Skip to content

Comment on How LLMs Work, Explained Without Mathparent

Comments

The problem here is that the actual Pyou isnt a distribution over all possible words.. that's just the observer's model of what's going on -- ie., your friend supposes that you could say anything.

Actually: (1) there's very very small number of possible words you are considering; (2) you arent considering 'words' at all, but future cognitive and sensory-motor states/vocalisation actions; (3) your vocalisation is moderated by a vast array of other causes/reasons (eg., being kind); etc.

The P(next word|previous, LLM-model, corpus...) for an LLM isnt an abstraction; it's actually implemented in its training.

The only sense in which, even in an instance, we seem to compute P(next|previous) is purely a radical abstraction which has nothing to do with any property we posess, but is an epistemic artefact from the outside.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.