The problem here is that the actual Pyou isnt a distribution over all possible words.. that's just the observer's model of what's going on -- ie., your friend supposes that you could say anything.
Actually: (1) there's very very small number of possible words you are considering; (2) you arent considering 'words' at all, but future cognitive and sensory-motor states/vocalisation actions; (3) your vocalisation is moderated by a vast array of other causes/reasons (eg., being kind); etc.
The P(next word|previous, LLM-model, corpus...) for an LLM isnt an abstraction; it's actually implemented in its training.
The only sense in which, even in an instance, we seem to compute P(next|previous) is purely a radical abstraction which has nothing to do with any property we posess, but is an epistemic artefact from the outside.
Comments
The problem here is that the actual Pyou isnt a distribution over all possible words.. that's just the observer's model of what's going on -- ie., your friend supposes that you could say anything.
Actually: (1) there's very very small number of possible words you are considering; (2) you arent considering 'words' at all, but future cognitive and sensory-motor states/vocalisation actions; (3) your vocalisation is moderated by a vast array of other causes/reasons (eg., being kind); etc.
The P(next word|previous, LLM-model, corpus...) for an LLM isnt an abstraction; it's actually implemented in its training.
The only sense in which, even in an instance, we seem to compute P(next|previous) is purely a radical abstraction which has nothing to do with any property we posess, but is an epistemic artefact from the outside.