Skip to content

Comment on How LLMs Work, Explained Without Mathparent

Comments

Are we afraid that this might happen to us too?

No, I'm more frustrated by the pseudoscience that models of frequency associations in text are explanations of people (or anything else). The choice isn't between a pseudoscientific behaviouralism where animals have no presence in the world, no mental faculties, and so on vs. "magic".

When I'm in a conversation I'm also selecting the optimal word from a predefined dictionary

Consider it this way: the probability dististribution over all possible world for you speaking is parameterized on space and time: Pyou(x, t; ...). And of the LLM generating text, Pllm(historical data).

So imagine plotting the live probability distributions of Pyou and Pllm for any given situation. As you think, imagine, move, recall, prefer, desire... the Pyou goes "wild" with dramatic discontinous shifts in distribution brought about by these causes.

Whereas the Pllm remains the same. It never changes. It never reacts to anything at all.

The whole distribution over all prior text tokens, Pllm is a stationary model of frequency associations. Yours is not. This makes all the difference in the world when claiming that Pllm somehow models, or is even relevant to, Pyou.

I think main tension here is that you think people confuse an LLM with a person having real world like experiences? I get that, and it's an interesting phenomenon, it's just a helpful way to talk to these models.

When I type in a query to ChatGPT, I know there's no real "you" in it, but it's just helpful for me to narrow down that Pllm probability space to give me a good enough answer I'm looking for. It's a helpful shortcut if I treat it like a person, but I should know it's a really good prediction machine which has access to most of the combined human knowledge in a latent space.

I rarely go "wild", people rarely go "wild" with their train of thoughts and actions. Everybody is moving on a more or less predefined set of rails. How long that rail is and where it goes depends on the person and situation.

Pyou is also limited in the sense that most of my outputs are history based, but all the sensory inputs I have gained over the years are now compressed into various neural circuits. I can act in surprising ways, but I'm more than a language model, although language is a huge part of me.

I'm more like a Pyou(xyz, t, Pllm, I) where "I" is the magic soup of human experience encoded in some gray matter.

You are right that most LLM-s are static, but we shouldn't ignore that complex systems (agents) can be built with LLM-s in which the language model is only an I/O layer to static knowledge. There could be other moving, dynamic parts which can be continuously updated and these parts can modify the behaviour of the system.

I think main tension here is that you think people confuse an LLM with a person having real world like experiences?

Many people do make this mistake, you see it all the time with people asking the LLM about itself or saying it is sentient based on text it outputs when asked about what it thinks etc.

Zoom in a bit. Freeze Pyou to a single conversation. A single spoken sentence, then another. There isn't enough time for emotions or imagination to shift. It's you in a particular situation, with some thought to express, words flowing out of your mouth.

I don't know about you, but to me this moment feels exaxtly like being an LLM.

I keep arguing that LLMs aren't similar to humans in entirety, but rather just to the "inner voice" - the bit that feeds your consciousness strings of words, which you utter, or consider, or send back if they make no sense.

The problem here is that the actual Pyou isnt a distribution over all possible words.. that's just the observer's model of what's going on -- ie., your friend supposes that you could say anything.

Actually: (1) there's very very small number of possible words you are considering; (2) you arent considering 'words' at all, but future cognitive and sensory-motor states/vocalisation actions; (3) your vocalisation is moderated by a vast array of other causes/reasons (eg., being kind); etc.

The P(next word|previous, LLM-model, corpus...) for an LLM isnt an abstraction; it's actually implemented in its training.

The only sense in which, even in an instance, we seem to compute P(next|previous) is purely a radical abstraction which has nothing to do with any property we posess, but is an epistemic artefact from the outside.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.