Point is that the simplest way to excel in next token prediction in the way human consider correct - which is rated by how people feel the predictor mimics a human understanding - is to actually have a world model and other components of human understanding.
This is a speculative theory for why a next token predictor might sound like it knows what it's talking about. Not something we actually know.
Comments
This is a speculative theory for why a next token predictor might sound like it knows what it's talking about. Not something we actually know.