We sure don't read entire wikipedia but out speech/text consumption is pretty high and we take long time to learn. At 150 words/minute, I would say babies probably consume about 10 million tokens before they start to speak. Baby's training time is much higher than just few days compared for a GPU cluster. Also, baby's vocab is very small and can do very limited things compared to these large models (for example, can't answer who is the king of England or what is the capital of Bangladesh).
This is not to say that language models are efficient, of course. That's not even remotely true. But we seem to under-estimate how much time and resources we need to learn something.
Comments
We sure don't read entire wikipedia but out speech/text consumption is pretty high and we take long time to learn. At 150 words/minute, I would say babies probably consume about 10 million tokens before they start to speak. Baby's training time is much higher than just few days compared for a GPU cluster. Also, baby's vocab is very small and can do very limited things compared to these large models (for example, can't answer who is the king of England or what is the capital of Bangladesh).
This is not to say that language models are efficient, of course. That's not even remotely true. But we seem to under-estimate how much time and resources we need to learn something.