There is no issue with AI ingesting data from itself in itself. Humans do it as well. That data might even be higher quality than human data. The scale at which humans produce data will most likely stay higher than AI data for a long time.
There is already bot data out there from lower quality AIs/bots, and chatGPT has ingested it.
LLMs are made to be good at some textual tasks, and not for what they're being used right now. They're not information stores, or Q/A. It only answers what a human is likely to answer.
This made something click for me. There really is an issue with humans ingesting our own data. Humans have ended up with many mutually unintelligible languages. If two groups of people train on their own data they gradually start diverging.
But this gets at the heart of the issue - separation. If we ensure that humans and AI are trained on roughly the same data then we will stay connected and be able to understand each other. We may even end up borrowing a few gpt-isms, and that's actually totally fine.
Comments
Sounds like a /r/showerthoughts post.
There is no issue with AI ingesting data from itself in itself. Humans do it as well. That data might even be higher quality than human data. The scale at which humans produce data will most likely stay higher than AI data for a long time.
There is already bot data out there from lower quality AIs/bots, and chatGPT has ingested it.
LLMs are made to be good at some textual tasks, and not for what they're being used right now. They're not information stores, or Q/A. It only answers what a human is likely to answer.
This made something click for me. There really is an issue with humans ingesting our own data. Humans have ended up with many mutually unintelligible languages. If two groups of people train on their own data they gradually start diverging.
But this gets at the heart of the issue - separation. If we ensure that humans and AI are trained on roughly the same data then we will stay connected and be able to understand each other. We may even end up borrowing a few gpt-isms, and that's actually totally fine.