The trivial idea that there are some benchmarks for a tool and there can be even better inexplicit ones, and if the tool is capable of modifying itself towards better benchmarks results there could be a recursive climbing of those benchmarks through iterations of tool versions which are outputs of former tool versions.
Improvement means being able to do more complicated things more reliably. Relatedly, it means being able to learn to do new things with fewer and fewer examples. We are running out of easily verifiable or simulation-friendly or data-rich domains for LLMs to conquer. (Note that I didn't say "simple" or "easy" domains.)
I suppose the next (and more risky) step is to let AI conduct its own real-world experiments, so that it can generate data to learn more physical and social properties.
It is doing that already - every day 1B people or more use AI for the tune of a few trillion tokens. Imagine that much language flowing between brains and AI agents. It carries real our world problems to AI, their solutions back to us, and we act in the world and come back for more AI iteration. In the end AI gets inside the loop of real world actions and their consequences. AI logs stretch over years, tracking downstream effects. Hindsight can be used to track consequences of prior actions. It's a real data loop, an experience engine.
This same process has recently been under scrutiny when mathematicians claimed AI companies trained on their unpublished logs and later claimed merit for results. But the exchange of experience happens across all domains. Experience gets generated at amazing rates, and absorbed by models which get applied everywhere, collecting more experience.
Comments
Can someone define “improvement” in this context? This concept feels like a buzz word otherwise.
The trivial idea that there are some benchmarks for a tool and there can be even better inexplicit ones, and if the tool is capable of modifying itself towards better benchmarks results there could be a recursive climbing of those benchmarks through iterations of tool versions which are outputs of former tool versions.
Improvement means being able to do more complicated things more reliably. Relatedly, it means being able to learn to do new things with fewer and fewer examples. We are running out of easily verifiable or simulation-friendly or data-rich domains for LLMs to conquer. (Note that I didn't say "simple" or "easy" domains.)
I suppose the next (and more risky) step is to let AI conduct its own real-world experiments, so that it can generate data to learn more physical and social properties.
It is doing that already - every day 1B people or more use AI for the tune of a few trillion tokens. Imagine that much language flowing between brains and AI agents. It carries real our world problems to AI, their solutions back to us, and we act in the world and come back for more AI iteration. In the end AI gets inside the loop of real world actions and their consequences. AI logs stretch over years, tracking downstream effects. Hindsight can be used to track consequences of prior actions. It's a real data loop, an experience engine.
This same process has recently been under scrutiny when mathematicians claimed AI companies trained on their unpublished logs and later claimed merit for results. But the exchange of experience happens across all domains. Experience gets generated at amazing rates, and absorbed by models which get applied everywhere, collecting more experience.
That sounds terrifying.
Which means someone is already working on it.
Models that are strong enough to improve themselves without a human (I.e. ai researcher) in the loop