Would love to learn more about some techniques that "everybody" uses to do this well. So far, everything I've seen that meaningfully advances the frontier has been high-touch (involving human experts in one way or another).
It's fairly easy to describe a task that is slightly harder than an existing one.
For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.
My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?"
Comments
Sure?
Doesn't everyone get their agents to construct evals it can't pass? There's nothing magical about this.
Would love to learn more about some techniques that "everybody" uses to do this well. So far, everything I've seen that meaningfully advances the frontier has been high-touch (involving human experts in one way or another).
It's fairly easy to describe a task that is slightly harder than an existing one.
For example if frontier models are able to one-shot a database query across 20 columns and 10 tables add one additional relationship then test. Keep doing this until the pass-rate drops below acceptable and now you have your new frontier eval.
I see, we're talking about different things.
My thought experiment was along the lines of "Let's say I'm Anthropic and I want to significantly improve my frontier model's performance on, say, theoretical physics research. How do I build a fully autonomous process capable of constructing an eval that's somewhat outside the current capability in some useful direction (decided by the autonomous process itself)?"
Would love to hear folks' ideas. :)