Skip to content

Comment on Self-Adapting Language Models

Comments

From Anthropic a couple days ago too, self finetuning:

https://arxiv.org/html/2506.10139v1

This is wild!

"when assessed by Claude 3.5 Sonnet’s production-grade RM, our unsupervised assistant policy wins 60% of head-to-head comparisons against the policy trained with the human-supervised RM." So now the models can even post-train the new models better than a human can

Everytop model in ARC AGI used a test time finery king approach. They they had one example pair though and would usually do transformations (color, mirroring, etc) of it for the finetuning, and that might have been coded by hand

Related ongoing thread:

Unsupervised Elicitation of Language Models - https://news.ycombinator.com/item?id=44276041

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.