Skip to content

Comment on Absolute Zero: Reinforced Self-Play Reasoning with Zero Data

Comments

Pretty sure OpenAI and/or DeepMind have already been doing something very similar for a while already, just without publishing it.

Agreed, it's a pretty obvious solution to the problems once you are immersed in the problem space. I think it's much harder to setup an efficient training pipeline for this which does every single little detail in the pipeline correctly while being efficient.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.