Skip to content

Comment on First Proof

Comments

At this time, the competition is soon finishing - with no models having succeeded. Given the incentives for top labs, and the short time needed for a successful automated solution, we can make a reliable upper bound on the capability of current models - better than any normal benchmaxed datasets.

What I would like to see is an easier version of this same format.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.