Skip to content

Comment on Quantitative AI progress needs accurate and transparent evaluation

Comments

For instance, if a cutting-edge AI tool can expend $1000 worth of compute resources to solve an Olympiad-level problem, but its success rate is only 20%, then the actual cost required to solve the problem (assuming for simplicity that success is independent across trials) becomes $5000 on the average (with significant variability). If only the 20% of trials that were successful were reported, this would give a highly misleading impression of the actual cost required (which could be even higher than this, if the expense of verifying task completion is also non-trivial, or if the failures to solve the goal were correlated across iterations).

This is a very valid point. Google and ChatGPT announced they got the gold medal with specialized models, but what exactly does that entail? If one of them used a billion dollars in compute and the other a fraction of that, we should know about it. Error rates are equally important. Since there are conflicts of interest here, academia would be best suited for producing reliable benchmarks, but they would need access to closed models.

Compute has been getting cheaper and models more optimised. So if models can do something it will not be long till they can do this cheap.

GPU compute per watt has grown by a factor of 2 in last 5 years

with specialized models
what exactly does that entail

Overfitting on the test set with models that are useless for anything else, that's what.

Arguably we're not that far from jumping up one level, having an AI agent that when encountering any new type of hard problem would train a new dedicated sub-AI to solve that.

Deep learning scientists automating themselves? Ironic.

Don't put Google and ChatGPT in the same category here. Google cooperated with the organizers, at least.

Also neither got a gold medal. Both solved problems to meet the threshold for a human child getting a gold medal but it’s like saying an F1 car got a gold medal in the 100m sprint at the Olympics.

The popular science title was funnier with a pun on "mathed" [1]

"Human teens beat AI at an international math competition Google and OpenAI earned gold medals, but were still out-mathed by students."

[1] https://www.popsci.com/technology/ai-math-competition/

Indeed, it’s like saying a jet plane can fly!

"Google F1 Preview Experimental beat the record of the fastest man on earth Usain Bolt"

Could you clarify what you mean by this?

Google's answers were judged by IMO. OpenAI's were judged by themselves internally. Whether it matters is up to the reader.

TheZvi had a summarization of this here: https://thezvi.substack.com/i/168895545/not-announcing-so-fa...

In short (there is nuance), Google cooperated with the IMO team while OpenAI didn't which is why OpenAI announced before Google.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.