Skip to content

Comment on Schema Harness Achieves ~99% on Arc‑AGI‑3 Publicparent

Comments

Benchmarks are meant to measure something, so can't be too hard else all measurements will be 0. At the same time the systems being tested - LLMs - are getting larger and more capable, at least in the narrow areas most benchmarks are focusing on, so all benchmarks will continually become saturated and need to be revised.

So, are you against all benchmarks or specifically ARC AGI? At least ARC AGI is trying to test for something a bit different and not play to the text prediction strength of LLMs. It should go without saying that no single test, or type of test, can claim to test for AGI or human level intelligence, which would require a suite of tests as broad and varied as the generality of intelligence you are trying to test for.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.