Skip to content

Comment on Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark

Comments

The blog post: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-...

Using Claude Opus 5, but it can use others:

AVO is also designed to operate across frontier models. While our full public-set result used Claude Opus 5, we additionally paired AVO with GPT-5.6 Sol on a challenging subset of games. In these limited experiments, Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions in matched-level comparisons. These preliminary results suggest complementary operating profiles across models, and we leave a broader systematic comparison to future work

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.