Skip to content

Comment on Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark

Comments

Is this AGI? I don't think I can score 100% on ARC AGI.

You'll find that those goalposts are very movable.

You could probably score 100% on ARC 3 if you were motivated enough. I find some of the current problems to be kind of like Chess - mechanically simple, and ~solvable, but it's difficult to force myself to think at length about a monotonous and meaningless problem. The machines do have an advantage on the "energy" front; they've become almost psychotically persistent (and don't get tired after too many prompts).

Anyway yes I think we've had AGI for a while now, even if the GI doesn't quite match up with what we expect from a human.

A 100% score means AI agents can beat every game as efficiently as humans. (0)

Yey, AGI is finally solved.

(0) https://arcprize.org/arc-agi/3

Is this AGI? I don't think I can score 100% on ARC AGI.

100% is some "RHAE" metric: its performance of median human first time seeing those problem.

It is until ARC-AGI-4. Maybe around 73 we will stop.

Yes, we’ve had AGI for years now.

Its interesting because I didn't think it was, but then reading the NVIDIA approach, this kind of loop plus generating a program to explain things. Maybe that is AGI? I don't know, but it seems like an additional layer that maybe is a fundamental shift in capabilities (kind of like reinforcement learning and COT was).

The capability for a man-made machine machine to solve problems drawn from arbitrary problem domains without domain specific pre-training: Artificial. General. Intelligence. AGI. It is what the term of art means.

Depends on which definition they’ll use today.

agi with a context of 250k-1mill tokens?

Could KITT? Cmdr. Data?

What's AGI, at the end of the day? Equivalence to the average human?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.