Skip to content

Comment on OpenAI o1 Results on ARC-AGI-Pub

Comments

Takeaway:

o1-preview is about on par with Anthropic's Claude 3.5 Sonnet in terms of accuracy but takes about 10X longer to achieve similar results to Sonnet.

Scores:

GPT-4o: 9%
o1-preview: 21%
Claude 3.5 Sonnet: 21%
MindsAI: 46% (current highest score)

The takeaway is also that o1-preview is a major improvement compare to GPT-4o.

Anthropic is just ahead.

It's a little embarrassing for OpenAI though?

I think they’d expect as much from Dario. He designed GPT3…

True. I've updated my post to include some of the scores.

How the hell is Anthropic this far ahead? I am yet impressed

There were rumors that 3.5 Sonnet heavily used synthetic data for training, in the same way that OpenAI plans to use o1 to train Orion. Maybe this confirm it?

who knows what MindsAI is?

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.