Comment on OpenAI o1 Results on ARC-AGI-PubComments−meowface1yTakeaway:o1-preview is about on par with Anthropic's Claude 3.5 Sonnet in terms of accuracy but takes about 10X longer to achieve similar results to Sonnet.Scores:GPT-4o: 9%o1-preview: 21%Claude 3.5 Sonnet: 21%MindsAI: 46% (current highest score)−GaggiX1yThe takeaway is also that o1-preview is a major improvement compare to GPT-4o.Anthropic is just ahead.−disgruntledphd21yIt's a little embarrassing for OpenAI though?−dr_dshiv1yI think they’d expect as much from Dario. He designed GPT3…−meowface1yTrue. I've updated my post to include some of the scores.−Alifatisk1yHow the hell is Anthropic this far ahead? I am yet impressed−krackers1yThere were rumors that 3.5 Sonnet heavily used synthetic data for training, in the same way that OpenAI plans to use o1 to train Orion. Maybe this confirm it?−attentive1ywho knows what MindsAI is?
Comments
Takeaway:
Scores:
The takeaway is also that o1-preview is a major improvement compare to GPT-4o.
Anthropic is just ahead.
It's a little embarrassing for OpenAI though?
I think they’d expect as much from Dario. He designed GPT3…
True. I've updated my post to include some of the scores.
How the hell is Anthropic this far ahead? I am yet impressed
There were rumors that 3.5 Sonnet heavily used synthetic data for training, in the same way that OpenAI plans to use o1 to train Orion. Maybe this confirm it?
who knows what MindsAI is?