Comment on OpenAI o1 Results on ARC-AGI-PubparentComments−riku_iki1yI think huge advantage is that they keep eval tests private, so corps can't finetune them to model and claim breakthrough, which possibly happened with many other benchmarks.
Comments
I think huge advantage is that they keep eval tests private, so corps can't finetune them to model and claim breakthrough, which possibly happened with many other benchmarks.