Comment on Top AI models fail at >96% of tasksparentComments−kolinko7moThey didn't test Opus at all, only Sonnet.One of the tasks was "Build an interactive dashboard for exploring data from the World Happiness Report." -- I can't imagine how Opus4.5 could've failed that.−codexonOP7moCheck the link to the study. It has been updated for Opus 4.5.
Comments
They didn't test Opus at all, only Sonnet.
One of the tasks was "Build an interactive dashboard for exploring data from the World Happiness Report." -- I can't imagine how Opus4.5 could've failed that.
Check the link to the study. It has been updated for Opus 4.5.