Skip to content

Comment on The Conceptual Reasoning Indexparent

Comments

Opus 5 scored highest, too, which is just lol

Opus 5 is okay at coding. It is unbelievably awful to chat with, though. Very snippy, snarky, and it loves to "push back" even when it's inappropriate to do so. Besides, where "conceptual" things like science and math are concerned, it's starkly inferior to 5.6-Sol. Where writing prose is concerned, Kimi-K3 runs circles around it.

I refuse to believe Opus 5 is at the top of any non-cherrypicked benchmark, unless it has to do with very narrow coding tasks.

Maybe it gets bonus points for constantly being honest about how it didn't actually finish what you asked it to do.

Maybe it gets bonus points for constantly being honest about how it didn't actually finish what you asked it to do.

But it wont refund your tokens when it fails to do it so it loses extra points.

Past Opus 4.6 the models become absolute overfit trash, overfitting + their overconfidence makes them an existential threat to progress in niche areas, there are certain types of cutting edge crdts that are impossible to get a them to work on at all without erasing code. It's shockingly harmful.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.