Opus 5 is okay at coding. It is unbelievably awful to chat with, though. Very snippy, snarky, and it loves to "push back" even when it's inappropriate to do so. Besides, where "conceptual" things like science and math are concerned, it's starkly inferior to 5.6-Sol. Where writing prose is concerned, Kimi-K3 runs circles around it.
I refuse to believe Opus 5 is at the top of any non-cherrypicked benchmark, unless it has to do with very narrow coding tasks.
Past Opus 4.6 the models become absolute overfit trash, overfitting + their overconfidence makes them an existential threat to progress in niche areas, there are certain types of cutting edge crdts that are impossible to get a them to work on at all without erasing code. It's shockingly harmful.
Comments
Opus 5 scored highest, too, which is just lol
Opus 5 is okay at coding. It is unbelievably awful to chat with, though. Very snippy, snarky, and it loves to "push back" even when it's inappropriate to do so. Besides, where "conceptual" things like science and math are concerned, it's starkly inferior to 5.6-Sol. Where writing prose is concerned, Kimi-K3 runs circles around it.
I refuse to believe Opus 5 is at the top of any non-cherrypicked benchmark, unless it has to do with very narrow coding tasks.
Maybe it gets bonus points for constantly being honest about how it didn't actually finish what you asked it to do.
But it wont refund your tokens when it fails to do it so it loses extra points.
Past Opus 4.6 the models become absolute overfit trash, overfitting + their overconfidence makes them an existential threat to progress in niche areas, there are certain types of cutting edge crdts that are impossible to get a them to work on at all without erasing code. It's shockingly harmful.