Skip to content

Comment on Claude: Elevated Error Rates for Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6parent

Comments

but they also burn more tokens per task, so in the end, Claude comes out as the more efficient one, despite giving you less tokens.

You've got it backwards. Opus is the token/money burning one https://deepswe.datacurve.ai/

Gpt 5.5 uses a third of the opus 4.8 tokens for the same task and scores higher. Glm 5.2 was worse in quality but used half the tokens - 5.3 is not tested yet but will be higher.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.