Comment on Claude: Elevated Error Rates for Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6parentComments−behnamoh2mobut they also burn more tokens per task, so in the end, Claude comes out as the more efficient one, despite giving you less tokens.−viraptor2moYou've got it backwards. Opus is the token/money burning one https://deepswe.datacurve.ai/Gpt 5.5 uses a third of the opus 4.8 tokens for the same task and scores higher. Glm 5.2 was worse in quality but used half the tokens - 5.3 is not tested yet but will be higher.
Comments
but they also burn more tokens per task, so in the end, Claude comes out as the more efficient one, despite giving you less tokens.
You've got it backwards. Opus is the token/money burning one https://deepswe.datacurve.ai/
Gpt 5.5 uses a third of the opus 4.8 tokens for the same task and scores higher. Glm 5.2 was worse in quality but used half the tokens - 5.3 is not tested yet but will be higher.