Comment on Claude Sonnet 5 – benchmark resultsComments−Tiberium2moSeems like the model is incredibly inefficient at max reasoning, and even at high/xhigh it uses far more tokens than other models, including Gemini 3.5 Flash, GLM 5.2 and so on. GPT 5.5's efficiency in tokens is still unmatched.See also: https://cursor.com/cursorbench−trentor2moSame with opus nothing above medium has a reasonable improvement for the tokens spent.
Comments
Seems like the model is incredibly inefficient at max reasoning, and even at high/xhigh it uses far more tokens than other models, including Gemini 3.5 Flash, GLM 5.2 and so on. GPT 5.5's efficiency in tokens is still unmatched.
See also: https://cursor.com/cursorbench
Same with opus nothing above medium has a reasonable improvement for the tokens spent.