Skip to content

Comment on DeepSWE Measuring frontier coding agents

Comments

e2e4OP

gpt-5.5xhigh leading benchmark, coincides with my recent experience. I've been opus 4.7 user but it burns tokens so quickly, so gave gpt-5.5xhigh (via codex) a try, quality was similar (if not better), and tokens lasted a lot longer.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.