Skip to content

Comment on Qodo CLI agent scores 71.2% on SWE-bench Verifiedparent

Comments

It's really not that hard to not build a custom bench setup to game the benchmark instead of just using your product straight out of the box, though.

Right, other than financial pressure. Which is, of course, immense.

Right. Building a custom setup is blatant- that will wildly overfit.

But let's say a group uses it as a metric as part of CI and each new idea / feature they create runs against SWE bench. Maybe they have parameterized bits and pieces they adjust, maybe they have multiple candidates datasets for fine tuning, maybe they're choosing between checkpoints.

This will also end up overfitting - especially if done habitually. It might be a great metric and result in a more powerful overall model. Or it might not.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.