Comment on Some critical issues with the SWE-bench datasetparentComments−siva71yo3-mini and gpt-4o are so piss poor in agent coding compared to claude that you don't even need a benchmark−jbellis1yo3-mini-medium is slower than claude but comparable in quality. o3-mini-high is even slower, but better.−danielbln1yClaude really is a step above the rest when it comes to agentic coding.−dr_kiszonka1yWhen I used it with Open Hands it was great but also quite expensive (~$8/hr). In Trea, it was pretty bad, but free. Maybe it depends on how the agents use it? (I was writing the same piece of software, a simple web crawler for a hobby RAG project.)
Comments
o3-mini and gpt-4o are so piss poor in agent coding compared to claude that you don't even need a benchmark
o3-mini-medium is slower than claude but comparable in quality. o3-mini-high is even slower, but better.
Claude really is a step above the rest when it comes to agentic coding.
When I used it with Open Hands it was great but also quite expensive (~$8/hr). In Trea, it was pretty bad, but free. Maybe it depends on how the agents use it? (I was writing the same piece of software, a simple web crawler for a hobby RAG project.)