Skip to content

Comment on Testing Moonshot AI's Kimi K3 Inside Claude Codeparent

Comments

There's a very interesting benchmark comparison below of Opus 4.7 run under three different harnesses : OpenCode, Cursor and Claude Code where it's not very close at all and Opus's native harness, Claude Code, performs worst of all three.

The pass@1 scores are 50/45/40 for OpenCode/Cursor/Claude Code respectively.

https://artificialanalysis.ai/agents/coding-agents#harness-c...

I've seen other benchmarks where Pi also outperforms Claude Code both in model performance and in much reduced token usage.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.