Skip to content

Comment on Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%parent

Comments

If this is counter-intuitive, a refresher on basic statistics and probability theory may be in order.

I'm not running "statistics". I'm running an individual run. I care about the individual quality of my run and not the general quality of the "aggregate".

The problem here is that the difference may not be immediately observable. Sure, if it doesn't give a correct answer, that's quickly catchable. If it costs me 10x the time, that's not immediately catchable but no less problematic.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.