Comment on Claude Opus 4.6 accuracy on BridgeBench hallucination test drops from 83% to 68%parentComments−coldtea4moWhich is a useless distinction. When we say models in this context we mean the whole LLM + infrastructure to serve it (including caches, etc).
Comments
Which is a useless distinction. When we say models in this context we mean the whole LLM + infrastructure to serve it (including caches, etc).