Right. But given all pairs of mainstream LLM combinations, it seems a model is more likely to say “yes I am X” when it is X than when it isn’t X, even if it still has a high chance of being wrong.
Which means you should (as a bayesian actor) update on it saying “I am X” as evidence it is X
Comments
Right. But given all pairs of mainstream LLM combinations, it seems a model is more likely to say “yes I am X” when it is X than when it isn’t X, even if it still has a high chance of being wrong.
Which means you should (as a bayesian actor) update on it saying “I am X” as evidence it is X