But if an LLM were to say, "I'm sad" or "I'm happy", does that actually correlate to anything? I'm OK with saying "The LLM was sad", if there is an internal state that leads to observable changes in behavior correlating with the kinds of changes in behavior humans have when they're sad. The question is, if the LLM says "I'm sad", is that because it has that internal state (self-reflection)? Or is it because that's the kind of thing a human would say in that context?
If they are not trained specifically to find those internal state correlates when introspecting, I would find it quite shocking to see that introspection is an emergent behavior of LLMs
Comments
If they are not trained specifically to find those internal state correlates when introspecting, I would find it quite shocking to see that introspection is an emergent behavior of LLMs