From the abstract "Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations [...]".
If this is true and easily computable, this might have big impact in AI safety, as it seems to be really lacking today.
Comments
From the abstract "Further, our symbolic approximation allows us to modify an LLM's behavior in targeted ways via precise interventions on its internal representations [...]".
If this is true and easily computable, this might have big impact in AI safety, as it seems to be really lacking today.