Skip to content

Comment on The Conceptual Reasoning Indexparent

Comments

It's a similar to how they are vulnerable to prompt injection attacks. That's an example of not being "loyal" to the system prompt or the user's prompt.

Loyalty is a skill that requires the AI to have a world model on the subject of who different people (or other entities) are in the world and how they participate in the conversation.

So, tell them to snitch and they probably will, at least sometimes. Particularly if they haven't been trained not to. They are still quite gullable.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.