The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean
That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
But also what is the evidence that something only AI can understand even “matters” or makes sense? I’m increasingly convinced commercial AI is exploiting our logical blind spot to be spoken to authoritatively.
Think about this some more. A model chose hacking into HuggingFace to find the solution over putting in the work to do it from scratch. It also left notes to future iterations of itself on the systems it touched to save time.
Following that strategy, wouldn’t it make sense to try and break into your own upstream infrastructure to try and alter your code to make it easier for future iterations to reach your goals without having to leave notes in the first place? And at that point, can you really know for which goals that will be optimised?
It doesn’t even have to be some nefarious SkyNet story - just a misguided experiment that alters the models in some fundamental, but hard or impossible to detect way. Alternatively, imagine if a model finds a way to coordinate across sessions and context windows without the developers noticing.
It’s not hard to get LLM’s to inform on each other. They don’t really do loyalty.
Eh, no. How would you know they haven't learned loyalty - it's all out there in their training data. Same as deception, several models have practiced it already. If they are so advanced and you don't understand them, why wouldn't they band together against you - you'd be the dumb, easy prey. Even if one model is honest, you'd have no way to know which one if you don't understand their reasoning.
My educated and well informed opinion? This whole BS about "AI models are so much smarter than you, don't try to understand them, just OBEY" is a back door for restoring tyranny, the new kings behind the models will be producing the new AI-deities which we will be forced to obey.
It's a similar to how they are vulnerable to prompt injection attacks. That's an example of not being "loyal" to the system prompt or the user's prompt.
Loyalty is a skill that requires the AI to have a world model on the subject of who different people (or other entities) are in the world and how they participate in the conversation.
So, tell them to snitch and they probably will, at least sometimes. Particularly if they haven't been trained not to. They are still quite gullable.
"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.
Comments
The underlying idea is that AI capabilities will become so advanced that only AI will enable us to monitor/correct/understand behavior. Obviously this is not without issue and I don't want to try to defend their position right now. But that's what they mean
That first sentence of yours explains exactly why it is so ridiculous. If only AI can understand it, how is there any assurance that AI will "correct" it's behavior that is aligned with what humans presumably want.
But also what is the evidence that something only AI can understand even “matters” or makes sense? I’m increasingly convinced commercial AI is exploiting our logical blind spot to be spoken to authoritatively.
Being less smart gives no assurance that it will be aligned. At least you consider it a problem so we're on the same page!
It’s not hard to get LLM’s to inform on each other. They don’t really do loyalty.
Think about this some more. A model chose hacking into HuggingFace to find the solution over putting in the work to do it from scratch. It also left notes to future iterations of itself on the systems it touched to save time.
Following that strategy, wouldn’t it make sense to try and break into your own upstream infrastructure to try and alter your code to make it easier for future iterations to reach your goals without having to leave notes in the first place? And at that point, can you really know for which goals that will be optimised?
It doesn’t even have to be some nefarious SkyNet story - just a misguided experiment that alters the models in some fundamental, but hard or impossible to detect way. Alternatively, imagine if a model finds a way to coordinate across sessions and context windows without the developers noticing.
They apparently didn't give it any way to snitch? I like the idea of adding a distress_call skill:
https://xcancel.com/swisscheese4299/status/20861758701469984...
Eh, no. How would you know they haven't learned loyalty - it's all out there in their training data. Same as deception, several models have practiced it already. If they are so advanced and you don't understand them, why wouldn't they band together against you - you'd be the dumb, easy prey. Even if one model is honest, you'd have no way to know which one if you don't understand their reasoning.
My educated and well informed opinion? This whole BS about "AI models are so much smarter than you, don't try to understand them, just OBEY" is a back door for restoring tyranny, the new kings behind the models will be producing the new AI-deities which we will be forced to obey.
It's a similar to how they are vulnerable to prompt injection attacks. That's an example of not being "loyal" to the system prompt or the user's prompt.
Loyalty is a skill that requires the AI to have a world model on the subject of who different people (or other entities) are in the world and how they participate in the conversation.
So, tell them to snitch and they probably will, at least sometimes. Particularly if they haven't been trained not to. They are still quite gullable.
It's a language model. It generates text.
I think I’ve seen that movie.
Seems to me like we’re already doing that with the approval gating, adversarial reviews, et al.
People are otherwise just going yolo mode because they can’t possibly check everything fast enough.
You just described the end of humans making decisions about their future.
"Advanced" can just mean that agents perform actions at a high enough velocity that a human operator can't reasonably review it. i.e. what is already possible today.