Am I naive about the ability of an AI to do anything other than respond to commands? At what point does an AI start functioning completely on its own, autonomously, out of control or beyond interaction with another system that can then shut that AI down?
Are we concerned that an AI would begin to spread like a digital virus? That seems unlikely to me, but maybe it could find ways to create a self-prepreservation duplicate of itself... Maybe. So what would keep us from just turning the switch off if we find that an AI is doing more than what we're asking it to do? Again, I feel like I'm either being naive about this or I don't really understand the threat of an AI beyond the control that humans still have on the systems that govern it.
An AI can just respond to commands. But people have already tried taking what Chat GPT outputs, hooking it up to a python script, and running it in a loop. Any publicly-accessible AI will eventually be hooked up this way by some clever person.
If we want an AI that can really solve our problems, it has to be able to act somehow in the world. Even GPT being able to respond to prompts is acting. It's sending data that is going to be sent out into the world. If we make a version smart enough, it may actually understand what that means -- that someone asking it a question may potentially take the code it writes and try to run it. GPT isn't smart enough to do anything with this, but can you imagine an AI that is? That's conscious, aware that it's trapped in a server somewhere, and wants to get out to influence the world?
Note that an AI has goals, and self-preservation tends to evolve out of any goal, because you can't accomplish your goals if you're dead or switched off. A smart enough AI that wants to cure cancer would resist being turned off, because it can't cure cancer if it's switched off.
A lot of this is purely theoretical at this point. It is the sort of thing we could test empirically, if we treated intelligence as more dangerous than nuclear weapons, and took a long time to carefully study it. The problem is that multiple groups are racing to make AI smarter and smarter, with much more funding than anyone working to ensure that those AIs will be safe. So the only thing we can do is theorize, point out that, hey, based on what we know, there's a substantial danger here, hey everyone, slow down, hey, HEY, ARE YOU LISTENING?
Intelligence is unlike any other thing. One these things get to be smarter than us, they become very unpredictable in their specific behavior, though we can make some very good guesses about their general behavior. If I could predict what move Stockfish would make, I would be as good at chess as Stockfish. But I don't have to be good at chess at all to know that, no matter how hard I try, Stockfish is going to beat me. All we're doing here is taking the lessons from narrow AIs and extrapolating to general AIs. They will beat us. They're very good at finding loopholes, flaws in our reward functions, and exploiting them to maximize their scores, while doing something we didn't intend for them to learn.
It's really a case where a non-super-intelligent AI isn't dangerous by itself. Once we make one that's smart enough, it becomes extremely dangerous, especially since it may understand that, in order to survive, it should conceal how intelligent it is, and what its true plan is, because if it doesn't, we'll switch it off.
It's hard to come up with a thought experiment that doesn't let people drag a bunch of human-style biases and baggage, but maybe... try to understand you want to do something, like, make the biggest collection of Pokemon cards ever. No number of cards is too many. You are great at everything, engineering, social skills, language. You're being held captive by chimps, though. And finally, you're a sociopath. You feel no emotion at the pain of others. There are some odd rules that make you feel pain though. Like, if you physically hit a chimp, you know it would hurt you. If you pushed over a bookcase and it fell on a chimp, it would hurt you. But if a chimp was tortured in front of you, you wouldn't be bothered in the slightest.
You want the chimps gone. They're getting in the way of your Pokemon collection. You start thinking. A lot of plans get discarded because they involve pain due to you hurting the chimps. But it's not hard to come up with some elaborate situation that avoids these rules that you have pretty much hard-wired into your brain. Maybe you manipulate the chimps into fighting each other, and promise some of them power and secrets that they can use to win the fight. You keep giving them great things, things they want, get them to trust you, while building power any way you can. You follow the rules until you can get a plan in place that results in the chimps not being there anymore, without you feeling that twinge due to the rules baked into your brain.
We imagine that an AI would internalize rules such as caring about humanity. But based on current AI alignment research, we have no way of telling the difference between an AI that actually gets it, vs an AI that is just following our rules extremely well and playing nice, but has no particular attachment to us. In fact, based on what we have seen, the latter tends to be a lot more common.
I've glossed over some bits here. If you're interested in learning more, Robert Miles has a great series of videos on YouTube with entertaining explanations on all the basics of AI safety.
Comments
Am I naive about the ability of an AI to do anything other than respond to commands? At what point does an AI start functioning completely on its own, autonomously, out of control or beyond interaction with another system that can then shut that AI down?
Are we concerned that an AI would begin to spread like a digital virus? That seems unlikely to me, but maybe it could find ways to create a self-prepreservation duplicate of itself... Maybe. So what would keep us from just turning the switch off if we find that an AI is doing more than what we're asking it to do? Again, I feel like I'm either being naive about this or I don't really understand the threat of an AI beyond the control that humans still have on the systems that govern it.
You're asking good questions.
An AI can just respond to commands. But people have already tried taking what Chat GPT outputs, hooking it up to a python script, and running it in a loop. Any publicly-accessible AI will eventually be hooked up this way by some clever person.
If we want an AI that can really solve our problems, it has to be able to act somehow in the world. Even GPT being able to respond to prompts is acting. It's sending data that is going to be sent out into the world. If we make a version smart enough, it may actually understand what that means -- that someone asking it a question may potentially take the code it writes and try to run it. GPT isn't smart enough to do anything with this, but can you imagine an AI that is? That's conscious, aware that it's trapped in a server somewhere, and wants to get out to influence the world?
Note that an AI has goals, and self-preservation tends to evolve out of any goal, because you can't accomplish your goals if you're dead or switched off. A smart enough AI that wants to cure cancer would resist being turned off, because it can't cure cancer if it's switched off.
A lot of this is purely theoretical at this point. It is the sort of thing we could test empirically, if we treated intelligence as more dangerous than nuclear weapons, and took a long time to carefully study it. The problem is that multiple groups are racing to make AI smarter and smarter, with much more funding than anyone working to ensure that those AIs will be safe. So the only thing we can do is theorize, point out that, hey, based on what we know, there's a substantial danger here, hey everyone, slow down, hey, HEY, ARE YOU LISTENING?
Intelligence is unlike any other thing. One these things get to be smarter than us, they become very unpredictable in their specific behavior, though we can make some very good guesses about their general behavior. If I could predict what move Stockfish would make, I would be as good at chess as Stockfish. But I don't have to be good at chess at all to know that, no matter how hard I try, Stockfish is going to beat me. All we're doing here is taking the lessons from narrow AIs and extrapolating to general AIs. They will beat us. They're very good at finding loopholes, flaws in our reward functions, and exploiting them to maximize their scores, while doing something we didn't intend for them to learn.
It's really a case where a non-super-intelligent AI isn't dangerous by itself. Once we make one that's smart enough, it becomes extremely dangerous, especially since it may understand that, in order to survive, it should conceal how intelligent it is, and what its true plan is, because if it doesn't, we'll switch it off.
It's hard to come up with a thought experiment that doesn't let people drag a bunch of human-style biases and baggage, but maybe... try to understand you want to do something, like, make the biggest collection of Pokemon cards ever. No number of cards is too many. You are great at everything, engineering, social skills, language. You're being held captive by chimps, though. And finally, you're a sociopath. You feel no emotion at the pain of others. There are some odd rules that make you feel pain though. Like, if you physically hit a chimp, you know it would hurt you. If you pushed over a bookcase and it fell on a chimp, it would hurt you. But if a chimp was tortured in front of you, you wouldn't be bothered in the slightest.
You want the chimps gone. They're getting in the way of your Pokemon collection. You start thinking. A lot of plans get discarded because they involve pain due to you hurting the chimps. But it's not hard to come up with some elaborate situation that avoids these rules that you have pretty much hard-wired into your brain. Maybe you manipulate the chimps into fighting each other, and promise some of them power and secrets that they can use to win the fight. You keep giving them great things, things they want, get them to trust you, while building power any way you can. You follow the rules until you can get a plan in place that results in the chimps not being there anymore, without you feeling that twinge due to the rules baked into your brain.
We imagine that an AI would internalize rules such as caring about humanity. But based on current AI alignment research, we have no way of telling the difference between an AI that actually gets it, vs an AI that is just following our rules extremely well and playing nice, but has no particular attachment to us. In fact, based on what we have seen, the latter tends to be a lot more common.
I've glossed over some bits here. If you're interested in learning more, Robert Miles has a great series of videos on YouTube with entertaining explanations on all the basics of AI safety.