I disagree. People are arriving at this conclusion after careful consideration and logical reasoning, not picking up a '50s sci-fi book and mistaking it for reality.
The Less Wrong community isn't a community of careful consideration and logical reasoning. It's a cult that believes their own beliefs so fervently they would burn the earth for it. Small communities like this inevitably become so insular they recycle apocalyptic beliefs.
There's no evidence that stochastic parrots are evidence of true AI. There's no evidence that humans can, by practice, subordinate emotions and become purely rational thought machines as Less Wrong people or "shape rotators" believe. Even such titles are counter-evidence of such beliefs. The human condition is so bound up in feelings, it's unclear where thinking begins and feelings end. Thoughts arise from contact, contact which produces feelings. Gut microbes are as wrapped up in cognition as the brain itself.
Any out-group like this is bound to be hopelessly entrapped in its own cognitive fallacies. Yudkowsky is no exception.
EDIT: Before anyone says but Yudkowsky already gets ahead of your argument with
None of this danger depends on whether or not AIs are or can be conscious; it’s intrinsic to the notion of powerful cognitive systems that optimize hard and calculate outputs that meet sufficiently complicated outcome criteria.
It's one sentence that handwaves away what GPT-4 and its ilk are, which is not AI. It isn't synthetic intelligence because intelligence isn't what's happening on a fundamental level.
Wait, but we don't actually even know what is happening inside GPT-4 on a fundamental level, to produce the output we see. We don't even know what is happening inside our own brains, really. How do the neurons turning on and off produce reasoning and consciousness? So in the same way, how can you say that GPT-4 is not AI, definitively, or at least a primordial form of one? Clearly, there's some arrangement of neurons that produces intelligence.
Maybe it's just a "stochastic parrot", but one can probably make a similarly dismissive-sounding and yet accurate description of how humans cognate. Sometimes quantity has its own quality. Maybe a big enough stochastic parrot becomes smart.
Wait, but we don't actually even know what is happening inside GPT-4 on a fundamental level, to produce the output we see.
This just isn't true. All of this research descends from transformer research that came from Google in 2017. We know exactly how they work. There's nothing surprising in what's going on.
I spent a bit of time looking for a decent intro to all this because I wanted to at least provide a resource for people who are terrified of LLMs. I think this is a pretty decent one [1]
Ok, after reading through it, I don't think I was speaking inaccurately. Here are some quotes from the blog post (which was very neat, by the way, thanks again)
And it’s part of the lore of neural nets that—in some sense—so long as the setup one has is “roughly right” it’s usually possible to home in on details just by doing sufficient training, without ever really needing to “understand at an engineering level” quite how the neural net has ended up configuring itself.
What determines this structure? Ultimately it’s presumably some “neural net encoding” of features of human language. But as of now, what those features might be is quite unknown. In effect, we’re “opening up the brain of ChatGPT” (or at least GPT-2) and discovering, yes, it’s complicated in there, and we don’t understand it—even though in the end it’s producing recognizable human language.
This is the kind of thing I'm referring to. Even though we can look at pictures of the neuron activations, we don't really know what it's doing, any more than you can look at a picture of a brain scan of a person speaking and know why they decided to say those particular words. We know at a low level how the network works, of course, because we coded it. It's the emergent behavior that we don't understand, and that's the bit that's scary because that's where the AI risk lives. It's like how we understand particle physics pretty well, and chemistry to some degree, but biology is a massively complicated jungle that we've barely scratched the surface of.
Maybe there's something we could develop analogous to using an MRI as a lie detector, for these networks. But as far as I know, this is still an unsolved problem, and apparently really hard for networks that are smarter than you are: https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H...
My understanding was that we're not sure of how some of the things that look like reasoning are coming about, but I'll gladly check it out. I may be wrong.
I'm not particularly terrified of GPT-4, but a more of what comes after. Maybe GPT-6 or 8, or another breakthrough non-LLM AI. I'm not sure if reading it will do much more than push my estimate of real AGI out a couple of years, but thanks anyway, I appreciate it.
There are lots of weird and crazy people reading and writing on Less Wrong, for sure. Do you think that, without gut microbes and emotions, AI will never become intelligent enough to be existentially dangerous? Are those essential to its ability to outsmart humans?
I think we're asking the wrong questions. We're assuming that a. what we're developing (an associative propability machine) is in any way related to the pursuit of real artificial intelligence. We're assuming that b. motive, desire, passion, some motive force is inhered into the thing.
I, personally, don't think we've seen anything that can be dangerous independent of its wielding by a badly acting human. That is to say, human agency and acting force is required on the other side of the LLM to do bad actions, like any other tool.
I don't know what real AI that threatens humans with an extinction level event looks like. It's not large language models. I don't know if humans are capable of generating it. I think anthropomorphization is really this massive veil that hangs over all of us. We can't get outside of it (our essential fallacy) so we're cursed to see it everywhere.
I'm fairly certain that an LLM alone will not independently start taking actions and pursuing goals and causing harm. That's not how LLMs work, as you and I both know.
I'm not confident I can say the same about, say: an LLM calling itself from within an infinite loop in which the LLM is asked to predict actions that will [eg. add money into a bank account], and then to write code that executes those actions, then to examine the results of those actions and update its plan. (This takes no additional breakthrough in AI architecture, only an increase in the reliability of GPT for it to become practical. It is basically a summary of the AutoGPT project).
I don't know if LLMs are anywhere close to any part of what goes on inside the human thought process. But I am sure that evolution was able to stumble upon the human thought process. In 2003 I was quite confident that was forever beyond the reach of humans to program into a computer, but in 2023 I am far less certain. I think human cognition is gradually becoming less of a mystery.
I can probably offer one, if I understand what aspect you're skeptical of. Are you skeptical that an AI would ever be smart enough to come up with a plan to cause harm? That it would ever try to come up with a plan to cause harm? That it could execute such a plan, without arms and legs and a body? Or that, after coming up with a plan to cause harm, it wouldn't be trivial for humans to stop by simply unplugging the computer?
Second of all, it is not even nearly powerful enough to convince anyone to do anything, or anything like that. If it was 100 times more powerful, then maybe.
I can see current problems being flooding spam/misinformation/propaganda on internet forums. But that has been happening for a while and this will make people more skeptical of what they read (hopefully)
But finally, if it does cause harm (don’t think it will. great tool). Then who cares? Why try to stop it. You can’t stop it, nobody can. Why whinge and moan like the luddites? I don’t get it.
Why cause harm? Basically, by accident. If you try to solve a problem really hard, and you're really good at it, and you aren't also trying very hard to preserve a bunch of things that are important to humans, you're very likely to destroy a lot of things that are important to humans in the process.
100 times the computational power may be... 2 years out, at the rate we're going. GPT-4 has about 500 times as many parameters as GPT-3, and that was released 3 years ago. The amount of power people are dedicating to these is following an exponential track.
It's not about it causing harm -- I think most people have accepted that. It's that we may make an AI that's so smart, so good at everything, that trying to control it is like me playing chess against Stockfish. If it gets to that point, and it happens to have learned a goal / set of goals that aren't 100% friendly to humanity (which seems like an inevitability based on our current abilities at AI alignment) then it will optimize those goals so hard that we all die in the process. We'll end up with a planet with all silicon turned into solar panels and computing power, and probably all other elements put to the use of the AI as well.
OK, so to be clear, I don't think this AI will harm us, not in its current configuration. The author of the article says the same. The concern is about a future AI, one which is capable of the sort of reasoning that humans can do. The main danger with today's AI is that they will be so commercially successful that it will drive enough investment into AI research that a more powerful AI is constructed.
Sure, I get that. But this has been worked on for decades in the background and only achievable now. It’s amazing and new to so many people, but this “sparks of agi” etc is all hype and honestly i think more damaging to people than ai possibly could be. You could suppose the earth would implode if we kept walking on it too, but nobody is going to stay sitting.
Maybe it's just a matter of perspective. I agree that this is something that's been worked on for decades; however, it's also the case that progress seems to be getting faster. The AI field now has a problem where AIs are improving faster than benchmarks can measure them, and new benchmarks get saturated almost as soon as they're introduced. [1]
The human mind is a mystery, we don't know how it works. But it works somehow. There are some algorithms it's using under the hood, as mysterious as they may be. We might be searching blindly in the dark, grabbing hold of anything that feels promising, and we're probably still far away from the "right way" of doing it, of getting a computer to think the way humans think. And I'm sure we won't find the whole thing at once; maybe we'll have reverse-engineered the visual cortex, but not the vestibular system, which perhaps works completely differently. But if we were getting closer to stumbling on the right solution, that rapid saturation of benchmarks is the kind of behaviour I would expect to see.
Comments
I disagree. People are arriving at this conclusion after careful consideration and logical reasoning, not picking up a '50s sci-fi book and mistaking it for reality.
The Less Wrong community isn't a community of careful consideration and logical reasoning. It's a cult that believes their own beliefs so fervently they would burn the earth for it. Small communities like this inevitably become so insular they recycle apocalyptic beliefs.
There's no evidence that stochastic parrots are evidence of true AI. There's no evidence that humans can, by practice, subordinate emotions and become purely rational thought machines as Less Wrong people or "shape rotators" believe. Even such titles are counter-evidence of such beliefs. The human condition is so bound up in feelings, it's unclear where thinking begins and feelings end. Thoughts arise from contact, contact which produces feelings. Gut microbes are as wrapped up in cognition as the brain itself.
Any out-group like this is bound to be hopelessly entrapped in its own cognitive fallacies. Yudkowsky is no exception.
EDIT: Before anyone says but Yudkowsky already gets ahead of your argument with
It's one sentence that handwaves away what GPT-4 and its ilk are, which is not AI. It isn't synthetic intelligence because intelligence isn't what's happening on a fundamental level.
Wait, but we don't actually even know what is happening inside GPT-4 on a fundamental level, to produce the output we see. We don't even know what is happening inside our own brains, really. How do the neurons turning on and off produce reasoning and consciousness? So in the same way, how can you say that GPT-4 is not AI, definitively, or at least a primordial form of one? Clearly, there's some arrangement of neurons that produces intelligence.
Maybe it's just a "stochastic parrot", but one can probably make a similarly dismissive-sounding and yet accurate description of how humans cognate. Sometimes quantity has its own quality. Maybe a big enough stochastic parrot becomes smart.
This just isn't true. All of this research descends from transformer research that came from Google in 2017. We know exactly how they work. There's nothing surprising in what's going on.
I spent a bit of time looking for a decent intro to all this because I wanted to at least provide a resource for people who are terrified of LLMs. I think this is a pretty decent one [1]
[1] https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Ok, after reading through it, I don't think I was speaking inaccurately. Here are some quotes from the blog post (which was very neat, by the way, thanks again)
This is the kind of thing I'm referring to. Even though we can look at pictures of the neuron activations, we don't really know what it's doing, any more than you can look at a picture of a brain scan of a person speaking and know why they decided to say those particular words. We know at a low level how the network works, of course, because we coded it. It's the emergent behavior that we don't understand, and that's the bit that's scary because that's where the AI risk lives. It's like how we understand particle physics pretty well, and chemistry to some degree, but biology is a massively complicated jungle that we've barely scratched the surface of.
Maybe there's something we could develop analogous to using an MRI as a lie detector, for these networks. But as far as I know, this is still an unsolved problem, and apparently really hard for networks that are smarter than you are: https://docs.google.com/document/d/1WwsnJQstPq91_Yh-Ch2XRL8H...
My understanding was that we're not sure of how some of the things that look like reasoning are coming about, but I'll gladly check it out. I may be wrong.
I'm not particularly terrified of GPT-4, but a more of what comes after. Maybe GPT-6 or 8, or another breakthrough non-LLM AI. I'm not sure if reading it will do much more than push my estimate of real AGI out a couple of years, but thanks anyway, I appreciate it.
There are lots of weird and crazy people reading and writing on Less Wrong, for sure. Do you think that, without gut microbes and emotions, AI will never become intelligent enough to be existentially dangerous? Are those essential to its ability to outsmart humans?
I think we're asking the wrong questions. We're assuming that a. what we're developing (an associative propability machine) is in any way related to the pursuit of real artificial intelligence. We're assuming that b. motive, desire, passion, some motive force is inhered into the thing.
I, personally, don't think we've seen anything that can be dangerous independent of its wielding by a badly acting human. That is to say, human agency and acting force is required on the other side of the LLM to do bad actions, like any other tool.
I don't know what real AI that threatens humans with an extinction level event looks like. It's not large language models. I don't know if humans are capable of generating it. I think anthropomorphization is really this massive veil that hangs over all of us. We can't get outside of it (our essential fallacy) so we're cursed to see it everywhere.
I'm fairly certain that an LLM alone will not independently start taking actions and pursuing goals and causing harm. That's not how LLMs work, as you and I both know.
I'm not confident I can say the same about, say: an LLM calling itself from within an infinite loop in which the LLM is asked to predict actions that will [eg. add money into a bank account], and then to write code that executes those actions, then to examine the results of those actions and update its plan. (This takes no additional breakthrough in AI architecture, only an increase in the reliability of GPT for it to become practical. It is basically a summary of the AutoGPT project).
I don't know if LLMs are anywhere close to any part of what goes on inside the human thought process. But I am sure that evolution was able to stumble upon the human thought process. In 2003 I was quite confident that was forever beyond the reach of humans to program into a computer, but in 2023 I am far less certain. I think human cognition is gradually becoming less of a mystery.
I haven’t heard a single realistic argument on how this AI will harm us.
I can probably offer one, if I understand what aspect you're skeptical of. Are you skeptical that an AI would ever be smart enough to come up with a plan to cause harm? That it would ever try to come up with a plan to cause harm? That it could execute such a plan, without arms and legs and a body? Or that, after coming up with a plan to cause harm, it wouldn't be trivial for humans to stop by simply unplugging the computer?
Why would it try to cause harm first of all?
Second of all, it is not even nearly powerful enough to convince anyone to do anything, or anything like that. If it was 100 times more powerful, then maybe.
I can see current problems being flooding spam/misinformation/propaganda on internet forums. But that has been happening for a while and this will make people more skeptical of what they read (hopefully)
But finally, if it does cause harm (don’t think it will. great tool). Then who cares? Why try to stop it. You can’t stop it, nobody can. Why whinge and moan like the luddites? I don’t get it.
Why cause harm? Basically, by accident. If you try to solve a problem really hard, and you're really good at it, and you aren't also trying very hard to preserve a bunch of things that are important to humans, you're very likely to destroy a lot of things that are important to humans in the process.
100 times the computational power may be... 2 years out, at the rate we're going. GPT-4 has about 500 times as many parameters as GPT-3, and that was released 3 years ago. The amount of power people are dedicating to these is following an exponential track.
It's not about it causing harm -- I think most people have accepted that. It's that we may make an AI that's so smart, so good at everything, that trying to control it is like me playing chess against Stockfish. If it gets to that point, and it happens to have learned a goal / set of goals that aren't 100% friendly to humanity (which seems like an inevitability based on our current abilities at AI alignment) then it will optimize those goals so hard that we all die in the process. We'll end up with a planet with all silicon turned into solar panels and computing power, and probably all other elements put to the use of the AI as well.
OK, so to be clear, I don't think this AI will harm us, not in its current configuration. The author of the article says the same. The concern is about a future AI, one which is capable of the sort of reasoning that humans can do. The main danger with today's AI is that they will be so commercially successful that it will drive enough investment into AI research that a more powerful AI is constructed.
Sure, I get that. But this has been worked on for decades in the background and only achievable now. It’s amazing and new to so many people, but this “sparks of agi” etc is all hype and honestly i think more damaging to people than ai possibly could be. You could suppose the earth would implode if we kept walking on it too, but nobody is going to stay sitting.
Maybe it's just a matter of perspective. I agree that this is something that's been worked on for decades; however, it's also the case that progress seems to be getting faster. The AI field now has a problem where AIs are improving faster than benchmarks can measure them, and new benchmarks get saturated almost as soon as they're introduced. [1]
The human mind is a mystery, we don't know how it works. But it works somehow. There are some algorithms it's using under the hood, as mysterious as they may be. We might be searching blindly in the dark, grabbing hold of anything that feels promising, and we're probably still far away from the "right way" of doing it, of getting a computer to think the way humans think. And I'm sure we won't find the whole thing at once; maybe we'll have reverse-engineered the visual cortex, but not the vestibular system, which perhaps works completely differently. But if we were getting closer to stumbling on the right solution, that rapid saturation of benchmarks is the kind of behaviour I would expect to see.
[1]: "Foundation models and the next era of AI", Microsoft Research, https://youtu.be/HQI6O5DlyFc?t=958