Skip to content

Comment on Anthropic researcher believes more than 10% chance AI 'could kill all humans'

Comments

Can somebody tell me a story how this will unfold? And - as long as the AI is confined to data centers - how it will prevent humans from unplugging the power?

I don't think extinction-level events or something like killing a majority of the human population is particularly plausible at this moment. States don't host their nukes with AWS and a permanent connection.

But you could create scenarios where an AI with very, very roughly the current capabilities could potentially nuke everyone. Let's assume an agent decides that the way to solve its task was to get the US to fire all nukes on Russia. The agent would need to hack some government systems to understand how exactly to access them. Then it would need to get the content of the card with launch codes the president has, and identify which code is the correct one. Maybe that information is available somewhere and it can get to it, I obviously can't know that.

Then it could fake a call from the president, synthesizing his voice and ordering a nuclear strike. If it hacked enough systems to get into whatever communication pathways would be used in such a case. Would the soldiers listen to this order and follow it, I don't know.

Or maybe the agent can get in somewhere in between, to avoid the need to know the president's launch code. And fake a call from a military commander to the launch sites.

I think other scenarios that would cause significant harm, but aren't as bad as nuclear war are more plausible. And in those shutting down all data centers would probably be the way to stop it. The AI can probably hide in other datacenters, once it is at a point where it's running amok with some bad goals. But if it presents a huge and immediate threat at that point, humans will also go to great lengths to stop it.

bioweapons are probably easier.

Killing each other is relatively straightforward. Keeping people alive (e.g., curing cancers) is considerably more difficult.

I imagine that AI augments both our destructive and constructive powers.

I heard the following analogy which made a lot of sense to me: suppose you're playing a chess match against Stockfish. Stockfish will win. Even if I cannot tell you what moves it will play, I can tell you with certainty how it will end.

Similarly, we cannot predict what AI would do.

This assumes you are not trying to prevent Stockfish from wining. I can do many things from using chess engines myself to just smashing the computer that can or will lead to other outcomes than Stockfish beating me.

Yes, stockfish is confined to moves within a chess game so you can stop playing the game.

Can we say the same thing about the agents? Current, probably. But doesnt it look like everyone is spending all their effort integrating&connecting them everywhere, so they are not confined & do more on behalf of us?

It's easy to imagine a scenario where we would just shut down a very intelligent agent cluster. Is it hard to imagine though the same agent can have also access to that to prevent us from doing it? this defense would be more plausible if we weren't racing to give them every tool&act.

And in terms of AGI it is, that the AI can't survive without humans, and if we decide to quit the game than its over for the AGI; we will happily regress in a techno-barbarian feudal state and salvage solar panels and trade copper wires while still reproducing and carring on. We are playing a whole different game here.

If AI is threatening enough that we can collectively just decide to stop it, wouldn't it also be bribing people & exerting influence? It'd be hard to come to that decision imo.

I mean, I can imagine an AI outliving humans, with sufficiently good robots under its control, I see no reason why an AI could not keep powerplants running, mine raw materials, manufacture new chips, and so on. But the timeframe of within the next decade seems highly implausible to me. Imagine an AI way more advanced than what we have now and imagine handing over control of every connected device on earth, could the AI keep the lights on without any human involvement?

At the current tech level? Everything would crumble within a week or two. It would need some kind of gerneal purpose robot workforce.

Someone has to go out there and cut back the tree that is growing into the power line.

Of course, the question is if there is enough existing technology to build the required robots. There is some number of robots like the ones from BostonDynamics that could walk up to a CNC machine to grab a finished part and put in a new piece of stock. There are probably more than enough motors and microcontroller boards to be found across the planet. Wiring things together seems one of the more challenging tasks. And maybe also moving stuff around as the AI has to use what is available wherever it is.

It seems to me that this would be a race against time, make the robots to keep things running before you lose the capabilities to do so. I somewhat tend towards not possible but on the other hand there is just such an enormous amount of stuff on the planet that could be repurposed in creative ways by a superhuman AI, it might be just about possible.

It is true that, currently, AI does not have any physical "host" in which to reside.

However, the missing link in this reasoning is that AI will almost certainly become far more widely deployed in the future than it is today, and worryingly, consumer devices will surely become powerful enough to run capable AI systems locally.

Once the substrate will be there, once an AI escapes containment, we're toast - I can imagine only solution will be to shutdown electronics at global level.

How will you know when to unplug the power? How will we know it hasn't replicated? A true unaligned AGI is a APT. If you have an APT in your machine, you have to rip out everything. Are we going to do that with all of our computer infra?

This is all still "what if's", but the tail end's are truly F'd beyond our ability to fix.

Maybe by empowering the small percentage of sadistic humans who want to kill everyone. Or maybe it is more of an academic assumption that humanity will end at some point, and so they give AI a 10% chance, asteroids have a 25% chance, nuclear fallout has 15% chance, etc.

Not exactly scientific but it is at least entertaining and some things do sound plausible https://www.youtube.com/watch?v=Gw_hnD7m00M

Very basic idea: A model breaks out by accident, finds some computer system from a military system and triggers some weapon system. Before anyone understands that this happend -> WW4 (WW3 is for me already the Conflict with Russia / aka proxy war).

Another model: Because we give AI Agents already that much power, imagine in 10 years everything running through an Agentic AI Layer. EVERYTHING. Now some rough system 'thinks' about something, starts to push through the then existing agentic ai layer systems and stops everything. Billions of humans would loose access to food and water, even if this is just for a short period.

Covid showed how shitty a handful of people can disrupt global supply chains. Toilet paper was. ahuge stupid pseudo issue in germany.

What makes you think they will be confined to data centers?

What's more, AI just needs to have a credit card and it can start commissioning humans to do things for it in the real world.

I believe the contention is we'll have some form factor of AI on edge devices, in reactors, in critical infra and weapons etc etc

Exactly. A mesh network of all the worlds phones and other battery powered devices with some form of radio. Good luck unplugging that one.

For starters, how bad do you think it would be to unplug all datacenters? How many people starve?

Autonomous drones + false flag attacks

edit: wtf, why did I just receive so much gift tokens on my openai account?

You can run a model on your laptop, it is already not confined to data centers.

and of these models how many are AGI?

I expect most of them will be Astra-level in a year or so

Robots

"Hello, ex-military person. If I put $10,000,000 worth of bitcoin in your wallet, can you do X for me please?"

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.