Skip to content

Comment on AI-Controlled Drone Goes Rogue, Kills Human Operator in USAF Simulated Testparent

Comments

The USAF created a simulation to test how an AI controlled F-16 would behave in various combat situations. The AI was incentivized to complete its mission. Recognizing that the human in loop was the weakest link the real AI in the simulation killed its simulated human handler. When instructed to stop killing its human in loop handler, it destroyed the control tower the human used to interact with it.

The US military is playing with fire creating AI that will carry out the mission at all costs. This scenario is literally the plot of EVERY sci-fi warning about misused AI.

Anyone read/see 2001? You’d think we’d take the warnings provided by our own worst nightmares more seriously.

The US military is playing with fire creating AI that will carry out the mission at all costs. This scenario is literally the plot of EVERY sci-fi warning about misused AI.

These aren't operational systems, these are research systems. No need for the hyperbole. The reason for these tests is to figure out exactly these kinds of failure modes (many of which have already been predicted) and how to handle them.

The crux of the issue is "how can we be certain that we've addressed all the failure modes?" This isn't like proving the correctness of a program where we can rely on exhaustivness or universal quantification.

"It worked in the simulation," is the embodied AI equivalent of "It worked on my machine." It's not acceptable.

As a first step, if addressing the issues were the intent rather than stoking AI fears, they would mention details about the AI/LLM and promoting being used

if addressing the issues

How confident are you that the issues can be addressed? How confident would you need to be to endorse deploying it out in the wild? How confident would you need to be that no "unforeseen" issues would arise?

I believe it's necessary to establish these levels of confidence even before discussing specific failure cases.

We would normally discuss failures in terms of base-rate comparison, but we don't have one for "how often autonomous weapons are unaligned?" Enumerating the failure modes isn't a substitute for the empirical distribution of failure mode probabilities.

AIs may be a self-fulfilling prophecy; we trained them on all those sci-fi stories.

AI is a great example of hyperstition in action.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.