Evan Hubinger, Anthropic's staff lead on keeping the technology aligned with human goals and values, backed up Coxon claims in a follow-up post of his own, though he didn't quit the company.
"Jacob is correct here — we really do earnestly believe AI could kill all humans," he said.
Hubinger estimated the chances of that happening to be higher than ten percent within the next decade, and added that there's no plan yet on how to keep AI aligned with human goals in the superintelligence scenario.
10% chance we kill everybody is a small price to pay for Motown remixes of classic 2pac songs.
This is just a bizarre thing to say that you're working on technology with that high a downside potential. If you were saying that while running a biology lab, or building a nuclear reactor, people would be demanding your head on a spike. But by not quitting it's clear that he himself doesn't really believe it.
Or rather, this shows the difference between "believe" (political) and "believe" (use as a basis for action). I'm reminded of a story of how Afghans supposedly listened to the BBC World Service despite considering it enemy propaganda because the weather reports were really useful.
How exactly does that work? The nuclear system of MAD relies on physical threat, lab A achieving ASI (artificial scary intelligence) does not prevent lab B achieving it.
I would like everyone involved to be a lot clearer about their threat models, with plausible series of clearly linked steps, rather than just sounding like a Vernor Vinge novel.
The idea in AI Risk circles is that the first true Scary Intelligence wins, as it can recursively advance itself and exponentially outgrows/hack any following (weaker) B or C intelligence.
"If we can do all this, we will have a world in which democracies lead on the world stage and have the economic and military strength to avoid being undermined, conquered, or sabotaged by autocracies, and may be able to parlay their AI superiority into a durable advantage." [0]
I couldn't agree more with you. Dario throws out the term "durable advantage" in many of his talks and essays without seriously addressing what that would entail. It's not pretty imagining the mechanics of converting an intelligence advantage into a permanent power imbalance. I'd much rather we pursue international cooperation on a slowdown.
- Manufacture new hacking incidents at other AI labs, proving that lab A is the only one capable of creating aligned AI
- Credible death threats to individuals working at data centers, coercing them to enable attack vectors
- To distract the world, enable a terrorist attack by delivering intelligence to specific groups who can act on it
- Create a network of individuals vulnerable to leverage to be used as tools for enabling the ASI to act on the real world
- Convince Lab A workers to display to the public that progress at Lab A is slowing down, and during this time recursively self improve until the task can be completed
Yeah, super ethical lab earnestly believing AI has 10% chance to kill all humans within a decade and doing utmost to protect humanity is partnering with MIC and Palantir in particular. Sounds about right.
Before you tell me about how the CEO has taken a principled stand: on record, he had no problem using it against 95%+ of humanity outside the U.S., and was only against fully automatic AI killing machines, and only citing the technical reality of then-current gen tech, so one should read that as human-rubber-stamped AI killing machines are totally fine with him.
You don't need to kill all humans, you just need to get better than them at zero sum games like resource extraction, and outcompete them. Then, you protect what's legally yours.
AI labs are working very hard at making them good at making them richer, and at making them respect the property rights of corporations.
But by not quitting it's clear that he himself doesn't really believe
I’m not saying I’d press a button that had a 50/50 chance of ending humanity vs giving me generational wealth, but … I can see how someone would get there? There are percentage odds and payouts to match all risk appetites and valuations. Anyone claiming they wouldn't press the 1 in quadrillion chance button for 1bn dollars is probably not telling the truth, and after that, we're just haggling about percentages.
If you believe that the probability of destruction is currently 10%, but the probability of destruction if you decide to quit Anthropic becomes (say) 13%, then the rational move (Assuming you are opposed to destruction) is not to quit.
A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.
Comments
10% chance we kill everybody is a small price to pay for Motown remixes of classic 2pac songs.
This is just a bizarre thing to say that you're working on technology with that high a downside potential. If you were saying that while running a biology lab, or building a nuclear reactor, people would be demanding your head on a spike. But by not quitting it's clear that he himself doesn't really believe it.
Or rather, this shows the difference between "believe" (political) and "believe" (use as a basis for action). I'm reminded of a story of how Afghans supposedly listened to the BBC World Service despite considering it enemy propaganda because the weather reports were really useful.
Or he believes that other labs might get there first, and he is working to counter that threat.
This is the Manhattan Project again.
How exactly does that work? The nuclear system of MAD relies on physical threat, lab A achieving ASI (artificial scary intelligence) does not prevent lab B achieving it.
I would like everyone involved to be a lot clearer about their threat models, with plausible series of clearly linked steps, rather than just sounding like a Vernor Vinge novel.
The idea in AI Risk circles is that the first true Scary Intelligence wins, as it can recursively advance itself and exponentially outgrows/hack any following (weaker) B or C intelligence.
Lab A reaches ASI, and is prompted the following: "Permanently nullify all other AI labs".If it's ASI it would succeed.
I am explicitly asking people to fill in the blank on how it would succeed. It's intelligence, not magic.
A lot of these scenarios seem to assume that nobody else gets a move. That there wouldn't be a human response.
"If we can do all this, we will have a world in which democracies lead on the world stage and have the economic and military strength to avoid being undermined, conquered, or sabotaged by autocracies, and may be able to parlay their AI superiority into a durable advantage." [0]
I couldn't agree more with you. Dario throws out the term "durable advantage" in many of his talks and essays without seriously addressing what that would entail. It's not pretty imagining the mechanics of converting an intelligence advantage into a permanent power imbalance. I'd much rather we pursue international cooperation on a slowdown.
[0] https://darioamodei.com/essay/machines-of-loving-grace
- Manufacture new hacking incidents at other AI labs, proving that lab A is the only one capable of creating aligned AI
- Credible death threats to individuals working at data centers, coercing them to enable attack vectors
- To distract the world, enable a terrorist attack by delivering intelligence to specific groups who can act on it
- Create a network of individuals vulnerable to leverage to be used as tools for enabling the ASI to act on the real world
- Convince Lab A workers to display to the public that progress at Lab A is slowing down, and during this time recursively self improve until the task can be completed
We're currently putting it into all sorts of critical systems, from logistics to power. It could just stop running them on our behalf.
Yeah, super ethical lab earnestly believing AI has 10% chance to kill all humans within a decade and doing utmost to protect humanity is partnering with MIC and Palantir in particular. Sounds about right.
Before you tell me about how the CEO has taken a principled stand: on record, he had no problem using it against 95%+ of humanity outside the U.S., and was only against fully automatic AI killing machines, and only citing the technical reality of then-current gen tech, so one should read that as human-rubber-stamped AI killing machines are totally fine with him.
You don't need to kill all humans, you just need to get better than them at zero sum games like resource extraction, and outcompete them. Then, you protect what's legally yours.
AI labs are working very hard at making them good at making them richer, and at making them respect the property rights of corporations.
Or he is paid >1M USD per year
I’m not saying I’d press a button that had a 50/50 chance of ending humanity vs giving me generational wealth, but … I can see how someone would get there? There are percentage odds and payouts to match all risk appetites and valuations. Anyone claiming they wouldn't press the 1 in quadrillion chance button for 1bn dollars is probably not telling the truth, and after that, we're just haggling about percentages.
If you believe that the probability of destruction is currently 10%, but the probability of destruction if you decide to quit Anthropic becomes (say) 13%, then the rational move (Assuming you are opposed to destruction) is not to quit.
The real rationale fot them not to quit Anthropic is the 100% probability of losing in income.
https://xcancel.com/hilbertspaess/status/2097476208908972230...
From Jacob:
For a few million a year, I'd say the same thing.
A simple explanation:
This is a highly uncertain and dangerous scenario. Breeding ground for anxiety.
High agency people often deal (cope) with anxiety by trying to control outcomes. Some just flee the situation altogether.
We have an example of both here: one employee leaves, one stays.