Saturday, September 12, 2026

Jacob Coxon Warns of Human Extinction and Triggers a Preference Cascade

CEOs of major AI labs, and employees of major AI labs, including OpenAI and Anthropic, often say they plan to build superintelligence soon, as in within a few years create AIs that are superior to humans at essentially all cognitive tasks.

They often warn that such AIs might kill everyone. Or that AIs might cause mass unemployment, cause cyberattacks across the internet, enable mass surveillance or risk causing any number of other highly bad things.

These warnings are consistently and directly against the interests of the labs. Yet the warnings have recently gotten a lot louder and more frequent. OpenAI has been practically screaming, for those with ears to listen, on many occasions.

A series of events, over two months and especially the last week or so, including internal observations of the pace of progress at OpenAI and also Anthropic, have freaked out everyone involved quite a lot more than they were already freaked out.

After all the events, plus statements by Dean Ball and Jakub Pachocki, we were seeing the beginnings of a preference cascade.

Then along came Jacob Coxon as the tipping point, and things took off.  [...]

Jacob Coxon Resigns From Anthropic In Protest And Sounds The Alarm 

Jacob Coxon spent the last three years doing pretraining research at both OpenAI and Anthropic. He has come to realize that everyone involved is being wildly irresponsible.

He warns us: They are racing straight to superintelligence and gambling with our lives. I agree with and strongly endorse his statement.

If anything he sounds like an optimist. He’s asking you to consider what the next few years will actually feel like, which means he thinks you have a few years left.
Jacob Coxon (former Anthropic and OpenAI, 160m+ views, September 8): I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.

Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.

The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible - but I hear the same people express fear privately. No other human activity poses this level of danger.

A common response is “if they truly believe this, why are they still building it?” At OpenAI, many have not deeply internalized the civilizational stakes. At Anthropic, the stakes are well-understood, but they are locked in a race to get there first - they believe no one else will act responsibly, so they must do it themselves, despite the risk.

Accepting this race and entering the “endgame” is a hubristic gamble that should not be launched from a private company’s Slack. Attempting to speedrun alignment should require extraordinary confidence that there are no better trajectories available.

I am optimistic about the potential for coordination. Warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. I don’t feel like we’re on track to prevent a global race, which may require costly actions such as a temporary ban on improving model capabilities.

If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?

Jacob Coxon (WSJ interview): We’re on track for a lot of the most aggressive of these scenarios where by the end of next year things could be out of control already.
If you want Jacob Coxon’s full views, I recommend his interview with Wired’s Maxwell Zeff. This thread has extensive quotes.

Here is Jacob Coxon doing a 5 minute interview with Anderson Cooper. He speaks well and plainly, and it is clear how much the events of the last two months have made it much easier to speak plainly to a civilian like Cooper about what is happening.

Here is Jimmy Kimmel doing four minutes on this. He gets it. How is this not the top news story on every site, indeed.

Yes, this is a common view, even if few have the courage to act.
Alex Turner: I left Google DeepMind in June. Jacob is right: many researchers believe they are building something that could kill everyone on the planet. It was literally my day job to think about how to stop that.
That tells you how bad Alex Turner thought DeepMind’s actions were with regard to the Department of War. His day job was that he got paid by Google to think about how to stop AI from killing everyone, and he felt morally obligated to quit in protest.

Derek Thompson here writes about this as part of AI Safety Is Having a Moment.

If you want to see the full list of lab employee quotes from the preference cascade, I compiled them into another post today. [...]

A few days ago, I had no idea who Jacob Coxon was, and I was not alone.

The timing of a preference cascade is difficult to predict. Once they start they can happen very quickly. You don’t want to talk until you are confident others will, and at some point the evidence that others will follow can snowball and it happens. A classic concrete example was replacing Biden in 2024.
Derek Thompson: what’s weird is that, in a way, this is breaking thru even more than huggingface! ... and it’s just a guy nobody had heard of reiterate a position that his CEO has said on podcasts 1,000 times
Why was Coxon able to set off a preference cascade? How did this one break through?

A confluence of factors, all of them downstream of the obvious actual reason, which is that there is a good chance that AI kills everyone soon.

by Zvi Mowshowitz, DWAV |  Read more:
[ed. Who's betting on reason and political courage to win the day? Uh, huh. First of all, agree on a temporary halt to recursive self-improvement (AIs improving AIs). Second, delay/suspend IPO's (even if most of the current economy is driven by AI spending and debt). Third, light a fire under politicans (it's been done successfully with data center buildouts). At this point, just one of these would be a big win. Here's Nate Soares, co-author of the book If Anyone Builds It, Everyone DiesA case for courage, when speaking of AI danger (LW). Also this:]
***
Terry Pratchett: “Some humans would do anything to see if it was possible to do it. If you put a large switch in some cave somewhere, with a sign on it saying 'End-of-the-World Switch. PLEASE DO NOT TOUCH', the paint wouldn't even have time to dry."
***
UPDATE: Dario Amodei (Anthropic) proposes a three-point plan. Original here (We Must Pace the Frontier.]