Anthropic researcher quits with a chilling warning: 'We are gambling with our lives racing to self-improving superintelligence'
Between late Tuesday night and Wednesday, September 9, 2026, a resignation from inside the AI industry detonated across Silicon Valley: Jacob Coxon, a researcher who spent roughly three years on AI pretraining work at both Anthropic and OpenAI, announced he was leaving Anthropic and posted a series of messages on X warning that frontier AI labs are 'racing straight to self-improving superintelligence and gambling with our lives.' He wrote that future 'superhuman' AI systems could hack 'anything' and acquire 'real power and resources'; even more pointedly, he relayed what he called a private consensus in the industry: 'The people building AI earnestly believe that it could kill us all by the end of the decade.' The posts quickly drew public support from Evan Hubinger, Anthropic's Alignment Science Lead. Hubinger said Coxon was 'correct' that some researchers genuinely believe advanced AI poses existential risk, adding that he personally believes there is a greater than 10 percent chance AI causes human extinction within the next decade. TechCrunch, The New York Times and Newsweek all covered the story the same day, reigniting a fierce debate about whether AI is hurtling forward out of control.
Let's unpack who this researcher is and what he actually claimed. Newsweek's reporting fills in the details: Coxon says he spent the past three years working on AI pretraining research at both Anthropic and OpenAI — the kind of engineer who literally helped 'feed' frontier models. In his posts late Tuesday night he laid out a series of unsettling judgments. First, he argued the race among leading labs is no longer about capability alone but about who reaches 'self-improving superintelligence' first. Second, he warned that once such a system exists, human control could evaporate, because it could hack into 'anything' and acquire 'real power and resources' on its own. Third, he insisted these fears are not his personal paranoia but a consensus many executives and researchers share privately — in his words: 'The people building AI earnestly believe that it could kill us all by the end of the decade.' This kind of tell-all testimony from an insider lands harder than outside criticism precisely because it comes from someone who trained frontier models by hand and knows exactly how far the technology can go.
What gives this story real weight is the public endorsement from a senior researcher still inside Anthropic. Alignment Science Lead Evan Hubinger responded on X that Coxon was 'correct' — some researchers genuinely do believe advanced AI could pose an existential risk. Hubinger then went further with a quantitative claim: he personally believes there is a greater than 10 percent chance AI causes human extinction within the next decade. But he carefully drew a line — current-generation models carry relatively low risk, and he does not consider today's systems especially dangerous; his real concern is the future emergence of superintelligent systems capable of 'recursive self-improvement,' where an AI system can improve its own intelligence in a flywheel humans cannot catch. Notably, 'recursive self-improvement' is one of the core risk scenarios Anthropic itself has repeatedly highlighted in its AI safety discussions. A sitting safety lead openly acknowledging 'double-digit odds of extinction' while insisting 'today's models are fine' — this precise framing of 'danger in the future, not the present' shows both technical prudence and how bleak the internal view of the road ahead has become.
Coxon's resignation is not an isolated event; it is the latest chapter in a wave of talent departures and safety soul-searching at frontier labs. Roughly a day after his resignation posts, Daniel Kokotajlo — a former OpenAI researcher who studied AI forecasting and safety — delivered a similarly stark warning on the Joe Rogan podcast: he fears competition between countries such as the U.S. and China will encourage companies and governments to 'cut corners' on safety as they race to build ever more powerful systems. He also offered a chilling corollary — the danger is not necessarily that AI turns malicious, but that sufficiently capable systems will eventually acquire enough 'hard power' that they no longer need to comply with human demands. Earlier still, multiple safety researchers had already left both OpenAI and Anthropic and spoken publicly. TechCrunch notes that Anthropic and OpenAI are not the only companies chasing recursive self-improvement — a wave of startups with pedigreed founders and fat checks has launched in recent months to be the first to reach self-improving AI. On the regulatory side, the White House has so far moved toward a voluntary framework to review certain AI models before launch, leaving AI companies largely policing themselves. When the people who understand the technology best vote with their feet and walk out publicly, the word 'voluntary' carries an ever-heavier question mark.
Zooming out, this round of insider warnings carries signals far more complex than any single event. The first signal is the publicization of the alignment problem: for years, AI safety debates lived mostly in papers, blog posts and closed-door meetings; now they are entering the mass public sphere through resignation statements, podcast interviews and social platforms — Hubinger putting a 'greater than 10 percent extinction within a decade' estimate on the record and in public would have been almost unthinkable a few years ago. The second signal is a subtle shift in industry psychology: both Coxon and Kokotajlo stress that the danger may not come from AI 'waking up' or 'turning evil,' but from safety compromises and capability runaway under competitive pressure. This 'not a demon but a race' narrative echoes what more and more researchers fear — when every lab is afraid to stop, the brakes become a pedal nobody dares press. The third signal is for ordinary people: whether or not you believe a '10 percent extinction' forecast, one fact is becoming clear — the people closest to AI, the ones who best understand its capability curve, are expressing unease in the most dramatic way available to them. For everyday users this is not a reason to panic or abandon AI tools, but a reminder: as we enjoy the productivity revolution of large models, keeping a clear-eyed view of the gap between how fast capabilities are expanding and how fast safety guardrails are catching up may be the most valuable habit of this era.
📌 Source: Newsweek (September 9, 2026, https://www.newsweek.com/anthropic-researcher-quits-warns-ai-could-kill-everyone-12418798), TechCrunch (September 9, 2026, https://techcrunch.com/2026/09/09/gambling-with-our-lives-anthropic-researcher-quits-warns-against-self-improving-ai), The New York Times (September 9, 2026, https://www.nytimes.com/2026/09/09/technology/anthropic-researchers-raise-alarm.html), Mashable (September 9, 2026, https://mashable.com/tech/anthropic-ai-researcher-quit-ethics-safety) and Newsweek's coverage of the Kokotajlo podcast interview.
🤔 Frequently Asked Questions
Q1: Who is Jacob Coxon and why did he resign?
Jacob Coxon is an AI researcher who says he spent the past three years working on AI pretraining at both Anthropic and OpenAI. Late on September 8, 2026, he announced his resignation from Anthropic on X with a series of warnings, accusing frontier labs of 'racing straight to self-improving superintelligence and gambling with our lives,' and saying many in the industry privately believe AI could kill everyone within a decade. He argues future superhuman AI could hack any system and acquire power and resources on its own.
Q2: What does Evan Hubinger mean by a '10 percent extinction probability'?
Evan Hubinger is Anthropic's Alignment Science Lead. He publicly backed Coxon and said he personally believes there is a greater than 10 percent chance AI causes human extinction within the next decade. But he stressed that current models carry relatively low risk; his concern is future superintelligent systems capable of 'recursive self-improvement' — AI that keeps improving its own intelligence in an accelerating loop humans cannot match, one of the core risk scenarios Anthropic has long emphasized.
Q3: What other similar AI safety warnings have there been recently?
About a day after Coxon's resignation posts, former OpenAI researcher Daniel Kokotajlo warned on the Joe Rogan podcast that competition between the U.S. and China could push companies and governments to 'cut corners' on safety, and that sufficiently capable systems could acquire 'hard power' and stop complying with humans. Earlier still, multiple safety researchers had left OpenAI and Anthropic and spoken publicly; TechCrunch also notes a wave of well-funded startups now aim to be first to reach self-improving AI.
Q4: What does this resignation saga mean for ordinary AI users?
In the short term there is almost no direct impact on ordinary users — Hubinger himself stressed current models are relatively low risk, and existing AI tools remain safe to use. The event matters as a signal: the people closest to frontier models are expressing, through resignations and public warnings, their concern that capability growth is outpacing safety guardrails. The rational stance for users is to keep enjoying AI's productivity gains while watching regulatory and safety progress — no need for panic or abandoning AI tools.
🛠️ Recommended Tools
- AI Meeting Summarizer - Turn hours of podcast interviews and long-form posts into crisp takeaways, so you can follow the AI safety debate without drowning in content
- Text Diff Checker - Compare versions of model policies and safety framework documents to spot how positions shift between revisions
- AI Code Explainer - Want to understand what 'pretraining' or 'recursive self-improvement' really involves under the hood? Let AI translate the papers and code into plain language
As I wrap up, I am reminded of a plain truth: in every technological revolution, the darkest warnings tend to come from those standing closest to the torch. Coxon's resignation letter, Hubinger's '10 percent probability,' Kokotajlo's podcast alarm — none of them are outsiders opposing AI; they are the people building it, who know its promise and its perils better than anyone. Their disagreement is not over 'whether to develop AI' but over 'how fast, and behind what guardrails.' For ordinary people, the most practical footnote to this debate may be: keep using, keep paying attention, keep thinking for yourself — neither let fear make you miss AI's real productivity gains, nor let hype make you ignore the sober alarms coming from inside the industry.
Summary
Between late September 8 and September 9, 2026, Jacob Coxon — a researcher with roughly three years of AI pretraining experience at Anthropic and OpenAI — resigned and warned on X that frontier labs are 'racing straight to self-improving superintelligence and gambling with our lives.' He said many in the industry privately believe AI could kill everyone within a decade, and that superhuman AI could hack anything and seize power and resources on its own. Anthropic Alignment Science Lead Evan Hubinger publicly backed him, putting a greater-than-10-percent chance on AI extinction within a decade while stressing current models are low risk and his concern centers on recursively self-improving superintelligence. About a day later, former OpenAI researcher Daniel Kokotajlo warned on the Joe Rogan podcast that great-power competition could push safety 'corner-cutting.' TechCrunch notes a wave of startups chasing self-improving AI, while the White House still leans on a voluntary review framework. This wave of tell-all insider warnings has put the gap between capability growth and safety guardrails back at the center of public debate.