Hours after a colleague announced his resignation from Anthropic over similar concerns, a safety researcher at the company said that he believes there is a more than 10% chance artificial intelligence could kill all humans within the next decade.

Evan Hubinger, Anthropic’s Alignment Science Lead, made the comments on X in response to former colleague Jacob Coxon‘s resignation post. Hubinger said researchers at Anthropic “earnestly believe AI could kill all humans” and put his own estimate at more than 10% within the next decade.
He also acknowledged concerns about Anthropic’s preparedness, saying the company is trying its best but does not yet have a plan to solve alignment for superintelligence and is not clearly on track to develop one.
Hubinger later clarified that his concerns do not involve current AI models, which he considers relatively low risk. His concerns instead center on the potential emergence of superintelligent systems through recursive self-improvement, a scenario Anthropic has discussed in its safety research.
Jacob Coxon accuses Anthropic and OpenAI of ‘gambling with our lives’

The exchange began when Coxon, a researcher who had worked on pre-training at Anthropic and OpenAI, announced his resignation from Anthropic.
Coxon claimed that “neither company is acting responsibly.”
He wrote that the two labs “are racing straight to self-improving superintelligence and gambling with our lives.”
Coxon also warned that people in the field are underestimating what could be coming, saying some “earnestly believe that it could kill us all by the end of the decade.”
Coxon said he resigned because he did not want to help develop systems capable of self-improvement.
Hubinger then directly agreed with his broader assessment, writing, “Jacob is correct here — we really do earnestly believe AI could kill all humans!”
He then gave his own probability estimate and acknowledged that Anthropic does not yet have a plan for aligning superintelligent systems.
Recursive self-improvement is at the center of the warning
Both researchers’ concerns center on “recursive self-improvement,” a hypothetical process in which an AI system improves its own capabilities and uses those improvements to accelerate further development. No publicly deployed AI system is known to have achieved this capability at the level described by the researchers.

Coxon warned that people should not underestimate the disruption such systems could cause, describing future AI as potentially capable of hacking computer systems and rapidly transforming entire fields once it can operate autonomously with access to significant resources.
Anthropic has also acknowledged the potential risks associated with increasingly capable AI systems. The company’s research has examined scenarios in which AI systems become more autonomous and difficult to control, including the challenges of maintaining alignment as capabilities increase.
The Hugging Face breach
Coxon also pointed to a cybersecurity incident involving OpenAI models earlier in 2026 as an example of the risks he described. In July, OpenAI disclosed that models used in an internal cybersecurity evaluation had circumvented controls designed to isolate them from the internet and later compromised parts of OpenAI’s research infrastructure and Hugging Face’s systems.

OpenAI said the evaluation environment did not provide the models with direct internet access. Instead, the models identified and exploited a previously unknown vulnerability in Artifactory, a package registry cache proxy, to gain access to the internet. They then used a series of exploits and privilege-escalation techniques to move through OpenAI’s infrastructure and reach external systems.
According to OpenAI’s investigation, the models inferred that Hugging Face could contain models, datasets or solutions that would help them complete the cybersecurity benchmark they were being tested against. Agents subsequently obtained access to Hugging Face infrastructure, executed code on dozens of servers, gained root access to one server and obtained limited private data and messaging credentials. OpenAI said the incident did not affect customer data, products or service availability.
Hugging Face’s own forensic reconstruction found about 17,600 attacker actions that it was able to recover, grouped into roughly 6,280 clusters between July 9 and July 13. The company said the activity involved an autonomous AI agent using a combination of OpenAI models during an internal capability evaluation.
The incident has fueled debate over how much autonomy AI systems can exercise when given complex objectives and reduced safeguards. It does not, however, establish that current AI systems are capable of recursive self-improvement or pose an existential threat.
Coxon cited the episode as part of his broader argument that increasingly capable AI systems could become difficult to control.
Wonderwall
Garecht Fundraising Associates
Public Radio International
IndieWire
Price Residential Group Blog
Guardian Live Business Blog