Anthropic Researcher Puts Catastrophic AI Risk Above 10 Per Cent This Decade
A senior safety engineer at the Claude maker has quantified existential danger from artificial intelligence, as a colleague resigns citing reckless development practices across leading labs

A Public Warning from Inside the Lab
A senior safety researcher at Anthropic has stated publicly that the probability of artificial intelligence causing human extinction exceeds 10 per cent before the end of this decade. The assessment, shared on social media, represents one of the most explicit quantifications of catastrophic risk to emerge from inside a leading AI laboratory. It arrives within hours of another Anthropic researcher, Jacob Coxon, announcing his resignation over what he describes as inadequate safety protocols at the company and its competitors.
Coxon, who has worked on training AI systems at both Anthropic and OpenAI, framed his departure as a response to what he characterises as a reckless development trajectory. In his resignation statement posted on X, he accused the two organisations of "racing straight to self-improving superintelligence and gambling with our lives." The language is stark, even by the standards of AI safety discourse, and it points to mounting internal tension over the pace at which frontier models are advancing relative to the safeguards designed to constrain them.
The 10 Per Cent Threshold
The unnamed senior researcher's probability estimate is significant not because it is new in substance, but because it comes from someone embedded in the safety apparatus of a company that positions itself as cautious. Anthropic has publicly committed to a research agenda centred on constitutional AI and interpretability, frameworks intended to make models more predictable and aligned with human intent. Yet the internal assessment suggests that even those efforts may be insufficient to address tail risks associated with systems that could recursively improve their own capabilities.
A 10 per cent probability of extinction is, in any other domain, a threshold that would trigger immediate regulatory intervention. In climate science, pandemic preparedness, or nuclear security, such a figure would be treated as intolerable. Within AI development, however, the response has been more ambiguous. The industry operates in a regulatory vacuum, with voluntary commitments substituting for enforceable standards. Anthropic, OpenAI, Google DeepMind, and other labs have published safety frameworks, but these remain self-imposed and subject to competitive pressure.
Self-Improving Systems and the Control Problem
The core anxiety articulated by Coxon centres on self-improving superintelligence, a scenario in which an AI system gains the ability to enhance its own architecture and capabilities without human oversight. This is not a distant hypothetical. Current large language models already exhibit emergent behaviours that were not explicitly programmed, and research into agentic AI, systems that can pursue goals across extended time horizons, is accelerating. The concern is that once a system crosses a certain capability threshold, it may become effectively uncontrollable.
Anthropic and OpenAI have both invested in alignment research intended to prevent such outcomes. Techniques like reinforcement learning from human feedback and red-teaming are designed to identify and mitigate dangerous behaviours before deployment. Yet these methods rely on assumptions about the behaviour of future systems that may not hold. A model that is aligned during training may behave unpredictably when deployed at scale, especially if it encounters edge cases or adversarial inputs not represented in its training distribution.
The resignation of a researcher who has trained models at two of the most prominent labs suggests that confidence in these safeguards is not universal, even among those closest to the work. Coxon's public statement is unusual. Most departures from AI labs are quiet, governed by non-disclosure agreements and professional norms that discourage airing internal disputes. His willingness to speak openly indicates a level of alarm that overrides those constraints.
The Competitive Dynamics of Frontier AI
The language of "racing" is central to Coxon's critique and reflects a broader concern about the structure of incentives in the AI industry. Anthropic was founded in part by former OpenAI researchers who left over disagreements about the company's direction, particularly its partnership with Microsoft and the perceived erosion of its original non-profit mission. The stated aim was to build a lab that prioritised safety over speed. Yet the dynamics of capital, talent, and public attention have pushed even Anthropic into a competitive posture.
Venture funding for AI companies has reached unprecedented levels, with Anthropic itself raising billions from investors including Google. The pressure to demonstrate progress, ship products, and justify valuations creates incentives that can conflict with cautious development. When multiple labs are working toward similar goals, the risk is that any one of them will cut corners to avoid being left behind. This is a classic collective action problem, and it is playing out in real time across San Francisco, London, and Beijing.
At DailyTechWire, we have tracked the funding rounds and product announcements that drive this cycle. Each new model release is accompanied by benchmarks showing incremental improvements in reasoning, coding, or multimodal understanding. What is less visible is the internal debate over whether those improvements are safe to deploy, and under what conditions. The public statements from Coxon and his colleague suggest that debate is intensifying.
What Comes Next
The immediate question is whether these warnings will have any impact on the trajectory of AI development. Historically, public resignations and whistleblower statements have had limited effect on the behaviour of large technology companies. The incentives are too strong, and the regulatory environment too permissive, for voluntary restraint to be a stable equilibrium.
There are, however, signs that the landscape may be shifting. Governments in the United States, the European Union, and the United Kingdom have begun drafting AI safety legislation, though implementation remains years away. International convenings, including AI safety summits hosted by the UK government, have brought together lab leaders, researchers, and policymakers to discuss risk mitigation. Whether these efforts will translate into binding constraints is uncertain.
For now, the burden of safety remains with the labs themselves. The public quantification of existential risk by a senior Anthropic researcher, combined with a high-profile resignation, adds pressure on the company and its peers to demonstrate that their safety commitments are more than rhetorical. The industry has long argued that it can self-regulate, that the people building these systems are best positioned to understand and manage the risks. The statements from inside Anthropic test that claim directly.
The 10 per cent figure will be debated, as all probability estimates of low-frequency, high-consequence events are. But the fact that it has been stated at all, by someone with access to the models and the internal deliberations, is itself a signal. It suggests that the distance between speculative risk and operational reality is narrower than many outside the labs have assumed.


