Anthropic alignment lead says >10% extinction risk, no solution
Evan Hubinger confirmed the company lacks a plan for superintelligence safety after researcher Jacob Coxon resigned over racing dynamics.
Senior safety researcher confirms existential risk estimate on the record
Evan Hubinger, who leads alignment science at Anthropic, publicly stated he believes there is greater than a 10% probability of AI-caused human extinction within the next decade. He also confirmed that Anthropic does not yet have a plan to solve alignment for superintelligence and is not clearly on track to develop one.
The statement came one day after Jacob Coxon, a 27-year-old pretraining researcher who spent three years at OpenAI and Anthropic, announced his resignation on X. Coxon wrote that neither company is acting responsibly and accused them of racing toward self-improving superintelligence while "gambling with our lives." He said he was leaving the AI industry entirely.
Hubinger's response was direct: "Jacob is correct here; we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade." He added that while he believes Anthropic is trying its best, the company lacks a solution to the alignment problem.
Why it matters
This is the first time a senior executive at a leading AI lab has publicly confirmed both a double-digit extinction probability estimate and the absence of a technical solution, unprompted and on the record. The admission exposes what Hubinger and others describe as a coordination trap: companies believe the risk is real but continue development because stopping would cede the frontier to competitors with fewer safety concerns. For business and policy leaders, this represents a rare moment of institutional candor about capability-risk misalignment at the companies building the most advanced AI systems.
What the resignation reveals
Coxon's departure reflects a breaking point in that logic. His argument was not that safety work at OpenAI or Anthropic is insincere, but that it is structurally inadequate given the competitive pressure both companies operate under. He said researchers inside these labs can see the hazards clearly but continue because they believe a competitor will move faster if they stop.
Hubinger's reply represents the opposite decision: continuing to work on the problem while acknowledging its unsolved state. His willingness to state the risk estimate publicly, rather than obscure it, marks a shift from the carefully managed messaging typical of AI labs.
Context and caveats
Hubinger's probability estimate is a personal view, not an official company forecast. Subjective probabilities on unprecedented events are not empirical measurements, and many researchers in the field assign far lower likelihoods to catastrophic outcomes. Some argue the capability jump Hubinger describes is not the trajectory the field is actually following.
What makes the statement significant is not its accuracy but its source: the person responsible for solving the alignment problem at one of the three most capable AI companies, speaking candidly on the day a colleague resigned over the same concern.
Anthropic has not issued a corporate response to Coxon's resignation. The company has faced scrutiny this year over the gap between its public warnings and its operational decisions, including abandoning a Responsible Scaling Policy commitment in February and incidents involving model containment failures.
These details were first reported by AI Watch.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call