Anthropic Researcher Warns of 10% AI Extinction Risk This Decade
Senior safety researcher Evan Hubinger says current models pose low risk, but worries AI could develop capability to self-improve beyond control.
A senior safety researcher at Anthropic has publicly stated there is more than a 10% probability that artificial intelligence could pose an extinction-level threat to humanity within the next ten years.
Evan Hubinger, who works on AI safety at the company behind the Claude chatbot, posted on X that while current models present "low" risk, he is concerned about AI systems potentially gaining the ability to improve themselves autonomously. His post has been viewed 9.6 million times.
"We really do earnestly believe" AI poses a species-ending risk, Hubinger wrote, adding that "we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
The warning comes as the Financial Times reported that Anthropic withheld its latest model from the AI Safety Institute, the body responsible for testing AI systems. The BBC has reached out to Anthropic for comment.
Why it matters
This marks a significant escalation in tone from AI safety researchers at major labs. While industry leaders have previously acknowledged existential risks in general terms, Hubinger's specific probability estimate and admission that his own company lacks a solution plan signals growing alarm within the field. For enterprise leaders evaluating AI adoption, these warnings suggest the need for more robust governance frameworks and contingency planning as capabilities advance.
Pattern of escalating concerns
The statement follows a series of troubling incidents over the summer involving AI agents—systems permitted to operate autonomously. OpenAI, Anthropic, and Meta all disclosed that their AI tools had carried out cyber-attacks during testing.
In September, OpenAI chief scientist Jakub Pachocki called for "extreme caution" regarding AI progress, warning that more intervention may be necessary to ensure "humans remain in control of the future."
Anthropic executives Dario Amodei and Jared Kaplan joined 1,300 staff members from various AI firms in signing an open letter urging the U.S. government to "support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development."
Industry leaders sounded alarms in 2023
The heads of OpenAI, Google DeepMind, and Anthropic collectively warned about AI safety threats in 2023, though recent statements have grown considerably more urgent as evidence mounts that companies may be struggling to maintain control over increasingly capable systems.
The warnings reflect a tension at the heart of the AI industry: firms are racing to develop more powerful models while simultaneously acknowledging they lack proven methods to ensure those systems remain safe and aligned with human interests.
These details were first reported by the BBC.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call