Anthropic researcher quits, warns AI labs racing toward risky superintelligence
Jacob Coxon's public resignation highlights growing internal concern that leading AI companies are prioritizing speed over safety in pursuit of advanced systems.
An artificial intelligence researcher has publicly resigned from Anthropic, one of the industry's leading AI safety companies, warning that both it and OpenAI are gambling with humanity's future in their rush to develop superintelligent systems.
Jacob Coxon announced his departure on Tuesday in a series of posts on X, stating bluntly that "neither company is acting responsibly." According to Coxon, the two organizations are "racing straight to self-improving superintelligence and gambling with our lives."
The warning from inside
Coxon's resignation is notable not just for its public nature but for the specificity of his concerns. He cautioned that the AI field is approaching "superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources" — capabilities he believes are being underestimated by those building them.
His critique distinguished between the two companies' approaches. He characterized OpenAI members as not having "fully internalized the stakes of artificial intelligence," while describing Anthropic as "locked in a race to get there first" — believing that no one else will act responsibly, so they must reach advanced AI themselves despite the risks.
Why it matters
This resignation comes from within Anthropic, a company explicitly founded on AI safety principles and positioned as a more cautious alternative to OpenAI. When safety-focused organizations face internal dissent over their risk management, it signals that the competitive dynamics of AI development may be overriding safety considerations even among the most conscientious players. For business leaders evaluating AI partnerships or investments, these warnings suggest that regulatory frameworks and safety standards are lagging behind the technology's advancement.
Senior researcher confirms the stakes
In a striking validation of Coxon's concerns, Evan Hubinger, Anthropic's Alignment Science lead, responded directly to the resignation posts, confirming the severity of the risk assessment.
"We really do earnestly believe AI could kill all humans!" Hubinger wrote. He added that he personally estimates the probability of this outcome at "greater than 10% within the next decade."
Hubinger acknowledged that while "Anthropic is trying its best," the company does "not yet have a plan to solve alignment for superintelligence and are not clearly on track to." Alignment refers to the technical challenge of ensuring advanced AI systems reliably pursue intended goals without causing unintended harm.
Coxon urged researchers at AI labs to exercise extreme caution over the coming years of development, as systems approach and potentially exceed human-level capabilities across domains.
The details of Coxon's resignation and the subsequent exchange were first reported by The Jerusalem Post.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call