AI Safety Evaluators Lack Power to Stop Model Releases
Anthropic and OpenAI propose embedding third-party reviewers inside their labs, but experts say the arrangement falls short of meaningful regulatory oversight.

Leading AI companies are proposing a new approach to model safety: embedding independent evaluators inside their labs with access to unreleased systems. But legal and regulatory experts warn the arrangement lacks the enforcement teeth necessary to function as genuine oversight.
Anthropic CEO Dario Amodei recently committed to providing third-party safety evaluators with access comparable to internal risk teams and the right to publish findings without editorial control. The proposal came after a former researcher resigned with warnings that frontier labs were racing toward systems they might not be able to control. OpenAI has made similar commitments but has not yet detailed how its framework will operate.
Amodei explicitly compared the arrangement to bank supervisors embedded within financial institutions. But Julie Andersen Hill, dean of the University of Wyoming College of Law and an expert on banking regulation, says the comparison breaks down under scrutiny.
The power gap
At major banks, government examiners maintain offices inside the institution with continuous access to systems and employees. Critically, they possess enforcement authority: they can order a bank to halt practices, restrict growth, force management changes, or in extreme cases shut down operations entirely.
Anthropic's proposed evaluators would investigate and report, but Amodei's plan grants them no comparable authority to prevent a model from being trained or deployed. "If you don't give them that kind of power, I don't know what they are doing," Hill told CNBC, which first reported the details.
Neither Anthropic's proposal nor OpenAI's existing framework gives outside evaluators independent authority to halt development or deployment decisions. The companies did not respond to requests for comment.
What evaluators actually see
Albert Ziegler, who leads AI evaluation at cybersecurity firm XBOW, said his team has received early access to unreleased models from Anthropic, OpenAI, and other developers. The work happens in XBOW's own environment, with findings typically shared back to the model providers.
The day-to-day reality is less dramatic than the existential framing suggests, Ziegler said. His team might find that a model produces errors under unusual inputs or that safety systems need more frequent intervention. "But the kind of insidious, catastrophic consequences produced by subterfuge combined with unprecedented abilities that people are afraid of — that's not something we've seen ourselves," he noted.
Evaluators can document risks a developer missed and "compel an informed decision before release," Ziegler said. But the final call remains with the company. "It's true that we don't have any veto power," he acknowledged.
Independence questions
Anthropic has pointed to the nonprofit Model Evaluation and Threat Research (METR) as a potential embedded evaluator. The organization recently reviewed cybersecurity incidents involving Anthropic's Claude model. A former Anthropic researcher, Joe Benton, recently left the company to join METR and work on embedded assessments.
METR says it accepts no cash payments from AI companies or their executives. However, in its own reporting, the organization acknowledged that some employees have strong social ties to AI company workers and share research facilities with lab employees. The connections illustrate how small and interconnected the frontier AI safety field remains.
The structural tension is familiar from other industries. Christina Ho, chief assurance officer at accounting firm Oath and former board member of the Public Company Accounting Oversight Board, noted that auditors face similar conflicts because clients pay for the work that must challenge them. But in financial auditing, criminal liability provisions help enforce independence.
Why it matters
The debate over AI safety evaluators reveals a fundamental question about how frontier AI development should be governed. If companies retain full control over training and deployment decisions while evaluators can only observe and report, the arrangement may provide transparency without accountability. Hill framed the test simply: "If you really believe that AI has the power to destroy society, then you have to have an independent supervisor that has the ability to pull the plug on it." Without enforcement authority, embedded evaluators may function more like internal compliance departments than independent regulators — offering the appearance of oversight without its substance. For business leaders evaluating AI partnerships and policymakers designing regulatory frameworks, understanding this distinction is critical.
The details were first reported by CNBC.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
