How to Design a U.S. AI Safety Regulator That Actually Works
As the White House reviews proposals for a FINRA-style AI oversight body, five critical design choices will determine whether it becomes trusted infrastructure or a failed experiment.

The White House is reportedly reviewing a proposal to create a U.S. AI oversight body modeled after the Financial Industry Regulatory Authority (FINRA), according to a Bloomberg report from mid-July. The framework, originally proposed by Google DeepMind CEO Demis Hassabis and developed with Treasury Secretary Scott Bessent's involvement, would require frontier AI companies to submit their models for safety testing up to 30 days before public release.
Under the proposal, these systems would be evaluated for dangerous cyber, biological, and deceptive capabilities. Participation would start voluntarily, then become mandatory for U.S. deployment once the process proves viable. While the institutional form appears sound, former government AI leaders warn that critical design questions remain unanswered—and getting them wrong could undermine the regulator's credibility from day one.
Five unresolved design questions
Vinh X. Nguyen, former chief AI officer at the National Security Agency, and Elham Tabassi, former chief AI advisor at the National Institute of Standards and Technology, identify five structural challenges that must be resolved:
Independence and funding conflicts. Industry-funded oversight raises immediate concerns about conflicts of interest. The issuer-pays credit-rating model that contributed to the 2008 financial crisis offers a cautionary parallel. While FINRA requires public governors to outnumber industry representatives, concerns about "composition drift"—where shifting conflicts gradually move the board away from public priorities—persist.
National security classification. Unlike FINRA examinations, AI evaluations will inevitably produce findings that are themselves national security assets. The oversight body will also become a high-value espionage target by holding prerelease access to every frontier lab's models. Mishandling classification could either turn the body into an intelligence subsidiary operating outside proper safeguards, or leave critical capabilities unexamined until adversaries discover them first.
International legitimacy. Allies in Brussels, London, Seoul, and Tokyo are unlikely to embrace an American industry-funded body as an arbiter of international standards. Public trust also requires examining concerns beyond what companies and security agencies prioritize—including jobs, privacy, and inequality.
Guaranteed access and talent. Without mandatory model access, an oversight body can only evaluate what companies allow it to see. Currently, access is granted at company discretion, under non-disclosure agreements, and on a release-by-release basis. The scarcest resource will be talent capable of testing frontier capabilities in scientifically valid ways—expertise that has increasingly consolidated within private companies.
Measurement science. No settled science of frontier AI assessment exists. Benchmarks quickly saturate, results carry unreported uncertainty, and point-in-time certification fails for systems that change between measurements. No broader accountability ecosystem exists to translate findings into practical deployment standards.
Design recommendations
The authors propose separating three functions—safety specifications, infrastructure, and evaluation—structurally and financially from the start. Safety standards should be funded exclusively by neutral, non-industry money. Industry resources should instead support shared testing infrastructure through organizations like MLCommons.
Evaluations themselves should be conducted by institutions with proven public-interest commitment, such as pairing RAND's national security expertise with civil rights organizations like the Center for Democracy & Technology. This structural separation addresses what board composition alone cannot deliver.
Other critical recommendations include binding funding commitments to guaranteed model access, designing classified information protocols before the first classified finding occurs, and creating stable career paths for evaluators to counter private-sector talent asymmetry. The authors also emphasize embedding allied governments and societal-impact evaluation into the architecture immediately, not as afterthoughts.
Why it matters
The institutional design choices made in the coming months could establish an AI oversight framework that persists for decades. A well-designed body could become a quality standard that American AI models carry into global markets as a competitive advantage. A poorly designed one—perceived as an industry stamp rather than independent oversight—will lack credibility at home and abroad, requiring costly rebuilding after its first crisis of confidence. With measurement science still immature and AI capabilities advancing rapidly, the window to establish durable, trusted infrastructure is narrow.
These details were first reported by Vinh X. Nguyen, Elham Tabassi, and Kat Duffy writing for the Council on Foreign Relations.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
