NTI and Concordia AI Form Global Group on AI Bio-Risk Testing
New international working group will establish standards for evaluating whether AI systems could enable biological threats.
The Nuclear Threat Initiative and Concordia AI have established an international working group to develop common standards for testing whether artificial intelligence systems could lower barriers to biological misuse, according to an announcement from NTI.
The AIxBio Technical Working Group on Evaluation Practice, launched with technical input from SecureBio, brings together experts from organizations across China, the European Union, the United Kingdom, and the United States that conduct biological capability evaluations of AI systems.
The evaluation challenge
Biological capability evaluations assess whether AI systems possess capabilities that could make it easier to create or deploy biological threats. Frontier AI developers increasingly use these assessments to inform decisions about safeguards, access restrictions, monitoring levels, and deployment strategies.
The challenge arises when evaluation results from one organization inform risk decisions at another. Current evaluation methods and reporting practices vary significantly across jurisdictions and institutions. Results reflect both the AI model's actual capabilities and the specific testing methodology used—but inconsistent reporting makes it difficult to distinguish between these factors.
Without standardized approaches, differences stemming from how systems are configured or how tests are designed may be incorrectly attributed to the underlying AI model itself.
What the working group will do
The group will compare published evaluation approaches across different jurisdictions and institutional contexts to establish a common technical foundation. Its work will produce shared terminology and a minimum reporting baseline for evaluation results.
Specific objectives include clarifying what biological capability evaluations actually measure and how evaluation design relates to the threat scenarios they represent. The group will identify conditions under which results can be meaningfully compared across different systems and settings, and distinguish between what findings reveal about tested capabilities versus what they indicate about actual biological risk.
The working group will also examine how different evidentiary standards and action thresholds affect how results are used in practice.
Why it matters
As AI capabilities advance rapidly, the lack of standardized evaluation methods creates real governance challenges. Governments and companies making high-stakes decisions about AI deployment need to understand whether evaluation results are comparable and what they genuinely indicate about risk. Without common standards, the AI safety ecosystem risks fragmentation, where evaluations conducted in one context provide limited useful information for decision-makers elsewhere. Establishing shared methodological foundations now could prevent costly misinterpretations as biological AI capabilities continue to evolve.
Outputs and participation
The working group will publish methodological findings, voluntary reporting conventions, and research priorities to support more consistent interpretation of evaluation evidence by AI developers, governments, and third-party evaluators.
The initiative complements the AIxBio Forum's existing Working Group on Horizon Scanning, Risk Assessment, and Evaluations by focusing specifically on assessment design and comparability.
Organizations and individuals working on biological capability evaluations can contact Helia Samani at samani@nti.org to discuss participation or share related work, according to the announcement.
Details were first reported by the Nuclear Threat Initiative.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call