Mathematicians Design New AI Benchmark for Research Progress
Harmonic and the American Institute of Mathematics are creating evaluations that measure whether AI helps solve real research problems, not just standardized tests.

Mathematicians take control of AI evaluation
Harmonic, an AI startup backed by Nvidia, has formed a partnership with the American Institute of Mathematics to develop a new benchmark designed by working mathematicians rather than AI researchers. The collaboration aims to measure whether AI systems actually help advance mathematical research, not just whether they can answer test questions correctly.
The partnership was disclosed exclusively to Axios.
Why it matters
AI companies have relied heavily on benchmark scores to demonstrate model improvements, but these evaluations are increasingly criticized as disconnected from real-world expert needs. By giving domain specialists direct control over evaluation design, this initiative represents a fundamental shift in how the industry measures progress—prioritizing practical research utility over performance on standardized assessments.
A benchmark built for working researchers
The American Institute of Mathematics and Harmonic will jointly create an open benchmark featuring mathematical problems selected by practicing mathematicians. The initial release includes more than 50 number theory problems that span from generating examples to tackling open research questions.
Unlike conventional AI evaluations, this framework assesses two dimensions: whether models produce correct answers and whether they help mathematicians make meaningful progress on challenging research problems. The benchmark materials and evaluation criteria will be released publicly, enabling academic and industry researchers to contribute both problems and results.
"Harmonic's goal is to amplify human creativity and discovery by designing AI tools that complement human insight," said Tudor Achim, Harmonic's CEO. AIM Executive Director Sergei Gukov described the partnership as "an important step" because it centers mathematicians' needs, adding that AIM intends to release additional benchmarks covering other research areas.
Industry rethinks evaluation standards
The partnership reflects broader dissatisfaction with existing AI benchmarks. OpenAI recently challenged widely used evaluations, arguing that many are fundamentally flawed. The company advocated for benchmarks "built by experienced software developers specifically to test model capabilities."
Companies now argue that next-generation evaluations should test how AI assists experts in solving real-world problems rather than measuring performance on standardized exams disconnected from practical application.
Harmonic has consistently maintained that mathematics provides an especially clear test of AI reasoning because many answers can be formally verified. The company was founded in 2023 and counts Robinhood CEO Vlad Tenev among its backers.
Domain experts redefine AI progress
The collaboration signals that experts in specific fields are claiming a larger role in defining what constitutes meaningful AI progress. Rather than allowing AI developers to set evaluation standards unilaterally, this approach positions domain specialists as the ultimate judges of whether systems deliver genuine value for their work.
These details were first reported by Axios.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call