Vals Raises $40M to Fix AI Benchmarking's Credibility Problem
The Andreessen Horowitz-backed startup tests models on real-world tasks rather than abstract knowledge, keeping test materials private to prevent gaming.

A two-year-old startup is positioning itself to overhaul how the AI industry measures model performance, addressing growing concerns that companies have learned to game traditional benchmarking systems.
Vals announced a $40 million Series A led by Andreessen Horowitz last month, following a seed round from 8VC and Bloomberg Beta. The San Francisco-based company, founded in 2024, has grown from eight employees at the start of this year to 25, while reporting revenue eight times higher than the previous year, according to TechCrunch.
The benchmarking credibility gap
Co-founder Rayan Krishnan, 25, identified a fundamental problem: legacy benchmarking systems weren't designed for modern AI capabilities, and their public availability allows companies to train models specifically to perform well on those tests. The result is metrics that may look impressive but don't necessarily reflect real-world performance.
"We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance," Krishnan told TechCrunch during a tour of the company's Folsom Street office.
Vals takes a different approach. Rather than measuring abstract intelligence through publicly available tests like bar exams, the startup evaluates models on their ability to complete complex, domain-specific tasks in fields including law, finance, coding, cybersecurity, and biosecurity. Critically, Vals keeps its test materials private to prevent optimization against known benchmarks.
Why it matters
As AI companies prepare for public offerings—Anthropic is reportedly planning an IPO this year, and OpenAI is expected to follow—standardized, credible performance metrics become essential infrastructure. Investors, regulators, and enterprise buyers need trustworthy data about what models can actually do. If benchmarks can be gamed, the entire market operates on unreliable information. Vals is betting that as AI becomes core economic infrastructure, companies will pay for rigorous third-party validation the same way students pay for standardized tests.
Business model and expansion
Companies pay Vals to test their models, a revenue structure Krishnan compares to students paying the College Board for SAT exams. While it may seem counterintuitive for companies to pay to potentially learn their models underperform, the evaluations help engineering teams identify weaknesses and guide improvement efforts. The assessments are also becoming decision factors for enterprises evaluating which AI models to deploy.
Vals is expanding its evaluation scope beyond traditional industries. The company now offers benchmarks on recursive self-improvement, mental health applications, and even how models apply international humanitarian law under the Geneva Convention. It recently launched a program providing model evaluations to federal agencies.
Krishnan, who previously interned at Palantir and worked at Microsoft and Stanford's AI lab as an undergraduate, expects Vals' benchmarking approach to become central to how public AI companies report performance in SEC filings and discuss investment strategies.
The company plans to relocate to larger offices and add 10 to 15 additional staff members as it scales.
These details were first reported by TechCrunch.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call
