AI Training Faces Data Scarcity as Models Approach Human Limits
Researchers are exploring synthetic data and elite expertise to push language models beyond general knowledge boundaries.
The artificial intelligence industry confronts an emerging challenge: what happens when machine learning models consume all available human knowledge and need to advance further?
Researchers are already tackling this question, according to work detailed by Northeastern University graduate student Harsh Raj, who has explored the data frontier during co-op positions at Bespoke Labs and Scale AI.
The elite knowledge problem
Large language models require massive training datasets to improve their capabilities. But as these systems approach the boundaries of general human knowledge, finding suitable training material becomes exponentially harder.
"It's very hard to collect data which trains the model to be better than humans, because there are very few humans who can create that data," Raj explained. The goal is creating AI that can reason at the level of Nobel Prize winners or Fields Medal mathematicians—but such individuals represent a vanishingly small pool of potential training sources.
Raj's research focuses on understanding where current models fail and identifying pathways to superhuman performance. At Scale AI, he investigates failure taxonomies—the specific reasons language models produce incorrect or inadequate responses—and proposes targeted improvements based on experimental work and scientific literature.
Synthetic data as a solution
One approach involves generating synthetic data: artificially created text or code that mimics high-quality human output without requiring actual human experts to produce it. During his work at Bespoke Labs, Raj helped create synthetic training material for AI agents in controlled learning environments.
The challenge lies in matching the quality of human-generated content. While computing power can be purchased, creating synthetic data that genuinely advances model capabilities remains experimental territory.
David Bau, a Northeastern professor of machine learning who secured a $9 million National Science Foundation grant in 2024 for AI infrastructure research, noted Raj's contributions to understanding the "black box" of language model reasoning. Raj developed methods to train models with structured internal thought processes, making their decision-making more interpretable.
Beyond individual expertise
The next frontier may involve training models on the collective knowledge of entire teams rather than individual experts. AI laboratories are exploring whether models can eventually automate scientific discovery itself—performing not just the work of a competent programmer but an entire startup's worth of specialized talent.
"You want to create data which is automating science, automating discovery," Raj said, though he dismissed concerns about AI surpassing human knowledge limits. "There are very few people on earth who actually have experience in that … so it's very hard to create data."
Why it matters
The data scarcity problem represents a fundamental constraint on AI advancement that computing power alone cannot solve. How the industry addresses this challenge will determine whether language models plateau at current capabilities or achieve the superhuman reasoning their developers envision. The solutions being explored—from synthetic data generation to targeting elite expertise—will shape the trajectory of AI development over the coming years.
These details were first reported by Northeastern Global News in a profile of Raj's research work.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call

