Science

AI models need 100,000x more data than children to learn language

Researchers are building baby-scale language models to understand how kids master speech with far less input than machines require.

Omega Editorial· August 24, 2026· 3 min read

Large language models have achieved remarkable fluency in human language, but they require an extraordinary amount of data to get there. Modern LLMs like Claude and GPT models train on trillions of tokens—roughly 100,000 times more language than a child experiences while learning to speak.

This stark contrast, known as the data efficiency gap, has become a central puzzle for both AI researchers and cognitive scientists. A child raised in a linguistically rich environment hears around 100 million words by their preteen years. By age 20, adding reading might bring that to 300 million words. Meanwhile, Meta's Llama 3.1 processed 15 trillion tokens during pretraining alone, and frontier models may train on ten times that amount.

The scale difference is almost incomprehensible. As Ethan Gotlieb Wilcox, a cognitive scientist at Georgetown University, puts it: "Claude has seen the amount of language that an entire city will experience in one generation." If printed on paper, the training data for a modern LLM would stack higher than the International Space Station. A child's 100 million words would reach just 20 meters.

Why it matters

The data efficiency gap has practical implications beyond academic curiosity. The internet's supply of easily available training data could run dry as early as the 2030s, creating a ceiling for traditional scaling approaches. Understanding how children learn language with minimal data could unlock more efficient AI training methods—critical for developing models in minority languages or training on video data. It also addresses fundamental questions about human cognition: whether we're born with language instincts or can learn purely from experience.

Testing theories with baby-scale models

Researchers have launched competitions like BabyLM to train language models on developmentally plausible datasets of just 100 million words, drawn from storybooks, dialogue, and transcripts of child-directed speech. These models are evaluated using the same grammar benchmarks psycholinguists use with humans.

The results have already challenged assumptions. Curriculum learning—starting with simple data and progressing to complex inputs, mimicking how adults talk to babies—was the most popular approach in early competitions but didn't work as well as expected. "It seems like these transformers don't really need to have their data ordered in such a way to learn effectively," says Aaron Mueller, a computer scientist at Boston University.

The 2024 winning model, GPT-BERT, trained on about 100 million words, beat Meta's Llama 2 70B—pretrained on 15,000 times more data—on one benchmark. Still, these baby-scale models can't match commercial LLMs in overall performance.

Learning through a child's eyes

Some researchers believe closing the data gap requires moving beyond text. Children learn through multiple senses, especially vision and hearing. Projects like SAYCam have equipped babies with headcams to capture their actual visual experience—revealing a world much more focused than adults assumed, with objects close-up and a "forest of knees."

Brenden Lake at Princeton trained a model on 61 hours of SAYCam footage that learned to identify objects and associate them with words, without the biases many theories suggest children need. "It turns out you can get a real start on language learning using a lot less than what a number of theories suggested," Lake notes, though the model still falls short of two-year-old capabilities.

The challenge of replicating human language acquisition remains unsolved, but researchers are making progress understanding what makes children such remarkably efficient learners.

These findings were first reported by MIT Technology Review.

#language models#cognitive science#child development#data efficiency#llm training#natural language processing

This is an original analysis by the Omega editorial team. Source reporting: AI Watch.

Want systems like this working for your business?

Book a Call

More in Science

Science· 3 min read

Singapore Deploys 16 Million Living Neurons in Data Center Prototype

A partnership between NUS Medicine, DayOne, and Cortical Labs has created the first biological computing server rack using lab-grown human neurons alongside silicon hardware.

Via AI Watch · Aug 24, 2026
Science· 3 min read

DOE Genesis Mission Awards UC San Diego $M for AI Physics Tools

SIDERIUS project will deploy intelligent agents to accelerate fundamental physics research and quantum computing applications.

Via AI Watch · Aug 21, 2026
Science· 3 min read

First Blinded AI Antibody Design Benchmark Shows Wide Variance

A prospective competition involving 29 organizations and 511 sequences reveals no dominant computational approach—and exposes a regulatory validation gap.

Via AI Watch · Aug 21, 2026