AI Pioneer Rich Sutton Calls Synthetic Training Data a Mistake
The Turing Award winner argues real experiential data—not algorithmically generated substitutes—is essential for advancing AI systems.

AI luminary challenges industry's data strategy
Rich Sutton, a Canadian computer scientist whose work laid foundations for today's AI systems, is pushing back against one of the industry's most popular solutions to the data scarcity problem. In a podcast episode released Tuesday by Sequoia, the Turing Award winner declared that relying on synthetic training data to continue scaling large language models represents "a big mistake."
Synthetic data—information artificially generated by algorithms rather than collected from real-world sources—has become an increasingly attractive option as AI companies exhaust readily available internet content. The approach includes computer-generated images for autonomous vehicle training and fabricated financial records for fraud detection systems.
Why it matters
As AI development costs soar and companies race to build more capable models, the choice between synthetic and real-world data could determine which approaches succeed. Sutton's critique carries weight given his foundational contributions to reinforcement learning, and his stance reflects a broader tension in the industry between scaling shortcuts and fundamental capability building.
The limits of manufactured information
Sutton argues that algorithmically produced data cannot adequately substitute for information drawn from actual experience. He pointed to human behavior as a domain where synthetic alternatives fall short, noting that "there's no way we can have synthetic data for other people's minds."
The physical world presents similar challenges. When discussing autonomous systems like drones, Sutton emphasized that simulations cannot capture the full complexity of real environments, including variables such as friction and motor wear. "The world is infinitely complex, and any simulation of it is like, microscopic," he said.
Tech giants hedge their bets
Despite the appeal of synthetic data, major technology companies appear to share some of Sutton's concerns. OpenAI has publicly sought large-scale proprietary datasets not available through standard web scraping. This week, Google agreed to pay $10 million for internal data and software from bankrupt Spirit Airlines, underscoring the premium placed on authentic, real-world information.
A new approach takes shape
Sutton's preferred alternative centers on what he calls real experiential data—information that AI agents gather by directly interacting with their environments, observing outcomes, and learning from consequences over time. This philosophy now underpins Oak Lab, a startup Sutton launched last month with former student Khurram Javed. The company is building agents designed to learn continuously from their own experiences rather than depending primarily on large pre-assembled datasets. Oak Lab has not yet disclosed funding details or investor information.
These details were first reported by Business Insider.
This is an original analysis by the Omega editorial team. Source reporting: AI Watch.
Want systems like this working for your business?
Book a Call