
Rich Sutton is one of the pioneers behind the key technologies driving today’s artificial intelligence boom. Now, however, he believes the path technology giants are taking to keep AI advancing has serious problems.

The Canadian computer scientist and Turing Award winner said on a Sequoia Capital podcast released Tuesday local time that the AI industry is increasingly relying on synthetic training data, which may be steering the entire industry in the wrong direction.
Asked whether synthetic data could help large language models continue to scale, Sutton said bluntly: “This is absolutely a major mistake. Maybe this will become the next important lesson.”
Synthetic data refers to information artificially generated by algorithms or AI models, rather than collected from raw data in the real world. As AI companies continue searching the internet and other new sources for training data, synthetic data is becoming increasingly attractive. For example, self-driving training can use computer-generated images of cars, while AI fraud-detection systems can use fabricated bank transaction records.
To some extent, technology giants also agree with Sutton’s view.
Technology companies are doing whatever it takes to find more real-world data. OpenAI previously put out a public call for large-scale proprietary datasets, seeking data that could not easily be obtained from the internet for training AI models. Earlier this week, Google agreed to pay $10 million (Note: approximately RMB 6,756.9 ten-thousand yuan at the current exchange rate) to acquire the internal data and software of the bankrupt Spirit Airlines, underscoring the growing value of proprietary real-world data for AI training.
On Tuesday’s podcast, Sutton said artificially generated data cannot adequately replace real-world data.
He used human behavior as an example. “We cannot manufacture synthetic data for other people’s thoughts,” he said.
Sutton favors using real experience data instead: information that AI agents obtain by interacting with real environments, observing what actually happens, and continuously learning from the results of their actions.
“You cannot use synthetic data to predict how a drone will interact with its environment in the real physical world,” Sutton said. He also pointed out that variables such as friction and wear exist in a robot’s motors. “This world is extraordinarily complex, and any simulation of it captures only a tiny part of the whole.”
Today, this philosophy has become the direction of a startup.
Last month, Sutton and his former student Khurram Javed co-founded Oak Lab. The startup is developing AI agents that can continuously learn from their own experiences, rather than relying primarily on enormous, pre-organized datasets.
Oak Lab has not yet disclosed its funding rounds or investor information.
