AI companies copying internet to train models biggest theft of labour in history: Microsoft executive
When AI companies like OpenAI, Microsoft and others started scraping vast amounts of internet data to train their models, not everyone inside the companies was happy about it. A Microsoft executive called the practice an "astonishing theft" and even described it as the "largest theft of labour in human history."

How do AI models like ChatGPT, Claude, Gemini know so much about the world? They were trained on vast amounts of data scraped from the web. And not all of that data was necessarily free to use. OpenAI, Anthropic and other AI labs have faced accusations from authors, publishers and news organisations over the use of their work to train AI models without permission or payment. And not everyone inside these companies was happy about AI scraping either. One Microsoft executive called it the “largest theft of labour in human history.”
The comment came from Brent Hecht, Microsoft’s director of applied science. His remarks came to light in unsealed court filings from the copyright case brought by The New York Times and other news organisations against Microsoft and OpenAI.According to the filing, while discussing the use of online content to train AI models, Hecht described the practice as the biggest theft in the history.
“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” he wrote. He also noted that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”
On the other hand, Microsoft and OpenAI have argued that their use of copyrighted material to train AI models is protected by fair use. The companies say their AI systems transform the material rather than simply republishing it. The New York Times and other news organisations meanwhile disagree, arguing that the companies’ AI products can substitute for the original news sources.
The filings also reveal how AI chatbots can change the way people access news. Instead of clicking on a news website, users can ask a chatbot and get the information directly in the AI platform. In his deposition, Microsoft CEO Satya Nadella acknowledged this, saying that talking to chatbots has “substituted giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
The news organisations say this could create a difficult cycle for publishers: AI models use content produced by news websites, while AI chatbots can also reduce the need for users to visit those websites. The filing cites Microsoft data showing that click-through rates for The New York Times and Daily News domains were 83 to 93 per cent lower on Copilot’s “answer engine” than on traditional Bing Search. The filing also cites OpenAI’s ChatGPT head Nick Turley describing the products as “largely substitutive” and saying they would become “more and more substitutive as they get better.”
The filings also include internal comments about how well AI models could handle news. OpenAI co-founder Greg Brockman wrote that the models were “particularly good at predicting text of news articles” and “very good at any news task”. He made these comments while discussing how the models performed on New York Times content.
The New York Times sued Microsoft and OpenAI in 2023, about a year after ChatGPT launched. Other news organisations, including the New York Daily News, The Intercept and the Center for Investigative Reporting, later joined the case. Microsoft and OpenAI have maintained their fair-use defence, while the news plaintiffs are now asking the court to rule in their favour on their copyright claims.

