Microsoft and OpenAI Executives Called AI Training Practices 'Theft'
Newly unsealed documents in The New York Times' three-year-old copyright lawsuit against OpenAI and Microsoft reveal internal admissions that executives privately viewed AI training practices as theft.
The companies allegedly scraped paywalled content undetected, built massive training datasets, and deliberately stripped copyright notices before content reached their models. This has put the underlying content supply chain at risk, according to Microsoft's director of Applied Science, Brent Hecht.
Hecht described declining click-through traffic to Times content as a 'doom loop,' with one instance seeing a 93% decline following Copilot's launch. In internal documents, he also called the scraping 'the largest theft of labor in human history.'
Microsoft CEO Satya Nadella testified that paywalled content should be licensed before use in training, and said he would have required OpenAI to retrain its models had he known paywalled material was involved.