OpenAI and Microsoft's AI Training Practices Exposed as Massive Theft
A court case between The New York Times and tech giants OpenAI and Microsoft has revealed internal admissions from top executives that threaten their legal defense of 'fair use'.
Microsoft's Director of Applied Science, Brent Hecht, described the data scraping required for AI training as an 'astonishing theft of unprecedented proportions' and potentially 'the largest theft of labor in human history.'
Hecht further noted that winning a fair use defense on these grounds would 'make a complete mockery of the idea of fair use'. OpenAI's Head of ChatGPT, Nick Turley, admitted in internal chats that their products are 'largely substitutive, period', posing an 'existential threat' to news publishers by diverting readers and cutting off traffic.
The court filings claim that OpenAI leadership used technical workarounds and 'hacks' to bypass paywalls to extract training data. Microsoft's CEO, Satya Nadella, testified that he would have forced OpenAI to retrain its models from scratch had he known paywalled content was being scraped.
Internal Microsoft presentations outlined a self-defeating 'doom loop', acknowledging that degrading or bankrupting the news industry would ultimately hurt the quality of the web and data supplies that future LLMs rely on. News outlets argue that these unsealed documents serve as a smoking gun, proving that tech executives privately knew their chatbots served as direct substitutes for traditional reporting, caused massive drops in referral traffic, and relied on unauthorized, systemic scraping of copyrighted material rather than legally protected transformation.