Microsoft Executive Describes AI Training Practices as 'Theft of Labor in Human History'
A recent unredacted filing in the copyright lawsuit between The New York Times and OpenAI has revealed internal documents that describe AI training practices as 'theft' by top Microsoft executive Brent Hecht. In the documents, Hecht writes that AI products pose an 'existential threat' to publications and that there is a 'real risk' of significant disruption to employment due to generative AI.
The lawsuit alleges that OpenAI and Microsoft used copyrighted material without permission to train their AI models. The companies allegedly obtained content by scraping it from the Bing Index, bypassing paywalls undetected, and stripping copyright notices from training data.
Internal documents show that OpenAI's mid-training datasets contain over 91,692 copies of works published by The New York Times, Daily News, and Center for Investigative Reporting. Microsoft's Copilot 'answer engine' caused click-through rates for The New York Times' domain to drop as much as 93% compared to traditional Bing search.
Micorsoft CEO Satya Nadella testified in a deposition that he would have required OpenAI to retrain its models if he had known they were scraped from behind paywalls. OpenAI employees allegedly came up with a plan to circumvent paywalls without detection, and the companies' actions cut against several pillars of the fair-use test.