Microsoft Exec Describes AI Training as 'Largest Theft of Labor in Human History'
A recent internal memo from Microsoft's Director of Applied Science, Brent Hecht, has shed light on the company's AI training practices. In the memo, dated January 2023, Hecht described the mass scraping of journalism as 'the largest theft of labor in human history.' The statement is not a plaintiff's accusation, but rather a candid assessment from within Microsoft itself.
The internal communications highlight the harm caused by AI training practices in language that would never be approved for public consumption. A separate document warned of a 'real risk' that generative AI could significantly disrupt the employment of those who generated the training data. This threat is directly linked to scraping, which poses an existential danger to journalists' jobs.
OpenAI's internal characterization of publishers facing an 'existential threat' from substitutive products echoes Hecht's concerns. The companies' own analysts predicted that AI products would suppress publisher traffic, weakening newsrooms' capacity to produce content. This creates a 'doom loop,' where AI models rely on degraded content for quality.
The filings also reveal that OpenAI bypassed paywalls to obtain content and stripped copyright notices from datasets before training. Microsoft CEO Satya Nadella agreed in sworn testimony that chatbot conversations can substitute for visiting original publisher sites. He also testified that he would have required OpenAI to retrain its models had he known otherwise.
Project Mango, a collaboration between Microsoft and OpenAI, assembled a training dataset containing at least 160,903 unique works from the news publishers involved. The evidence presented points toward licensing frameworks as a more durable path forward rather than relying on legal maneuvering.