Microsoft Execs Slam Web Scraping for LLM Training: 'Largest Theft of Labor in Human History'
A recently unsealed court document has revealed that executives from OpenAI and Microsoft expressed concerns about the use of web scraping to train large language models (LLMs). Dr. Brent Hecht, director of Applied Science at Microsoft, described it as 'the largest theft of labor in human history,' which could create a 'doom loop.'
Nick Turley, an executive from OpenAI, also referred to the practice as an 'existential threat' to publishers.
The comments were made as part of a lawsuit launched by The New York Times against OpenAI and Microsoft in 2023. The lawsuit alleges that big tech companies broke copyright law by scraping millions of news articles off the internet without permission or compensation to train advanced LLM systems.
The use of web scraping has raised concerns about the impact on journalism, with some executives acknowledging the potential danger to news gathering and the role of AI models in substituting for human work. The court documents also revealed that OpenAI's president, Greg Brockman, responded 'ah nice' when informed of a new paywall hack.