OpenAI and Microsoft Accused of Knowingly Using Stolen News Content
New court filings in The New York Times' lawsuit against OpenAI and Microsoft reveal that the companies' executives were aware of the risks of mass use of copyrighted materials to train artificial intelligence models. According to the plaintiff, OpenAI's intermediate training datasets contained over 91,692 copies of material from The New York Times, Daily News, and the Center for Investigative Reporting.
One dataset, compiled using Common Crawl, contained more than two million documents from nytimes.com. Microsoft internal memo from January 2024 shows that the Copilot system had reduced the number of visits to The New York Times website by 93% compared to traditional Bing search.
Microsoft CEO Satya Nadella testified that material available behind a paywall should be licensed by anyone using it to train models or supplement responses. He also stated that he would have required OpenAI's models to be retrained had he known that paywalled data was being used.