OpenAI Accused of Stealing Millions of News Articles to Train ChatGPT
The New York Times has accused OpenAI and its investor Microsoft of stealing news content to train their AI models. The NYT alleges that OpenAI scraped content from over 10 million articles, with nearly a third coming from the NYT itself.
Brent Hect, Director of Applied Science at Microsoft, described the alleged theft as 'an astonishing theft of unprecedented proportions' and possibly the 'largest theft of labor in human history.'
The NYT sued OpenAI and Microsoft three years ago for allegedly stealing copyrighted material to train ChatGPT. Several other news publishers have joined the lawsuit, including Ziff Davis, Mother Jones, The Intercept, and multiple local newspapers.
Mircosoft claims Hect's statements are just 'one employee's individual perspective' and do not represent the company's views. OpenAI has signed content licensing deals with many news publishers since launching ChatGPT in 2022, but concerns persist that generative AI software is reducing traffic to internet news sites.
The plaintiffs have requested a summary judgment in their favor, which would prevent the case from going to trial. A ruling is not expected until 2027.