Microsoft Defends Copilot Against Copyright Claims with AI Training Data Analysis
Microsoft is fighting copyright claims from publishers and authors over its AI chatbot Copilot. In new legal filings, Microsoft argues that using copyrighted content for AI training datasets should be considered fair use.
The company provided 8.2 million chat logs to an expert hired by news publishers as part of the lawsuit's discovery process. The analysis found that only 59,545 conversations contained at least 16 words in common with news content used to ground the AI model.
Microsoft claims that this shows Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original.
The company argues that while systems like Copilot rely on using copyrighted material, the resulting systems are used for significantly different purposes than the original. Microsoft says that the fact that they sometimes reproduce sections of text 'hardly undermines the transformative purpose of LLM training.'