Microsoft Says Copilot Training on Copyrighted Books is Fair Use
Microsoft has moved for summary judgment in a consolidated copyright litigation against book authors and news publishers. The company argues that training its large language model, Copilot, on copyrighted books is fair use as a matter of law. Microsoft claims that an expert's review of 8.2 million conversations found only 24 responses containing at least 30 words matching the authors' asserted works.
The expert, Dr. Shawn Shan, used an adversarial extraction protocol to test Copilot's ability to reproduce book passages. In a laboratory setting, he fed the model verbatim book passages hundreds of words long and found that fewer than 1 percent produced any 30-word match.
In a real-world setting, Dr. Shan reviewed 8.2 million conversations from Microsoft's Copilot product and identified 24 responses containing 30 matching words - a rate described as .00029 percent. For 202 of the 212 asserted books, he found no regurgitation in the logs at all.
Microsoft also cited the authors' own sales evidence, stating that they are selling just as many books as they would have had ChatGPT and Copilot never been released. The company argues that training an LLM on copyrighted books serves a technological purpose: building a model that generates natural-language responses to prompts.