Microsoft and OpenAI's AI Training Data Raises Red Flags
Newly unsealed court documents reveal internal concerns at Microsoft and OpenAI about using millions of news articles to train AI systems, including ChatGPT. One Microsoft executive described the practice as 'the largest theft of labour in human history', while an OpenAI executive warned that AI products posed an 'existential threat' to publishers.
The disclosures emerged from a copyright lawsuit brought by The New York Times against Microsoft and OpenAI, which centers on allegations that the companies used copyrighted news content without permission. The companies have argued that the legal doctrine of fair use protects their use of the material.
Internal documents show Brent Hecht, Microsoft's director of applied science, expressed concern about the scale at which AI companies were collecting online content. He described it as 'a product that destroys its supply chain', referring to AI systems' dependence on content produced by publishers and journalists.