Microsoft has submitted evidence in ongoing lawsuits, arguing that its AI tool Copilot rarely copies large portions of news articles or books, countering allegations from The New York Times and authors that the company unlawfully uses their content to train AI models.
- Copilot chat logs show minimal text matching with NYT and authors’ works.
- Microsoft claims AI training use qualifies as fair use.
- The NYT contests Microsoft’s findings, highlighting business harm.
What happened
Microsoft provided 8.2 million chat logs from its AI system Copilot as part of discovery in copyright infringement lawsuits filed by The New York Times, authors, and other publishers. The logs were chosen specifically to include references likely connected to publisher websites and copyrighted content. Upon expert analysis, only a very small number of conversations actually contained significant overlap with copyrighted news articles or books, with 59,545 instances containing 16 or more matching words, and only 24 responses containing 30 or more matching words in author-related data.
Microsoft uses these findings to argue that its AI chatbot rarely reproduces large, substantive portions of copyrighted works. The company contends that this limited overlap should not amount to infringement. The New York Times strongly disagrees, asserting that Microsoft and OpenAI's use of their materials in training AI has substituted for original journalism and harms their business model.
Why it matters
The case centers on whether using copyrighted content as training data for AI systems like Copilot constitutes fair use, a key legal question with broad implications for AI development and content creators. Microsoft's argument relies on the claim that AI transforms the original material and does not simply replicate or replace it, supporting a fair use defense that could allow extensive use of protected works in AI training.
On the other hand, publishers and authors view this practice as unauthorized copying that threatens their livelihood and the sustainability of journalism and publishing industries. A ruling against Microsoft and OpenAI could impose stricter limits on AI training datasets and reshape how AI companies source information, potentially increasing costs or limiting AI capabilities.
What to watch next
The judge overseeing these consolidated cases may issue a summary judgment ruling to decide the dispute early without a full trial. Microsoft is currently pushing for such a ruling in its favor, seeking to end the case based on the limited evidence of direct copying. Meanwhile, the New York Times and other plaintiffs are pushing back, seeking to continue litigation to hold Microsoft and OpenAI accountable for what they characterize as theft of content.
The outcome could set a precedent for how courts handle AI training and copyright issues in the future. Additionally, public interest is heightened by the Trump administration filing a statement supporting OpenAI’s position, signaling government attention on the implications of AI technology and copyright law. Stakeholders in media, technology, and policy will be closely monitoring any decisions and subsequent legal developments.