A recently unsealed 92-page legal brief in the ongoing copyright litigation against OpenAI includes stark warnings from a Microsoft research scientist about the profound risks posed by AI training on publisher content. The brief spotlights internal admissions that AI systems may irreversibly disrupt the economic foundations of news and media providers.
- Microsoft data shows up to 93% drop in click-throughs from AI-powered search.
- AI training datasets contain tens of thousands of copyrighted articles from major publishers.
- Internal documents warn AI is disrupting both employment and economic models for news media.
What happened
A 92-page brief filed in the consolidated OpenAI copyright litigation was unsealed, revealing detailed internal Microsoft documents and statements from its own scientist, Dr. Brent Hecht. Hecht and others describe the use of content from news publishers to train AI as a massive, unprecedented form of labor theft. The brief quotes Hecht calling the practice possibly the largest theft of labor in human history.
The documents disclose Microsoft’s recognition that AI-powered products like Copilot and Bing Chat have caused severe declines in web traffic to publishers, with click-through rates falling 51% to 93%. It also reveals concerns that AI models replicate patterns from their training content without compensating the original creators, threatening the economic foundation of media suppliers.
Why it matters
The brief reveals a paradox within Microsoft and OpenAI’s AI strategy: their products depend heavily on content created by news organizations and publishers, but simultaneously undermine the economic incentives that sustain that content. This ‘doom loop’ risks destabilizing the very ecosystem that fuels AI development.
Acknowledgments from multiple internal sources underscore the existential threat that AI poses to journalists and publishers, potentially disrupting employment and destroying supply chains of labor and creativity. These concerns also question the boundaries of fair use in AI training, which could reshape content licensing and compensation laws going forward.
What to watch next
The ongoing court case will be critical in determining how copyright law is interpreted in the context of AI training datasets, especially for copyrighted news content. Outcomes could mandate changes in how tech companies source, use, and compensate for the creative labor embedded in AI training materials.
Publishers and media companies continue to monitor and challenge AI platforms as they deploy increasingly substitutive models that replicate human-generated content. Future regulation or settlement agreements may emerge to address the economic disruption and clarify rights and responsibilities for AI-generated content.