A federal court has approved a $1.5 billion settlement in favor of authors and publishers whose copyrighted works were unlawfully used by Anthropic PBC to develop its AI model, marking the largest copyright class action settlement to date.

  • Settlement covers half a million pirated books used to train AI.
  • Authors and publishers to receive $3,000 per infringed work.
  • Court ruling leaves AI training legality unresolved overall.

What happened

A federal judge approved a $1.5 billion settlement to compensate authors and publishers whose creative works were used without authorization by Anthropic PBC in training its AI chatbot Claude. The disputed materials, nearly 500,000 books, were allegedly sourced from pirate websites such as Library Genesis and Pirate Library Mirror. The ruling requires Anthropic to destroy all infringing copies of the content.

Judge Araceli Martínez-Olguín deemed the settlement fair and adequate given the complex, novel copyright claims involved. While a previous judge ruled that training AI on copyrighted content can constitute fair use, obtaining these materials through piracy did not fall under such protection and thus constituted infringement.

Why it matters

This $1.5 billion settlement is the largest copyright class action payout in history, underscoring the significant legal and financial risks for AI developers who rely on copyrighted works obtained through unauthorized means. It sends a strong signal about respecting copyright boundaries in AI training data acquisition.

The case highlights ongoing legal ambiguity over the treatment of copyrighted content in AI training processes. Though some courts have recognized fair use for training purposes, the provenance of training data remains a crucial factor. This settlement may influence how other major AI firms approach data sourcing amid numerous pending lawsuits.

What to watch next

With this precedent set, key AI companies like Google, Meta, OpenAI, and others still face pending litigation over their use of copyrighted materials for model training. Observers will closely monitor how courts balance fair use and copyright infringement claims going forward and whether further settlements follow.

Anthropic’s deputy general counsel noted that the settlement was reached after a ruling affirming trainability of books under fair use but emphasized that reliance on pirated content was the infringing factor. Industry stakeholders should watch regulatory and judicial developments that may clarify the boundaries for ethical and legal AI training practices.

Source assisted: This briefing began from a discovered source item from SiliconANGLE. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings