A hidden Airtag placed in a rare book has exposed Amazon’s practice of buying rare books in bulk, scanning them for AI training, and then destroying the physical copies, raising ethical concerns among booksellers and the AI community.

  • Amazon uses rare books to create proprietary AI training datasets.
  • Books are physically damaged and destroyed after scanning.
  • The practice intensifies ethical debates around AI data sourcing.

What happened

Amazon’s AI development team has been revealed to purchase rare books in bulk, only to physically dismantle these volumes after scanning their contents for use in training advanced AI models. This was confirmed when a bookseller planted a hidden Airtag in a rare book purchased by Amazon, tracking the item to an AI training facility in Las Vegas. There, a specialized team was found tearing books from their spines to digitize pages before discarding the damaged texts.

This finding confirms suspicions held by booksellers and industry observers that AI firms have been systematically sourcing rare and older books with unique ISBNs, aiming to build expansive and exclusive datasets. Amazon’s practice contrasts with competitors like Anthropic and xAI, who have publicly asserted they avoid using rare or antique books for training purposes.

Why it matters

The destruction of rare books by Amazon to create AI training data raises significant ethical and cultural concerns. Booksellers worry that valuable literary and historical works—especially those in lesser-known languages or editions—are being sacrificed without regard to their inherent worth. This contributes to a broader debate over how AI companies source data and the impact of those methods on cultural heritage.

Compounding the issue is Amazon’s apparent systematic approach to scanning, where workers are trained to first scan ISBN barcodes, suggesting a methodical attempt to capture every printed book available. The approach could give Amazon a competitive advantage by incorporating unique and difficult-to-obtain texts, while also underscoring the tension between innovation in AI and preservation of rare cultural artifacts.

What to watch next

Industry and regulatory scrutiny is likely to increase around data acquisition practices employed by AI companies, especially those that involve the physical destruction of rare materials. Public backlash and pressure from booksellers and cultural institutions may prompt calls for more ethical frameworks or legal constraints on how training data is gathered.

Amazon’s next steps—whether it continues to prioritize rare physical materials despite controversy, or adapts its methods in response to growing criticism—will be closely monitored. Additionally, rival AI firms’ approaches to sourcing data from books versus alternative text sources could influence broader industry norms regarding intellectual property, preservation, and ethical AI training.

Source assisted: This briefing began from a discovered source item from Ars Technica Tech Policy. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings