In the rush to enhance artificial intelligence models, some US-based AI companies are reportedly buying rare and fragile books in bulk, only to destroy the physical copies after scanning pages for data extraction. This practice has sparked resistance from booksellers and library preservationists alarmed at the loss of irreplaceable cultural artifacts.

  • AI companies destroying rare books to quickly collect training data
  • Non-destructive scanning methods exist but are slower and costlier
  • Internet Archive uses meticulous manual scanning to preserve rare volumes

What happened

Several AI firms have been quietly purchasing rare and antique books in large quantities, only to dismantle them for rapid scanning and digital data extraction. The process often involves cutting up books to feed pages into scanners, resulting in physical destruction of the original copies. This aggressive approach stems from the high demand for long-form, quality text data to train increasingly sophisticated AI language models.

Book dealers and preservation advocates have voiced alarm about this trend, emphasizing the cultural loss and irreversible damage inflicted on these fragile works. Despite public concerns, the urgency in AI development incentivizes looking for the fastest and cheapest data acquisition strategies, often sidelining the value of the physical books themselves.

Why it matters

The destruction of rare books to fuel AI development raises significant ethical and cultural preservation questions. These old and sometimes unique volumes contain historical knowledge and craftsmanship that cannot be replaced once lost. The contrast between destructive scanning and available non-damaging technologies underlines a tension between innovation speed and heritage conservation.

Notably, Google patented a non-destructive scanning technology in 2009, but companies largely opt for faster destructive methods due to cost and scale considerations. Organizations like the Internet Archive demonstrate that painstaking manual scanning can safeguard fragile rare books, though this approach requires more time and resources and may not be favored by profit-driven AI firms.

What to watch next

How the AI industry responds to growing criticism regarding rare book destruction will be important to monitor. Potential regulatory interest or public pressure may push for adoption of less harmful scanning methods or establish guidelines to protect cultural artifacts involved in data sourcing. Advances in technology could also enable better scanning solutions that strike a balance between efficiency and preservation.

In parallel, the strategies of libraries, archives, and preservation entities like the Internet Archive will be critical in setting standards and providing models for ethical digitization. Their ongoing efforts to scan precious collections carefully may serve as a benchmark for responsible data collection practices within the evolving AI training ecosystem.

Source assisted: This briefing began from a discovered source item from Ars Technica Tech Policy. Open the original source.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings