According to a recent review by Digital Trends Computing, booksellers worldwide have reported unusual large-scale orders for old and out-of-print books, prompting theories that AI companies might be acquiring physical copies for training their language models. While no definitive proof links these purchases directly to AI firms, reports indicate a growing pattern of mysterious orders that could be connected to AI data sourcing needs.

  • Unusual bulk orders of old books seen globally, sparking AI sourcing theories
  • Older printed books valued for authentic, non-AI-generated training data
  • Processes may involve scanning and recycling physical books after digitization

Product angle

The source review from Digital Trends Computing observes a notable trend among booksellers globally where large orders targeting old and out-of-print books are linked to possible AI training needs. The reported activity includes purchases through opaque aliases and full-price payments without bulk discount negotiations, indicative of organizational motivations beyond typical collectors or resellers. This suggests that AI companies might be investing in physical book collections to build extensive, human-origin content libraries for language model development.

While no direct testing or confirmation from SignalDesk exists, public domain information points to companies like Anthropic engaging in book acquisitions for AI training, sometimes destroying the physical copies post-scanning. An earlier attempt by ISBNdb to offer printed book sourcing for AI datasets was retracted, underscoring the challenges and ethical concerns around digitizing and repurposing physical literary material for AI research.

Best for / avoid if

This phenomenon in AI data sourcing is best suited for organizations requiring vast and diverse human-authored textual corpora with minimal AI-generated content interference, such as language model developers focusing on training data purity. Companies with capabilities or partnerships to handle physical book acquisition and digitization might find opportunities in such a niche market segment.

Conversely, this approach is ill-advised for buyers or stakeholders concerned with cultural heritage preservation or those reliant on retaining physical rare or antiquarian texts, as the process reportedly includes spines being cut and books recycled. Collectors or institutions focused on conservation should avoid participating in or supporting bulk book liquidation suspected to aid AI training to prevent loss of irreplaceable literary artifacts.

Pricing and alternatives to check

The review notes that buyers in these bulk book orders typically pay full retail prices rather than seeking discounts, suggesting a premium is accepted for securing rare and varied materials. Although no pricing plans or models are directly available, the underlying transaction patterns hint at a willingness to invest significantly in obtaining hard-to-source content for AI training purposes.

For those exploring alternatives, traditional digital dataset providers or licenses for existing e-book collections may offer less intrusive options. Additionally, some AI firms utilize synthetic data generation or web scraping aligned with copyright compliance to avoid the complexities and ethical questions raised by physically sourcing and destroying books. Exploring these established or emerging alternatives remains critical to balancing AI progress with preservation concerns.

Source assisted: This briefing began from a discovered source item from Digital Trends Computing. Open the original source.
Review disclosure: Review-watch pages are buyer briefings unless clearly labelled as hands-on SignalDesk reviews. Affiliate, sponsor or free-access relationships should be disclosed on the page. Read the review methodology.
How SignalDesk reports: feeds and outside sources are used for discovery. Public briefings are edited to add context, buyer relevance and attribution before they are published. Read the standards

Related briefings