BotBeat
...
← Back

> ▌

AnthropicAnthropic
INDUSTRY REPORTAnthropic2026-07-21

AI Companies Turn to Old Printed Books as 'Clean' Training Data, Creating Hidden Market

Key Takeaways

  • ▸ISBNdb facilitates bulk book acquisitions for AI companies, with orders ranging from 1,000 to 1 million books per transaction, targeting pre-2022 publications guaranteed free of AI-generated content
  • ▸Anthropic and Google have been caught acquiring and destroying millions of printed books for training data, exposed through copyright lawsuits from authors and publishers
  • ▸AI companies use strict NDAs with ISBNdb to conceal their book acquisition operations and avoid reputational damage from book destruction
Source:
Hacker Newshttps://www.404media.co/ai-companies-are-buying-tons-of-old-books-because-theyre-free-of-ai-slop/↗

Summary

As AI companies face the challenge of AI-generated content contaminating their training datasets, a new market has quietly emerged: ISBNdb, a book database company, now offers bulk book sourcing services to AI labs seeking pre-2022 publications guaranteed free of AI-generated text. Books published before the generative AI boom are particularly valuable because they avoid 'model collapse,' a phenomenon where models trained on AI-generated content produce progressively worse results. According to ISBNdb's marketing materials, printed books offer 'curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate.'

The scale of these operations was exposed through copyright lawsuits. Internal documents revealed Anthropic planned to acquire and scan millions of printed books from marketplaces like Better World Books, while Google faces similar legal action for training Gemini on copyrighted books. ISBNdb now facilitates bulk orders ranging from 1,000 to 1 million books per transaction, helping AI companies systematically digitize physical books while maintaining anonymity through legally binding non-disclosure agreements.

The practice reveals an uncomfortable infrastructure behind AI advancement: companies are destroying cultural artifacts to improve their models. ISBNdb's own marketing acknowledges the reputational risk, noting that 'AI company destroys two million books' generates no sympathy. Industry insiders report historic sales spikes in book marketplaces since April, suggesting the practice is accelerating despite growing legal and ethical scrutiny.

  • Model collapse—where training on AI-generated text produces degraded model performance—is the primary driver for AI companies seeking high-quality human-written training data from physical books

Editorial Opinion

The emergence of ISBNdb as a book-sourcing infrastructure layer exposes the uncomfortable reality that scaling AI models now requires destroying existing cultural artifacts at scale. While companies defend this as technically necessary—pre-2022 books do offer genuine advantages over contaminated web data—the reliance on NDAs and the acknowledgment of 'optics problems' suggests companies know their practices are indefensible in public discourse, even if lawyers argue they're legally defensible. The copyright lawsuits are just the beginning; this practice will ultimately force a reckoning between AI companies' data hunger and society's expectations around intellectual property and cultural preservation.

Large Language Models (LLMs)Generative AIMarket TrendsEthics & Bias

More from Anthropic

AnthropicAnthropic
FUNDING & BUSINESS

Anthropic Reaches Record $1.5 Billion Copyright Settlement Over Claude Training Data

2026-07-21
AnthropicAnthropic
POLICY & REGULATION

University of Tennessee Sues Anthropic Over Neural Network Technology

2026-07-21
AnthropicAnthropic
RESEARCH

Research Shows AI Advice Suppresses Critical Thinking, Even When Accuracy Drops

2026-07-21

Comments

Suggested

CloudflareCloudflare
PRODUCT LAUNCH

Cloudflare R2 Reaches General Availability, Eliminating Egress Fees for Object Storage

2026-07-21
MetaMeta
PRODUCT LAUNCH

Meta's AI Models Power First Wave of Genesis Mission Projects

2026-07-21
Google / AlphabetGoogle / Alphabet
RESEARCH

Google Begins Pre-Training Gemini 4, Next-Generation LLM in Development

2026-07-21
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us