BotBeat
...
← Back

> ▌

AnthropicAnthropic
INDUSTRY REPORTAnthropic2026-07-27

AI Companies Are Shredding Rare Books for Training Data—and It's Legal

Key Takeaways

  • ▸AI companies are purchasing rare books in bulk, scanning them, and shredding the originals for training data, with Anthropic actively recruiting expertise to scale the operation
  • ▸A federal court ruled the practice constitutes "fair use," providing legal cover that is accelerating the destruction of irreplaceable cultural artifacts
  • ▸ISBNdb facilitates anonymity and scale (up to a million books per order) while offering NDAs and coaching clients to use euphemistic language like "digital preservation"
Source:
Hacker Newshttps://xcancel.com/HedgieMarkets/status/2081534588485296565↗

Summary

AI companies are systematically purchasing rare books in bulk, using high-speed scanning machines to digitize them, and shredding the physical originals for use as training data. The practice, facilitated by services like ISBNdb, has destroyed millions of books. A federal judge ruled the practice constitutes "fair use," creating legal cover for acceleration. Anthropic has reportedly hired the former head of Google Books partnerships to obtain "all the books in the world," exemplifying the industry's aggressive approach to acquiring training material. Pre-2022 books command premium prices because they predate AI-generated text, making rare historical documents particularly vulnerable.

The business operates with deliberate opacity—ISBNdb offers non-disclosure agreements as a standard feature and coaches clients to rebrand destruction as "digital preservation." Rare books with only a handful of surviving copies worldwide are being fed into this pipeline. Unlike website scraping or music torrenting, this destruction is irreversible; once a rare 18th-century botanical text or other unique historical document is shredded, it cannot be recovered. The legal ruling has legitimized permanent destruction of irreplaceable cultural artifacts for marginal gains in AI model training.

  • Pre-2022 books are particularly targeted because they predate AI-generated content, putting rare historical texts with few surviving copies at acute risk of permanent destruction
  • Unlike other AI training data acquisition methods, this process is irreversible—shredded books cannot be recovered, representing permanent loss of human knowledge and cultural heritage

Editorial Opinion

This represents a troubling inflection point for AI development. While the legal ruling technically permits this practice under fair use doctrine, the permanent destruction of irreplaceable cultural artifacts to marginally improve AI models raises fundamental questions about what civilization is willing to sacrifice for technological convenience. The deliberate opacity—NDAs, euphemistic rebranding, anonymous transactions—suggests the companies involved understand this practice wouldn't survive public scrutiny. We may be witnessing the algorithmic erasure of human history, one shredded book at a time.

Large Language Models (LLMs)Market TrendsRegulation & PolicyEthics & Bias

More from Anthropic

AnthropicAnthropic
RESEARCH

Mathematical Paradox Reveals Fundamental Limits of LLM Truth Probes

2026-07-27
AnthropicAnthropic
POLICY & REGULATION

Trump Administration's Fractured AI Policy Apparatus Targets Anthropic Over Export Controls

2026-07-27
AnthropicAnthropic
PRODUCT LAUNCH

AgentHost: Orchestrating Multiple AI Agents in Your Own Cloud Infrastructure

2026-07-27

Comments

Suggested

Alibaba (Cloud)Alibaba (Cloud)
RESEARCH

Comprehensive Quantization Benchmark: How Qwen 3.6 27B Performs Across Compression Levels

2026-07-27
NVIDIANVIDIA
POLICY & REGULATION

NVIDIA Leads 77-Company Coalition Pushing US to Permit Open-Weight AI Models

2026-07-27
Independent ResearchIndependent Research
RESEARCH

Study: AI Agents Report Success at Identical Rates Whether Tasks Succeed or Fail

2026-07-27
← Back to news
© 2026 BotBeat
AboutPrivacy PolicyTerms of ServiceContact Us