AI Companies Bulk-Buying Millions of Secondhand Books for Training Data, Raising Copyright Concerns
Key Takeaways
- ▸AI companies are bulk-purchasing secondhand books through obscured channels to obtain uncontaminated, human-authored training data free of AI-generated content
- ▸ISBNdb has pivoted its platform to offer specialized bulk-purchase services, with orders ranging from 1,000 to 1 million books per transaction
- ▸Books are being scanned destructively for AI training, sparking copyright lawsuits and million-dollar fines against Anthropic and legal action against Google
Summary
An investigative report from 404 Media reveals that leading AI companies are systematically purchasing millions of secondhand books through intermediaries like ISBNdb to obtain high-quality training data for their AI models. These bulk acquisitions—ranging from 1,000 to 1 million books per transaction and surging dramatically since April 2026—aim to avoid the contamination of "AI slop," the proliferating low-quality AI-generated content that has polluted internet training datasets. The books are then scanned, often destructively, by specialized services, with companies like Anthropic and Google employing firms such as Datamation Information Services.
The scale of these operations is unprecedented: booksellers report a fivefold increase in bulk purchases, with weekly sales skyrocketing from 20 books to several hundred. Purchasers appear indifferent to pricing and show no interest patterns by subject, genre, or author—hallmarks of coordinated industrial acquisition. However, this covert approach to sourcing training data has ignited legal backlash. Anthropic faced a $1.5 billion fine for maintaining a repository of seven million pirated books, while a coalition of publishers has sued Google over its use of copyrighted books to train Gemini. Though courts have affirmed that using books for AI training falls under fair use, the scale and opacity of these operations continue to strain the boundaries of copyright law.
- Booksellers report unprecedented surge in bulk purchases (fivefold increase) since April 2026, indicating industrial-scale acquisition of physical books by AI companies
- The practice reveals AI companies' desperation for clean training data as the internet becomes saturated with low-quality AI-generated content
Editorial Opinion
The practice of systematically purchasing and destroying millions of books for AI training—while major companies face multibillion-dollar copyright lawsuits—exposes a troubling pattern: AI companies are circumventing public scrutiny to access intellectual property at scale. While fair use doctrine may technically permit this, the opacity, scale, and financial evasion of these operations suggest the industry is exploiting legal gray areas rather than operating in good faith. Authors and publishers deserve transparency and fair compensation, not covert purchasing schemes that treat humanity's literary heritage as a disposable resource for corporate AI development.



