AI Companies Race to Acquire Old Books to Escape AI-Generated Training Data
Key Takeaways
- ▸AI companies are actively buying older, pre-AI books to secure training data free of AI-generated content
- ▸Concern about 'AI slop' — lower-quality AI text — polluting training datasets and degrading new models
- ▸Data sourcing is becoming a premium business, with companies investing heavily in high-quality, curated datasets
Summary
AI companies are increasingly purchasing large collections of older books published before the widespread generation of AI content, seeking to source high-quality training data untainted by artificial content. The trend reflects growing concerns about 'AI slop'—lower-quality AI-generated text—contaminating public datasets and degrading the performance of new models trained on internet-sourced data. By acquiring out-of-print and rare books, major AI developers aim to access pristine, human-authored content while avoiding the recursive problem of training models on AI-generated outputs. This shift signals a market willingness to pay premium prices for curated, pre-AI-era datasets as data quality becomes a critical competitive factor in generative AI development.
- The market recognizes that training on AI-generated outputs creates a recursive degradation problem
- Older published works represent valuable, human-authored content that pre-dates AI generation era
Editorial Opinion
This trend exposes a fundamental weakness in the AI development supply chain: the increasingly problematic quality of freely available training data. As AI-generated content floods the internet, companies are forced to pay significant sums for the privilege of accessing older books—essentially paying to escape the consequences of their own technology. It's a cautionary tale about the long-term sustainability of training-data sourcing and raises important questions about content scarcity and data quality in an AI-saturated world.



