Study Finds Scientific Literature's Quality Issues Harm LLM Training
Key Takeaways
- ▸A 2024 multi-institution study found that removing ArXiv, PhilPapers, and NIH ExPorter from LLM training data improved model performance and reduced toxic outputs
- ▸Scientific literature contains a problematic mix of honest work, half-truths, methodological games, and outright fraud that are difficult to distinguish
- ▸Scientific metadata like citations and authorship have been corrupted by gaming and are unreliable indicators of quality
Summary
A 2024 research study involving teams from MIT, Cornell, Carnegie Mellon, Google, and OpenAI has found that scientific literature is fundamentally problematic as training data for large language models. Rather than improving LLM performance, papers from sources like ArXiv, PhilPapers, and NIH ExPorter actually decreased model accuracy on academic questions and increased toxic outputs—removing these datasets improved overall model performance. The researchers found that scientific literature is scattered across multiple quality dimensions, containing a minefield of half-truths, convenient omissions, methodological games, and outright fraud that are nearly indistinguishable from genuine findings.
The core issue goes beyond individual bad actors. Honest scientists often dress up their work with superlatives and adopt popular but incorrect theories to compete for journal spots against fraudsters who simply fabricate data. Combined with gamed citation metrics and authorship information, the result is a corrupted signal that actively undermines LLM training. While removing PubMed data slightly reduced performance, the broader finding that wholesale removal of large scientific corpora improved model accuracy suggests that pre-slop scientific literature represents a unique training liability.
The findings raise important questions for AI-assisted science efforts, which currently lack access to the informal social networks and personal reputation channels that human scientists use to identify reliable research. Without these verification mechanisms, AI systems trained on scientific literature inherit its biases and corrupted signals, potentially perpetuating misconduct and poor methodology.
- AI systems for science lack access to the informal social networks and reputation channels humans use to verify research reliability
- These findings suggest that scientific literature quality is a significant bottleneck for AI-assisted research efforts
Editorial Opinion
This research reveals a fundamental tension: the scientific literature that should be humanity's most reliable knowledge base is corrupted precisely because it's the field's primary currency. The solution isn't simply filtering out bad papers—it's recognizing that AI systems need access to the informal social verification networks that human scientists rely on. Whether that means surveilling scientific practice or building better tools for transparent peer review, the status quo of training AI on unfiltered scientific literature is clearly insufficient.


