Gentoo Bugzilla Taken Offline by AI Bot Scraper Overload
Key Takeaways
- ▸AI bot scraping has become so aggressive that it can render critical open-source infrastructure inaccessible to its intended users
- ▸Many AI companies and research labs are collecting training data without respecting rate limits, robots.txt, or attempting coordination with project maintainers
- ▸Open-source projects have limited defenses against volumetric scraping and are beginning to shut down services rather than continue the drain
Summary
The Gentoo Project was forced to take its Bugzilla instance offline due to an overwhelming surge of scraping bots collecting data for AI training. The relentless volume of automated requests made the bug tracker inaccessible to actual Gentoo developers and users, forcing the project to shut down the service rather than allow its infrastructure to be consumed by AI data collection efforts.
The incident highlights a growing friction point between AI companies' voracious appetite for training data and the sustainability of open-source infrastructure that the tech industry depends on. Multiple AI labs and researchers training large language models have been aggressively scraping publicly available content without adequate rate limiting, robots.txt compliance, or coordination with site operators. Gentoo's decision to take Bugzilla offline represents an inflection point where a major open-source project was forced to choose between being accessible to its community or being accessible to AI training pipelines.
This incident has reignited important conversations about data collection ethics, the responsibility of AI companies to be good actors on shared internet infrastructure, and whether new technical standards or regulations are needed to balance AI training needs with the viability of critical open-source projects.
- The incident raises urgent questions about AI training data collection ethics and whether industry self-regulation is sufficient
Editorial Opinion
This is a wake-up call for the AI industry. The behavior described—wholesale scraping of infrastructure that many depend on—is extractive and unsustainable. While training data is essential for AI development, the current approach treats open-source projects as free resources to be harvested without consent, throttling, or consideration for their operational needs. If we continue down this path, we risk damaging the very open-source ecosystem that modern AI systems are built on. The industry needs to establish clear norms around respectful data collection, including rate limiting, robots.txt compliance, and direct engagement with project maintainers.



