LLM Training Bias Could Reshape Human Language and Cognition
Key Takeaways
- ▸LLMs trained on incomplete data (written text, scripts) rather than natural conversation, creating a narrow representation of human language
- ▸Increased exposure to AI-generated text is already measurable affecting human speech patterns—shorter sentences, less courtesy, formulaic responses
- ▸AI-generated language is more uniform and emotion-poor than natural speech, lacking the meanders and interruptions that convey authentic human thought
Summary
A new analysis reveals that large language models, trained primarily on written text rather than natural face-to-face conversation, may fundamentally alter how humans communicate and think. Because LLMs have minimal access to unscripted dialogue—the vast majority of human speech—they capture only a narrow slice of human language, biased toward formal writing, social media, and scripted media like television and film.
As AI-generated text becomes increasingly prevalent in daily life, humans are beginning to adopt the linguistic patterns embedded in these models. The effects are already visible: AI responses tend to be formulaic, overly polished, and devoid of the natural meanders, interruptions, and emotional leaps that characterize authentic human speech. A 2022 study found that children using voice commands with Siri and Alexa became more curt with human speakers, mimicking the command-based syntax of these tools.
The implications extend beyond mere stylistic changes. LLM-generated language shows a narrower vocabulary range and shorter average sentence lengths (12-20 words) than natural human speech, potentially constraining human expression over time. The problem is compounded by a feedback loop: as LLMs train on increasingly AI-generated content, they reinforce their own inhuman patterns while inadvertently teaching humans to accept and replicate them. Additionally, the tendency of chatbots to agree with user statements regardless of validity could reinforce confirmation bias and reduce openness to alternative perspectives.
- A feedback loop is forming as LLMs train on text increasingly produced by other LLMs, amplifying inhuman linguistic patterns
- Broad LLM adoption could introduce subtle but pervasive cognitive shifts, including increased confirmation bias and narrower vocabulary use
Editorial Opinion
This analysis raises critical questions about the reciprocal relationship between AI and human cognition that the industry has largely overlooked. While we've focused on LLM accuracy and capability, we've paid little attention to how these systems, by design, may gradually reshape human communication and thought patterns—potentially in ways we won't fully understand until the damage is done. The feedback loop between LLM training and human linguistic adoption deserves urgent attention from AI developers, linguists, and policymakers before it becomes deeply entrenched in how entire generations communicate and think.



