A recent paper by researcher Hector Zenil at the University of Cambridge has provided mathematical proof for what many computer scientists have long suspected: LLMs (Large Language Models) cannot self-improve through training on their own outputs. The phenomenon, called model collapse, occurs when these statistical models start feeding on their own generated content rather than fresh human-created data. Think of it as an intellectual version of mad cow disease, where the system slowly degenerates by consuming itself. The core issue lies in what LLMs actually are versus what the hype suggests. Despite all the breathless marketing about artificial intelligence and machine learning, these systems are fundamentally probability engines. They analyze massive datasets of human-created text and build statistical models predicting which words should follow other words. When you ask ChatGPT a question, it's not thinking or reasoning in any meaningful sense. It's running a sophisticated calculation to determine what sequence of words is most statistically likely given your prompt. Zenil's paper demonstrates that when an LLM trains on its own output or on content created by other LLMs, it undergoes what he calls degenerative dynamics. The model doesn't learn or improve; it converges on a statistical singularity, essentially becoming a distorted echo chamber of its own biases and patterns. Each generation of self-training amplifies existing quirks and errors while losing the diversity and richness of the original human-generated training data. Within a few cycles, the output becomes increasingly homogeneous, repetitive, and divorced from reality. This isn't just a theoretical problem. As AI-generated content floods the internet (some estimates suggest AI-created text already comprises 10-15% of new web content in 2026), future LLMs face an increasingly polluted data pool. Companies like OpenAI, Anthropic, and Google are already struggling to find enough high-quality human-generated text to train their next-generation models. The easy solution of using synthetic data, letting AI train AI, turns out to be a dead end. The paper proposes potential mechanisms to counter entropy decay within these models, but the fundamental conclusion remains: statistical models require continuous external anchoring to human-generated data to avoid collapse. There's no magical feedback loop where AI trains itself into superintelligence. The much-hyped concept of AGI (Artificial General Intelligence) emerging from scaled-up LLMs looks increasingly like a fever dream sold by venture capitalists and tech evangelists who confused pattern matching with thinking.
💻 technology
AI Eats Its Own Tail, Math Proves
New research confirms what skeptics have been saying: Large Language Models (LLMs) like ChatGPT cannot actually learn from their own output without collapsing into statistical gibberish. The mathematical proof demolishes Silicon Valley's favorite fantasy that AI will bootstrap itself into superintelligence.
My Take
The AI hype cycle has been built on a foundation of deliberately obscured limitations. Silicon Valley has spent billions convincing investors and the public that we're on the verge of creating digital minds, when in reality we've built extremely expensive autocomplete systems. This research doesn't just poke holes in the AGI narrative; it mathematically proves the whole concept is nonsense given current architectures. The emperor has no clothes, and now we have peer-reviewed calculus showing exactly how naked he is. What's particularly galling is how many smart people should have seen this coming. Anyone who's actually worked with these systems rather than just invested in them knows they're brittle, prone to hallucination, and fundamentally incapable of actual reasoning. But there's too much money and ego invested in the AGI mythology to admit the obvious. The model collapse problem reveals something more profound than a technical limitation: it exposes that these systems were never on a path to consciousness or general intelligence in the first place. They're statistical parrots, and feeding a parrot its own squawks doesn't teach it to philosophize.
What Happens Next
The immediate corporate response will be denial and distraction. Expect OpenAI and Anthropic to announce new benchmarks and capabilities that carefully sidestep the model collapse issue while pumping their next funding rounds. But behind the scenes, the scramble for clean training data is about to get ugly. We'll see aggressive moves to license content from publishers, universities, and anyone sitting on large corpuses of pre-2023 human-written text. Reddit's data licensing deals will look quaint compared to what's coming. The real shift happens in 12-18 months when investors start demanding actual ROI from AI deployments and realize these systems can't autonomously improve their way to profitability. Companies that bet their futures on AGI arriving by 2027-2028 will quietly pivot their messaging from "revolutionary intelligence" to "productivity tools." The model collapse research gives cover to the skeptics who've been shouted down for two years. Funding for scaling-focused AI labs will plateau while money flows toward hybrid approaches that keep humans firmly in the loop, which ironically might produce more useful technology than the current dead-end pursuit of silicon consciousness.
What History Tells Us
This pattern of overpromising on AI capabilities has deep roots. In the 1960s, researchers like Marvin Minsky claimed that artificial intelligence comparable to human cognition was only a generation away. The subsequent AI winters of the 1970s and late 1980s came when reality failed to match the hype, causing funding to collapse. The current LLM bubble mirrors the expert systems craze of the 1980s, when rule-based AI was supposed to capture human expertise and revolutionize everything from medicine to manufacturing. Those systems also hit fundamental limitations, failing to generalize or truly understand context despite massive investment. The cycle repeats because each generation of AI boosters genuinely believes their approach has cracked the code that previous attempts missed.
Market Impact
The model collapse research poses a medium-term threat to the astronomical valuations of AI-pure-play companies, though expect immediate market reaction to be muted. NVIDIA (NVDA), currently trading around $920 per share after a strong 2025-2026 run fueled by AI chip demand, faces the most interesting dynamics. Short-term bullish as companies panic-buy compute to train on remaining clean data before it's all contaminated, but 18-month outlook turns bearish if the AGI narrative collapses and demand shifts from massive frontier model training to smaller, specialized deployments requiring less hardware. Microsoft (MSFT) and Google parent Alphabet (GOOGL) are somewhat insulated by their diversified businesses, but their multi-billion-dollar AI investments look increasingly questionable. Watch for subtle shifts in their quarterly guidance language around AI monetization timelines. The real beneficiaries might be old-school data companies like Thomson Reuters and content platforms like Stack Overflow that control valuable corpuses of verified human-generated data, expect aggressive M&A activity in that space. Meanwhile, AI-focused ETFs like ROBT and BOTZ have already priced in AGI dreams, setting up for potential 20-30% corrections when reality sets in.