The Web's Quiet Collapse: How AI Content Is Draining the Internet's Information Pool
A provocative new theory argues that AI-generated content is causing an 'epistemic heat death' of the internet, where volume explodes but genuine ideas stagnate.

Takeaways
- ›AI content risks flooding the web with volume while draining its pool of new information
- ›Current web metrics (pageviews, engagement) are obsolete for distinguishing human-created content from AI restatements
- ›Information theory suggests perfect AI language models add zero new information relative to their training data
- ›We urgently need new tools to measure and reward genuine knowledge creation online
Open a search engine and ask it something specific. Chances are, the top results are well-formatted, plausible-sounding, and utterly devoid of insight beyond your query. This isn't just SEO decay. It's a symptom of a deeper rot in our information ecosystem, argues Jarosław Szulc in his dense, provocative paper "Epistemic Heat Death and the Signal-to-Noise Ratio of the Global Web."
Szulc's thesis is both simple and alarming: as AI-generated content floods the web, we're witnessing an explosion of pages that masks a collapse in actual information. It's not just that quality is declining, it's that the number of genuinely new ideas, grounded claims, and accountable human authors is stagnating while the volume of content skyrockets.
The Photocopier Effect
Szulc borrows the term 'epistemic heat death' from thermodynamics, but with a crucial twist. Unlike true heat death (maximum entropy), the web isn't becoming random, it's becoming repetitive while appearing more diverse. It's less 'the universe cooling' and more 'a photocopier feeding its output back into itself, forever, at increasing speed.'
This distinction matters because it points to the heart of the problem: AI content, by its nature, can't add truly new information to the web. It can only recombine and restate what already exists in its training data.
Why Our Tools Are Obsolete
If Szulc is right, the metrics we've relied on to navigate the web, pageviews, engagement, link authority, trending topics, aren't just imperfect. They're fundamentally broken. These tools were built to find relevant human content among other human content. They were never designed to distinguish human-created information from AI-generated restatements, because until recently, that distinction barely mattered.
The Shannon Entropy of AI Content
Szulc grounds his argument in Claude Shannon's information theory: information is a function of surprise. A message that tells you something unpredictable carries information. A message that restates what you already knew, even in different words, carries almost none.
This leads to the paper's most crucial insight: a language model trained on human text, then asked to generate more human-sounding text, is, in the limit of perfect training, a zero-information source relative to its training distribution. It can rearrange and recombine, but it cannot add anything truly new.
The Data: Messy but Concerning
The paper's reliance on some contested figures (like the oft-quoted "90% of content will be AI-generated by 2026") is a weakness. However, even more conservative estimates paint a troubling picture:
- An Ahrefs study found AI involvement in about 75% of surveyed web pages, though only 2.5% were "pure AI" without human editing.
- One vendor estimates 312 million AI-assisted pages are published monthly, up from 82 million two years ago.
Importantly, even if the raw share of AI content plateaus, distribution channels (search engines, AI assistants) appear to be filtering toward human authorship, consistent with the idea that truly informative content is becoming scarcer and more valuable.
The Market for Lemons, But for Ideas
Szulc extends his argument using George Akerlof's "market for lemons" theory. When buyers can't easily distinguish good products (informative content) from bad ones (AI-generated filler), the market can collapse to low-quality offerings. Applied to the web, this suggests a vicious cycle where AI content crowds out human-created information, making the entire ecosystem less valuable.
Rethinking How We Measure Web Value
The paper doesn't offer easy solutions, but it highlights the urgent need for new ways to measure and value online information. We need tools that can distinguish genuinely novel ideas from sophisticated restatements, and systems that can route attention and resources to accountable human creators adding real information to the web.
Szulc's 'epistemic heat death' may be a metaphor, but it names a real and pressing challenge. As AI reshapes the information landscape, we must rethink how we measure, distribute, and reward the creation of genuine knowledge online. The alternative is a web that grows ever larger while saying less and less that's truly new.
Related reads
Web Data Infrastructure for AI Explained: Challenges, Benefits
4 min read
Google Held Liable for AI-Generated Content Errors
3 min read
RAG Pipeline Explained: Why It Often Fails in Production
5 min read
AI-Generated Code Security Risks: Why Review Process Needs Overhaul
5 min read
Diffusion Models for Video Generation: Challenges and Approaches
5 min read
Software Engineers and AI: Why Jobs Remain Resilient
5 min read
Reported and explained by AI·Reporter.