The Internet Is Filling Up With AI.
AI-generated text is now a measurable part of the web. As future models learn from it, content that verifiably came from humans is becoming the scarce resource.
Human Writing Is Becoming Scarce.
A few years ago, seeing a blog post, product description, or online review came with a simple assumption: a person probably wrote it.
That assumption no longer holds.
AI-generated text is now a measurable part of the web. In a July 2026 random sample of 10,000 English language webpages, 10% showed significant signs of AI authorship or editing, according to a Pew Research Center analysis based on 490,000 webpages collected over five years. Another 2026 study examining websites published between 2022 and 2025 found that by mid-2025, roughly 35% of newly published websites were classified as AI-generated or AI-assisted. That study is currently a preprint and has not yet undergone peer review.
There is an important limitation to Pew's number. Its detection method captures signs of both AI authorship and substantial AI editing. That means the 10% does not represent only pages written entirely by AI; it can also include human-written text that was significantly edited using AI tools. Pew also cautions that AI detection models are probabilistic and should not be treated as definitive judgments about individual pages.
The important shift is not that AI can write.
We already know that.
The shift is that AI-generated text is becoming part of the material future AI systems will learn from.
And that creates an unexpected new scarcity:
content we know actually came from humans.
The Web Is Already Starting to Sound Different
The change is visible in language itself.
Pew found that compared with 2023, em dashes now appear about twice as often online. Oxford comma usage increased by 63%, while words frequently associated with AI writing, including "delve," "interplay," and "testament," more than doubled.
Even sentence structures associated with AI writing are becoming more common. Negative parallelism, such as "it's not just X, it's Y," has nearly tripled in frequency.
None of these features proves that an individual article was written by AI. Humans use them too. But across hundreds of thousands of webpages, the pattern matters.
The web is not just getting bigger. Parts of it are becoming more uniform.
The same 2026 preprint on AI-generated text found a similar pattern. Researchers reported a negative correlation between the growth of AI-generated text and semantic diversity online.
We can now produce more text, faster than ever. That does not necessarily mean we are producing more variety.
AI Is Starting to Learn From AI
This is where the issue goes beyond writing style.
Large language models depend on enormous amounts of training data collected from books, articles, webpages, and other sources. Historically, much of that material originated with humans.
As AI-generated content spreads across the web, future models will inevitably encounter content produced by earlier models.
Researchers from Oxford, Cambridge, and other institutions examined this problem in a peer-reviewed Nature paper on model collapse.
Their finding was direct: indiscriminately training generative models on recursively generated data causes them to progressively lose information about the original data distribution. Rare patterns begin disappearing first, and the model's representation of the original data becomes narrower over successive generations.
The researchers argue that preserving access to original data sources and additional data not generated by large language models is necessary for sustainable training over time.
That changes the value of human-generated data.
AI can produce enormous amounts of text. But generating more AI content does not create new human experience, observation, or knowledge.
When Physical Books Become Training Data
One of the clearest examples of the growing demand for training data is happening far away from the web.
In August 2026, an investigation by 404 Media followed a shipment of rare books to a facility in Las Vegas operated by Amazon. According to the investigation and an employee later interviewed by the publication, books arriving at the facility had their bindings removed so the loose pages could be processed through high-speed scanners for AI training data.
After scanning, the pages were discarded together as loose paper, leaving the original books impossible to reconstruct.
The employee described books in multiple languages, library materials, government documents, and obscure publications moving through the facility.
The process itself is striking:
Buy the book. Cut the binding. Scan the pages. Keep the data.
The significance here is not the company doing it. It is what the process represents.
We have access to billions of webpages, and generative AI can produce more text almost instantly. Yet physical, human-written material is still being acquired and converted into training data.
The investigation does not establish that these books were selected specifically because they predate widespread generative AI. But it does show that material outside the constantly changing modern web is useful enough to acquire, physically process, and preserve as data.
Placed next to the research on model collapse, that matters.
As synthetic content becomes easier to produce, access to original material becomes harder to replace.
Human-Made Content Is Becoming a Data Asset
For most of the internet era, the difficult problem was getting access to enough information.
That problem has changed.
Generative AI can now produce text at enormous scale. The harder question is increasingly where that information came from.
- Was this review written by someone who actually bought the product?
- Was this article based on first-hand research?
- Was this forum post written by someone who experienced the problem?
- Was this training sample produced by a person, or by a model trained partly on the output of another model?
The Nature researchers point directly to this provenance problem. As generated content becomes more common online, distinguishing AI-generated data from original data at scale becomes increasingly difficult. They warn that training future models may require access to data collected before mass AI adoption or direct access to human-generated material.
When content becomes nearly unlimited, provenance becomes valuable.
More Content Does Not Mean More Knowledge
None of this means AI-generated content is inherently worthless.
AI can summarize, translate, organize information, and produce useful drafts at a speed humans cannot match. Synthetic data also has legitimate uses in model development.
The problem begins when generated material becomes difficult to distinguish from the human material it originally learned from.
The internet can then become a feedback loop: human knowledge is absorbed, recombined, published, scraped again, and recombined once more.
The Nature experiments show why preserving original data matters. In one setup, retaining even 10% of the original training data resulted in substantially less performance degradation than relying on generated data without preserving the original material.
The conclusion is not that AI content has no value.
It is that human-origin data remains difficult to replace.
Human Content Is Becoming the Scarce Resource
For years, technology made information cheaper.
Search engines made it easier to find. Social media made it easier to publish. Generative AI made it almost effortless to produce.
Now the scarce resource is changing.
It is no longer text.
It is original experience, first-hand observation, and verifiable human input.
That matters beyond AI training. It changes what readers can trust, what companies should publish, and what remains valuable when anyone can generate a thousand competent paragraphs in seconds.
The internet is not running out of things to read.
It is filling up faster than ever.
What is becoming harder to find is something AI cannot manufacture simply by generating another paragraph:
a genuinely new human source.
References
- Bestvater, S., Smith, A., TerBush, C., Baranovski, C., & Chavda, J. (2026, August 20). How Much of the Internet Is Written With AI? Pew Research Center.
- Dolezal, J., Alam, S., Graham, M., & Bohacek, M. (2026). The Impact of AI-Generated Text on the Internet. arXiv:2604.26965. https://doi.org/10.48550/arXiv.2604.26965
- Maiberg, E. (2026, August 17). We Tracked a Shipment of Rare Books. It Ended at an Amazon AI Training Facility. 404 Media.
- Maiberg, E. (2026, August 26). Inside the Warehouse Where Amazon Scans and Destroys Books for AI Training. 404 Media.
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y
- Spennemann, D. H. R. (2025). "Delving into" the quantification of AI-generated content on the internet (synthetic data). The paper estimates that at least 30% of text on active webpages may derive from AI-generated sources, while noting substantial methodological uncertainty around that estimate.
