If the web’s signal-to-noise ratio keeps getting worse, publishers that depend on search traffic may be hit first.
An article published by TechFlowPost cites a new Pew Research Center study built on Common Crawl web archives. The research team randomly sampled 10,000 English-language web pages from each of 49 crawls conducted between January 2021 and July 2026, producing a dataset of roughly 490,000 pages. Each page was then scanned with an AI detection model called Open Pangram.
AI involvement is rising in newly published pages
The study’s findings can be reduced to two headline numbers. In the full July 2026 snapshot, about 10% of English-language web pages showed what the study called significant signs of AI writing or deep AI-assisted editing. If the sample is limited to pages published after ChatGPT launched on Nov. 30, 2022, the share jumps to 35%.
The article notes that 35% does not mean one-third of the entire internet was written by AI, since the web still contains a large stock of pages created before 2022. What the figure does show is the direction of change in newly added content: AI participation in writing is moving from exception to routine practice.
AI traces vary sharply by domain
The distribution is uneven across different parts of the web.
- Among .com domains, about one in 10 pages in the 2026 sample showed AI-writing features.
- For .org domains, the figure was 4.6%.
- For .edu and .gov domains, both were about 1%.
On that basis, the share seen in .com pages was more than double the level for .org and about 10 times the level for .edu and .gov. The article says commercial websites are absorbing AI writing far faster than academic and government institutions.
Language patterns are shifting too
Pew also broke down what it described as the linguistic fingerprints of AI writing. Compared with 2023, newly published pages in 2026 used em dashes at twice the previous rate. Use of the Oxford comma rose 63%. Words commonly favored by AI models, including delve, interplay, tapestry, and pivotal, more than doubled. The frequency of the construction “it’s not just X, it’s Y” was close to triple.
None of those markers alone is enough to prove AI authorship, the article says. Taken together at scale, though, they point to a clear shift in the texture of new web language toward the distribution typically seen in AI-generated text.
On the other side of the market, machines are reading too
The Pew report focuses on who is writing. The TechFlowPost article places it next to another set of data to complete the picture.
On June 3, 2026, Cloudflare CEO Matthew Prince said on social media that AI bot traffic had surpassed human traffic for the first time, accounting for 57.4% of global web HTTP requests. He said he had expected that crossover to arrive at the end of 2027, but it came 18 months earlier.
HUMAN Security’s 2026 report pointed in the same direction. Across all of 2025, AI-driven automated traffic grew at eight times the pace of human traffic. Within that category, traffic from agentic AI, defined in the article as AI acting on a user’s behalf, was up nearly 8,000% year over year.
Put together, the supply side and demand side tell a broader story: roughly one-third of newly added pages involve AI in the writing process, while more than half of the “readers” are now machines. The article argues that an infrastructure built for human communication is turning into an information pipeline where machines increasingly write for other machines.
Recursive training and model collapse
The piece says this is not only a story about content quality. It also touches the foundation of the AI industry itself.
In 2024, Nature published a paper from research teams at Oxford and Cambridge showing that AI models can experience model collapse during recursive training, meaning the use of AI-generated data to train the next generation of AI. Under that process, outputs gradually drift away from the distribution of real-world data, rare patterns in the tail disappear, and after several generations the resulting content becomes increasingly homogeneous or even meaningless.
Researchers at Epoch AI have predicted that high-quality human-written text suitable for AI training could be exhausted between 2026 and 2032.
The cycle described in the article is straightforward: AI models ingest the human internet and generate large volumes of new pages; those pages are indexed by search engines and archived by crawlers such as Common Crawl; the next generation of models is then trained on that material. With each turn of the loop, the share of “AI writing AI” rises, while signals rooted in first-hand human experience are diluted.
Search-dependent publishers may feel the pressure first
If the web’s signal quality keeps deteriorating, the article says, content producers that rely on search distribution are likely to face the earliest impact.
As search results fill up with homogenized AI-rewritten pages, readers may find it harder to tell the original from the copy. Google is already using AI Overview in place of some clickable search-result links. A separate Pew study released in July found that only 20% of users considered AI search summaries “very useful,” while just 6% said they “very much trust” them.
In that environment, the article argues, media outlets able to provide first-hand reporting from the scene, exclusive access to data and documents, opinions and judgments traceable to identifiable human sources, and editorial taste that readers are willing to pay for may command a scarcity premium.
What counts as a media moat is changing
The article closes by saying the metric for media competitiveness is quietly shifting. Volume is no longer much of a barrier. AI can produce 1,000 articles in a day. The real moat is verifiable human originality: whether the information in a story was reported by a journalist or generated by a model, and whether an editorial judgment came from industry experience or from prompt assembly.
On an internet where 35% of new pages already carry AI traces, the ability to answer those questions well is itself a competitive advantage.


