One in Three New Web Pages Likely Written by AI, Pew Researchers Find
Analysis of nearly 500,000 pages reveals the scale of machine-generated content reshaping the internet since late 2022

The Scale of Machine Authorship
More than one-third of web pages published since late 2022 display characteristics consistent with AI authorship or substantial machine editing, according to data from Pew Research. The finding underscores how rapidly generative models have moved from novelty to infrastructure across the commercial web.
The research drew on Common Crawl archives, examining close to half a million English-language pages spanning five years. Detection technology from Open Pangram flagged content likely produced or heavily reshaped by large language models. In a July 2026 snapshot of 10,000 randomly sampled pages, roughly 10% showed significant AI fingerprints. But that figure includes pages predating the November 2022 release of ChatGPT, when such tools were not yet widely accessible.
Filtering for recency changes the picture. Among pages published only after ChatGPT became available, 35% bore markers of machine generation. The difference illustrates how quickly adoption has accelerated: the web's newer layer is fundamentally different in composition from what came before.
Domain Patterns and Institutional Resistance
Commercial domains are driving the shift. Pages on .com addresses showed AI authorship signatures at roughly ten times the rate of .edu or .gov sites, both of which hovered near 1%. Non-profit .org domains registered a 4.6% rate, still well below the commercial web but higher than government or academic publishers.
The divergence suggests institutional caution. Universities and government agencies face reputational and legal constraints that discourage wholesale adoption of machine-generated text, particularly in contexts requiring accountability. Commercial publishers, under pressure to scale content for search visibility and advertising inventory, face different incentives.
At DailyTechWire, we have tracked similar patterns in Southeast Asian e-commerce platforms, where product descriptions and category pages increasingly rely on template-driven LLM output. The economics are straightforward: human writers cost more and produce slower. For many commercial publishers, the trade-off has already tipped.
Stylistic Tells and Detection Limits
Beyond algorithmic detection, Pew researchers noted rising frequency of stylistic patterns associated with generative models. Use of em dashes, Oxford commas, and rhetorical structures like "it's not X, it's Y" have all increased in prevalence since late 2022. These are not definitive proof, but they align with the syntactic preferences encoded in training data and reinforcement tuning.
Detection tools like Pangram operate probabilistically and can misclassify human-written text as machine-generated, particularly when authors use formal or templated language. At scale, however, the directional signal is likely robust. The researchers acknowledged the margin of error but argued the trend is clear enough to warrant attention.
The bigger challenge lies in pages that blend human and machine contributions. An editor who drafts an outline, feeds it to a model, then revises the output has produced something neither fully human nor fully synthetic. Current detection methods struggle with this hybrid zone, which is likely where much professional content now lives.
Bots Reading Bot-Written Pages
The Pew data arrives weeks after Cloudflare reported that automated traffic had surpassed human browsing, a threshold the infrastructure provider had not expected to cross so soon. Taken together, the findings sketch a feedback loop: bots crawling pages written by other bots, with search engines and recommendation algorithms mediating the exchange.
This has implications for training data. If future models ingest web snapshots heavily populated by earlier model outputs, the risk of model collapse or stylistic homogenization grows. Researchers have warned that recursive training on synthetic data can degrade performance, particularly for tasks requiring nuance or factual grounding.
It also reshapes the economics of search engine optimization. If a significant share of new pages are machine-generated and optimized for algorithmic rather than human readers, the web's informational value to end users may decline even as its volume expands. Publishers chasing traffic may find themselves competing primarily with other automated systems rather than with human attention.
What Comes Next
The 35% figure is a snapshot, not a ceiling. Adoption of generative tools continues to accelerate, particularly in markets where content production costs are high relative to revenue. We have seen this in India's edtech sector, where course descriptions and marketing pages are increasingly templated through LLMs, and in Indonesia's travel aggregators, where destination guides are assembled from model output with minimal human review.
Regulatory responses remain fragmented. The European Union's AI Act includes transparency requirements for synthetic content, but enforcement mechanisms are still taking shape. In the United States, no comparable federal framework exists. Academic and government institutions have moved more cautiously, but commercial incentives point toward further automation.
The question is not whether machine-generated content will continue to grow, but whether the web's architecture can adapt without collapsing into a low-signal environment where bots primarily serve other bots. For publishers, the challenge is to deploy generative tools in ways that preserve editorial value rather than simply inflating page counts. For platforms, it is to build discovery systems that can surface human insight in a sea of synthetic text. Neither problem has an obvious solution yet.


