Roughly one in three English-language web pages published since ChatGPT arrived carries detectable fingerprints of machine writing, according to research the Pew Research Center released on Thursday. The finding lands as the open web absorbs a wave of generative tooling that has been easier to observe anecdotally than to measure.
Pew drew its sample from Common Crawl, the public web archive, pulling close to half a million English-language pages spanning roughly five years. Crucially, the window opens about two years before ChatGPT shipped in November 2022, giving the researchers a pre-generative baseline to measure against rather than a single snapshot with nothing to compare it to.
How the Measurement Works
Detection ran through technology from Pangram, using an open-weight classifier that returns a score between 0 and 1 for a block of text. A zero indicates writing the model reads as fully human; a one indicates text it reads as fully machine-generated. Pew applied that scoring across the corpus rather than hand-auditing pages.
The headline number depends heavily on which pages you count. In a random sample of 10,000 pages captured in a July 2026 crawl, only about 10% showed significant signs of AI authorship. But a random slice of the web is mostly old web. Pages published years before generative writing tools existed cannot have been written by them, and they drag the average down.
Once Pew restricted the sample to pages carrying post-ChatGPT publication dates, the share jumped to 35%.
Where the Machine Writing Lives
The distribution is lopsided by domain. Commercial .com addresses showed signs of machine authorship at roughly ten times the rate of .edu and .gov addresses, both of which came in near 1%. Nonprofit .org domains landed in between at 4.6%.
That gradient tracks incentives more than technology. Commercial publishers face volume pressure that universities and government agencies largely do not, and the cheapest way to meet a content quota is now a prompt.
Pew also tracked stylistic markers that have become informal shorthand for machine prose, including rising use of em dashes, Oxford commas, and the rhetorical construction that sets up a negation before delivering a substitute. Those markers climbed over the same period.
The Caveats Are Real
Detection tools of this class misfire in both directions. They flag careful human writing as synthetic and clear synthetic text as human, and Pew acknowledges the limitation directly. What survives the noise is the direction of travel rather than any single percentage.
The study also arrives alongside a separate signal from the infrastructure layer. Cloudflare recently reported that automated traffic had overtaken human traffic on the web, and reached that crossover earlier than the company had projected. Read together, the two datasets describe a web where bots increasingly read pages that other bots wrote.
What It Changes
Earlier work had suggested the shift was less dramatic, with machine-written articles approaching rough parity with human ones in some samples without displacing them. Pew's numbers do not contradict that so much as sharpen it: the growth is concentrated in a specific, commercially motivated slice of the web.
The practical stakes sit downstream. Search ranking systems, retrieval pipelines, and the training corpora for the next generation of models all draw from this same pool. If a third of new commercial pages are machine-authored, the question of what counts as a primary source gets harder to answer every quarter.






