AI Newsway

Pew: 35% of Web Pages Published After ChatGPT Show Signs of AI Authorship

A Common Crawl analysis of roughly half a million pages finds machine writing concentrated on .com domains, while .edu and .gov barely register

|3 min read0
AI Summary
Pew Research Center reported Thursday that 35% of English-language web pages published after ChatGPT's November 2022 launch show detectable signs of AI authorship, based on nearly half a million Common Crawl pages scored by a Pangram classifier. The rate falls to about 10% across a random 2026 sample, and .com domains show machine writing at roughly ten times the near-1% rate of .edu and .gov sites. Detection errors mean the trend, not the exact percentage, is the durable finding.
A desktop keyboard sits idle as new research finds a third of recently published web pages show signs of machine authorship rather than human typing.
A desktop keyboard sits idle as new research finds a third of recently published web pages show signs of machine authorship rather than human typing.

Roughly one in three English-language web pages published since ChatGPT arrived carries detectable fingerprints of machine writing, according to research the Pew Research Center released on Thursday. The finding lands as the open web absorbs a wave of generative tooling that has been easier to observe anecdotally than to measure.

Pew drew its sample from Common Crawl, the public web archive, pulling close to half a million English-language pages spanning roughly five years. Crucially, the window opens about two years before ChatGPT shipped in November 2022, giving the researchers a pre-generative baseline to measure against rather than a single snapshot with nothing to compare it to.

How the Measurement Works

Detection ran through technology from Pangram, using an open-weight classifier that returns a score between 0 and 1 for a block of text. A zero indicates writing the model reads as fully human; a one indicates text it reads as fully machine-generated. Pew applied that scoring across the corpus rather than hand-auditing pages.

The headline number depends heavily on which pages you count. In a random sample of 10,000 pages captured in a July 2026 crawl, only about 10% showed significant signs of AI authorship. But a random slice of the web is mostly old web. Pages published years before generative writing tools existed cannot have been written by them, and they drag the average down.

Once Pew restricted the sample to pages carrying post-ChatGPT publication dates, the share jumped to 35%.

Where the Machine Writing Lives

The distribution is lopsided by domain. Commercial .com addresses showed signs of machine authorship at roughly ten times the rate of .edu and .gov addresses, both of which came in near 1%. Nonprofit .org domains landed in between at 4.6%.

That gradient tracks incentives more than technology. Commercial publishers face volume pressure that universities and government agencies largely do not, and the cheapest way to meet a content quota is now a prompt.

Pew also tracked stylistic markers that have become informal shorthand for machine prose, including rising use of em dashes, Oxford commas, and the rhetorical construction that sets up a negation before delivering a substitute. Those markers climbed over the same period.

The Caveats Are Real

Detection tools of this class misfire in both directions. They flag careful human writing as synthetic and clear synthetic text as human, and Pew acknowledges the limitation directly. What survives the noise is the direction of travel rather than any single percentage.

The study also arrives alongside a separate signal from the infrastructure layer. Cloudflare recently reported that automated traffic had overtaken human traffic on the web, and reached that crossover earlier than the company had projected. Read together, the two datasets describe a web where bots increasingly read pages that other bots wrote.

What It Changes

Earlier work had suggested the shift was less dramatic, with machine-written articles approaching rough parity with human ones in some samples without displacing them. Pew's numbers do not contradict that so much as sharpen it: the growth is concentrated in a specific, commercially motivated slice of the web.

The practical stakes sit downstream. Search ranking systems, retrieval pipelines, and the training corpora for the next generation of models all draw from this same pool. If a third of new commercial pages are machine-authored, the question of what counts as a primary source gets harder to answer every quarter.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.
AI & Machine Learning

A Researcher Quit Anthropic Over Extinction Risk. His Safety Lead Agreed in Public.

Jacob Coxon resigned from Anthropic over extinction risk. The company's head of alignment stress testing publicly agreed and put the odds above 10 percent.

Seung Jung8 days ago
OpenAI Says It Hit Its Automated Research Intern Goal. The Caveats Are in the Data
AI & Machine Learning

OpenAI Says It Hit Its Automated Research Intern Goal. The Caveats Are in the Data

OpenAI says agents now run 3.1 workdays of effort per human workday in its research org, but most long successful tasks still need human intervention.

Seung Jung6 days ago
Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier
AI & Machine Learning

Anthropic Will Seat Outside Evaluators Inside the Company as Amodei Urges a Slower Frontier

Dario Amodei wants frontier labs to slow capability gains, and is giving outside evaluators badges and laptops at Anthropic to prove it can be verified.

Seung Jung5 days ago
A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds
AI & Machine Learning

A 25-Turn Pressure Test Shows LLMs Fold While Their Reasoning Holds

A new arXiv benchmark called SPINE argues with models for up to 25 turns and finds collapse rates rise with conversation length for all seven systems tested.

Seung Jung4 days ago
An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.
AI & Machine Learning

An AI Cracked a 373-Year-Old Cipher. Then Someone Checked the Microfilm.

Vals AI reported Claude Fable 5.1 solved a 373-year-old cipher in 44 minutes. An independent replication reports 8 of 64 letters match — chance level.

Seung Jung4 days ago
OpenAI's 10,000-Agent Navier-Stokes Proof Lands With an Asterisk
AI & Machine Learning

OpenAI's 10,000-Agent Navier-Stokes Proof Lands With an Asterisk

OpenAI says 10,000 agents found a singularity in the Navier-Stokes equations, verified in Lean. It won't claim the Clay prize, and a credit fight has erupted.

Seung Jung7 days ago