Researchers who pretrained 800 language models report that AI-generated web text helps or hurts a model depending entirely on how data-starved that model is, and that the standard scaling laws used to plan training runs cannot represent the difference. In a paper posted to arXiv on September 30, they also quantify how much of the open web is already machine-written: after FineWeb quality filtering, 27.5% of tokens in June 2026 web data were labeled AI-generated, rising to 31.1% by August.
Key takeaways
- After FineWeb filtering, the AI-generated share of web tokens rose from 27.5% in June 2026 to 31.1% in August, as measured by the Pangram classifier.
- For data-starved models, adding AI tokens first lowers loss on human text, then saturates, then reverses into harm; for models with large human-text budgets, AI tokens raise loss almost immediately.
- The authors' proposed scaling law splits benefit and harm into separate terms, predicts models up to 3.6x larger with 41% lower error than the best existing law, and collapses back to Chinchilla when no AI text is present.
Why wild AI text is not synthetic data
Most prior work on training models with machine output studies either deliberate synthetic data generation or model-collapse scenarios, where a model is fed its own outputs in a loop. The paper argues that what actually shows up in pretraining corpora is neither. Wild AI text comes from many different models, it was written to be read by humans rather than to serve as training fuel, and it arrives unlabeled β mixed into crawls with no marker distinguishing it from human prose.
That matters because quality filtering does not remove it. The 27.5% and 31.1% figures are measured after FineWeb's quality pipeline has run, which means the AI-written share survives exactly the screen most large language model teams rely on. Pew's finding that 35% of post-ChatGPT web pages show signs of AI authorship described the same trend from the publishing side; this paper measures what lands in the training set and then tests what it does there.
What 800 pretraining runs showed
The team varied the ratio of added AI tokens to human tokens across those 800 models and fit scaling laws to held-out losses on both human-written and AI-written text. Two regimes emerged.
When a model is data-starved β more compute than clean human text to spend it on β adding AI tokens initially lowers loss on human text. The benefit saturates as the ratio climbs, then turns negative. When a model is trained on a high budget of human text, there is no honeymoon at all: AI tokens raise loss almost immediately, while the same number of fresh human tokens keeps pushing loss down.
The practical consequence is that there is no single answer to what an AI token is worth. Its value depends on the alternative, and the sign flips. That is precisely what the Chinchilla-style laws from Hoffmann et al. (2022) cannot express: they model loss as a smooth function of parameters and tokens, treating all tokens as interchangeable units of the same good.
A scaling law that allows a negative token
The authors propose a replacement with separate benefit and harm terms, so the marginal value of an AI token can change sign as the mix shifts. Crucially, the law reduces to Chinchilla when the AI share is zero, which makes it a strict extension rather than a competing formalism.
Fit on smaller models, it predicted the effect of AI text on held-out human-text loss for models up to 3.6x larger with 41% lower error than the best existing law, measured across all AI ratios. That extrapolation property is the point: labs fit scaling laws on cheap small runs specifically to decide how to spend a far larger budget, and a law that misprices contaminated data misprices the whole run.
What the paper tells practitioners to do
Three recommendations follow. Filter AI text when the target distribution is human text. Repeat human text before expanding the dataset with AI-generated web text β a direct challenge to the reflex of simply crawling more. And report validation loss on human and AI text separately, because a single blended number hides the tradeoff entirely.
The authors are careful about the inverse case: AI text remains valuable when the target is AI text. A model meant to operate on machine-written input is not harmed by training on it. The damage is specific to the mismatch between what you train on and what you intend to model.
To support replication, the team released WildAI, an 83-billion-token corpus labeled for AI authorship, topic and format, along with all 800 models and the code. The AI-authorship labels come from Pangram Labs, whose co-founders Max Spero and Bradley Emi are among the paper's seven authors alongside Jenna Russell, Ben Glickenhaus, Katherine Thai, John Wieting and Mohit Iyyer β a detail worth noting, since the measured contamination rate depends on that classifier's accuracy.
Outlook
If the AI share of filtered web text keeps climbing at roughly a point a month, the window for treating fresh crawls as a clean human-text reservoir is closing fast. The paper's framing turns that from a vague worry into a budgeting problem with a formula attached: data curation stops being hygiene and becomes a term in the scaling equation. Expect detection quality itself to become contested infrastructure, because every lab's compute plan now depends on a classifier's verdict about who wrote what. Treating AI web text as synthetic data with a known provenance was never accurate; the honest position is that nobody knows the mix in a 2026 crawl without measuring it.
FAQ
Does training on AI-generated web text always hurt a model?
No. The paper finds the effect depends on how much human text the model already has. Data-starved models see loss on human text fall at first when AI tokens are added, before the benefit saturates and reverses. Models trained on large human-text budgets are hurt almost immediately.
Why do Chinchilla scaling laws fail here?
Chinchilla-style laws treat tokens as interchangeable and model loss as a smooth function of parameters and data volume, so they cannot represent a token whose marginal value changes from positive to negative. The authors add separate benefit and harm terms, and their law reduces to Chinchilla when there is no AI text.
What is WildAI?
WildAI is the 83-billion-token corpus the authors released alongside the paper, labeled for AI authorship, topic and format. All 800 pretrained models and the training code were released with it, so other researchers can refit the scaling law or test alternative filtering strategies.






