Three of the largest crawler operators on the internet have agreed to a set of conditions written by a network provider. Cloudflare announced on September 15 that Apple, Google and Microsoft meet, or have committed to dates for meeting, a designation it calls Accountable - and alongside it shipped a setting that lets a site refuse AI training without dropping out of search.
Key takeaways
- Seventeen percent of Cloudflare sites already block AI training, while under one percent block search - a gap mixed-use crawlers had made impossible to act on.
- Earning Accountable status requires training and summary opt-outs, URL-level reporting on what was used, and a guarantee that refusing training will not move search rankings.
- Microsoft's robots.txt support lands in early 2027, so the new setting conveys nothing to Bing until then.
The bind mixed-use crawlers created
A mixed-use crawler fetches a page once and puts it toward two ends: building a search index, and training a model. Publishers had no way to split those. Turning the crawler away to protect content from training also removed the pages from search, which for an ad-funded or subscription business is the expensive half of the bargain.
The numbers show how lopsided the preference is. Almost nobody turns down search, and roughly one site in six has reached for some mechanism against training. Cloudflare's new option, Disallow AI Training, writes a no-training directive into robots.txt through Bot Preference Sync while leaving Applebot, Bingbot and Googlebot free to index.
Why a file alone could not settle it
Publishing a directive and having it obeyed are different things. A robots.txt cannot identify who is fetching a page, determine the purpose, or stop a crawler that disregards it. Sitting in front of the origin, Cloudflare can do all three, then publish what each operator actually did on its Radar service - which is the enforcement the designation leans on.
Training-only crawlers were never the hard case. Amazon, Anthropic, Meta and OpenAI run separate bots for search and training, so the training one can be blocked with no cost to discoverability. Those operators are also counted as Accountable.
Where each operator stands
Google honors a Disallow on Google-Extended, offers a generative-results toggle in its webmaster portal, and says URL-level transparency tools are weeks away. Apple respects Applebot-Extended and reads nosnippet for summaries, but has no inspection tool yet and showed Cloudflare one in progress for next year. Both companies state that refusing training does not affect search rank.
Microsoft is the gap. Bing accepts a NOARCHIVE meta tag today, and its robots.txt handling of a no-training preference is targeted for early 2027. Site owners who want out of Bing training before then are pointed at NOARCHIVE or the content removal tools.
Summaries, not training, are the next argument
Cloudflare separates the two questions deliberately. Training asks whether content can build a model; summarization decides whether anyone visits the site at all. Its own figures make that tension concrete: more than half of consumers read search summaries, and those readers are over 40% more likely to stop searching afterward - yet visitors who do arrive from AI search convert at three to five times the rate of ordinary search traffic.
A site-wide yes or no is too blunt for that tradeoff, so the stated goal for early next year is control over how much content a summary may draw on, set once at the network instead of operator by operator, expressed through the emerging ai-prefs standard.
For Cloudflare, which has argued the web needs rebuilding around agent traffic rather than human traffic, writing the rules that Apple, Google and Microsoft then sign is the more durable win. The controls are live on every plan today.
FAQ
Does disallowing AI training remove my site from Google Search?
No. Accountable mixed-use crawlers such as Googlebot and Applebot keep crawling for search under the new setting, and both Google and Apple have stated that opting out of training does not affect search ranking. Choosing Block instead now stops those crawlers entirely, search included.
Which crawlers does Disallow AI Training stop?
All training crawlers except Accountable mixed-use ones, including the training-only bots run by Amazon, Anthropic, Meta and OpenAI. Blocking those carries no search penalty because each company operates a separate crawler for indexing.
Is the setting available on free plans?
Yes. Cloudflare says the controls reach all customers on all plans, configured per domain in Security Settings. Domains onboarded from September 15 are offered preset configurations that differ depending on whether the site serves advertising.






