AI Newsway

Apple, Google and Microsoft Sign On to Cloudflare's Crawler Rules

A new Disallow AI Training setting ends the choice between search visibility and model training

|4 min read0
AI Summary
Cloudflare launched a Disallow AI Training setting on September 15, 2026, letting sites refuse AI model training while staying indexed by crawlers that serve both search and training. Apple, Google and Microsoft meet its new Accountable designation, which requires opt-outs, URL-level reporting and a promise that refusing training will not hurt search rank. Microsoft's robots.txt support arrives in early 2027, and summary-level controls are promised for next year.
The entrance to a Cloudflare office, where the company drafted the crawler conditions Apple, Google and Microsoft have now agreed to meet
The entrance to a Cloudflare office, where the company drafted the crawler conditions Apple, Google and Microsoft have now agreed to meet

Three of the largest crawler operators on the internet have agreed to a set of conditions written by a network provider. Cloudflare announced on September 15 that Apple, Google and Microsoft meet, or have committed to dates for meeting, a designation it calls Accountable - and alongside it shipped a setting that lets a site refuse AI training without dropping out of search.

Key takeaways

  • Seventeen percent of Cloudflare sites already block AI training, while under one percent block search - a gap mixed-use crawlers had made impossible to act on.
  • Earning Accountable status requires training and summary opt-outs, URL-level reporting on what was used, and a guarantee that refusing training will not move search rankings.
  • Microsoft's robots.txt support lands in early 2027, so the new setting conveys nothing to Bing until then.

The bind mixed-use crawlers created

A mixed-use crawler fetches a page once and puts it toward two ends: building a search index, and training a model. Publishers had no way to split those. Turning the crawler away to protect content from training also removed the pages from search, which for an ad-funded or subscription business is the expensive half of the bargain.

The numbers show how lopsided the preference is. Almost nobody turns down search, and roughly one site in six has reached for some mechanism against training. Cloudflare's new option, Disallow AI Training, writes a no-training directive into robots.txt through Bot Preference Sync while leaving Applebot, Bingbot and Googlebot free to index.

Why a file alone could not settle it

Publishing a directive and having it obeyed are different things. A robots.txt cannot identify who is fetching a page, determine the purpose, or stop a crawler that disregards it. Sitting in front of the origin, Cloudflare can do all three, then publish what each operator actually did on its Radar service - which is the enforcement the designation leans on.

Training-only crawlers were never the hard case. Amazon, Anthropic, Meta and OpenAI run separate bots for search and training, so the training one can be blocked with no cost to discoverability. Those operators are also counted as Accountable.

Where each operator stands

Google honors a Disallow on Google-Extended, offers a generative-results toggle in its webmaster portal, and says URL-level transparency tools are weeks away. Apple respects Applebot-Extended and reads nosnippet for summaries, but has no inspection tool yet and showed Cloudflare one in progress for next year. Both companies state that refusing training does not affect search rank.

Microsoft is the gap. Bing accepts a NOARCHIVE meta tag today, and its robots.txt handling of a no-training preference is targeted for early 2027. Site owners who want out of Bing training before then are pointed at NOARCHIVE or the content removal tools.

Summaries, not training, are the next argument

Cloudflare separates the two questions deliberately. Training asks whether content can build a model; summarization decides whether anyone visits the site at all. Its own figures make that tension concrete: more than half of consumers read search summaries, and those readers are over 40% more likely to stop searching afterward - yet visitors who do arrive from AI search convert at three to five times the rate of ordinary search traffic.

A site-wide yes or no is too blunt for that tradeoff, so the stated goal for early next year is control over how much content a summary may draw on, set once at the network instead of operator by operator, expressed through the emerging ai-prefs standard.

For Cloudflare, which has argued the web needs rebuilding around agent traffic rather than human traffic, writing the rules that Apple, Google and Microsoft then sign is the more durable win. The controls are live on every plan today.

FAQ

Does disallowing AI training remove my site from Google Search?

No. Accountable mixed-use crawlers such as Googlebot and Applebot keep crawling for search under the new setting, and both Google and Apple have stated that opting out of training does not affect search ranking. Choosing Block instead now stops those crawlers entirely, search included.

Which crawlers does Disallow AI Training stop?

All training crawlers except Accountable mixed-use ones, including the training-only bots run by Amazon, Anthropic, Meta and OpenAI. Blocking those carries no search penalty because each company operates a separate crawler for indexing.

Is the setting available on free plans?

Yes. Cloudflare says the controls reach all customers on all plans, configured per domain in Security Settings. Domains onboarded from September 15 are offered preset configurations that differ depending on whether the site serves advertising.

How do you feel about this article?

SJ

Discussion

Sign in to post
Loading...

Related articles

OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months
Tech & Business

OpenAI Says Its Own Models Helped Tape Out Jalapeño in Nine Months

AI-generated kernels beat OpenAI expert-written versions by up to 1.8x, as Jalapeño posts its first InferenceX benchmark results.

Seung Jung22 days ago
Starcloud Raised $250M for Orbital AI Compute. Its Hardest Problem Is Booking a Rocket.
Tech & Business

Starcloud Raised $250M for Orbital AI Compute. Its Hardest Problem Is Booking a Rocket.

The Series A extension doubles Starcloud to a $2.3B valuation, but Falcon 9 winds down in 2028 and Starship has yet to prove rapid reuse.

Seung Jung25 days ago
Feds Warn AI-Written Exploits Are Already Hitting Critical Infrastructure Controllers
Tech & Business

Feds Warn AI-Written Exploits Are Already Hitting Critical Infrastructure Controllers

Five US agencies say attackers are using AI coding assistants and the open source snap7 library to build tools that read and write Siemens S7 controller memory.

Seung Jung28 days ago
Google Hands Publishers a Button to Claw Back AI-Era Search Traffic
Tech & Business

Google Hands Publishers a Button to Claw Back AI-Era Search Traffic

Google now lets publishers embed a Preferred Sources button on their own sites, letting readers pin them across Search, Discover, and News in one tap.

Seung Jung27 days ago
Judge Voids Pentagon Supply Chain Risk Label on Anthropic as Unlawful Retaliation
Tech & Business

Judge Voids Pentagon Supply Chain Risk Label on Anthropic as Unlawful Retaliation

A 59-page order found the Pentagon retaliated against Anthropic for criticizing the government, violating the First and Fifth Amendments.

Seung Jung20 days ago
AWS Buys DuckLabs, and DuckDB's Extension Stack Is About to Open Up
Tech & Business

AWS Buys DuckLabs, and DuckDB's Extension Stack Is About to Open Up

AWS will acquire DuckLabs in early September. DuckDB stays MIT-licensed under an independent foundation, and the extension stack is set to open up.

Seung Jung20 days ago