Updated 16 Sep 2026: Adds Cloudflare's Sep 15 Disallow AI Training and Accountable mixed-use crawler model, clarifying how publishers can preserve search access while expressing no-training preferences and how Block semantics now apply to mixed-use crawlers.

Key details

  1. Disallow AI Training publishes the applicable no-training preference through Bot Preference Sync.
  2. Accountable mixed-use crawlers remain allowed for Search under Disallow AI Training.
  3. Cloudflare identifies Applebot, Googlebot and Bingbot as Accountable mixed-use crawlers.
  4. Cloudflare's Accountable criteria include training opt-out, AI-summary opt-out or a commitment to it, URL-level transparency or a commitment to it, and assurance that training opt-out does not harm traditional search.
  5. Block and Block on pages with ads now apply to mixed-use crawlers, so those settings can affect Search as well as Training.
  6. Training-only crawlers from Amazon, Anthropic, Meta and OpenAI can be blocked without the same search-discoverability trade-off because those operators separate Search and Training crawlers.
  7. For existing domains using granular controls, previous Training Block or Block on pages with ads selections migrate to Disallow AI Training.
  8. For new ad-supported domains, Cloudflare recommends Search Allow, Training Disallow AI Training and Agent Block on pages with ads.
  9. Bing's domain-level robots.txt no-training support is targeted for early 2027; until then Cloudflare notes that Disallow AI Training does not automatically convey that preference to Bing through robots.txt.

What builders should take away

  1. If you want search visibility but not AI training, review whether Disallow AI Training better matches your intent than a full Training Block.
  2. Treat Block as an access-control decision: for mixed-use crawlers it can remove traditional search crawling too.
  3. Recheck existing Cloudflare zones after the September 15 migration so you understand how legacy settings were translated.
  4. Publishers using Bing should note the current gap in domain-level robots.txt no-training support and use Microsoft's available controls where required.
  5. Separate Search, Training and Agent policy decisions instead of treating every automated visitor as one class.

What changed

Cloudflare has expanded the September 15 rollout with a Disallow AI Training setting and an Accountable designation for crawler operators. Disallow AI Training publishes a no-training preference through Bot Preference Sync while allowing Accountable mixed-use crawlers to continue serving Search. Cloudflare says Apple, Google and Microsoft meet or have time-bound commitments to meet the Accountable requirements. At the same time, the semantics of Block and Block on pages with ads have changed: those stronger settings now apply to mixed-use crawlers as well as training-only crawlers. Existing granular Training Block selections are migrated to Disallow AI Training to preserve their practical intent, while new ad-supported domains are recommended to allow Search, disallow Training and block Agents on pages with ads.

Why it matters

This removes an important false binary for publishers. A crawler such as Googlebot, Applebot or Bingbot may serve both Search and AI-related purposes, so simply blocking the crawler can sacrifice discoverability. Cloudflare now offers a preference layer intended to separate those uses where operators support it. Builders still need to understand the enforcement boundary: Disallow AI Training is partly an expressed preference for Accountable mixed-use crawlers, whereas Block is an edge access control that can remove the crawler entirely. The distinction matters for SEO, AI training policy and publisher monetisation.

Disallow AI Training is deliberately different from Block

Cloudflare now separates a content-use preference from an access denial. Disallow AI Training publishes a no-training preference and leaves Accountable mixed-use crawlers available for Search. Block is the stronger network control: it stops the crawler, including its search function when the same bot serves multiple purposes.

Accountable is a capability-and-commitment designation

Cloudflare's Accountable label is meant to tell publishers which operators provide, or have made time-bound commitments to provide, meaningful separation and transparency. The requirements cover training opt-out, AI-summary controls, URL-level visibility and an assurance that opting out of training will not reduce ordinary search ranking. Cloudflare currently places Apple, Google and Microsoft in this group for their mixed-use crawlers.

The September 15 migration changes existing settings too

Cloudflare is migrating previous granular Training Block selections to Disallow AI Training, preserving the intended no-training behavior without unnecessarily blocking mixed-use search crawlers. Legacy Block AI configurations are also mapped into the new Search, Training and Agent controls.

Not every crawler can honor the preference in the same way yet

The mechanism depends partly on crawler-operator support. Google and Apple already expose separate training preference mechanisms; Microsoft is targeting domain-level robots.txt support for Bing in early 2027. Cloudflare can still edge-block crawlers, but the ability to preserve Search while refusing Training depends on the operator honoring the relevant preference.

What to watch next

  • Whether Microsoft ships Bing's promised domain-level robots.txt training opt-out in early 2027.
  • Google's planned additional URL-level Google-Extended transparency tools.
  • Cloudflare's planned controls for how much publisher content may appear in AI summaries.
  • Whether the Accountable designation expands to more mixed-use crawler operators and whether commitments are delivered on schedule.
  • Adoption of open preference standards such as ai-prefs.

Still unclear

  • The Accountable designation includes both capabilities available now and time-bound commitments, so not every listed feature is implemented by every operator today.
  • Bing's robots.txt no-training support is not expected until early 2027, leaving a temporary difference between the intended policy and what can be conveyed automatically.
  • Crawler classifications and operator commitments may evolve, changing which controls are available or how Cloudflare applies them.

Sources

Direct reading behind this dossier.

4 sources

Discussion

Discussion is reader-contributed. Comments are not part of the BTN dossier or its editorial evidence.

0 visible comments

Join the discussion

Keep comments useful and relevant. Reader contributions may be moderated and are not BTN editorial evidence.

Sign in to comment