# Cloudflare can now keep robots.txt aligned with AI bot policy

Cloudflare’s new Bot Preference Sync mirrors zone-level Search, Agent and Training choices into robots.txt, reducing a long-standing mismatch between stated crawler preferences and edge enforcement while preserving existing disallow rules.

Bot Preference Sync turns Cloudflare’s AI crawler controls into both an enforcement policy and a published robots.txt signal. It is rolling out to every plan, with different defaults for publishers and other sites and important limits around custom per-bot rules.

- Status: Active
- Published: 2026-08-23T20:25:55+12:00
- Updated: 2026-08-23T20:25:55+12:00
- Categories: Cloud & Infrastructure, Marketing & Distribution, GEO & AI Search, Edge & CDN
- Tags: AI search, Cloudflare
- Canonical HTML: https://beyondthe.news/dossiers/cloudflare-bot-preference-sync-robots-ai-crawlers

## What changed

Cloudflare announced Bot Preference Sync on August 21, 2026. The feature uses a site’s existing zone-level AI bot settings for Search, Agent and Training traffic to generate or update corresponding robots.txt directives, while keeping the edge enforcement policy and the public preference signal aligned. It will be available on every Cloudflare plan. For existing robots.txt files, Cloudflare prepends the generated block and preserves the site’s existing Disallow directives. New customers will have sync enabled by default, while existing users of Cloudflare’s legacy managed robots.txt feature will be prompted to review and confirm their preferences when the new feature launches.

## Why it matters

Publishers and other site owners increasingly have two separate control planes for AI crawlers: a voluntary robots.txt signal and an enforceable CDN or bot-management rule. Those can drift apart, creating ambiguity for compliant crawlers and operational mistakes for site owners. Bot Preference Sync makes one category-level policy feed both layers. The practical consequence is easier separation of search discovery, agent access and model training, especially for sites that want to remain discoverable while declining training. The limitation is equally important: Cloudflare is synchronizing category-level policy, not arbitrary custom per-bot exceptions.

## One policy now drives two different control layers

Cloudflare’s AI bot settings can Allow, block on ad-supported pages or block Search and Agent traffic, while Training can be disallowed. Bot Preference Sync translates those category choices into robots.txt while edge controls continue to provide actual enforcement where configured. The feature therefore reduces policy drift without pretending that robots.txt itself is an access-control mechanism.

## Disallow Training is designed to preserve search for cooperating mixed-use crawlers

Cloudflare says a Training disallow writes a no-training preference into robots.txt. Mixed-use crawlers that meet its transparency criteria can still access content for search indexing while honoring the training preference. Cloudflare’s criteria include respecting a no-training signal, offering an AI-summary opt-out, giving site owners visibility into training/search usage and demonstrating that declining training does not hurt traditional search results.

## Existing files and custom rules need deliberate review

Generated directives are prepended to an existing robots.txt, so existing Disallow entries remain. But Bot Preference Sync does not read or reproduce individual custom rules with more complex exceptions. Sites with negotiated crawler access, granular bot rules or other hand-maintained logic should review the generated result and may need to disable sync.

## Publisher onboarding gets a different default

Cloudflare says new publisher or ad-supported customers can choose an onboarding option that defaults Training to Disallow while leaving search available. Other new customers start without category blocks or disallows. Those defaults are editable, but they make business model part of the initial crawler-policy configuration.

## Key details

- Cloudflare announced Bot Preference Sync on August 21, 2026 and says it will roll out to all plans in the following week.
- The feature mirrors zone-level Search, Agent and Training policies into robots.txt.
- Generated directives are prepended to an existing robots.txt and preserve existing Disallow directives.
- Bot lists are updated from Cloudflare’s BotBase/verified-bot classifications.
- The sync works at category level and does not automatically reproduce individual custom bot rules.
- New publisher/ad-supported customers can default Training to Disallow; other new customers start with no blocks or disallows.

## Builder takeaways

- Treat robots.txt and edge blocking as different mechanisms even after enabling sync: one states a preference, the other can enforce it.
- If AI-search visibility matters, explicitly decide Search, Agent and Training policy separately instead of applying one blanket crawler rule.
- Audit the resulting robots.txt after rollout, especially if the site already has hand-written crawler directives or commercial exceptions.
- Sites with per-bot custom rules should not assume Bot Preference Sync understands them; either keep the sync off or maintain a documented reconciliation process.
- Publisher teams should coordinate SEO, editorial and infrastructure ownership so a dashboard policy change does not unexpectedly alter discoverability or training preferences.

## What to watch

- The exact general-availability date and migration prompt for existing managed-robots.txt users.
- Which mixed-use crawlers satisfy Cloudflare’s transparency requirements and how often that list changes.
- Whether major AI crawler operators consistently honor the generated no-training preferences.
- Operational edge cases when category-wide sync interacts with custom robots.txt logic or negotiated crawler exceptions.

## Uncertainties

- Cloudflare’s claim that cooperating mixed-use crawlers can preserve search visibility while respecting no-training preferences depends on crawler behavior outside Cloudflare’s control.
- The feature had been announced but was still due to roll out in the following week at publication time.
- Cloudflare does not claim that robots.txt can technically prevent non-compliant crawlers from accessing content.

## Sources

- [Say it once: introducing Bot Preference Sync](https://blog.cloudflare.com/bot-preference-sync/) — Cloudflare · primary · 2026-08-21T00:00:00+12:00. Primary announcement describing rollout, defaults, category policies, robots.txt behavior and transparency criteria.

