How Bot Preference Sync Keeps robots.txt Current
Cloudflare announced Bot Preference Sync on its official blog on August 22, 2026, describing a feature that automatically keeps a website's robots.txt file aligned with the AI bot policies an owner has already configured in the Cloudflare dashboard. The tool covers three bot categories - Search, Agent, and Training - so one dashboard choice can update the rules that govern all three crawler types at once. Cloudflare generates or updates the robots.txt file based on the site's zone-level configuration, meaning the settings already stored for that domain drive the resulting text file automatically.
New directives are prepended to the existing robots.txt file, which preserves any Disallow rules an owner had already written by hand. Cloudflare also keeps the file current by drawing on its own BotBase bot-classification list, so entries update automatically as new bots are added to blocked or disallowed categories. The feature can be switched on or off at any time, and it is available to every Cloudflare customer, from the Free tier through Enterprise, rolling out within the week following the announcement.
What Happens Automatically for New and Ad-Monetized Sites
Brand-new Cloudflare customers start with no bot blocks at all under this system, leaving the entire policy choice to the site owner from the first setup. Sites that identify themselves as ad-monetized, meaning they tell Cloudflare they monetize pages with ads, can instead select a preset that sets the Training category to Disallow by default. That single distinction already shows how a site's stated business model, and not a deliberate technical decision, can steer its AI-training exposure before the owner writes a single rule.
The Training category carries four distinct configuration options: Allow, Block on pages that serve ads, Block everywhere, or Disallow, which writes a no-training preference directly into the robots.txt file. Search and Agent traffic remain separately controllable from Training, so an owner can keep a site fully open to search indexing while restricting or blocking crawlers that harvest content for AI-training datasets. Because these choices sit inside a dashboard rather than a manually maintained file, changing a policy takes effect the next time Cloudflare syncs the robots.txt output.
From Manual Dashboard Toggle to Automatic Enforcement
Cloudflare first gave site owners granular control over AI bot traffic on July 1, 2026, letting them separately allow or block crawlers across the same three categories, Search, Agent, and Training, from the dashboard. Bot Preference Sync, announced seven weeks later on August 22, is the automation layer that keeps the resulting robots.txt file matched to those dashboard choices. Owners no longer need to hand-edit a static text file every time a preference changes; Cloudflare handles that step automatically.
The shift moves control of a website's AI-crawling permissions out of a document most owners never open and into a product Cloudflare is actively building out over successive releases. Once the default behavior of robots.txt is generated by dashboard settings and not by whoever last edited the file, the practical AI-training policy for a large share of the web starts to depend on what Cloudflare ships as its defaults. A mixed-use crawler that does both search-indexing and AI-training must meet four specific criteria, including respecting robots.txt preferences and providing URL-level visibility into what was crawled and why, to avoid being blocked once an owner disallows Training.
The New Checkbox Decision for EU and UK Site Owners
Any EU or UK business running a website now faces an AI-training decision that resolves to a dashboard checkbox. Most site owners never open their Cloudflare bot settings page, which means the platform's chosen defaults, such as the ad-monetized preset that disallows Training, will end up setting policy for a large number of sites regardless of whether the owner ever makes an active choice. For the first time, staying silent produces a specific, automatically enforced outcome.
Owners who want a deliberate policy, whether that means allowing AI-training crawlers, blocking them everywhere, or restricting them only on ad-monetized pages, can now set that policy in minutes through the same dashboard used for Search and Agent traffic. The four Training options give enough range to match most editorial or commercial positions without touching a text file directly. Businesses that skip the settings page are not opting out of a decision; they are simply accepting whichever default Cloudflare or their own site profile assigns to them.
Read next: Cloudflare's New Browser Cuts AI Costs, Not Time | The Spec Moved Three Revisions Past the Product



