Cloudflare Gives Publishers Default Block on AI Training Crawlers
The infrastructure giant's second Content Independence Day introduces granular bot controls that could reshape how AI companies access the web.
Cloudflare announced July 1, 2026, that all customers can now independently control three categories of AI bot traffic—Search, Agent, and Training—marking the company's second Content Independence Day and a significant escalation in the standoff between publishers and AI scrapers.
The new controls, live immediately in zone Settings for all existing customers, let website owners block Training and Agent crawlers while keeping Search bots allowed. The distinction matters: multi-purpose crawlers from Google, Apple, and Microsoft that combine search indexing with model training will be blocked entirely if a site owner selects the Training block, according to Cloudflare's documentation.
For new domains onboarding to Cloudflare, the default flips entirely. Starting September 15, 2026 (for new domains onboarding to Cloudflare; existing customers can opt in or out), Training and Agent bots will be blocked by default on pages displaying ads, while Search crawlers remain permitted.
The move builds on Cloudflare's first Content Independence Day in July 2025, which offered a one-click "Block AI Bots" option. That blunt instrument has now evolved into surgical controls that force AI companies to declare their intentions—or lose access to a significant portion of web domains.
"The deal between crawlers and website owners that had held up for 30 years — we crawl you, and you get referrals — was no longer true," Cloudflare wrote in its announcement. The company frames the update as a response to an existential threat: AI models now consume content for training without generating the referral traffic that once justified open crawling.
The three-category taxonomy creates immediate pressure on crawlers that blur the line between search and training. Googlebot, Applebot, and BingBot all perform dual functions under current configurations. Site owners who block Training will see these crawlers rejected entirely, even if Search remains allowed—a technical reality that could accelerate the industry's move toward separated crawling functions or compensated content agreements.
The timing coincides with growing publisher frustration over uncompensated model training. By giving default protection to new ad-monetized sites, Cloudflare is betting that the economic model of the open web requires explicit consent for AI training use, not just for search indexing.
What happens next depends on how major AI companies respond. They can separate their search and training crawlers, negotiate access deals with publishers, or accept reduced training data from a significant portion of the web. Cloudflare's move suggests the third option may be the default outcome.