If your website runs behind Cloudflare — and roughly 20% of all websites do — something important is changing on September 15, 2026, and most small business owners have not heard about it.
On July 1, 2026, Cloudflare announced it was retiring its familiar single-toggle "Block AI Bots" switch and replacing it with three independent controls, each targeting a different type of crawler: Search, Agent, and Training. The new defaults take effect September 15. For many sites, these changes are straightforward. But for a specific group — sites on Cloudflare that carry advertising and previously enabled the old "Block AI Bots" toggle — the update contains a serious hidden risk: blocking Training crawlers under the new rules can also block Googlebot, Applebot, and BingBot. (Cloudflare Blog, July 1 2026)
What Are the Three Crawler Types?
Cloudflare's new system separates bot traffic into three categories:
- Search crawlers: Bots that index your site so it appears in search results. Googlebot, Bingbot, and Applebot historically fall here — but see the important caveat below.
- Training crawlers: Bots that scrape content to train AI language models. Examples include CCBot and similar scrapers used by AI companies for dataset building.
- Agent crawlers: Bots that pull information at query time to power AI assistant answers — the kind of bot that allows ChatGPT or Perplexity to answer a real-time question using your page content.
You can set each toggle independently. Want to be indexed by search engines and cited by AI assistants but not used for model training? In theory, you can allow Search and Agent while blocking Training. That sounds logical. Here is where it gets complicated.
The Multi-Purpose Crawler Problem
Googlebot, Applebot, and BingBot are not single-purpose crawlers. They crawl for search indexing, but they also gather data used for AI-related purposes — placing them in the "mixed-use" or multi-purpose category under Cloudflare's new classification system.
Under the September 15 rules, any crawler that serves more than one purpose is treated under the most restrictive rule that applies to it. This means that on pages where Training is blocked, Google's and Bing's multi-purpose crawlers can be caught by that block — even if you explicitly allow Search crawlers. (Search Engine Journal, July 2026)
The consequence: a site that wants to keep Google indexing while blocking AI training scrapers may inadvertently prevent Googlebot from crawling its pages at all. This is not theoretical — it is the documented behavior of the new policy for sites on ad-monetized pages.
Who Is at Risk?
You are most at risk if any of these apply to your site:
- Your site runs behind Cloudflare (free or paid plan)
- Your site carries advertising (display ads, affiliate links, or sponsorships)
- You previously enabled the "Block AI Bots" toggle that Cloudflare launched in July 2025
- You are a new Cloudflare customer setting up a new domain after September 15
Small business owners who set up Cloudflare once and never revisited the security settings are in the most precarious position. The new defaults are being applied automatically — this is not an opt-in change.
The GEO Dimension: Agent Crawlers Are Your Citation Pipeline
There is a second implication beyond search indexing that matters specifically for businesses focused on AI visibility. Agent crawlers — the bots that power ChatGPT's real-time search, Perplexity's answers, and Google's AI Mode — fall into their own separate category under the new system. If you block Agent crawlers, you are effectively cutting off the pipeline that allows AI tools to find and cite your business when a user asks a relevant question.
Given that AI-referred traffic converts at significantly higher rates than standard search traffic — Previsible's 2026 study found ChatGPT visitors convert at roughly 15.9% versus Google organic's ~1.76% — inadvertently blocking Agent crawlers is not a neutral choice. It is actively closing the highest-value channel before it has a chance to send you customers.
Three Things to Do Before September 15
- Audit your Cloudflare settings. Log in to your Cloudflare dashboard, navigate to Security → Bots, and look for the AI Crawler controls. Note which toggles are currently enabled and compare against the new three-category system. If you enabled the old "Block AI Bots" toggle in 2025, review its status — it is the highest-risk configuration under the September 15 changes.
- Decide on each category intentionally. The recommended configuration for most small businesses focused on both search ranking and AI citation: allow Search, allow Agent, block Training. This keeps Googlebot indexing normally and keeps AI assistant citation pipelines open, while preventing bulk training scrapes. Make sure your Cloudflare settings match this intent after September 15.
- Check your robots.txt for AI crawlers. Cloudflare is not the only place crawler access is controlled. If your robots.txt disallows GPTBot, PerplexityBot, ClaudeBot, or Google-Extended, those bots will not crawl your site regardless of Cloudflare settings. Verify that named AI crawlers are explicitly allowed — or at minimum not blocked — in your robots.txt.
The Bigger Picture
The Cloudflare change reflects a broader shift underway across the web: the era of treating all bots the same is ending. Search bots, training scrapers, and AI agents are meaningfully different in what they do with your content and what they give back to your business. Having a clear, intentional policy for each type — rather than a single all-or-nothing toggle — is now a real operational requirement, not just a technical detail for large publishers.
For small businesses in New York City's competitive local markets, the practical checklist is short: make sure Google can still find you, make sure AI assistants can cite you, and make sure you are not inadvertently doing both at once in a way that blocks one or the other.
Sources
- Your site, your rules: new AI traffic options for all customers — Cloudflare Blog, July 1 2026
- Cloudflare's AI Crawler Rules Can Block Googlebot — Search Engine Journal
- New options to manage AI traffic — Cloudflare Changelog, July 2026
- Cloudflare's new policy pushes AI companies to pay for publishers' content — TechCrunch
