Does Cloudflare block Googlebot? How the Block AI bots rules changed, and how to check yours

The short answer: the old tick box does not take you out of Google, but Block does

"Nothing, in almost every case. Your current settings carry over on their own."

Cloudflare, on what changes for existing sites

If your site runs through Cloudflare and you once ticked "Block AI bots", Google's search crawler can still reach you. The warning that the old tick box would catch Googlebot no longer describes what happens. For existing sites, the answer is "Nothing, in almost every case. Your current settings carry over on their own."

The setting that does remove Google is the one now simply called Block. Cloudflare says Block "will stop Applebot, Bingbot, and Googlebot from reaching your site", and that includes search. If you want to stop AI training but stay in search, the setting to choose is Disallow AI Training, not Block.

None of this touches you unless your site actually uses Cloudflare. If you are not sure whether it does, ask your web developer or host before you change anything.

What was forecast in July, and what actually happened

  1. Said

    1. Cloudflare announces that from 15 September, multi-purpose crawlers (Googlebot, Applebot, Bingbot) will be blocked for anyone blocking Training, including through the legacy "Block AI bots" service.

      Superseded Cloudflare blog
  2. Done

    1. Cloudflare publishes a new post. The plan is superseded and current settings carry over.

      Cloudflare blog
    2. Search Engine Journal confirms: "This differs from the plan I mentioned back in July."

      Search Engine Journal
  3. Planned, not shipped

    1. "In the weeks to come"

      Google URL-level transparency tools for Google-Extended.

      Cloudflare blog
    2. "By early next year"

      A single Cloudflare control over AI-summary inclusion.

      Cloudflare blog
    3. "Targeted for early 2027"

      Microsoft support for the robots.txt no-training preference.

      Cloudflare blog
Source: Cloudflare blog, Search Engine Journal

In July, Cloudflare announced that from 15 September, multi-purpose crawlers such as Googlebot, Applebot and Bingbot would be blocked for anyone blocking AI training, including anyone using the legacy "Block AI bots" service. For a small business that had ticked the box to keep AI scrapers out, that would have meant losing Google too.

That plan did not go ahead. On 15 September Cloudflare published a new post, and its earlier Bot Preference Sync post now sends readers there for "how the Block and Disallow AI Training settings now work together". Search Engine Journal's coverage that day said it plainly: "This differs from the plan I mentioned back in July."

So if a July article, email or forum thread is still nagging at you, set it aside. What follows is the position since 15 September.

What your old setting turned into

If your site was...Search becomesTraining becomesAgent becomes
Never on the detailed controls, old "Block AI bots" off Allow Allow Allow
Never on the detailed controls, old setting on Block or Block on pages with ads Allow Disallow AI Training Block on pages with ads
Already on the detailed controls, with Training on Block or Block on pages with ads Unchanged Disallow AI Training Unchanged
New domain from 15 September, no ads Allow Allow Allow
New domain from 15 September, funded by ads Allow Disallow AI Training Block on pages with ads
"Block AI Bots" and Managed Robots.txt are deprecated. Bot Preference Sync replaces Managed Robots.txt.
Source: Cloudflare blog

Cloudflare now splits its AI crawler controls into three settings: Search, Training and Agent. Where you land depends on what you had before.

New domains added from 15 September get presets. A site that does not run ads gets Allow for all three. An ad-funded site gets the same mix as the second case above.

The old tick box itself is on its way out. "Block AI Bots" and Managed Robots.txt are deprecated, and Managed Robots.txt is replaced by Bot Preference Sync, which writes your choices into your robots.txt file.

The four Training settings, and which one shuts out Google

AllowCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotTraining-only crawlersAmazonAnthropicMetaOpenAIAll pass DisallowAI TrainingCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotrobots.txt:no trainingTraining-only crawlersAmazonAnthropicMetaOpenAISearch stays open Block onpages with adsCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotAdpages:stoppedTraining-only crawlersAmazonAnthropicMetaOpenAISearch stops on ad pages BlockCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotTraining-only crawlersAmazonAnthropicMetaOpenAISearch stops too AllowCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotTraining-only crawlersAmazonAnthropicMetaOpenAIAll pass DisallowAI TrainingCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotrobots.txt:no trainingTraining-only crawlersAmazonAnthropicMetaOpenAISearch stays open Block onpages with adsCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotAdpages:stoppedTraining-only crawlersAmazonAnthropicMetaOpenAISearch stops on ad pages BlockCLOUDFLAREOutsideOn your siteMixed-use crawlersGooglebotBingbotApplebotTraining-only crawlersAmazonAnthropicMetaOpenAISearch stops too
Source: Cloudflare blog

The Training setting now has four options: Allow; Disallow AI Training; Block on pages with ads; Block. Only two of them matter for staying in Google.

Disallow AI Training is the middle path. It publishes a no-training preference in robots.txt. "Accountable" mixed-use crawlers stay allowed for search. Every other training crawler is blocked, including training-only crawlers from Amazon, Anthropic, Meta and OpenAI. In Search Engine Journal's words, Cloudflare lets sites disallow AI training without blocking Googlebot.

Block and Block on pages with ads are different. Both now apply to mixed-use crawlers, the ones that crawl for search and for AI at the same time. Block keeps Googlebot, Bingbot and Applebot off your site altogether.

Training settingGooglebot (Google Search)Other mixed-use crawlers (Bingbot, Applebot)Training-only crawlers (Amazon, Anthropic, Meta, OpenAI)Good fit for
Allow Allowed Allowed Allowed Happy for AI training
Disallow AI Training Stays in for search Accountable mixed-use crawlers stay in for search. Bing gets no robots.txt no-training signal yet Blocked Stop AI training, stay in search
Block on pages with ads Applies to mixed-use crawlers on pages with ads Applies on pages with ads Blocked Ad-funded publishers who accept the search trade-off
BlockLeaves search Stopped, search included Stopped, search included Blocked Only if leaving search is intended
Source: Cloudflare blog, Search Engine Journal

For most small businesses the choice is between Allow and Disallow AI Training. Block makes sense only if leaving search is what you actually want.

How rare blocking search is on Cloudflare's network

Cloudflare's network only, not the whole web

Share of sites on Cloudflare

Block Search bots under 1% Restrict AI training 17% 0% 50% 100%

Cloudflare press release

Mixed-use crawlers

36.6% of verified crawler traffic

Cloudflare press release

Source: Cloudflare

On its own network, Cloudflare says fewer than 1% of sites block Search bots, while 17% restrict AI training. Mixed-use crawlers make up 36.6% of verified crawler traffic it sees. These figures describe sites on Cloudflare, not the whole web.

Our reading: the gap between those two shares is the whole story. Far more owners wanted to keep AI training off their work than wanted to leave search, and the July plan would have tied those two wishes together.

Stopping AI training and staying in Google are separate decisions

  • Googlebot (robots.txt token)

    Controls Google Search, including Discover and all Search features.

    Google Search Central

  • Google-Extended (robots.txt token)

    Controls use for Gemini training and grounding. Google: it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal".

    Google Search Central

  • Search Console AI Overviews and AI Mode control

    Separate from Google-Extended, and does not affect training, per Search Engine Journal's reading.

    Search Engine Journal

  • Bing

    Cloudflare points to Bing's own controls: the NOARCHIVE meta tag, or the Block URLs and Content Removal tools in Bing Webmaster Tools. robots.txt support is "targeted for early 2027".

    Cloudflare blog

Source: Google Search Central, Search Engine Journal, Cloudflare blog

Google keeps these apart in its own crawler rules. Googlebot is the token that affects Google Search, including Discover and all Search features. A separate token, Google-Extended, manages whether your content may be used to train Gemini models and for grounding. Google says Google-Extended "does not impact a site's inclusion in Google Search nor is it used as a ranking signal".

There is a third switch, and it is easy to confuse with the other two. According to Search Engine Journal's reading of Google's help page, Google-Extended is separate from the Search Console control for AI Overviews and AI Mode, and that control does not affect training.

At Search Engine Hub, our advice is to make these as two deliberate choices. First: do you want to be found in Google? For almost every business, yes. Second: are you comfortable with your content training AI models? That one is yours to weigh, and saying no to training does not require saying no to search.

Bing works differently, for now

Disallow AI Training does not yet send Bing a no-training preference through robots.txt. Microsoft's support is "targeted for early 2027". Until then, Cloudflare points to Bing's own controls: the NOARCHIVE meta tag, or the Block URLs and Content Removal tools in Bing Webmaster Tools. Microsoft says NOARCHIVE does not affect search ranking.

Two more changes are promised but are plans, not features you can switch on. Cloudflare says Google intends to add URL-level transparency tools for Google-Extended "in the weeks to come", and Cloudflare is aiming for a single control over AI-summary inclusion "by early next year".

A 10-minute check if your site uses Cloudflare

This check tells you whether Google can reach your pages today. It tells you nothing about where you will rank, and that is fine: reachability comes first.

  1. Confirm you are on Cloudflare. If you do not know, ask your web developer or host. If you are not, the rest of this section does not apply.
  2. Read your Cloudflare AI crawler settings. Look at what Search and Training say now. If you want to stay in Google, Search should read Allow, and Training should not read Block.
  3. Open your robots.txt file. Type your address followed by /robots.txt into a browser. Bot Preference Sync keeps your existing robots.txt content and adds Cloudflare's text above it, so expect new lines at the top and your own lines underneath.
  4. Check Crawl Stats in Search Console. The Crawl Stats report shows server responses and host status. Google says it is aimed at advanced users and available only for root-level properties, so you may need to open the domain-level property rather than a single folder.
  5. Run a URL Inspection live test. On your home page and one page that brings in work, the Live Test shows "Crawl allowed?" and "Page fetch" for the current version of the page.
A small-business back office in an Australian home or shop: a laptop open on a timber desk beside a coffee mug and a mobile phone, an Australian three-pin power point on the wall behind, late afternoon light through a window.
Illustration: a laptop open on a desk in a small home office.

If you want to go further in the same sitting, our 20-minute website health check covers the rest of the basics.

Why your robots.txt file has to load cleanly

Counts as successful200403404410 Counts as unsuccessful4295xx 0 to 12 hoursGoogle stops crawling After that, up to 30 daysGoogle uses the last goodcopy of robots.txt Counts as successful200403404410 Counts as unsuccessful4295xx 0 to 12 hoursGoogle stops crawling After that, up to 30 daysGoogle uses the last good copy of robots.txt
Source: Google Search Console Help

Step 3 matters beyond Cloudflare. Google treats robots.txt as successful only if it returns 200, 403, 404 or 410. A 429 or any 5xx error counts as unsuccessful. When that happens, Google stops crawling for the first 12 hours, then uses the last good copy of the file for up to 30 days.

In practice: if your robots.txt loads normally in a browser and Crawl Stats shows no host problems, this part is in order. If it throws an error, raise it with whoever looks after your site.

What it comes down to

  • Block stops Applebot, Bingbot and Googlebot, search included.

    Cloudflare blog

  • Disallow AI Training: accountable mixed-use crawlers stay in for search; training-only crawlers blocked.

    Cloudflare blog

  • Google-Extended does not impact inclusion in Google Search.

    Google Search Central

  • robots.txt: 200, 403, 404, 410 count as successful; 429 or 5xx means 12 hours with no crawling, then up to 30 days on the last good copy.

    Google Search Console Help

Source: Cloudflare blog, Google Search Central, Google Search Console Help

The old tick box no longer takes you out of Google. Block does. Disallow AI Training is the setting for stopping AI training while staying in search, and Google-Extended is Google's own separate switch for training that leaves search alone. Bing still needs its own setting, usually NOARCHIVE.

If your traffic looks odd and you want to rule things out in order, our guide on how to tell whether your SEO is actually working is a good next read.

Your next step

Spend ten minutes on the check above this week. If something in your Cloudflare settings, robots.txt or Search Console does not look the way it should, or you simply want someone to look at it with you, send Search Engine Hub your question through our contact page and we will give you a straight answer on your site.

Sources