Cloudflare flipped a switch in 2026. Every new domain that connects to its network now blocks AI crawlers by default. GPTBot, ClaudeBot, PerplexityBot, and a growing roster of lesser known bots hit a wall the moment a site goes live, unless the owner turns them back on. If your site already sits behind Cloudflare, the company is not flipping that switch retroactively. Instead, it is asking every account holder to set an explicit policy, crawler by crawler, before September 15, 2026. Miss that date, and Cloudflar’ own fallback rules make the decision for you.

I bring this up because it is the first question on nearly every client call this month. Did Cloudflare just cut us off from ChatGPT? The honest answer is maybe, and you will not know until you check. Cloudflare sits in front of an enormous share of the sites that AI answer engines crawl, train on, and cite. A default change at that scale reshuffles who shows up in ChatGPT, Perplexity, and Google AI Overviews answers, and who quietly disappears from all three at once.

Is Your Site Even on Cloudflare

Before any of this applies to you, confirm your site actually runs through Cloudflare. A large share of the web sits behind Cloudflar’ network without the site owner ever logging into a Cloudflare account directly, because a web host, a platform like Shopify or Webflow, or an agency set it up during the original build and never mentioned it again.

The fastest check takes under a minute. Look up your domai’ nameservers through any public WHOIS lookup tool. If the nameservers end in cloudflare.com, your site sits behind Cloudflar’ network, whether or not you personally hold the login. A second method: view your sit’ HTTP response headers in a browse’ developer tools. A header labeled cf-ray or cf-cache-status confirms Cloudflare is sitting in front of your server.

If your site is not on Cloudflare, this specific deadline does not apply to you, though the underlying question still does. Whatever CDN or hosting platform you use likely has its own bot management layer, and the same conflict between an edge level rule and your robots.txt can exist there too. If your site is on Cloudflare and you have never logged into the dashboard yourself, this is the week to ask whoever manages your hosting for access, or to have them run through the checklist below on your behalf.

What Cloudflare Changed

The control panel for this decision is called AI Crawl Control. It replaced Cloudflar’ earlier AI Audit tool and gave every domain on the network a dedicated switch for each AI crawler category: training crawlers like GPTBot and ClaudeBot, retrieval crawlers like PerplexityBot, and a fast growing list of agent and assistant bots that fetch pages on behalf of a live chat session happening somewhere else on the internet.

For any domain added to Cloudflare after the rollout, every crawler category defaults to blocked. The owner has to open AI Crawl Control and turn crawlers back on, one category at a time, or leave the block in place. For domains that were already on Cloudflare before the change, the default did not flip automatically. Cloudflare is instead prompting those accounts, through dashboard banners and account notifications, to make an explicit choice by September 15, 2026.

After that date, any account that never set a policy inherits whatever Cloudflare applies as its global fallback, not a choice the site owner actually made. That is the part worth sitting with. A decision this consequential for your AI visibility should not be settled by a platform default you never reviewed.

Cloudflar’ AI crawler policy is now a setting you own, not a background assumption. Every domain on the network needs an explicit answer, allow, block, or charge, for GPTBot, ClaudeBot, PerplexityBot, and every other AI crawler that matters to your category.

Why Cloudflare Made AI Crawlers Opt In

Cloudflare framed the change as correcting an imbalance. For years, AI crawlers were treated the same as search engine crawlers. Site owners who wanted their pages found in Google left the gate open, and AI companies walked through the same open gate. As a handful of AI crawlers began accounting for a larger share of total crawl traffic on some sites, while sending back comparatively little referral traffic in return, publishers started asking for a different arrangement. Some blocked AI crawlers outright. Some pursued direct licensing deals with AI companies. Cloudflar’ answer was to build the negotiation into the infrastructure layer itself, rather than leaving each publisher to fight the same battle alone.

Instead of one blunt setting, allowed or blocked, Cloudflare gave every site owner three explicit choices for each AI crawler it can identify. Allow the crawler, the way you always could. Block the crawler outright. Or charge the crawler a fee for every successful fetch, through a feature Cloudflare calls pay per crawl.

The timing is not a coincidence. Cloudflare has spent the past several product cycles building infrastructure around bot identification, from its existing bot management tools to signed, cryptographically verifiable crawler agents. AI Crawl Control is the customer facing surface of that infrastructure. Once a network can reliably tell a real GPTBot request from a spoofed one, offering a paid access tier becomes possible with real confidence behind it. Pay per crawl could not have existed as a credible product before that verification layer was in place.

The Three Settings: Allow, Block, Charge

Each setting produces a different outcome for your AI visibility, and the right one depends entirely on what your content is for.

Setting What It Does Best Fit
Allow The crawler fetches your pages the way it would have before the default change. Your content stays eligible for AI training data and real time citation. Marketing sites, services businesses, and any brand trying to earn AI citations.
Block Cloudflare denies the request at the edge before it reaches your server. The crawler never sees the page. Publishers whose content is the product, or any site with a deliberate licensing position.
Charge Cloudflare returns a payment required response. Crawlers enrolled in the pay per crawl marketplace pay the price and proceed. Others are turned away. Publishers with expensive to produce content who want compensation without a full block.

Allow

Allow is the setting most business sites should choose, and it works exactly as the name implies. The crawler fetches the page, reads the content, and treats it the same as it would have before Cloudflare changed the default. Your content stays eligible for training data and for the kind of real time retrieval that gets a page cited in a Perplexity or ChatGPT answer the same week it is crawled.

One caution: setting Allow in AI Crawl Control is necessary but not sufficient. Your robots.txt file still needs to agree with the setting, and so does any bot management rule sitting further down your configuration. A mismatch between layers is one of the most common ways sites end up invisible to AI crawlers without anyone noticing. More on that below.

Block

Block does what it says. Cloudflare intercepts the request at the edge, before it ever reaches your origin server, and the crawler receives nothing. No training data, no retrieval, no citation eligibility. For a site that wants to protect content with standalone commercial value, a paid research report or a subscription news archive, this is a defensible position. For a site whose content exists to generate inquiries, Block quietly removes the brand from every AI answer that content could have earned.

Charge: Cloudflar’ Pay Per Crawl

Charge sits between the other two. Instead of an open door or a locked one, the site owner sets a price for access. When a participating crawler requests a page under this setting, Cloudflar’ edge returns an HTTP 402 Payment Required response along with the price. A crawler enrolled in Cloudflar’ pay per crawl marketplace pays that price automatically and receives the page. A crawler that has not enrolled, or that declines the price, is turned away and treated under the sit’ fallback rule instead.

Search Crawlers vs Training Crawlers: Why the Category Matters

AI Crawl Control does not force one setting across every AI bot. It groups crawlers into categories, and a site owner can set a different policy for each one. This distinction matters more than it looks like it should.

Training crawlers, like GPTBot and ClaudeBot, gather content that shapes what a model knows the next time it is updated. The effect shows up on a delay. Block a training crawler today, and the consequence surfaces months from now, when a model update rolls out without your brand in its training set. Retrieval crawlers, like PerplexityBot and OAI-SearchBot, fetch live content to answer a specific question right now. Block a retrieval crawler and the consequence is immediate. Your page cannot be cited in tomorro’ answer, even if it was cited yesterday.

Some site owners land on a split policy: allow retrieval crawlers because the payoff is immediate and visible, while blocking or charging training crawlers because that payoff is harder to see and further away. This is a reasonable position to hold on purpose. It is a much weaker position to hold by accident, because a default handled one category differently than you assumed it would.

How Pay Per Crawl Actually Works

Pay per crawl only functions for crawlers Cloudflare can verify. Any bot can claim to be GPTBot by setting its user agent string to say so. Cloudflare does not take that claim at face value. It verifies crawler identity cryptographically before applying whichever policy the site owner picked, which means the price, the block, or the allow only ever reaches the crawler it was actually meant for. A spoofed bot claiming to be GPTBot gets treated under your default rule for unverified traffic, not under the specific policy you set for the real thing.

That verification step matters more than it sounds like it should. Our field guide to GPTBot, ClaudeBot, and PerplexityBot covers a related problem: user agent strings are case sensitive and easy to fake, and a site that relies purely on string matching in robots.txt can be fooled or can accidentally block the crawler it meant to allow. Cloudflar’ cryptographic check closes that gap at the edge, at least for the crawlers it has agreements with.

For most AEO Hunt clients, pay per crawl is not the setting to reach for. It is built for publishers whose content carries a price on its own, independent of any inquiry it might generate. A services business, a local provider, or a SaaS company selling a product that is not the content itself gets more value from citation than from a per fetch fee that most AI crawlers will simply decline by routing around the priced page.

The marketplace piece is still young. Most AI companies have not broadly enrolled every one of their crawlers to pay for access on a per fetch basis, which means a Charge setting configured today often behaves closer to a soft block for whichever crawlers have not signed up yet. That will likely shift as more AI companies opt in, but a Charge setting should be judged on how it behaves right now, not on a hypothetical steady stream of future micropayments.

Why This Overrides Your Robots.txt

Robots.txt is a request. A well behaved crawler reads it and honors what it says. Cloudflar’ AI Crawl Control setting is enforcement, applied at the network edge before the request ever reaches your server or your robots.txt file. If the two disagree, the edge wins every time.

This is the same conflict our crawler field guide flags as a CDN configuration problem: a bot management layer can drop a crawle’ request before robots.txt is ever checked, which means your carefully written allow rules accomplish nothing if the CDN sitting in front of your server has already decided otherwise. Cloudflar’ new default just turned that occasional conflict into the norm instead of the exception. If you set Allow in your robots.txt for GPTBot but never touched AI Crawl Control after your domain onboarded, the domain default of Block is the setting that actually governs what happens.

Robots.txt is a request. Cloudflar’ edge is enforcement. When the two disagree, the edge wins, and the crawler never learns what your robots.txt said.

The Decision Framework for AEO

The right setting comes down to one question: is your content the product being sold, or is it the vehicle that earns attention and inquiries for something else you sell? Most businesses working with an AEO agency fall firmly into the second category, and the framework below reflects that split.

A marketing site, a local service business, or a SaaS company with a content led blog should default to Allow across every AI crawler that matters to its category. The content exists to be found and cited. Blocking the crawlers that make citation possible works against the entire point of publishing it.

A news publisher, a paid research firm, or any site where the content itself is the thing customers pay for should weigh Block or Charge deliberately, understanding exactly what citation visibility they are trading away in exchange for protecting that value. Neither choice is wrong for that kind of business. What is wrong is making the choice by accident, because a default flipped and nobody checked.

Business Type Recommended Default Why
Services business or local provider Allow Content exists to generate inquiries. Citation is the goal, not a risk.
SaaS company with a content led blog Allow Growth depends on being named when buyers ask AI engines for recommendations.
News publisher or paid research firm Block or Charge Content has standalone commercial value. Free training access can undercut subscription revenue.
Marketplace or community platform with user generated content Allow public pages, block gated ones Mixed content value needs a mixed policy, not one setting for the whole domain.

Notice that the fourth row is not a single setting at all. A platform with both public listings and gated member content should not apply one blanket policy to the whole domain. AI Crawl Control allows enough granularity to keep the pages that earn citations open while protecting the ones that should not be handed to a training pipeline for free. Treating the decision as domain wide when your content is not uniform is its own kind of mistake.

What to Check in Your Cloudflare Dashboard Before September 15

Here is the full process, and it takes less time than most site owners expect once they know where to look.

  1. Open AI Crawl Control for every domain you manage. Multi domain accounts need a policy set per domain, not a single account wide setting.
  2. Review the current per crawler policy. Check the existing setting for training crawlers, retrieval crawlers, and agent or assistant bots. Note which categories already have an explicit setting and which are sitting on a default you never touched.
  3. Cross check against your robots.txt. Compare the Cloudflare policy against your robots.txt rules for the same crawlers. Since Cloudflar’ edge decision overrides robots.txt, any contradiction between the two means your robots.txt is not doing what you think it is doing.
  4. Decide allow, block, or charge for each category. Use the product versus vehicle question above. Most marketing and services sites should land on allow.
  5. Set explicit rules instead of leaving categories on default. Save a deliberate setting for every crawler category relevant to your business, rather than letting any of them ride on Cloudflar’ fallback.
  6. Confirm the change in your logs. After saving, check Cloudflar’ bot analytics dashboard and your server access logs. Confirm the crawlers you allowed are actually reaching your pages and the ones you blocked are not.

Common Mistakes in Cloudflar’ AI Crawl Control

Reviewing crawler settings across client accounts this year turned up the same handful of errors, over and over.

Setting the account, not the domain. A Cloudflare account managing several properties needs a policy set on each domain individually. Configuring AI Crawl Control on one site does nothing for the others sitting on the same account.

Forgetting subdomains. A blog on a subdomain, a docs site, or a careers page hosted separately can carry its own zone and its own crawler policy, distinct from the main domain. Reviewing the root domain and assuming every subdomain inherited the same setting is a common gap.

Contradicting the robots.txt on file. As covered above, Cloudflar’ edge setting overrides robots.txt. Teams that update one and forget the other end up with a robots.txt that reads like an open door while Cloudflare quietly keeps it shut.

Confusing a staging zone with production. Some accounts run a staging or development zone through Cloudflare with a locked down default that made sense during a build phase. If that zone was later promoted to production without a fresh review, AI crawlers may be blocked from the live site because of a setting meant for a version nobody was supposed to see.

Assuming an agency or developer already checked. Plenty of site owners assume whoever built or maintains their site reviewed this setting as part of routine upkeep. Most did not, because the default change is recent enough that it has not made it into most standard maintenance checklists yet. Ask directly instead of assuming.

What Happens If You Do Nothing

Doing nothing before September 15 does not necessarily mean your site goes dark to every AI crawler overnight. What it means is that Cloudflar’ own global fallback policy takes over for any crawler category you never explicitly set, and that fallback was not written with your specific content, your specific category, or your specific citation goals in mind. It was written to be a reasonable default for the entire Cloudflare network at once.

This mirrors a pattern that shows up constantly in AEO audits: a technical blocker sitting quietly in a configuration nobody has opened in months, invisible to anyone who is not specifically looking for it. A robots.txt disallow rule left over from a redesign. A CDN bot rule inherited from a security review. Cloudflar’ new default is the newest version of that same pattern, just applied to every domain on the network at once instead of one site at a time.

For a business that depends on AI citation to reach buyers who ask ChatGPT or Perplexity before they ask anyone else, letting a platform wide default make that call is a real cost, and one that is easy to avoid. The review takes minutes. The consequence of skipping it compounds every month a crawler stays blocked without anyone noticing.

Where llms.txt Fits Into This Decision

Setting Allow across your AI crawlers opens the gate. It does not tell the crawler what is behind it. That is the job of an llms.txt file, a plain text file at your domain root that gives AI systems a curated map of your organization, your key pages, and the topics you cover.

The two decisions pair naturally. A site that sets Allow in AI Crawl Control but has no llms.txt is handing crawlers an open door into a building with no directory. A site that publishes a clear llms.txt but blocks the crawlers that would read it has built a directory nobody can enter. Get both right, and a crawler that reaches your domain knows immediately what it is looking at and where the content that matters actually lives.

How This Connects to How AI Engines Choose Sources

None of this matters if the crawler never reaches your content in the first place. Our breakdown of how ChatGPT, Perplexity, and Google AI Overviews choose sources covers the mechanics each engine uses to decide which pages earn a citation: training data quality, retrieval freshness, entity recognition, and formatting that makes a page easy to extract from. Every one of those mechanics assumes the crawler was able to reach the page at all.

A Cloudflare default set to Block skips the entire evaluation. It does not matter how well structured your FAQ sections are, how clean your schema markup is, or how directly your first paragraph answers the query someone is likely to ask. If GPTBot never gets past Cloudflar’ edge, none of that work has a chance to be judged, because the source selection process this article describes never starts.

The AEO Hunt Recommendation

For the large majority of businesses working with an AEO agency, the recommendation is straightforward: set Allow across GPTBot, OAI-SearchBot, ClaudeBot, anthropic-ai, PerplexityBot, and Google-Extended in AI Crawl Control, then confirm robots.txt agrees with that setting. Your content exists to generate inquiries. AI citation is one of the channels that generates them now, alongside organic search and paid media, and closing that channel by accident because a dashboard default changed underneath you is not a strategy anyone chose on purpose.

Publishers whose content carries independent commercial value have a real decision to make between Block and Charge, and that decision deserves its own analysis of licensing risk against citation value. But that is a narrow slice of the businesses reading this. For everyone else, the deadline is simple: go check the setting before September 15, confirm it matches what your robots.txt already says, and stop leaving your AI visibility to a platform default you never reviewed.

If you want a second set of eyes on your Cloudflare configuration alongside your robots.txt, schema, and entity signals, a full AI Visibility and AEO audit checks all of it in one pass and tells you exactly where the gate is closed.