Summarize This Article With AI
For most of the web’s history, a bot was either welcome or it was not.
You allowed it or you blocked it. There was no in between.
Cloudflare just tore up that binary. The company now lets site owners sort AI crawlers into three distinct categories based on what they actually do, Search, Agent, and Training, and manage each one separately. Even sites on the free tier get the controls. And buried in the announcement is a default change that could quietly block Googlebot on a lot of websites.
This is one of those updates that sounds like plumbing and turns out to be strategy.
Article Summary
- Cloudflare now sorts AI crawlers into three categories: Search, Agent, and Training.
- Site owners can allow, charge, or block each category independently, even on the free plan.
- From September 15, 2026, Training and Agent bots will be blocked by default on ad-supported pages.
- Because Googlebot is a mixed-use crawler, blocking Training can unintentionally block it too.
- Cloudflare data shows AI training now drives the majority of crawler requests on its network.
- The move pressures AI companies to separate their crawlers by purpose.
- Search Everywhere Optimization depends on understanding which crawlers you actually want reaching your content.
What Did Cloudflare Actually Launch?
The core of the update is a shift from asking whether a bot is “AI” to asking what the bot is doing on your site.
Cloudflare now groups automated traffic into three behaviors you can control independently.
Search covers crawlers that index your content so they can answer questions about it later, the kind of activity you expect to earn referral traffic in return. Agent covers automated activity acting in real time on a person’s behalf, such as chat fetch bots like ChatGPT-User or browser-use agents like Gemini or Claude driving a browser. Training covers crawlers that take your content to train or fine-tune a model, where your data is permanently absorbed into the AI itself.
For each category, you can choose to allow it, block it, or block it only on pages that display ads.
That granularity is the whole point. Previously, protecting your content from AI training often meant reaching for a blunt “block all bots” instrument that also cut off the crawlers driving your search visibility. Now you can welcome the crawlers that send you traffic while blocking the ones that simply take, and you can do it without touching your search presence.
Crucially, Cloudflare made these controls available to everyone, including free-tier customers.
The September 15 Default That Could Block Googlebot
Here is the part that deserves your full attention.
Starting September 15, 2026, Cloudflare is changing the default settings. For new domains, new sites added by existing customers, and all existing free-tier users who have not adjusted their settings, Training and Agent crawlers will be blocked by default on pages that display ads. Search crawlers stay allowed.
On its own, that sounds reasonable. The complication is what it does to mixed-use crawlers.
Googlebot is not a single-purpose bot. It crawls for both search indexing and AI training in one combined crawler, and so do Applebot and BingBot. Cloudflare applies the most restrictive rule to any multi-purpose crawler, which means a site that blocks Training to protect its content from AI models could unintentionally block Googlebot along with it.
That is a genuine risk worth flagging to anyone running a Cloudflare-protected site.
A network-level block is not the same as a polite line in robots.txt that Google can choose to ignore. If Googlebot gets caught by a Training block on your ad-supported pages, that is your search indexing on the line. The good news is that the controls are entirely in your hands. Site owners can review or change these settings in the Cloudflare dashboard any time before the deadline, and Cloudflare says it will keep notifying customers ahead of the date.
If you use Cloudflare, this is your reminder to check those settings well before September.
Why Cloudflare Is Doing This
The motivation becomes obvious once you look at the traffic data.
Cloudflare says AI training now accounts for the majority of crawler requests on its network, a sharp rise from roughly 20 percent in spring 2025. Daily AI agent requests, meanwhile, increased by more than 1,700 percent over the year. The old symbiotic deal, where crawlers took your content and sent you visitors in return, has broken down badly.
Cloudflare’s earlier research put numbers on that imbalance.
The company measured crawl-to-referral ratios ranging from 118 to 1 all the way up to nearly 50,000 to 1. In the worst cases, an AI crawler scraped a site tens of thousands of times and sent back a single human visitor. That is not a relationship. That is extraction with extra steps.
The new categories are Cloudflare’s attempt to rebalance that.
By pushing AI companies to separate their crawlers by purpose, Cloudflare is trying to force transparency into a system that has become deliberately opaque. If a company runs one crawler that indexes for search, acts as an agent, and harvests training data all at once, website owners have no way to allow the useful behavior while refusing the rest. Splitting those functions apart gives the choice back to the publisher.
Cloudflare is also opening a path to payment, with a Monetization Gateway that will let publishers charge automated systems directly for access. That is a longer story, but it points in the same direction. Content owners want to be compensated when their work creates value.
Why This Matters for Search Everywhere Optimization
This is where the plumbing becomes strategy.
Search Everywhere Optimization is about being visible, credible, and correctly represented wherever people discover and validate your brand. More and more, that discovery is mediated by AI systems reading your content long before a human ever sees a result. Which crawlers you allow, and which you block, now directly shapes where you can show up.
That makes crawler management a visibility decision, not just a security one.
Block too broadly, and you risk shutting yourself out of the AI platforms and search experiences you actually want to appear in. Allow everything, and you hand your content to training crawlers that absorb your work and send nothing back. The sweet spot sits in the middle, and Cloudflare’s three categories finally give you the vocabulary to find it.
Think about what each category means for your visibility.
Search crawlers are usually the ones you want to welcome, because they feed the indexes that surface your brand and send referral traffic your way. Agent crawlers are the emerging layer, the bots fetching your content to answer a question for a real person in real time, and blocking them may mean vanishing from AI-assisted journeys your customers are already taking. Training crawlers are the ones most publishers feel comfortable restricting, because they offer the least direct return. The right mix depends on your business, but you cannot make that call intelligently without seeing the categories clearly.
This is the same principle that runs through all of Search Everywhere Optimization. You cannot optimize what you cannot see, and you cannot protect what you cannot control. Cloudflare’s update adds a valuable layer of visibility over the AI crawlers shaping your brand’s presence, sitting alongside the technical SEO fundamentals that keep your content accessible to the right systems and closed to the wrong ones.
The brands that win the next phase of search will be the ones treating crawler access as a deliberate strategy rather than a default they never reviewed.
What Should Businesses Do Next?
If you use Cloudflare, start with the calendar.
Mark September 15 and review your AI crawler settings well before then. Decide, deliberately, which of the three categories you want to allow, charge, or block. Pay particular attention to the mixed-use crawler issue, because if you block Training on ad-supported pages without adjusting your settings, you may catch Googlebot in the same net and damage your search visibility.
Then think about your content strategy alongside your crawler strategy.
Which AI platforms do you actually want reaching your content? If AI-driven discovery matters to your brand, blocking every agent and search crawler is self-defeating. If protecting proprietary content from model training matters more, the new controls let you draw that line precisely. Most businesses will want a blend, and the point is to choose it on purpose.
Finally, treat this as part of a wider habit rather than a one-time task.
The crawler landscape is shifting fast, with new AI operators appearing constantly and existing ones changing behavior. Review your settings regularly, cross-reference them against your analytics and server logs, and keep asking the same question. Are the systems shaping my brand’s visibility the ones I have actually chosen to let in?
Robots.txt was always a polite request. Cloudflare just handed publishers something closer to a real decision.
Ready to Build Visibility Everywhere People Search?
AI crawlers are reshaping how your brand is discovered, and the businesses that win are the ones who understand exactly which systems are reaching their content and why. Visibility across Google, AI Overviews, AI Mode, and beyond starts with knowing who is reading your pages and choosing who you let in.
At SEO Sherpa, we help brands build the technical foundations, content authority, and cross-platform visibility needed to be found, trusted, and chosen across the modern search journey.
Book a free discovery call with our team today to find out how Search Everywhere Optimization can help your business increase its visibility, strengthen its authority, and turn more searches into customers.




















Leave a Reply