AI Crawler Control is Becoming Part of SEO: What Cloudflare Changed and How Website Owners Can Protect Their Content
Artificial intelligence is reshaping not only search results, content creation, and e-commerce operations. It is gradually altering the very architecture of the open web—specifically the rules by which automated systems access websites.
Previously, the relationship between website owners and search engines was quite straightforward: a search crawler scanned pages, added them to the index, and then the search engine directed visitors back to the site.
The website provided content; the search engine provided traffic.
Today, this model is no longer universally applicable. AI services can read a page, use its information to generate a direct answer, and bring virtually no users back to the original source.
The editors at Sitezilla reviewed Cloudflare's article "Your site, your rules: new AI traffic options for all customers", published on July 1, 2026. In it, the company proposes moving away from the simplistic binary of "AI bot or not" toward more granular control based on exactly what an automated system does with website content.
We break down what exactly has changed, why managing AI crawlers is steadily becoming a core component of SEO, and what content access policies online stores, corporate sites, and content projects should implement.
The Old "Content for Traffic" Exchange No Longer Works as It Used To
For decades, website owners essentially agreed to an informal exchange:
the search crawler gets access to pages;
the search engine indexes the information;
the user sees the page in search results;
the website gets a referral, an ad impression, a lead, or a sale.
Classic SEO, content marketing, informational portals, affiliate websites, and a major portion of e-commerce are built entirely on this model.
Cloudflare believes that the rise of generative search has disrupted this balance. AI crawlers can download massive volumes of content, while referral traffic from these AI services remains disproportionately low.
According to Cloudflare's observations shared during their inaugural Content Independence Day, the ratio of crawls to referrals for major AI operators ranged from roughly 118 crawls per single referral to nearly 50,000 crawls per referral. This means a server might serve content tens of thousands of times to an automated system before the website receives even a single visitor.
For website owners, this creates several immediate problems:
The AI bot consumes server resources and bandwidth;
The content can be used to generate a direct response on a third-party service;
The user gets the answer without ever opening the original source;
The website misses out on page views, ad impressions, leads, or sales;
It is difficult for the author to determine exactly who used their materials and for what purpose.
In other words, the issue is not just about using text for training models. The real problem is the lack of a transparent exchange of value.
Not All AI Bots Perform the Same Job
One of the core premises of Cloudflare's new approach is moving away from labeling all automated AI traffic with a single term.
The term "AI bot" has become far too broad.
One bot might collect information for a search index. Another might load a page directly at a user's request. A third might build datasets to train the next version of a large language model.
The consequences for website owners are fundamentally different in each case.
Cloudflare proposes evaluating automated traffic based on three questions:
What is the bot doing on the site?
What information does it store?
How can it reuse or display the content?
For practical management, Cloudflare defines three primary classes: Search, Agent and Training.
Search: Crawling for Search and Answers
The Search category includes systems that collect or index information to later use it for answering user queries.
According to Cloudflare's logic, a website owner who permits this crawling should expect at least one of two forms of value:
referral traffic to the site;
another fair form of compensation.
For an online store, search engines' access to products, categories, brands, and helpful articles remains crucial. Completely blocking all AI-enabled systems could hurt the store's visibility in the next generation of search.
One must consider that modern search platforms increasingly do more than just sort links; they generate direct answers themselves. As a result, the boundary between traditional search and AI search is becoming increasingly blurred.
Agent: Actions on Behalf of a Specific User
The Agent category covers automated systems that open a website in real-time to execute a specific user task.
For example, an AI agent might:
read a page linked by the user;
find product specifications;
check availability;
compare multiple models;
fill out a form;
add a product to the cart;
help complete an order.
In this scenario, a real person who expects a specific outcome is driving the bot's actions.
Cloudflare specifically classifies chatbots that load a page at a user's request and browser agents that control a web interface to execute a specific operation under the Agent category.
For e-commerce, this is a potentially valuable traffic source. In the future, AI agents could become a new sales channel, not only recommending products but also facilitating parts of the customer journey.
However, agent access requires oversight. There is no need to allow an automated system to trigger internal search thousands of times, cycle through all product filters, create empty shopping carts, or strain the checkout flow.
Training: Using Content to Train Models
The Training category refers to crawling where information is collected to train or fine-tune AI models.
In this scenario, the content may be integrated into a dataset and used to improve the future capabilities of the system.
This access does not necessarily provide the website with:
referral traffic;
brand mentions;
backlinks to the source;
leads;
direct revenue.
This is why many publishers and owners of unique databases are highly cautious regarding Training crawlers.
Cloudflare's stance is not that training must be universally blocked. Instead, the company aims to empower owners to choose their own terms: allow access, block it, or demand payment.
Cloudflare's Major Shift: Granular Rules Instead of Global Blocking
Previously, Cloudflare offered a Block AI Bots feature, allowing users to toggle a single switch to block crawlers associated with scraping data for model training.
However, this blanket blocking proved too rigid.
Consequently, Cloudflare introduced separate management for three distinct scenarios:
Category | What It's Used For | Potential Value for the Site |
|---|---|---|
Search | Indexing and generating answers | Visibility, mentions, referrals |
Agent | Executing tasks for real users | Potential leads, engagement, or sales |
Training | Training or fine-tuning models | Indirect benefits or compensation for access |
For each behavior type, Cloudflare allows users to select one of the core policies:
block site-wide;
block only on pages with ads;
allow (do not block).
The new settings are available through the AI bot policy management section in Cloudflare Security Settings, including for free-tier customers.
What Changes on September 15, 2026
Cloudflare announced that on September 15, 2026, it will update the default settings for new domains connecting to the platform.
On pages where Cloudflare detects the presence of advertisements:
Training bots will be blocked by default;
Agent bots will be blocked by default;
Search crawlers will remain allowed.
The underlying logic is that ad-supported pages are designed for human attention. If a bot scrapes the information without driving a user to the site, the publisher potentially loses advertising revenue.
At the same time, Search remains permitted, as search functions are most directly tied to potentially bringing audiences back to the site.
Another important update relates to multi-purpose bots. Some crawlers may simultaneously use content for both search and model training.
Cloudflare plans to apply the most restrictive of the established rules to these bots. For instance, if an owner allows Search but blocks Training, a mixed-use crawler with both functions may be blocked.
In Cloudflare's documentation, Googlebot, Applebot, and Bingbot are cited as examples of multi-purpose systems. Therefore, before activating aggressive Training blocks, website owners must verify that this rule won't inadvertently disrupt standard search indexing.
Allowing Access Doesn't Mean Letting Them Do Anything with Your Content
Cloudflare is also developing the concept of content use—the level of permissible utilization of content after it has been crawled.
Three levels are proposed:
Immediate
The system can interact with the page in real-time but must not store or reuse the retrieved content.
This mode is suitable for an agent executing a single user-requested task.
Reference
The system can index the information, show brief snippets, and link back to the primary source.
Cloudflare considers this level to be the default standard option.
Full
The system can generate comprehensive summaries and reproduce information much more broadly.
This is the most permissive mode, requiring careful consideration by owners of unique proprietary content.
As such, a future access policy might not simply read "allow AI search," but rather:
Allow search AI systems to index pages, quote short snippets, and mandate attribution to the source, but prohibit full content reproduction.
This is much more precise than simple User-Agent whitelisting or blocking.
New Signals in robots.txt
Cloudflare is testing the Content Signals extension, which allows website owners to state their preferences directly in the robots.txt file.
An example of such a policy:
User-agent: *
Content-Signal: search=yes,ai-train=no,use=reference
Allow: /
This directive translates to:
Crawling for search is permitted;
Utilization for AI training is not desired;
Content usage is limited to linking, indexing, and brief snippets (reference level).
At the same time, Cloudflare explicitly states that Content Signals in robots.txt express the site owner's preferences but do not act as technical enforcement. For true blocking, security rules, Bot Management, WAF, or other server-side mechanisms must be deployed.
This distinction is fundamentally important.
A robots.txt file is a declaration of guidelines for cooperative bots. Malicious or misidentified crawlers can simply ignore it.
What is Pay Per Crawl?
Cloudflare proposes a third alternative between free access and total blocking: Pay Per Crawl.
This model allows website owners to set a price for every successful page retrieval by an AI crawler.
The basic flow works as follows:
the crawler requests protected content;
Cloudflare signals that access requires payment;
the server returns an
HTTP 402 Payment Requiredstatus code and specifies the price;the crawler can accept the terms;
upon verification, it receives the page content;
the transaction is logged for subsequent billing and settlement.
Cloudflare acts as the technical platform and the Merchant of Record—handling the commercial clearing between the content owner and the crawler operator.
For any given crawler, owners can select one of three actions:
Allow—grant free access;
Charge—require payment;
Block—completely deny access.
Cloudflare's 2026 documentation describes setting domain-level pricing, configuring advanced rate rules, and monitoring monetized requests. The minimum base rate noted by Cloudflare is $0.01 USD per successful crawl. Feature availability and specific settings should be verified inside the user's Cloudflare dashboard.
Pay Per Crawl is unlikely to become an overnight windfall for typical e-commerce sites. However, the emergence of such a mechanism is milestone-worthy: access by automated systems to web content is finally being treated as a distinct economic transaction rather than an unconditional right given to any bot.
Why Completely Blocking All AI Bots Can Be a Mistake
At first glance, complete exclusion of AI traffic seems like the simplest solution.
However, for most commercial websites, this strategy might be overly aggressive.
Users are already looking for products, services, tutorials, and recommendations using AI engines. If a site is entirely locked down to search and discovery systems, its products risk being omitted from generative answers or represented using outdated, third-party data.
For small businesses, there is a dual risk:
Allowing unrestricted access and letting content be scraped without limits;
Blocking everything and losing visibility in emerging search channels.
Cloudflare itself acknowledges that global blocking is not a one-size-fits-all solution. That is why the company is pivoting to granular management of Search, Agent, and Training instead of a blanket "block all automation" policy.
In Sitezilla's view, the right approach is not a global ban, but a tailored access policy based on page types and crawler behavior.
Sitezilla's Recommended Policy for Online Stores
Below is a practical model that can serve as a foundation for e-commerce projects.
This is not a generic official Cloudflare recommendation, but rather a practical interpretation of the new features from Sitezilla.
Page Type | Search | Agent | Training |
|---|---|---|---|
Homepage | Allow | Allow Verified | Selective |
Product Categories | Allow | Allow | Selective or Paid |
Product Detail Pages | Allow | Allow | Selective |
Brand Pages | Allow | Allow | Selective |
Helpful Articles/Blog | Allow | Allow | Block, License, or Monetize |
Proprietary Research | Allow with Restrictions | Selective | Block or Require Payment |
Internal Search | Restrict | Restrict | Block |
Filter/Faceted Pages | Restrict | Restrict | Block |
Shopping Cart | Block Indexing | Allow Only Verified Agents (As Needed) | Block |
Checkout Page | Block | Strictly Control | Block |
User Account/Profile | Block | Only Authorized | Block |
Admin Panel | Block | Block | Block |
APIs & Service Routes | Custom Rules | Only Authorized | Block |
Product Feeds | Allow Specific Systems | Optional | Selective |
Product Pages and Categories
These pages must remain accessible to Search crawlers. They are what allow AI search engines to understand:
What the store is selling;
What specifications the product has;
What the current price is;
Whether the product is in stock;
What shipping and payment methods are available;
How one product differs from another.
For these pages, structured data is critical—including valid schema markup for Product, Offer, BreadcrumbList, Brand, and up-to-date pricing.
Blog Articles
Blogs should be open to search, but you do not necessarily have to permit unrestricted use for model training.
For standard SEO articles, you can apply a use=reference policy: indexing and short citations are allowed, provided there is attribution.
For unique proprietary research, custom databases, benchmarks, guides, and premium expert content, you might consider blocking Training or establishing paid access.
Filters and Internal Search
AI bots should not be allowed to run wild generating thousands of parameter combinations.
For example:
/category?color=red&size=42&sort=price&page=17Mass crawling of these URLs:
Creates duplicates;
Wastes server resources;
Pollutes analytics;
Increases database load;
Does not necessarily generate meaningful search value.
For filters, canonical tags, strict indexing rules, parameter handling constraints, and rate limiting are essential.
Administrative and Personal Sections
The administration panel, user profile, order history, service APIs, and internal routes must be sealed off not only from AI bots but from all unauthorized automated systems.
A robots.txt file is not enough here. You need:
Robust authentication;
WAF (Web Application Firewall) rules;
Rate limiting;
Session verification;
API protection;
Blocking of suspicious automated requests.
AI Crawlers Should Be Analyzed, Not Just Blocked
A single Cloudflare configuration is not a cure-all.
Website owners need visibility into which automated systems are visiting their site, how many pages they load, and whether they contribute any value.
To assist with this, Cloudflare is developing BotBase and Attribution Business Insights. BotBase classifies known crawlers based on behavior, while Attribution Business Insights monitors bot traffic volume, bandwidth consumption, search-to-referral ratios, and current policy outcomes (allowed or blocked).
In Sitezilla's opinion, a similar report should be integrated into your in-house web analytics stack.
What an "AI Crawlers" Report Should Track
For every crawler, it is advisable to log:
Bot name or recognized operator;
User-Agent string;
Verification status (verified vs. unverified);
IP addresses and countries of origin;
Total request count;
Unique URL count;
Volume of data transferred;
Average response time;
HTTP status codes (200, 301, 404, 403, 429, and 500);
Most frequently accessed pages;
Recrawl frequency;
Referrals generated by the corresponding AI service;
Crawl-to-referral ratio;
Leads or orders generated;
Estimated infrastructure/server overhead cost.
It is crucial to track operators of search engines, AI platforms, and agents without relying solely on the User-Agent header. Since headers can be easily spoofed, robust decision-making requires IP validation, reverse DNS, Web Bot Auth, or Cloudflare's proprietary classification.
How to Evaluate the Value of a Specific Crawler
Suppose that over a month, a specific bot:
Made 100,000 requests;
Downloaded 20 GB of data;
Frequently hit computationally expensive faceted filter pages;
Created noticeable database (e.g., MySQL) overhead;
Brought in exactly 2 visitors;
Generated zero sales.
Such an exchange is clearly disadvantageous for the store.
On the other hand, another operator might:
Make 2,000 requests;
Crawl only core product detail pages;
Drive 200 visitors to the site;
Generate several assisted conversions.
It is logical to keep the second bot whitelisted, whereas the first one should be subject to rate limiting, outright blocking, or a monetization rule.
Thus, access decisions must be driven by real-world behavior and metrics rather than company names or the broad "AI" label.
A Step-by-Step Plan to Implement an AI Crawler Policy
Step 1. Audit Your Automated Traffic
Analyze server logs to identify:
Which bots are visiting the site;
Which URLs they are crawling;
How frequently they return;
What errors they generate;
Which ones consume the most system resources.
It is recommended to analyze a window of at least 30 days.
Step 2. Segment Your URLs by Purpose
Group all application routes into distinct buckets:
Public commercial content;
Informational content;
Unique proprietary/creative assets;
Interactive/user-facing pages;
Personal profile pages;
Service routes;
APIs;
Parametric, technical, and faceted URLs.
Step 3. Assess the Value of Each Group
For each segment, answer the following:
Is search visibility required?
Can an AI agent drive a customer to this page?
Is utilization for model training permitted?
Does the page contain proprietary commercial data?
Does crawling this segment cause significant server load?
Can access to this content be monetized?
Step 4. Configure Distinct Policies
Avoid applying a single blanket action to the entire domain.
A baseline configuration might look like this:
Search—allow on public-facing pages;
Agent—allow verified systems with rate limiting applied;
Training—block, or allow highly selectively;
Internal pages—block for all automated systems;
Premium proprietary content—license or monetize access.
Step 5. Configure robots.txt and Content Signals
The robots.txt file should outline the site's overall policy but should not be relied upon as the sole line of defense.
For robust enforcement, pair it with Cloudflare rules, WAF configurations, server-side limits, and active bot verification.
Step 6. Track Referrals from AI Services
In your analytics platform, categorize and group referrals from:
AI search engines;
Chatbots;
Browser assistants;
Standard search engines;
Unidentified referrers.
Over time, this will reveal which AI platforms actually drive genuine user engagement.
Step 7. Review the Policy Monthly
The ecosystem is evolving at a breakneck pace. A bot utilized solely for training today might function as a search agent tomorrow, or vice versa.
Therefore, these rules should not be treated as a set-and-forget task.
SEO is Morphing into SEO, AEO, and GEO Simultaneously
Classic SEO optimizes a website to perform on search engine results pages (SERPs).
AEO (Answer Engine Optimization) helps content surface within direct, zero-click answers.
GEO (Generative Engine Optimization) focuses on ensuring that generative systems accurately interpret your brand, products, facts, and original sources.
Cloudflare explicitly highlights this shift from traditional SEO to the worlds of AEO and GEO, where analyzing positions and organic traffic is no longer sufficient; analyzing AI crawler behavior is now equally important.
For website owners, this introduces a new operational imperative:
Control content access;
Maintain visibility for valuable AI systems;
Protect unique proprietary assets;
Measure AI referrals;
Build a machine-readable site structure;
Verify the accuracy of brand information across models;
Prevent automated scraping from creating unnecessary resource overhead.
AI bot control is no longer solely the responsibility of system administrators or security teams. It is now a core issue of SEO, marketing, content strategy, and website economics.
Sitezilla's Position
Sitezilla advises against turning on a global block for all AI bots on commercial websites without prior analysis.
Conversely, complete openness is also not an optimal strategy.
A rational policy must be selective:
Public products and categories—accessible to Search;
Verified agents acting in the buyer's interest—allowed under restrictions;
Model training—strictly on the owner's explicit consent;
Unique, high-cost content—protected, licensed, or monetized;
Faceted filters, internal search, and technical URLs—restricted;
Shopping cart, user accounts, APIs, and admin routes—sealed off;
All bot activity—independently logged and analyzed.
The core principle can be stated simply:
You do not need to block all AI. You need to understand who is visiting your site, what they are doing, what value they extract, and what they return to the content owner.
Conclusion
Cloudflare is effectively proposing a new paradigm of relations between websites and automated systems.
Instead of unconditional access or complete prohibition, owners should have the tools to:
Allow constructive search crawling;
Serve agents acting for real users;
Block model training;
Restrict the methods and scope of content reuse;
Demand payment for crawling access;
Access transparent analytics on crawls and referrals.
For website owners, this means that robots.txt configurations, SEO, and bot protection can no longer be addressed in silos.
In the coming years, the competitive edge will belong not to websites that simply opened or closed their doors to AI, but to those that have built a structured control loop:
allowing the constructive, blocking the parasitic, measuring the impact, and setting explicit terms for content usage.
This is the central theme of Content Independence Day: your site, your content—your access rules.