Marketing

Cloudflare Blocks AI Crawlers by Default: What Publishers Need to Know Before September 15

By Post For Success · Jul 16, 2026 · 9 min read
A translucent digital gateway filtering streams of automated crawler bots, allowing some through and blocking others

On July 1, 2026, Cloudflare — the network that sits in front of roughly a fifth of the web — announced the biggest change yet to how AI bots are allowed to read your site. From September 15, 2026, Cloudflare will block "training" and "agent" AI crawlers by default on any page that shows ads, while still letting genuine search crawlers through. AI companies have until that date to cleanly separate the bots they use for search from the ones they use for training and automated agents — or risk being blocked across a huge slice of the open web.

If you publish content and monetise it with advertising — or plan to — this is not an infrastructure footnote. It changes the default answer to a question every content owner now faces: should AI systems be allowed to take your work for free? Here is exactly what changed, who it affects, and what to do before the deadline.

What Cloudflare actually announced

Cloudflare is splitting AI bot traffic into three named categories and letting site owners set a policy for each one independently. The distinction matters, because "AI bot" has always been a blunt label that lumped very different behaviours together.

CategoryCloudflare's definitionDefault on ad pages (new sites)
SearchBehaviour that collects or indexes your content so it can answer questions about it later — and link back to you.Allowed
AgentAutomated behaviour acting in real time on a person's behalf to get something done right now.Blocked
TrainingA crawler taking your content to train or fine-tune a model — permanently absorbing it into the model's architecture.Blocked

The logic is that search sends you traffic — a bot indexes your page, then a human clicks through to it — while training and agentic scraping usually do not. A model that swallows your article to answer a user inside a chatbot rarely sends that reader back to your site. Cloudflare's position, in CEO Matthew Prince's words: "Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge."

The September 15 deadline and what "ad pages" means

Two things happen on September 15, 2026. First, for all new domains onboarding to Cloudflare, the Training and Agent categories will be blocked by default — but only on pages that display advertising. Search stays allowed. Second, AI companies are on notice: multi-purpose crawlers that blend search with training will be treated as training bots unless they separate the two.

Why single out ad-supported pages? Cloudflare's framing is that an ad is a signal of intent: "An ad is a signal that a website owner meant for a person to land there and see it — something monetizable that fuels the business." A page carrying ads is, almost by definition, one where a human visit has commercial value — exactly the value that uncompensated AI scraping erodes. Pages without ads are not force-blocked; the stricter default is reserved for the monetised web.

One consequence worth flagging: because major crawlers such as Googlebot, Applebot and Bingbot are multi-purpose, a site that opts to block Training may also block those bots if the AI companies behind them have not separated their crawlers by the deadline. That is precisely the pressure Cloudflare is trying to apply.

Who is affected — and who is not (yet)

The default change is deliberately scoped so it does not silently rewrite the rules for millions of established sites overnight. It applies to:

  • All new domains onboarding to Cloudflare after September 15, 2026.
  • New sites set up by existing Cloudflare customers.
  • All existing free-tier customers.

Existing paid customers are not switched over automatically — but every Cloudflare customer can set these preferences manually today in their Security settings, choosing to allow or block Search, Agent and Training independently. In other words, the block is now the default for new and free sites, and a one-click option for everyone else. If you would rather keep letting AI training bots in — some publishers do, in exchange for licensing deals — you can opt out before the deadline.

Pay Per Crawl: from blocking to charging

Blocking is only half of Cloudflare's play. The other half is a marketplace called Pay Per Crawl, which lets a site charge AI bots a fee for access instead of simply turning them away. Rather than a binary "allow or block," a publisher can set a price and let AI companies decide whether the content is worth paying for.

Cloudflare says this is now evolving into a broader "Pay Per Use" model — the idea being to charge AI companies when your content actually creates value (for example, when it appears in an AI-generated answer or search result), not merely when a bot fetches the page. Early partners named alongside the launch include You.com and Ceramic.ai. The company also points to a striking inefficiency it says justifies the whole effort: more than half of AI crawler traffic is re-fetching pages that have not changed, burning bandwidth for no new information.

For small publishers, Pay Per Crawl is unlikely to become meaningful revenue on its own any time soon. But it establishes something that did not exist before: a price and a permission layer between AI companies and the open web. That is the structural shift, even if the cheques are small at first.

What this means for your SEO and content strategy

It is easy to read "Cloudflare blocks AI crawlers" and panic that your search visibility is about to vanish. It is not — as long as you understand the distinction. Here is how to think about it.

1. Search crawlers are still allowed — protect that

The default keeps Search bots (including AI search that cites and links back) flowing. Blocking Training and Agent does not remove you from Google, Bing, or AI search results that link to sources. If anything, Cloudflare is trying to preserve the search-and-cite model that still sends you traffic. Do not reflexively block everything — check that legitimate search and AI-citation crawlers remain allowed, because appearing in AI search results is still a growth channel, not a threat.

2. Decide your stance on training deliberately

You now have a real choice: block AI training entirely, allow it, or charge for it. Publishers with distinctive, high-value content may prefer to block or monetise; sites that benefit from broad AI exposure may keep it open. Make it a decision, not a default you never noticed — especially if you run ads, where the stricter default now applies.

3. Watch the traffic mix, not just rankings

As AI answers absorb more clicks, standalone chatbots are becoming their own referral channel — one where ChatGPT alone drives over 90% of trackable AI referral traffic. If you block agent and training bots but keep search open, you are betting that citation-with-a-link beats silent absorption. Track how much traffic actually arrives from AI sources so you can tell whether that bet is paying off.

4. Keep your content machine-readable where you want it read

For the crawlers you do welcome, the fundamentals still apply: server-rendered HTML, clean structure, and clear, quotable passages make you easier to index and cite. Blocking the bots you don't want does not help the ones you do — that is still won with good on-page optimization.

The bigger picture: a permission layer for the web

For twenty-five years the open web ran on an honour system: robots.txt politely asked bots to behave, and most search engines complied because indexing sent traffic back to publishers. Generative AI broke that bargain — it takes the content but often keeps the reader. Cloudflare's move is an attempt to rebuild the bargain at the network level, where a polite request becomes an enforced default and, increasingly, a priced transaction.

Whether that reshapes the AI-content economy or simply nudges it depends on what the AI companies do before September 15. But the direction is clear: the era of free, unlimited, unattributed scraping of ad-supported content is ending, and "who gets to read your site, and on what terms" is now a setting you control — not an assumption you inherit.

FAQ

Does Cloudflare's change remove my site from Google or AI search?

No. The default blocks only the Training and Agent categories on ad-supported pages. Search crawlers — including AI search that indexes and links back to your content — remain allowed by default, so your visibility in Google, Bing and citation-based AI answers is not affected by this change.

When does the new default take effect and who does it apply to?

September 15, 2026. The default block on Training and Agent bots applies to all new domains onboarding to Cloudflare, new sites created by existing customers, and all existing free-tier customers. Existing paid customers are not switched automatically but can set the same policy manually in their Security settings at any time.

What counts as an "ad page" for the stricter default?

Cloudflare applies the tougher default to pages that display advertising, on the reasoning that an ad signals the page was built for a human visit with commercial value. Pages without ads are not force-blocked, though you can still choose to block AI bots on them manually.

What is Cloudflare's Pay Per Crawl?

Pay Per Crawl is a marketplace that lets publishers charge AI bots a fee for access instead of simply blocking them. Cloudflare is extending it into a "Pay Per Use" model that aims to charge AI companies when your content creates value — such as appearing in an AI answer — rather than only when a bot fetches the page.

Should small publishers block AI training bots?

It depends on your goals. If your content is distinctive and you monetise with ads, blocking or charging for training access protects value that scraping erodes. If broad AI exposure helps you, you may keep it open. The key change is that it is now an explicit, controllable decision rather than an unmanaged default.

← More in Marketing