Future TechnologyFuture Technology
Security

Cloudflare just gave every website a switch to cut AI off from its content

· 5 min read · By Future Technology

Key takeaways

  • From September 15, 2026, Cloudflare will block AI training and agent crawlers by default on any site that runs ads, while still allowing search crawlers through
  • The change hits new Cloudflare customers, new sites from existing customers, and every site on the free tier automatically
  • Cloudflare now splits bots into three categories: Search, Agent, and Training, and judges multi-purpose bots on all their behaviors
  • Site owners can opt out of the new defaults any time before the deadline in their Cloudflare security settings

Cloudflare sits in front of roughly one fifth of the web, and starting September 15, 2026, it is going to start saying no on behalf of millions of sites that never explicitly asked it to. AI crawlers that scrape pages to train models, or that act live on someone's behalf, will be blocked by default on any site running ads. Search crawlers, the ones that index a page so a chatbot can cite it later, stay allowed.

If you run a website and do nothing between now and September, this happens to you automatically.

What Cloudflare actually changed

Cloudflare split bot traffic into three buckets. Search crawlers collect content to answer questions about it later, the way a search engine indexes a page. Agent crawlers act in real time on a person's behalf, think a chatbot fetching a page while you're mid-conversation with it. Training crawlers take content specifically to train or fine-tune a model, no immediate user request attached.

The new default blocks Training and Agent traffic on ad-supported pages. Search stays on, because blocking it would make a site invisible to the AI-powered search tools that are increasingly how people find things.

Mixed-use crawlers get judged on everything

The catch is multi-purpose bots. A crawler that does both Search and Training gets judged on everything it does, so if you have told Cloudflare to block Training, a mixed-use bot like Googlebot can get swept up in that block even though part of its job is indexing you for search. Cloudflare has said site owners can fine-tune this in their settings, but the default is blunt.

Why this is happening now

The framing from Cloudflare and most of the coverage around it is straightforward: AI companies have been quietly free-riding on content that publishers spent money to produce, and the training crawlers rarely paid for what they took. This policy is less about blocking AI outright and more about forcing a negotiation. Block by default, and AI companies that actually want the content have to come back and license it, or build a relationship with the publisher instead of just scraping.

TechCrunch reported the policy as explicitly designed to push AI companies toward paying publishers for their content, rather than treating the open web as a free training set. That's the commercial angle. The security angle, per Cloudflare's own writeup, is that it also gives smaller sites a default posture they'd otherwise have to configure themselves, since most site owners have never touched a robots.txt file, let alone a bot management dashboard.

Who this actually affects

If you are running a personal blog, a niche publication, or any ad-supported site on Cloudflare's free tier, you are getting this by default with no action required. If you are a big enterprise customer with an existing Cloudflare relationship, it is opt-in rather than forced on you retroactively, though new sites you spin up will inherit the new defaults too.

For anyone who's been through the SharePoint mess from earlier this month, this fits a pattern: the infrastructure layer of the internet increasingly makes security and access decisions for site owners by default, and you have to actively dig into settings to change them, in either direction. We covered the actively exploited SharePoint RCE a few days ago, same theme: default postures matter more than most people think, because most people never change them.

It is also worth reading alongside the wider AI industry conversation. The global race for training data, which we touched on in our look at how AI agents actually work, points to a bottleneck where training data matters as much as compute. Cloudflare's move directly targets that bottleneck.

What site owners should do before September 15

Check your Cloudflare security settings now rather than waiting. If you want AI companies to keep training on your content for free, you'll need to explicitly opt out of the new block. If you'd rather force a licensing conversation, the new default does that work for you. Either way, the deadline is fixed, and doing nothing is itself a choice here.

If you are running your own stack outside Cloudflare, this is also a good moment to check whether your own site allows Training crawlers by default. Publishers curious about running things themselves rather than depending on a single infrastructure provider might find our review of the self-hosted AI assistant OpenClaw a useful next read, same underlying question of who controls the defaults.

The bottom line

Cloudflare isn't banning AI training, it's changing who has to ask permission first. From September 15, the assumption flips: AI companies now have to show up and negotiate instead of quietly scraping, at least on the huge slice of the web that sits behind Cloudflare. Whether that actually gets small publishers paid, or just adds friction that big AI labs route around, is the thing worth watching over the next few months.

Read next

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.