Cloudflare's AI crawler wall went live, and Googlebot is caught in it
Key takeaways
- From 15 September, Cloudflare blocks Training and Agent category crawlers by default on any page showing ads, for new domains and all existing free customers.
- Googlebot crawls for both search indexing and AI training under one user agent, so blocking Training can also block it unless you opt out.
- Search crawling stays allowed by default, and site owners can override the defaults in either direction.
Cloudflare's crawler policy, announced back in July, took effect on 15 September. New domains, new sites from existing customers, and every existing free tier customer now block Training and Agent category crawlers by default on any page that displays ads. Search crawling stays allowed.
The mixed use bot problem
The complication is bots that do more than one job. Googlebot crawls for search indexing and for AI training through a single user agent, so it gets handled under whichever rule is most restrictive. A site that blocks Training therefore also blocks Googlebot on those pages unless the owner explicitly opts out.
That is a quiet way to lose search visibility while believing you only turned away scrapers. The failure mode is not an error message. It is a traffic graph that starts sloping down a few weeks later, by which point the cause is hard to see.
Why it matters
This is the largest default-on change to how AI companies reach the open web so far, and it moves the burden. Bots now have to declare what they are for. Publishers get a lever they did not previously have, which is the point: the policy is aimed at making AI companies negotiate for content rather than take it.
For site owners the practical effect is smaller and more urgent. The defaults changed under you, they are adjustable in either direction, and the documentation explains how. That makes this a settings check rather than a crisis, as long as you do the check. If you run anything on the Cloudflare free tier, this week is the right time, not the week after you notice.
What to watch next
Two things. Whether Google splits its crawler into separate agents for search and training, which would dissolve the trap entirely and which publishers have been asking for. And whether other CDNs follow with defaults of their own, because a fifth of the web moving at once tends to set the norm for everyone else. Related reading on the governance side: the Open Secure AI Alliance and the competitive pressure between the major labs that makes training data worth arguing over.