If you run a website, you probably want Google to crawl it. You probably don’t want every AI company on the planet helping itself to your content to train its models.
Until now, separating those two things hasn’t always been easy. Cloudflare is trying to fix that.
On September 15, Cloudflare launched a new Disallow AI Training setting that lets website owners continue welcoming traditional search crawlers while telling AI training crawlers to stay away.
And this isn’t just a polite request.
Cloudflare says training-only crawlers from companies including Amazon, Anthropic, Meta and OpenAI can be blocked outright, while mixed-use crawlers such as Googlebot can still crawl the site for search. Google already lets publishers restrict certain Gemini training and grounding uses through Google-Extended without affecting their Google Search ranking or inclusion.
Cloudflare is going even further with new ad-supported sites. Its recommended settings now essentially say: let search engines in, block AI training, and block AI agents on pages where ads are detected.
There are limitations. Robots.txt ultimately relies on bots behaving themselves, and blocking Google-Extended does notremove your content from Google’s AI Overviews or AI Mode. Those are considered part of Google Search.
But the bigger message is pretty clear.
For years, websites let bots crawl their content because there was something in return: traffic. AI threatens to rewrite that deal by consuming content without necessarily sending the visitor back.
Cloudflare is giving publishers a simple response:
Keep the search traffic. Give the AI training bots the middle finger.






