Cloudflare’s September 15 AI Crawler Defaults: What Publishers Must Check
Cloudflare’s September 15 defaults can block multipurpose crawlers when Training is blocked. Audit ad pages, verified bots, and final responses before the change.
Published August 10, 2026: Cloudflare says that on September 15 it will change the defaults applied to new domains. On pages that display ads, Training and Agent crawlers will be blocked by default while Search crawlers remain allowed. The consequential detail is how Cloudflare says it will treat a crawler assigned to more than one behavior.
Cloudflare says the most restrictive applicable rule will win. A customer who blocks Training can therefore block a multipurpose crawler that also performs Search. Cloudflare names Googlebot, Applebot, and BingBot as examples. Before September 15, publishers should save their current rules, identify pages that display ads, and test the verified search crawlers they cannot afford to block.
What Cloudflare announced
The change applies to new domains onboarding to Cloudflare, not automatically to every existing zone. Cloudflare separates three AI-related behaviors: Search, Agent, and Training. Its announced ad-page defaults allow Search and block Agent and Training.
That classification is behavior-based, not a promise that one user-agent always has one purpose. Cloudflare says a bot can be multipurpose and that the most restrictive applicable setting will decide whether it is allowed. The result can be broader than a publisher expects from a label such as “block training.”
| Behavior | Announced default | Publisher check |
|---|---|---|
| Search | Allowed | Confirm verified search crawlers return the intended status and HTML |
| Agent | Blocked | Document whether any required service is classified here |
| Training | Blocked | Check multipurpose crawlers affected by the restrictive rule |
Audit before September 15
- Export the current configuration. Record the zone, active bot products, legacy Block AI Bots state, custom WAF rules, exceptions, and owner.
- Map ad-bearing templates. Include article, category, search, tool, and archive pages; do not assume the homepage represents the site.
- List essential crawlers. Start with Googlebot and Bingbot, then add services required for previews, monitoring, feeds, or commerce.
- Verify identity. A user-agent string can be spoofed. Use the operator’s documented IP or reverse-DNS method and Cloudflare’s verified-bot signals where appropriate.
- Test the final request path. Check status, challenge page, response headers, robots controls, canonical, and visible main content from representative URLs.
- Set an alert. Watch verified-bot 403s, challenge responses, crawl-rate changes, and Search Console or Bing crawl anomalies after the release date.
The existing Cloudflare Googlebot 403 incident guide shows why an intended AI control must be checked at the delivered response. Use the broader Search, Agent, and Training bot policy to record the decision rather than leaving it inside a dashboard toggle.
Run a small response matrix
Choose at least one ad page and one non-ad page from every important template. Record the same fields before and after the change: URL, crawler identity, behavior classification, applied rule, HTTP status, challenge state, robots directive, canonical, body length, and whether the main content is present.
A passing browser visit is not a passing crawler test. A human session may carry cookies, JavaScript, or challenge clearance that the crawler does not have. Compare raw responses and Cloudflare event logs, then reproduce the result without the authenticated browser state.
If a required search crawler is blocked, narrow the rule or add a verified exception whose scope and owner are documented. Do not whitelist an arbitrary user-agent string, and do not remove all bot controls merely to repair one classification conflict.
What the announcement does not prove
- Blocking a Training-classified request does not prove that the operator would have used the response for model training.
- Allowing a Search-classified crawler does not guarantee discovery, indexing, ranking, citation, or referral traffic.
- A Cloudflare verified-bot classification does not replace the bot operator’s own documentation about intended use.
- The September defaults do not remove the publisher’s responsibility to review existing and custom rules.
Treat this as a release change with a rollback path. Preserve the old settings, name the person who can reverse the change, and run the technical SEO launch checklist on any production adjustment.
Ask a question or join the discussion