Cloudflare AI Crawl Control

Live

Live controls for classifying and blocking Search, Agent, and Training bots, alongside managed preference signals and monetization tools.

Website
blog.cloudflare.com/control-content-use-for-ai-training
Latest tracked evidence
Jul 01, 2026Search, Agent, and Training bot controls launched
Last checked
Jul 15, 2026
Profile updated
Jul 15, 2026
Primary approach
Tollgate
Practical force
Access controlIt can control the protected access path, but cannot guarantee control of downstream copies or alternate routes.
Also uses
Technical blocking / Preference signal
Pipeline
Collect / Retrieve
Status rationale
Publicly available or deployable in the latest catalog review.

Enforcement

Metered or authenticated access. Conditions access on payment, authentication, identity, or usage limits.

It can control the protected access path, but cannot guarantee control of downstream copies or alternate routes.

Catalog status describes public availability, not legal validity, adoption, or proven effectiveness.

What it is

Cloudflare’s controls distinguish Search, Agent, and Training traffic and let site owners apply edge blocking by use case. Its dashboard and BotBase provide visibility into known bots, managed robots.txt publishes Content Signals, redirects can steer verified training crawlers toward canonical content, and Pay Per Crawl supports monetized automated access.

Cloudflare AI Crawl Control conditions automated access through authentication, payment, or rate controls for web content across the collection and retrieval stages. It also incorporates technical blocking and preference signaling. The mechanism controls a protected access route; it does not determine how acquired material is used or whether alternate routes remain available. Public materials describe a currently available initiative; the newest dated source in this profile is “Search, Agent, and Training bot controls launched” (July 1, 2026). These details describe the published mechanism and evidence, not a finding about legal validity, adoption, or effectiveness.

Limitations

Cloudflare's robots.txt Content Signals express preferences and do not issue blocks directly. Its edge security rules can block traffic that Cloudflare classifies as automated, while the effectiveness of preference signals still depends on crawler behavior.

Evidence trail

Adoption signals

Examples

Real-world robots.txt blocklistSource

Cloudflare's Robotcop post uses this abbreviated news-site policy as a concrete AI crawler blocklist example.

User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Bytespider
Disallow: /
Managed Content Signals exampleSource

Cloudflare's July 2026 update adds a reference-use preference to its managed robots.txt signal.

User-Agent: *
Content-Signal: search=yes, ai-train=no, use=reference
Allow: /