Enforcement
Metered or authenticated access. Conditions access on payment, authentication, identity, or usage limits.
It can control the protected access path, but cannot guarantee control of downstream copies or alternate routes.
Catalog status describes public availability, not legal validity, adoption, or proven effectiveness.
What it is
Cloudflare’s controls distinguish Search, Agent, and Training traffic and let site owners apply edge blocking by use case. Its dashboard and BotBase provide visibility into known bots, managed robots.txt publishes Content Signals, redirects can steer verified training crawlers toward canonical content, and Pay Per Crawl supports monetized automated access.
Cloudflare AI Crawl Control conditions automated access through authentication, payment, or rate controls for web content across the collection and retrieval stages. It also incorporates technical blocking and preference signaling. The mechanism controls a protected access route; it does not determine how acquired material is used or whether alternate routes remain available. Public materials describe a currently available initiative; the newest dated source in this profile is “Search, Agent, and Training bot controls launched” (July 1, 2026). These details describe the published mechanism and evidence, not a finding about legal validity, adoption, or effectiveness.
Limitations
Cloudflare's robots.txt Content Signals express preferences and do not issue blocks directly. Its edge security rules can block traffic that Cloudflare classifies as automated, while the effectiveness of preference signals still depends on crawler behavior.
Evidence trail
- Jul 01, 2026 · Primary sourceSearch, Agent, and Training bot controls launched
- Apr 17, 2026 · Primary sourceRedirects for AI Training launched
- Sep 24, 2025 · Primary sourceContent Signals Policy launched
- Aug 28, 2025 · Primary sourceAI Crawl Control general availability announced
- Jul 01, 2025 · Primary sourceAI Audit and marketplace features launched
Adoption signals
Users
3.8M+ domains on managed robots.txt
- Content Signals Policy launchedSep 24, 2025
Data volume
1B+ HTTP 402 responses/day across Cloudflare customers
The company-wide figure is not specific to AI Crawl Control or paid crawling.
- AI Crawl Control general availability announcedAug 28, 2025
Examples
Cloudflare's Robotcop post uses this abbreviated news-site policy as a concrete AI crawler blocklist example.
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: anthropic-ai
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: Bytespider
Disallow: /Cloudflare's July 2026 update adds a reference-use preference to its managed robots.txt signal.
User-Agent: *
Content-Signal: search=yes, ai-train=no, use=reference
Allow: /