Cloudflare MCP Detection vs Crawler Control: A 24-Request Test

Separate Cloudflare Gateway MCP policy from public crawler controls with a 24-request fixture and a 145,325-request origin-log field audit.

Sonar separates an MCP Gateway lane from a public crawler and robots control lane while a security operator watches.

Updated August 29, 2026: Cloudflare Gateway can identify Model Context Protocol traffic crossing a managed network and expose a beta experimental.is_mcp selector in HTTP policies. That signal can help allow, block or isolate agent-to-tool connections. It is not a public-web crawler classification and does not replace robots.txt, AI Crawl Control, bot verification, WAF rules or origin authorization.

Direct answer: write the policy from the traffic path, not from the word “AI.” MCP governance protects a user or agent connecting to tools and data. Crawler governance protects a public website from automated retrieval. They may share vendors and HTTP, but they have different directions, identities, evidence and enforcement points.

What Cloudflare actually released

Cloudflare’s August 12 changelog says Gateway detects MCP requests by inspecting protocol-specific headers and payload characteristics. The new Is MCP selector is available in HTTP policies as experimental.is_mcp. Cloudflare marks the selector beta and says it may change before general availability.

The same release added an AI security report under Insights & Logs with MCP request volume, unique users, unique MCP servers and policies targeting MCP traffic. A separate Traffic Source selector identifies how traffic reached Gateway, including mcp_portal, device client, mesh, WAN and proxy endpoint. These are observations inside a configured Cloudflare One traffic path—not a census of every agent or bot on the Internet.

Documented capability versus inference to avoid
Documented capabilityWhat it supportsWhat it does not establish
experimental.is_mcpGateway classified the request as MCPThat the user was authorized for the tool or data
Traffic SourceHow traffic entered Cloudflare, such as an MCP portalThat the destination approved the action
AI security dashboardObserved MCP usage in the governed environmentOff-network use or public crawler activity
Allow, block or isolate ruleNetwork policy outcomeApplication-level scope, consent or data safety

Separate four control planes

Choose the control that can actually observe and stop the traffic
Traffic classTypical directionUseful controlsDecisive evidence
MCP agent-to-toolUser or agent network → MCP serverGateway policy, portal, identity, server authorization, tool scopesGateway event plus authenticated server decision
Public search or answer crawlerOperator infrastructure → public siterobots.txt, AI Crawl Control, WAF, verified-bot rules, origin controlsVerified bot identity and edge/origin request
Training crawlerModel operator → public siteRobots preference, behavior policy, WAF or crawl-control enforcementDocumented identity, behavior category and request outcome
Authenticated application callApproved client → private serviceAuthentication, authorization, rate limits, audit trailPrincipal, scope, resource and allow/deny result

A Gateway rule can block MCP traffic that bypasses an approved portal, but it does not tell Googlebot whether it may crawl an article. A robots.txt rule expresses a site owner’s crawl preference, but it does not decide whether an employee’s agent may call a payroll tool. Keep both policies when both paths exist.

Original field audit: why an origin log cannot prove MCP

Search Engine Answer inspected an archived production HTTPS access log before publishing this revision. The log covered 145,325 requests from July 31, 2026 at 12:37 UTC through August 29 at 12:04 UTC. We retained only aggregate counts and the log-field shape; no raw IP address, URL token or user identifier was copied into this article’s evidence package.

The ordinary Apache-style record exposed the request line, HTTP outcome, referrer and user agent. It did not include Cloudflare Gateway’s experimental.is_mcp result, Traffic Source, authenticated Gateway user, MCP server identity or tool-call authorization. Therefore this log can show that a URL was requested, but it cannot establish that the request was an MCP session, which portal carried it, what tool was called or whether an internal user was allowed to receive the result.

Build an MCP policy without removing server authorization

  1. Inventory the route. Record approved clients, portals, users, groups, servers, tools, data classes and environments.
  2. Choose the objective. Decide whether the rule should report, block unknown MCP, require an approved portal or isolate risky destinations.
  3. Scope narrowly. Combine experimental.is_mcp with Traffic Source, identity group, destination and environment where the product supports it.
  4. Observe first. Capture a representative period and identify legitimate automation, emergency access and non-MCP HTTP sharing the same destinations.
  5. Test fixtures. Exercise known MCP and ordinary HTTP requests before broad enforcement.
  6. Keep application authorization. The MCP server must still decide whether the authenticated principal may call the specific tool and access the requested record.
  7. Publish exception and rollback rules. Assign an owner, expiry, approval trail and a clear condition for reverting the policy.

Network admission and tool authorization answer different questions. “This looks like MCP from an approved portal” is not equivalent to “this user may invoke delete_customer against this tenant.” Log both decisions without placing secrets or sensitive payloads in the policy record.

Download the 24-request control fixture

The MCP traffic control fixture (CSV) contains 12 known-MCP cases and 12 non-MCP controls. It varies client approval, request shape, destination, portal use and authentication. The ordinary controls include browser traffic, APIs, webhooks, public crawlers, a non-MCP event stream and long polling so a broad detector does not win by matching every unusual HTTP request.

The file is a preflight template, not a Cloudflare accuracy benchmark. Replace generic clients and destinations with your own authorized fixtures, leave secrets out, and fill observed_is_mcp, action and evidence fields from a controlled run.

Calculate outcomes from the fixed denominators
OutcomeDenominatorRelease interpretation
Known MCP detected12 MCP fixturesRecall for the tested clients and shapes only
Known MCP missed12 MCP fixturesCritical-client misses block enforcement
Ordinary HTTP marked MCP12 non-MCP controlsFalse positives must be explained and reduced
Ordinary HTTP ignored12 non-MCP controlsConfirms separation within this fixture set

Twenty-four requests can expose a configuration error; they cannot estimate population-wide accuracy. Expand the matrix for every production client, protocol version, transport, proxy, on-ramp and failure mode before organization-wide enforcement.

Build the public-crawler policy separately

Cloudflare’s current bot taxonomy distinguishes behavior such as Search, Agent and Training, and also labels operation as Direct or Intermediary. Its Verified bot process uses identity mechanisms such as Web Bot Auth signatures, published IP lists with stable user agents or reverse DNS. That evidence belongs to public-site bot policy—not to the Gateway MCP selector.

For crawler preferences, publish robots.txt and verify that it returns the intended content. Cloudflare notes that robots compliance is voluntary; use AI Crawl Control, bot solutions, WAF or origin controls when the business requirement is technical enforcement. Decide separately whether to allow search discovery, user-directed agent access and model training. A single “block AI” label is too coarse for those different jobs.

The AI crawler guide covers robots and operator identities. The signed-agent verification guide explains why a claimed user agent alone is not identity.

Preserve a joinable evidence model

Do not ask one log to answer every question
Evidence sourceCan establishCannot establish alone
Gateway eventMCP classification, traffic source and network action when recordedTool-level permission or final data use
MCP server auditPrincipal, tool, resource, authorization and application resultOff-path traffic not reaching that server
Edge bot recordVerified identity or behavior classification and edge actionHuman intent behind an intermediary agent
Origin access logRequest, status, bytes, referrer and user agent when configuredGateway-only detector fields or payload meaning
robots.txt snapshotPublished crawler preference at a point in timeCompliance or enforcement outcome

Use a correlation identifier that is safe to retain across Gateway and server logs. Keep raw payloads out of the editorial worksheet; store only the minimum evidence needed to reproduce the decision, under the organization’s access and retention policy.

Release gates and rollback signals

  • Known production MCP clients are detected across every supported request shape and on-ramp.
  • Comparable non-MCP HTTP is not swept into the rule.
  • Approved users and portals continue to reach allowed servers.
  • Unapproved direct routes are blocked or isolated as designed.
  • The MCP server independently enforces tool and data authorization.
  • Public crawler controls are tested through their own robots, bot and origin evidence.
  • An operator can explain every alert and exception without exposing sensitive payloads.
  • A critical workflow block, classification regression or authorization bypass triggers rollback.

Limit: Cloudflare documents the beta capability, not an accuracy rate for every MCP implementation. Test your own clients, transports, encrypted paths and server controls. Recheck the selector name and availability before each release while it remains beta.

Primary documentation

Keep learning

Continue this topic

Community discussion

Discuss: Cloudflare MCP Detection vs Crawler Control: A 24-Request Test

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.