Cloudflare Web Search API Makes Crawler Rules Part of the Contract

Cloudflare's Web Search API requires provider crawler identity, robots.txt compliance and source links. Here is what publishers and AI teams can verify.

Sonar the Answer Whale checks identity, robots rules and source links across three search-provider streams

Cloudflare launched a Web Search API on October 2, 2026, with an unusual product requirement: search providers must identify their crawlers and respect publishers’ robots.txt rules, and return links to the sources behind their results. The first providers are Ceramic.ai, Exa, and Linkup. Developers reach them through Cloudflare AI Gateway rather than integrating three separate search endpoints.

This is useful infrastructure news, but it is not proof that publishers will receive more referrals or citations. The real change is contractual and operational. Crawler identity, publisher controls, source links, request logging, and provider choice are being placed in one auditable path.

The contract is the product difference.

Search APIs usually compete on index size, latency, freshness, and retrieval quality. Cloudflare adds a publisher-facing condition to that list. Its announcement says participating search providers must meet Cloudflare’s Verified Bot standards, follow robots.txtand supply source links with results.

What the Web Search API contract covers, and what still needs testing
LayerCloudflare statesEvidence a team should retainNot established by launch
Crawler identityProviders must meet Verified Bot standardsVerified user agent, published IP or identity method, server-log sampleEvery crawl is correctly classified by every firewall
Publisher rulesProviders must respect robots.txtRule in force, request time, affected bot, response codeA page will appear after access is allowed
Source linksResults include links to supporting sourcesReturned URL, title, rank, query, provider, timestampThe downstream AI answer will cite or send traffic to the source
Gateway controlsAI Gateway provides logs, billing, access controls and provider routingRequest ID, provider, latency, cost, retention setting, statusAll providers return equivalent results for the same query
Data handlingCloudflare marks provider features such as zero-data-retention support and BYOKAccount setting, provider policy, key ownership, contract versionA single label covers every downstream processor or use

Discovery, retrieval, and citation remain separate events

A source link in a search API response is valuable provenance. It still sits early in a longer chain. A provider discovers a page, decides that it matches a query, returns it to an application, and the application may then read, summarize, cite, omit, or replace it. A publisher referral happens only if a person or agent follows the link.

  1. Access: the provider’s crawler can request the page.
  2. Retrieval: the page enters a result set for a defined query.
  3. Selection: the calling application chooses the result for further use.
  4. Attribution: the output presents a source link or other visible credit.
  5. Referral: a user follows that attribution to the publisher.

Do not collapse those states into a single “AI visibility” number. The AI crawler guide explains the access layer. This API adds better evidence for the retrieval layer, while citation and referral still need their own observations.

A provider comparison should use one query panel.

The launch gives teams a practical way to compare three search providers without rebuilding the surrounding gateway. A defensible comparison keeps the prompt panel, locale, date, result count, and scoring rules fixed. It should not cherry-pick one provider’s best query and another provider’s worst.

Minimum fields for a reproducible provider test
FieldWhy it matters
Query and intent classSeparates navigational, current-event, local, technical, and research needs
Provider and model routePrevents a blended gateway result from being assigned to the wrong index
Run time and localeFreshness and regional results can change quickly
Returned URLs and ranksMakes overlap, domain diversity, and source reuse measurable
HTTP status and latencyDistinguishes retrieval quality from request failure or timeout
Source-link completenessChecks whether each substantive result can be traced to a publisher URL
Independent relevance scoreStops result count from becoming a proxy for usefulness

SearchEngineAnswer has not run that provider panel because the required account access was not available for this report. The table is a test protocol, not a set of performance results. Cloudflare’s launch post identifies the participating providers and their contract conditions, but it does not publish a common relevance benchmark.

When results are available, calculate domain overlap and unique-source coverage separately. High overlap can indicate agreement, but it can also show that every provider leans on the same small group of domains. Unique-source coverage reveals whether a provider expands the evidence set. Relevance scoring then determines whether that expansion is useful, not just different.

What publishers can verify now?

  • Keep named AI and search crawler rules explicit instead of relying on one blanket directive.
  • Confirm the effective robots.txt response at the public hostname, including CDN and bot-management behavior.
  • Record verified crawler requests separately from unverified user agents that copy the same name.
  • Use canonical, indexable source URLs and stable anchors so a returned link lands on the intended evidence.
  • Measure source inclusion, visible citation, and referral as separate outcomes.
  • Retain the provider, query, time, and returned URL when investigating an AI answer.

For application teams, the immediate benefit is routing and observability. AI Gateway can centralize credentials, spend controls, request logs, and provider selection. For publishers, the important signal is that access rules and source links are written into the provider relationship. Both are improvements in accountability. Neither guarantees ranking, citation, or traffic.

The next evidence to watch

Three measurements would make this launch more consequential: a documented common-query comparison across providers, crawler-log evidence showing consistent rule enforcement, and downstream applications that preserve source links in the final answer. Until those exist, treat the Web Search API as a cleaner retrieval contract rather than a new organic visibility channel.

Primary source: Cloudflare, “Introducing Cloudflare’s Web Search API”.

Keep learning

Continue this topic

Community discussion

Discuss: Cloudflare Web Search API Makes Crawler Rules Part of the Contract

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.