Cloudflare AI Search Is GA: What the New Retrieval Pricing Buys

Cloudflare AI Search is generally available with image embeddings, OCR, hybrid retrieval and usage pricing. This worksheet separates the cost and quality decisions.

Sonar the Answer Whale routes images, scanned pages and text through keyword and semantic retrieval lanes

Cloudflare AI Search became generally available on October 1, 2026. The release adds native image processing, document OCR, larger file support, and a new retrieval-based pricing model that begins billing on November 1. The practical question is no longer only whether the product can index a knowledge base. Teams now need to decide which retrieval path a question requires and what each path costs at their expected volume.

Short answer: exact identifiers and quoted strings belong on the full-text lane; conceptual questions and visual similarity belong on the semantic lane. Hybrid search combines vector and keyword retrieval and is billed at the semantic-query rate.

What changed at general availability

Cloudflare says AI Search can now create embeddings for images, generate image captions, extract text from supported documents with optical character recognition, and ingest files up to 10 MiB. The product combines semantic vector retrieval with full-text retrieval and can use both in a hybrid request.

Match the reader’s question to the retrieval lane
Question patternFirst lane to testReasonFailure to watch
Order ID, SKU, error code, quoted phraseFull textExact tokens carry the meaningVector similarity can blur a precise identifier
Concept, paraphrase, broad troubleshooting questionSemanticMeaning can survive different wordingNear-topic passages may outrank the decisive sentence
Screenshot, diagram, product imageImage-aware semanticVisual embeddings and captions create searchable signalsGenerated captions can omit a critical detail
Scanned PDF or photographed pageOCR plus semantic or full textText must first be extracted from pixelsOCR errors become retrieval errors
Mixed or uncertain intentHybridCombines exact and conceptual recallSemantic-rate billing and harder ranking diagnosis

The published prices turn architecture into a budget

Cloudflare lists base ingestion at $0.75 per million tokens. Image processing adds $0.50 per million input tokens. Stored index data costs $2 per GB-month. Semantic retrieval costs $0.75 per 1,000 queries, while full-text retrieval costs $0.10 per 1,000 queries. The monthly free allocation includes 5 million ingested tokens, 10 GB of storage, 1,000 semantic queries, and 1,000 full-text queries.

Illustrative monthly calculations using Cloudflare’s published prices
Workload assumptionCalculation after free allowanceIllustrative charge
4 million text tokens ingestedBelow the 5 million-token allowance$0.00 base ingestion
12 million text tokens ingested(12M − 5M) × $0.75/M$5.25 base ingestion
100,000 semantic queries(100,000 − 1,000) ÷ 1,000 × $0.75$74.25
100,000 full-text queries(100,000 − 1,000) ÷ 1,000 × $0.10$9.90
20 GB-month stored(20 GB − 10 GB) × $2$20.00

These are arithmetic examples, not a quote for a particular site. They exclude image-token volume, repeated re-ingestion, application-model costs, gateway fees, taxes, and any provider-specific charges outside AI Search. The official billing meter, not a content word count, determines the final amount.

A cost worksheet that exposes the real levers

  1. Inventory inputs. Count text tokens, image inputs, documents, and the resulting stored index size.
  2. Separate retrieval types. Estimate full-text requests separately from semantic requests, including hybrid and vector search in the semantic meter.
  3. Subtract the monthly allowance. Apply free units to the correct meter rather than to one blended total.
  4. Add re-indexing. Include content churn, deletion, scheduled refreshes, and the first import.
  5. Measure useful retrieval. Divide spend by supported answers or completed tasks, not by raw requests alone.

The wide price gap between semantic and full-text retrieval is a design signal. Sending every identifier lookup through semantic search can increase cost and reduce precision. Sending every natural-language question through full-text can miss paraphrases. Route obvious cases first, then reserve hybrid search for questions where the added recall changes the answer.

A useful budget has three query scenarios instead of one average: routine traffic, a launch-day peak, and a failure case in which retries multiply requests. Add a cache-hit assumption only after measuring it. If 40 percent of repeated questions are served from an application cache, that may reduce retrieval calls, but it also creates a freshness obligation. Record the cache key, source version, and expiry so cost savings do not turn into stale answers.

Compare hybrid, semantic-only, and full-text-only retrieval in the test environment and calculate useful source chunks per dollar. Cloudflare lists hybrid search in the semantic query meter, not as the sum of a semantic and a full-text charge. A cheaper lane that misses the decisive evidence is not efficient. A more expensive lane that adds duplicate passages is not automatically better.

Multimodal retrieval needs its own quality checks.

Native image support and OCR expand the searchable corpus, but they also add intermediate transformations. A diagram may receive a generated caption. A scanned page becomes OCR text. The retriever then ranks those derived representations, not the original visual meaning directly.

  • Keep the original file and a hash so an extracted chunk can be traced back to its source.
  • Sample OCR output by document type, language, layout, and scan quality.
  • Review image captions for omitted labels, quantities, and relationships.
  • Store page or region coordinates when the source format supports them.
  • Test retrieval with exact questions and paraphrases before connecting it to an answer model.

This is especially important for publishers and support teams that expect an answer to carry a reliable citation. The AI citation failure-layer guide explains why a source can be available but still disappear between retrieval and final output.

A release gate for production retrieval

Do not approve the system on latency alone
GatePass evidence
CoverageKnown-answer questions retrieve the decisive source across text, image, and scanned-document sets
TraceabilityEvery returned chunk maps to a stable source URL or file location
PrecisionExact identifiers are not displaced by semantically similar noise
FreshnessChanged and deleted content follows a measured update window
CostSemantic, full-text, storage, and ingestion meters reconcile with the workload
Failure handlingOCR errors, oversized files, empty results, and provider outages produce visible states

SearchEngineAnswer did not run a live Cloudflare AI Search quality benchmark for this report. The pricing math uses Cloudflare’s published rates, and the retrieval recommendations are an implementation framework to test. Retrieval accuracy, latency, and index behavior can vary with corpus, chunking, language, and configuration.

Primary source: Cloudflare, “Cloudflare AI Search is now generally available”.

Keep learning

Continue this topic

Community discussion

Discuss: Cloudflare AI Search Is GA: What the New Retrieval Pricing Buys

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.