Cloudflare AI Search Is GA: What the New Retrieval Pricing Buys
Cloudflare AI Search is generally available with image embeddings, OCR, hybrid retrieval and usage pricing. This worksheet separates the cost and quality decisions.
Cloudflare AI Search became generally available on October 1, 2026. The release adds native image processing, document OCR, larger file support, and a new retrieval-based pricing model that begins billing on November 1. The practical question is no longer only whether the product can index a knowledge base. Teams now need to decide which retrieval path a question requires and what each path costs at their expected volume.
Short answer: exact identifiers and quoted strings belong on the full-text lane; conceptual questions and visual similarity belong on the semantic lane. Hybrid search combines vector and keyword retrieval and is billed at the semantic-query rate.
What changed at general availability
Cloudflare says AI Search can now create embeddings for images, generate image captions, extract text from supported documents with optical character recognition, and ingest files up to 10 MiB. The product combines semantic vector retrieval with full-text retrieval and can use both in a hybrid request.
| Question pattern | First lane to test | Reason | Failure to watch |
|---|---|---|---|
| Order ID, SKU, error code, quoted phrase | Full text | Exact tokens carry the meaning | Vector similarity can blur a precise identifier |
| Concept, paraphrase, broad troubleshooting question | Semantic | Meaning can survive different wording | Near-topic passages may outrank the decisive sentence |
| Screenshot, diagram, product image | Image-aware semantic | Visual embeddings and captions create searchable signals | Generated captions can omit a critical detail |
| Scanned PDF or photographed page | OCR plus semantic or full text | Text must first be extracted from pixels | OCR errors become retrieval errors |
| Mixed or uncertain intent | Hybrid | Combines exact and conceptual recall | Semantic-rate billing and harder ranking diagnosis |
The published prices turn architecture into a budget
Cloudflare lists base ingestion at $0.75 per million tokens. Image processing adds $0.50 per million input tokens. Stored index data costs $2 per GB-month. Semantic retrieval costs $0.75 per 1,000 queries, while full-text retrieval costs $0.10 per 1,000 queries. The monthly free allocation includes 5 million ingested tokens, 10 GB of storage, 1,000 semantic queries, and 1,000 full-text queries.
| Workload assumption | Calculation after free allowance | Illustrative charge |
|---|---|---|
| 4 million text tokens ingested | Below the 5 million-token allowance | $0.00 base ingestion |
| 12 million text tokens ingested | (12M − 5M) × $0.75/M | $5.25 base ingestion |
| 100,000 semantic queries | (100,000 − 1,000) ÷ 1,000 × $0.75 | $74.25 |
| 100,000 full-text queries | (100,000 − 1,000) ÷ 1,000 × $0.10 | $9.90 |
| 20 GB-month stored | (20 GB − 10 GB) × $2 | $20.00 |
These are arithmetic examples, not a quote for a particular site. They exclude image-token volume, repeated re-ingestion, application-model costs, gateway fees, taxes, and any provider-specific charges outside AI Search. The official billing meter, not a content word count, determines the final amount.
A cost worksheet that exposes the real levers
- Inventory inputs. Count text tokens, image inputs, documents, and the resulting stored index size.
- Separate retrieval types. Estimate full-text requests separately from semantic requests, including hybrid and vector search in the semantic meter.
- Subtract the monthly allowance. Apply free units to the correct meter rather than to one blended total.
- Add re-indexing. Include content churn, deletion, scheduled refreshes, and the first import.
- Measure useful retrieval. Divide spend by supported answers or completed tasks, not by raw requests alone.
The wide price gap between semantic and full-text retrieval is a design signal. Sending every identifier lookup through semantic search can increase cost and reduce precision. Sending every natural-language question through full-text can miss paraphrases. Route obvious cases first, then reserve hybrid search for questions where the added recall changes the answer.
A useful budget has three query scenarios instead of one average: routine traffic, a launch-day peak, and a failure case in which retries multiply requests. Add a cache-hit assumption only after measuring it. If 40 percent of repeated questions are served from an application cache, that may reduce retrieval calls, but it also creates a freshness obligation. Record the cache key, source version, and expiry so cost savings do not turn into stale answers.
Compare hybrid, semantic-only, and full-text-only retrieval in the test environment and calculate useful source chunks per dollar. Cloudflare lists hybrid search in the semantic query meter, not as the sum of a semantic and a full-text charge. A cheaper lane that misses the decisive evidence is not efficient. A more expensive lane that adds duplicate passages is not automatically better.
Multimodal retrieval needs its own quality checks.
Native image support and OCR expand the searchable corpus, but they also add intermediate transformations. A diagram may receive a generated caption. A scanned page becomes OCR text. The retriever then ranks those derived representations, not the original visual meaning directly.
- Keep the original file and a hash so an extracted chunk can be traced back to its source.
- Sample OCR output by document type, language, layout, and scan quality.
- Review image captions for omitted labels, quantities, and relationships.
- Store page or region coordinates when the source format supports them.
- Test retrieval with exact questions and paraphrases before connecting it to an answer model.
This is especially important for publishers and support teams that expect an answer to carry a reliable citation. The AI citation failure-layer guide explains why a source can be available but still disappear between retrieval and final output.
A release gate for production retrieval
| Gate | Pass evidence |
|---|---|
| Coverage | Known-answer questions retrieve the decisive source across text, image, and scanned-document sets |
| Traceability | Every returned chunk maps to a stable source URL or file location |
| Precision | Exact identifiers are not displaced by semantically similar noise |
| Freshness | Changed and deleted content follows a measured update window |
| Cost | Semantic, full-text, storage, and ingestion meters reconcile with the workload |
| Failure handling | OCR errors, oversized files, empty results, and provider outages produce visible states |
SearchEngineAnswer did not run a live Cloudflare AI Search quality benchmark for this report. The pricing math uses Cloudflare’s published rates, and the retrieval recommendations are an implementation framework to test. Retrieval accuracy, latency, and index behavior can vary with corpus, chunking, language, and configuration.
Primary source: Cloudflare, “Cloudflare AI Search is now generally available”.
Keep learning
Continue this topic
Next in this topic
How to Test Tavily’s Language Boosting and Strict Filtering
Earlier in this topic
How to Use Generative AI for Website Content: A 30-Point Audit Checklist
Tools & Workflows
Ask a question or join the discussion