GPT-6 Prompt Caching: The Real Cost Math for AI Search Audits

Prompt caching can materially reduce repeated AI-search evaluation costs, but only when the shared prefix survives tool, schema and configuration changes.

Sonar the Answer Whale arranges stable prompt cards in a cache while one reordered tool card becomes a miss

Repeated AI-search audits are a perfect prompt-caching workload: the system instructions, scoring rubric, tool definitions and output schema remain stable while the query, brand or market changes. OpenAI’s GPT-6 caching update makes that design materially cheaper, but only when the shared prefix is truly identical.

The practical lesson is not “turn caching on.” Prompt caching is automatic. The work is to keep reusable instructions at the beginning, move changing data to the end and measure whether the cache survives the tool and configuration choices used by the audit.

Start with the bill, not the headline discount

OpenAI says cached GPT-6 input can receive discounts of up to 90%. “Up to” matters. A 90% cached-input discount does not mean a 90% lower total job cost because fresh input, cache writes and output tokens still count. The percentage saved depends on the share of input that can be reused.

Use this simple planning model:

Effective input cost = fresh-prefix cost + cached-prefix cost + changing-suffix cost.

Suppose an audit request contains 20,000 input tokens: a 16,000-token stable rubric and a 4,000-token changing case. If the 16,000-token prefix hits a tier with a 90% discount, the input-equivalent billed amount becomes 1,600 plus 4,000, or 5,600 tokens. That is a 72% input-side reduction, before output. If only half the supposed prefix matches, the saving is much smaller.

This is an illustrative calculation, not a quote for a specific model or account. Apply the current price card and dashboard data to your own token mix. Our SEO reporting workflow also separates acquisition cost from interpretation cost because long analytical output can remain the dominant expense.

Design a stable prefix for audit work

A strong reusable prefix contains rules that should not vary between rows in a test:

  • the evaluator’s role and decision standard;
  • definitions for citation, mention, recommendation and factual support;
  • the response schema and allowed labels;
  • tool definitions in a fixed order;
  • examples that apply to every case;
  • instructions for uncertainty and missing evidence.

Place volatile material after that block: the current query, captured answer, source list, locale, device, timestamp and brand under test. A timestamp near the top of every prompt can invalidate what should have been a large shared prefix. The same is true when a client ID, random request identifier or reordered JSON object appears before stable instructions.

Four ways an evaluation silently loses its cache

  1. Tool order changes. Reconstructing tool definitions from an unordered collection can produce a different prefix even when the tools are functionally identical.
  2. Schemas drift. Adding an optional property or rewriting a description changes the serialized bytes used for matching.
  3. Examples are personalized too early. A brand-specific example inside the shared rubric splits one evaluation family into many prefixes.
  4. Configuration is rewritten. OpenAI says reasoning effort can be changed by appending a configuration_update without breaking the cache. Rebuilding the earlier request is less cache-friendly.

OpenAI also recommends using allowed_tools or tool_choice: none when the available tool set changes, rather than changing the definitions themselves. This is valuable for AI visibility tests that sometimes browse and sometimes score an already captured answer.

Measure one evaluation family at a time

Do not average every production request into one cache-hit rate. Separate workloads by stable prefix: citation audits, answer-quality grading, source classification and report generation are different families. Record input tokens, cached input tokens, output tokens, latency, model, prompt version and tool-set version for every run.

OpenAI introduced a Prompt Caching Dashboard, a diagnostic tool for identifying where prefixes break, explicit cache breakpoints and prewarming. Those features make a three-step validation possible:

  1. run one cold request and confirm the intended breakpoint;
  2. repeat the same request inside the eligible window and confirm a hit;
  3. change only the dynamic suffix and verify that the shared prefix remains cached.

The documented reuse window is 30 minutes for eligible shared prefixes. That creates an operational choice: batch comparable evaluations closely enough to reuse the prefix, but do not bunch them so aggressively that rate limits or downstream review quality suffer.

What the published customer result does and does not prove

OpenAI reports that one customer moved an evaluation cache-hit rate from 83% to 91%, reduced cache writes by roughly two-thirds and cut inference cost by 36%. This is useful evidence that prompt organization can matter in a real evaluation workload. It is not a guaranteed outcome for every audit system.

The result combines a particular prompt shape, traffic pattern and token mix. A team with short prompts or mostly novel context may save far less. A team with a very large fixed rubric may save more on input while seeing little change in total spend because outputs dominate.

A one-day implementation plan

Choose the highest-volume evaluation family. Save ten real requests. Diff the serialized prefixes byte for byte. Move dynamic values below a named breakpoint, freeze tool ordering and version the stable rubric. Then rerun the ten requests in a controlled batch and compare cached-input share, total cost and answer quality.

Quality belongs in the comparison. A cheaper evaluation is not an improvement if a shortened rubric changes judgments. Keep a fixed validation set and compare labels before and after the prompt reorganization.

The primary documentation is OpenAI’s GPT-6 prompt caching update. Treat its discount ceiling and customer example as inputs to your measurement plan, not as a forecast. For the broader workflow, use our AI search analytics guide to keep prompt version, source capture and result interpretation in the same audit record.

Keep learning

Continue this topic

Community discussion

Discuss: GPT-6 Prompt Caching: The Real Cost Math for AI Search Audits

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.