Gemini Grounding: Trace Citations and Missing Source Evidence
Trace Gemini Search calls, source records and support mappings across two API surfaces, then stop structured output when provenance evidence is missing.
Direct answer: Preserve the Gemini API surface before you preserve its citations. The Interactions API records Google Search calls, results and inline URL annotations as timeline steps. generateContent returns separate grounding metadata, chunks and support mappings. They require different parsers and audit records.
This guide provides a 33-field citation trace for either surface. It keeps the raw answer, search activity, source URL, support review, rendering check and billing note connected without calling one grounded answer a Google ranking result. The current Google documentation was retrieved August 29, 2026.
The API surface determines the evidence shape
The old article described one generic set of search calls and citation spans. Google’s current documentation exposes two contracts. The newer Interactions API represents execution as ordered steps. The established generateContent surface returns a candidate plus groundingMetadata. A field with a similar name in one surface should not be assumed to share the other’s indexing, streaming or rendering behavior.
Store api_surface, API version and SDK language/version in every row. Freeze the raw response before extracting convenience properties such as combined output text. Google notes that a convenience output can omit earlier text blocks when non-text content or tools interrupt the sequence.
Interactions uses an observable step timeline
| Step or block | Current documented field | Audit use |
|---|---|---|
| Search call | google_search_call with executed query content | Preserve ordered retrieval activity. |
| Search result | google_search_result and search suggestions | Preserve the returned search layer. |
| Model output | Text content in a model_output step | Keep the exact displayed-source parent text. |
| URL citation | Annotation with URL, title, start and end indices | Connect one text range to one source record. |
| Terminal state | Interaction and step status | Separate completed, failed and partial traces. |
The timeline makes the call and result visible as separate events. That does not prove the source supports every generated claim. It proves what the API returned for this interaction under the recorded model, tools, project and time.
generateContent uses grounding metadata
| Field | What it records | Implementation boundary |
|---|---|---|
webSearchQueries | Queries used in the grounding flow | Preserve order and empty entries. |
searchEntryPoint | Rendered search-suggestion content | Follow Google’s display requirements. |
groundingChunks | Returned web source records | Keep stable indices across streamed chunks. |
groundingSupports | Segments mapped to chunk indices | Validate segment and referenced indices. |
groundingChunkIndices | Sources associated with one support segment | Association is not independent fact checking. |
Google’s API schema states that streamed grounding chunk indices refer across all responses and that the client must accumulate chunks in the same order. Dropping an early chunk can make a later support point to the wrong source even when the rendered answer still looks plausible.
Preserve the raw response before rendering
Save the complete response or interaction in a restricted evidence store. Hash the exact answer text and record the part or block index used for the citation. Then extract the fields into the public-safe trace. Do not retain credentials, authorization headers, private prompts or account identifiers in the downloadable CSV.
Transformations can break offsets. Trimming whitespace, normalizing Unicode, changing line endings, merging parts, inserting Markdown or converting HTML can shift the cited range. Render from the frozen source text or create and test an explicit offset map. Store the transformed text separately rather than overwriting the raw answer.
Validate indices under the correct contract
Record the offset-unit contract you actually implemented. Google’s generateContent schema describes segment start and end indices as byte offsets within a part, while the Interactions guide demonstrates slicing its annotation range through the client library. Do not copy one parser into the other surface without a fixture.
Test ASCII, accented characters, emoji, right-to-left text, line breaks and code. Reject negative, reversed, empty or out-of-range indices. Detect overlaps and adjacent citations explicitly. Compare the extracted cited text with the source response and keep a rendering-check result for every row.
Keep search queries in order
One prompt can lead to multiple Google Search queries. Preserve each query in order and record the total, including empty-query count where the response exposes it. A collapsed “search used” flag cannot explain why a subtopic was retrieved or why two runs produced different source sets.
Queries can contain sensitive user context. Store raw values only when the editorial or operational decision requires them and policy permits retention. The public trace can use a fixture ID, safe query class and a restricted raw-evidence reference. Do not publish confidential prompts to make a method appear more reproducible.
Resolve each source URL
Save the returned title and URL, then open the link and record the resolved destination and redirect state. A temporary or intermediary URL can hide the source owner. A page can also change after the model retrieved it, so the current destination is not proof of the exact content available at response time.
Deduplicate canonicals only after preserving the original record. Two returned URLs that resolve to one page are still two response records. An inaccessible source should remain in the trace with its failure status rather than being silently removed from the denominator.
Audit the source-claim relationship
A URL annotation or grounding support reports the association produced by the API. It does not certify that the source directly supports every statement in the mapped range. Open the source, locate the relevant passage and classify support as direct, qualified, conflicting, absent or inaccessible.
Use the citation-ready passage test for the claim review and the evidence-led publishing guide for source hierarchy. Keep the model’s association and the human review in separate fields so a later audit can disagree without rewriting the raw trace.
Stream without losing step or chunk order
The Interactions API streams specialized events for steps and deltas. generateContent streams response chunks whose grounding indices may depend on the accumulated order. Build separate assemblers and assert one terminal state. A UI that shows fluent text can still duplicate a delta, omit a source or attach a support record to the wrong part.
Record the first event time, terminal time and any incomplete state in the private trace. Re-run fixtures after SDK changes. Do not compare streaming and non-streaming citation counts unless the same assembly and inclusion rules are declared.
Separate one citation from search visibility
A grounded API response is an observation for one prompt, model, project, tool configuration and time. It is not a Google Search ranking, an aggregate Gemini citation share or evidence that the page will recur. A repeated visibility study needs a frozen prompt family, run count, environment and classification rubric.
Report citations, source support and referrals separately. The API trace does not show downstream visits unless another system records them. For measurement design, use the AI visibility measurement crosswalk.
Record the billing unit beside the run
Google currently documents query-based billing for Google Search grounding with Gemini 3: multiple executed queries in one prompt can count as multiple tool uses, while empty queries are ignored for that count. The page states that Gemini 2.5 and older models use per-prompt billing for this tool.
Pricing and model support can change. Store a billing-unit note and a link or private cost reference rather than hard-coding a price into the template. A query count is not the same as a final invoice line, and a cost observation does not prove citation quality.
Run a small parser fixture matrix
| Fixture | Expected parser behavior | Stop condition |
|---|---|---|
| ASCII answer | Extract every citation range exactly. | Any mismatch or missing source. |
| Unicode answer | Use the documented offset contract for that surface. | Range shifts after rendering. |
| Multiple queries | Preserve order and billable-unit note. | Queries collapse into one flag. |
| Streaming answer | Accumulate steps or chunks without duplication. | Index references change after assembly. |
| Redirecting source | Keep returned and resolved URLs. | Source owner becomes ambiguous. |
Structured JSON needs a provenance-completeness gate
Google now documents structured output with built-in tools as a preview for Gemini 3 models, including Google Search grounding. A response can satisfy the requested JSON schema while still lacking the source records needed to inspect its claims. Schema validity and provenance completeness are separate release decisions.
A Google AI Developers Forum thread contains user reports of groundingChunks or all groundingMetadata disappearing in some grounded responses, including structured-output workflows. Google has not confirmed a platform-wide incident in the located documentation, and SearchEngineAnswer has not reproduced the report. Treat it as a regression case to test, not a finding.
Record four states: source evidence present and resolvable; search activity present but source evidence absent; no search activity observed; request failed or incomplete. Do not collapse the middle two states into “not grounded.” They imply different debugging paths.
Fail closed when the answer outruns the evidence
function releaseGroundedJson(trace) {
if (!trace.schemaValid) return "hold_invalid_json";
if (!trace.searchInvoked) return "label_ungrounded";
if (!trace.sourceUrls.length) return "hold_missing_provenance";
if (!trace.supportMappings.length) return "hold_unmapped_claims";
return "review_sources_before_release";
}
This gate does not guess a URL from model text or replace missing metadata with a domain remembered by the model. It preserves the answer and tool evidence for diagnosis, then stops a source-dependent workflow before unsupported output reaches a reader.
Compare API surface and response mode before blaming the model
| API surface | Ordinary text | Structured JSON | Evidence to compare |
|---|---|---|---|
| Interactions | Record search-call, search-result and URL-annotation steps. | Validate the JSON and preserve the same step timeline. | Search invoked, result returned, annotation present, URL resolvable. |
generateContent | Record queries, chunks and support mappings. | Validate the JSON and preserve grounding metadata outside it. | Queries present, chunks present, supports mapped, URL resolvable. |
Use the same prompt classes, model version, SDK version and region in each cell. Repeat every cell enough times to report counts rather than a screenshot. A useful result separates valid JSON, search invocation, source evidence and claim support instead of reporting one pass rate.
Download the 33-field citation trace
Download the Gemini Google Search citation trace. Use one row per citation or grounding-support record. Replace both rows marked EXAMPLE-REMOVE, keep raw responses private and declare whether the row came from Interactions or generateContent.
The trace is an inspectable parser and evidence method, not a Gemini benchmark. It preserves enough context to reproduce the extraction, review the source and explain the editorial or release decision.
Current sources and limits
Primary sources: Google’s Grounding with Google Search guide, structured-output guide, Interactions migration guide, legacy grounding guide, and Gemini API changelog. Retrieved September 14, 2026.
SearchEngineAnswer did not execute live Gemini calls for this update. The new section is a release gate derived from the documented response contracts and a community regression lead, not a measured reliability result. Follow the Tools & Workflows archive for related implementation methods.
Keep learning
Continue this topic
Next in this topic
Gemini API Key Migration: Build a 30-Field Cutover Register
Earlier in this topic
Gemini Business Profile Changes: An Approval and Audit Workflow
Tools & Workflows
Ask a question or join the discussion