AI Search Citation Drift: Test Web and App Separately
Measure web/app citation drift with paired prompts, repeated runs, preserved URLs, overlap distributions, answer-use labels, change control, and a reusable observation ledger.
Direct answer: an AI search platform’s web and app clients can return different citation sets for the same prompt. That makes interface a measurement variable, not a presentation detail. A defensible visibility study records client, account state, prompt state, time, run number, displayed citations, resolved URLs, and answer use before comparing results.
The strongest current evidence is a July 2026 Chinese-language study covering four generative-search platforms, eight platform-interface combinations, 614 queries, and three replications per query-platform-interface combination. The authors produced 214,119 raw records and 160,860 cleaned citation records. Their result is not a universal benchmark for every platform or language, but it is enough to invalidate a common shortcut: pooling web and app observations under one platform label.
The decision this guide helps you make
Use this method when a dashboard reports that a brand “lost citations,” when two analysts cannot reproduce each other’s results, or when an app test contradicts a web test. The immediate question is not which interface is correct. It is whether the apparent change survives a controlled comparison within each interface.
Keep three effects separate: cross-interface drift, repeat-to-repeat instability inside one interface, and true change over time. They require different denominators and different corrective actions.
What the study established—and what it did not
| Finding | Reported result | Boundary |
|---|---|---|
| Design | 8 platform-interface combinations, 614 queries, 3 replications | Four Chinese-language platforms in the authors’ test period |
| Citation data | 214,119 raw records; 160,860 cleaned citation records | Cleaning depends on the authors’ URL and matching rules |
| Brand selection | 8.3% overall | Presence in the citation pool did not imply answer exposure |
| Contact carry-through | 12.4% of retrieved sources containing contact details contributed them to answers | Retrieval and answer use are different stages |
| Interface result | Source sets differed systematically between app and web | No universal effect size for other products or languages |
The paper also reports fitted half-lives of about 39 days for high-timeliness queries and 68 days for low-timeliness queries among cited pages with publication dates. Approximately 13% of brand exposures did not match the contemporaneous citation pool, while about 71% of contact-information exposures did not match crawled body text. These observations are useful for designing audits, not for declaring that unmatched output is fabricated. Other sources, incomplete captures, extraction limits, and generation behavior remain plausible explanations.
Define the observation unit before collecting answers
One observation should represent one prompt run in one declared environment. Do not store only the final answer and platform name. Freeze the fields below before the first run so analysts cannot quietly repair missing context after seeing results.
| Layer | Required fields | Why it matters |
|---|---|---|
| Environment | Platform, product, web/iOS/Android, app or browser version | Separates client effects from platform-level claims |
| Session | Account tier, signed-in state, location, language, conversation state | Captures personalization and context differences |
| Prompt | Exact text, prompt class, fixture ID, run number | Makes repeated comparison possible |
| Output | Answer text or hash, displayed citations, resolved URLs, visibility state | Preserves what the user actually saw |
| Use | Brand mention, claim support, contact use, uncertainty | Prevents citation presence from standing in for answer use |
Build a balanced prompt fixture
Start with 20 to 50 prompts if the test is manual. Include stable factual, comparative, navigational, local, and time-sensitive intents. Keep wording identical across clients. Record any interface that rewrites, expands, or refuses the prompt. If a product offers multiple research modes, treat each mode as another environment rather than mixing it into a client average.
Three repetitions per cell are a practical floor, not a guarantee of stability. Randomize the run order where possible. Run paired web and app tests inside a short time window so a news event or product release is less likely to explain the difference.
Preserve displayed and resolved citations
Save the label and URL exactly as displayed before normalization. Then follow redirects and store the resolved URL separately. Normalize tracking parameters and obvious URL variants for comparison, but retain the original so a reviewer can reconstruct the user experience. Record whether a citation appeared inline, in a source drawer, behind an expansion, or only after another interaction.
A screenshot can help with presentation evidence, but it is not the dataset. Screenshots are difficult to compare at scale and may omit redirect targets. Pair them with structured rows in the downloadable interface-drift observation ledger. Every included row is marked EXAMPLE-REMOVE and must be deleted before a real study.
Calculate overlap without hiding zeros
For each paired run, calculate Jaccard similarity as the size of the source-set intersection divided by the size of the union. A score of 1 means identical normalized source sets; 0 means no overlap. If both sets are empty, report an explicit empty-pair status instead of silently assigning 1 or dropping the row.
| Metric | Question | Do not substitute it for |
|---|---|---|
| Jaccard overlap | How similar are two citation sets? | Answer accuracy or source quality |
| Within-client repeat overlap | How stable is one interface? | Cross-interface difference |
| Brand selection rate | How often did a cited-source brand reach the answer? | Citation share |
| Claim support rate | How often did a cited passage support the answer claim? | Link presence |
| Contact carry-through | Did retrieved contact information reach the answer? | Lead quality or conversion |
Separate interface drift from ordinary instability
Compare web runs with other web runs and app runs with other app runs before comparing web with app. If each client is highly unstable, low cross-interface overlap may not be a client effect. A useful report shows the distribution of paired web/app overlap alongside both within-client distributions, not one blended average.
Stratify by prompt class. Stable factual questions may converge while time-sensitive or local prompts diverge. Report medians, quartiles, empty-result rates, and the number of valid pairs. The broader review of 45 GEO studies explains why repeated runs and explicit pipeline stages matter.
Do not confuse retrieval, citation, and absorption
A URL can be retrieved but not displayed. A displayed citation can fail to support a specific statement. A page can be cited without its brand, statistic, or contact detail being used. Code each stage separately. Our citation absorption guide supplies the claim-level distinction.
Use an adjudication field for ambiguous cases. A second reviewer should inspect disagreements involving brand identity, syndicated copies, paraphrase support, or contact details. Preserve the reviewer and decision date.
Add change control to longitudinal tracking
Record app version, browser, account tier, mode, location, and test date at every wave. When a platform changes its citation presentation, create a new measurement epoch. Do not join incompatible capture methods into an unbroken trend line. Keep a fixture version and note prompt additions, deletions, or wording changes.
Commercial tools should expose equivalent controls. The AI visibility tools buyer guide lists evidence questions to ask before accepting a platform score.
API repetition is another drift layer
A September 2026 study by Prefer adds a useful control case. It ran 80 questions three times across ChatGPT, Gemini, Claude and Perplexity APIs in one 16-minute window. The reported repeat-run source overlap ranged from 35% for Gemini and 37% for ChatGPT to 68% for Claude and 96% for Perplexity.
Because the collection window was short, the differences are unlikely to be explained only by a changed news cycle. They show that ordinary repeated retrieval can create substantial source-set movement before web-versus-app differences are considered.
Add a within-interface repeat baseline to every paired client test:
- Run each prompt at least three times in the web interface.
- Run the same repetitions in the app inside the same short window.
- Calculate overlap within web and within app before calculating cross-interface overlap.
- Flag an interface effect only when the paired gap is larger than ordinary repeat-to-repeat movement.
The Prefer test used APIs, US-English questions and a topic mix weighted toward AI search, GEO and AEO. It does not estimate consumer app drift, but it demonstrates why a single run cannot serve as the within-client baseline. See the aggregate study.
Write decision rules before reading the result
- Investigate capture quality when missing fields exceed the declared threshold.
- Do not call an interface effect when within-interface instability is of similar magnitude.
- Escalate a business-impact review only when the citation shift also changes answer use or an observable outcome.
- Repeat the fixture after a documented product or interface change.
- Keep a “no conclusion” state when the sample or pairing is insufficient.
Publish a reviewable result
A useful report names platforms and interfaces, dates, language, location, account state, prompt fixture, repetitions, capture method, URL normalization, empty-set handling, overlap distributions, within-client stability, answer-use labels, missing data, and limitations. Attach the ledger or a privacy-safe extract. State whether conclusions concern the tested clients or the platform generally.
This article does not provide a universal drift threshold. It translates one preprint’s strong measurement warning into a reusable protocol. The source is Zhen et al., “What Do Chinese-Language Generative Search Engines Cite and Surface?”, submitted July 17, 2026. Treat its figures as descriptive findings under the authors’ design.
The practical next step
Run one small paired fixture before buying more tracking coverage. If the result changes materially when web and app are separated, add interface to every collection schema, dashboard filter, alert, and editorial interpretation. If it does not, keep recording the field anyway: the absence of drift in one wave is evidence, not permission to stop measuring it.
Keep learning
Continue this topic
Next in this topic
AI Overview Click Study: What the 1% Means
Earlier in this topic
ChatGPT Outbound Clicks: 5.2% Is Not CTR
Research
Ask a question or join the discussion