The Same AI Search Engine Can Cite Different Sources on Web and App

A 160,860-record Chinese-language study shows why interface belongs in every AI visibility observation.

Sonar the Answer Whale follows one violet query into web and app branches that return visibly different source-card bundles.

Direct answer: a 2026 study of four Chinese-language generative search platforms found that the web and app interfaces of the same platform systematically returned different source sets. Interface is therefore part of the observation—not a cosmetic detail to omit from an AI visibility report.

The study covered eight platform-interface combinations, 614 queries and three replications per combination. It produced 214,119 raw records and a cleaned citation-level dataset of 160,860 records.

Interface is a measurement variable

Visibility tools often store platform, prompt and date. This paper shows why that schema is incomplete. A web session and an app session can differ in retrieval infrastructure, account context, location signals, source display and product constraints. Two observations with the same platform name are not necessarily repeats of the same system state.

Add these fields to every observation:

  • platform and product name;
  • interface: web, iOS, Android or another client;
  • account and subscription state;
  • location and language;
  • exact prompt and conversation state;
  • timestamp and replication number;
  • answer text, displayed citations and resolved URLs;
  • whether the citation was visible, expandable or only discoverable after interaction.

Without those fields, a team can interpret interface drift as ranking volatility or a failed optimization.

Selection and answer use also diverged

Bounded findings from the cleaned citation-level dataset
FindingObserved resultBoundary
Brand selection8.3% of brands in the citation pool were surfaced in answersPool presence did not guarantee answer exposure
Contact carry-through12.4% of retrieved sources containing contact information contributed it to answersRetrieval did not guarantee absorption
Freshness half-lifeAbout 39 days for high-timeliness and 68 days for low-timeliness queriesModelled among cited pages with dates
Unmatched exposure~13% of brand exposures and ~71% of contact exposures lacked the expected contemporaneous matchPossible other sources, extraction limits or generation effects

These results reinforce the distinction between citation count and citation absorption. A source can be retrieved without its brand or contact information reaching the answer. An answer can also surface information that the researchers could not match to the contemporaneous citation pool or crawled body text.

A cross-interface test you can run

  1. Freeze 20 to 50 prompts across navigational, factual, comparative and time-sensitive intents.
  2. Run each prompt three times on web and app in a short controlled window.
  3. Normalize citation URLs without discarding the original displayed URL.
  4. Calculate source-set overlap with Jaccard similarity: intersection divided by union.
  5. Record brand mention, passage use and contact-information use separately from citation presence.
  6. Repeat on a fixed schedule and preserve every zero.

Report the distribution, not only the average. A platform can have high overlap on stable factual prompts and low overlap on current or local questions. Also separate interface drift from run-to-run instability inside one interface.

Our review of 45 GEO studies explains why repeated runs and clearly defined pipeline stages are necessary. The AI visibility tools buyer matrix shows what evidence to demand from commercial trackers.

Limits before generalizing

This was a Chinese-language study of four mainstream platforms and their tested web/app interfaces in July 2026. It does not prove the same magnitude of drift for ChatGPT, Gemini, Perplexity, English-language queries or future product versions. The cleaned dataset also depends on URL resolution, crawling and matching rules.

The transferable conclusion is narrower and useful: record the interface, replicate observations, and do not collapse distinct clients into one platform score.

Primary source

Zhen et al.: What Do Chinese-Language Generative Search Engines Cite and Surface?

Status: 49-page arXiv preprint submitted July 17, 2026. The figures above are descriptive findings from the authors’ stated design, not universal benchmarks.

Community discussion

Discuss: The Same AI Search Engine Can Cite Different Sources on Web and App

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.