|

ChatGPT and Gemini Shared No Source Domains in 76.7% of Product Checks

A product-advice preprint found no shared displayed source domains in 76.7% of matched ChatGPT and Gemini checks. What the test means for AI visibility audits.

Sonar the Answer Whale compares two different product recommendations emerging from the same search prompt.

A product recommendation is not a stable inventory of what an AI assistant “knows.” In a September 2026 preprint, ChatGPT and Gemini displayed no source domains in common in 76.7% of matched shopping-answer comparisons. Even within one service, repeating a question changed the visible sources and products. For retailers and publishers, a single prompt screenshot is therefore a poor substitute for a repeatable visibility audit.

The study, “If I Had to Buy Just ONE…”, examines commercial advice rather than all search behavior. Its most useful contribution is not a league table of winning brands. It shows exactly which parts of a shopping answer can vary: the response, the products named, the websites displayed and the path between a search result and a final recommendation.

What the researchers actually compared

The authors built ConsumerQ from 2,528 user-written commercial-advice questions drawn from two public chat datasets. They manually selected 117 broad questions about physical products for controlled testing. Each was repeated three times across five conditions: logged-out ChatGPT and Gemini consumer interfaces, their corresponding APIs, and Google AI Overviews. The result was 1,755 attempted observations and 1,536 substantive answers.

The controlled runs took place September 4–5, 2026, using residential IP addresses in the Netherlands, fresh browser profiles and an English-language locale. This matters. The findings describe that test environment and those product questions, not a census of every account, market or future model version.

The source divergence is larger than a different ranking order

For a matched question and repetition, the average overlap between source domains displayed by ChatGPT and Gemini was 5.4%. In 76.7% of those comparisons, the two interfaces shared no displayed source domain at all. The study used a set-overlap measure, so a low percentage means different domain sets, not a measured decline in traffic to any particular publisher.

Repeats changed the answer within a service too. Average displayed-domain overlap across the three repetitions was 26.0% for ChatGPT, 29.8% for Gemini and 45.9% for Google AI Overviews. Product overlap was also limited: 0.178 for ChatGPT, 0.287 for Gemini and 0.421 for AI Overviews on the study’s measure. A retailer could be cited in one capture and absent in the next without changing its own page.

One practical consequence is that a campaign report must distinguish “appeared at least once” from “appeared consistently.” The former is an opportunity signal. The latter requires repeated, comparable observations. Neither, by itself, demonstrates revenue or even a visit.

Why an API dashboard can disagree with the consumer app

The researchers also compared consumer-interface and API answers. Average source-domain overlap was 12.0% for ChatGPT and 14.8% for Gemini. In 60.9% of matched ChatGPT pairs and 43.4% of matched Gemini pairs, there was no shared source domain. That is a warning for anyone treating an API-based prompt tracker as a direct record of what a shopper saw in the public app.

It is not proof that the interface alone caused the difference. The logged-out ChatGPT model version could not be verified, while the API used an approximation; retrieval settings and presentation can differ as well. A responsible dashboard should label its collection surface and model configuration rather than merging the two into one “ChatGPT visibility” score.

The wording can make a recommendation sound more certain than its evidence

Among product-recommending answers, ChatGPT used first-person preference language in 79% of cases, compared with 7% for Gemini and 2% for Google AI Overviews. That stylistic difference does not establish that one system tested the products more thoroughly. A conversational “I would buy” can be useful to a reader, but it should not be interpreted as first-hand product use.

For editorial teams, the distinction is especially important. A cited review, a merchant listing and an assistant’s preferred product are three different things. If a page earns a citation, inspect what claim the citation supports before announcing that the engine recommends the brand.

A small audit that can reveal the real pattern

Choose ten queries that reflect distinct buying decisions, not ten phrasings of the same one. Capture each query three times per service over several days. Keep the country, language, account state, device and interface fixed within a comparison. Save the full answer and every visible source URL. Resolve redirects, then record both the displayed domain and the final destination.

For each answer, mark four separate events: the brand is mentioned; a product is recommended; a source is cited; and a visit is actually referred. Our AI visibility reporting guide explains why those cannot be collapsed into one score. Also record the product model or version. “Nike” appearing in an answer is not the same observation as a particular shoe being advised for a particular use.

Then look for disagreement, not only share of voice. Which queries produce stable recommendations but unstable source links? Which domains appear only through merchant pages? Does the assistant cite a manufacturer for specifications but another publisher for judgment? Those splits reveal where content, distribution and measurement are actually operating.

What this changes, and what it does not

The study supports a narrow but important conclusion: product-answer visibility depends on the engine, interface and repetition in the tested setting. It does not show that one publisher lost search traffic, that an answer drove a purchase, or that a specific optimization tactic caused inclusion. The paper is a preprint, and its product-question sample was selected for analysis rather than drawn as a representative survey of all shoppers.

The useful next step is an instrumented test. Measure a fixed prompt panel alongside referral logs, product-page analytics and actual sales where available. Publish both positive and absent captures. The reference point is not one persuasive screenshot. It is a documented distribution of answers that another person could repeat.

Keep learning

Continue this topic

Community discussion

Discuss: ChatGPT and Gemini Shared No Source Domains in 76.7% of Product Checks

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.