AI Search Engines Cited Different Webs in a 960-Answer Study
The same 80 questions produced different search activation rates, source volumes and citation markets across four AI engines. A separate longitudinal panel confirms a sharp Perplexity and Gemini split around video.
Four AI search engines received the same 80 questions three times. They returned 8,579 cited URLs, but they did not search at the same rate or draw from the same source market. Nearly three quarters of the 1,329 cited sites appeared in only one engine.
The clearest contradiction appeared around video
In Prefer’s September 13 snapshot, Gemini cited YouTube in 113 of 240 answers. Perplexity cited it in none of 240. That contrast is difficult to explain as a general judgment about video quality because the same question set produced both results.
A separate Epovest longitudinal study found the same directional split over time. Perplexity’s share of sourced answers containing a video citation fell from 61.5% on August 10 to 0.7% on August 17. On nine travel trackers, 39 of 54 answers cited video on August 10 and none of 54 did one week later. From August 18 onward, only two of 276 tracked answers cited video. Epovest measured Gemini at 46.5%.
The two studies used different question sets and denominators, so their percentages should not be merged. Their agreement is narrower but still useful: both observed a near-zero Perplexity video presence after mid-August and a much higher Gemini rate.
The 960-answer snapshot held the questions still
Prefer published the study and its downloadable dataset under a CC BY 4.0 license. The test contained 80 questions across ten AI-search topics, three fresh runs per question, and four engines: ChatGPT using gpt-5.6-luna, Perplexity Sonar Pro, Gemini 3.8 Flash and Claude Sonnet 5.
Collection ran from 17:56 to 18:12 UTC on September 13 through the DataForSEO LLM Responses API. Web search was requested on every call, but the API did not report that every engine actually searched. These were API answers rather than captures from the consumer applications.
| Engine | Answers that searched | Total cited URLs | Sources per answer | Unique cited sites |
|---|---|---|---|---|
| ChatGPT | 216 of 240, 90.0% | 731 | 3.05 | 218 |
| Perplexity | 240 of 240, 100% | 4,674 | 19.48 | 748 |
| Gemini | 213 of 240, 88.8% | 2,074 | 8.64 | 580 |
| Claude | 192 of 240, 80.0% | 1,100 | 4.58 | 333 |
A source in this dataset is a unique cited URL inside one answer. The site-level counts remove www and merge subdomains only for Reddit, YouTube, LinkedIn and Wikipedia. Other subdomains remain separate because they can represent distinct properties.
Search activation controlled whether citations could appear
The four engines produced 960 answers. They searched on 861 and skipped search on 99. Every answer that skipped search contained zero cited sources.
Once search ran, a citation appeared in 858 of 861 answers. All searched Perplexity, Gemini and Claude answers included at least one source. ChatGPT had three searched answers with no source. This does not prove that search activation alone determines publisher visibility, but it identifies an upstream gate that citation-only monitoring cannot see.
The distinction supports the lifecycle described in our comparison of how conversational agents search the web. A missing citation can mean the engine never searched, searched but did not retrieve the domain, retrieved it without selecting it, or selected it without displaying it. Those are different failures.
The source markets separated after search
| Engine | YouTube | Mean site overlap between repeated runs | ||
|---|---|---|---|---|
| ChatGPT | 0 | 16 | 0 | 36.7% |
| Perplexity | 0 | 98 | 73 | 95.8% |
| Gemini | 113 | 93 | 0 | 35.1% |
| Claude | 0 | 0 | 2 | 67.6% |
The repeated-run overlap is Jaccard similarity between pairs of runs for the same question when both runs cited at least one site. It measures short-window consistency, not long-term stability. Perplexity repeated nearly the same cited-site sets, while ChatGPT and Gemini changed much more between fresh calls made seconds apart.
The large difference in source volume matters too. Perplexity exposed about 6.4 times as many cited URLs per answer as ChatGPT. That does not establish that Perplexity’s evidence was better. It means a raw citation count gives engines with larger visible source lists more opportunities to credit a domain.
The long tail diverged while the citation head converged
Across all four engines, 966 of the 1,329 cited sites appeared in one engine only. Another 205 appeared in two engines, 129 in three, and 29 in all four.
That looks like extreme fragmentation until citation frequency is considered. Of the 25 most-cited sites, 23 appeared in three or four engines. Among sites cited in ten or more answers, only 14 of 170 were exclusive to one engine. Much of the one-engine-only count came from the tail: 338 exclusive sites appeared in just one answer.
The useful reading is not “every engine uses a different internet.” Popular sources overlapped, while the long tail and platform mix differed sharply. A publisher can therefore face both a shared competitive head and an engine-specific discovery problem.
Build the denominator ledger before the visibility score
A cross-engine report should preserve at least five counts for every engine and prompt group:
- answers requested;
- answers that activated web search;
- answers that returned any visible source;
- answers citing the measured domain or platform;
- unique cited URLs and sites.
Keep the model identifier, API or consumer surface, geography, timestamp and run number beside those counts. Otherwise, a change in search activation can be misread as a source-quality change, and a longer citation list can look like greater visibility even when the selection rate is unchanged.
Keep crawler access as a separate diagnostic layer too. Our AI crawler guide maps which agents request pages, but a successful crawl does not prove that an answer engine searched, retrieved or cited the page for a specific question.
Download the cross-engine source-market comparison CSV
The file separates the Prefer snapshot from the Epovest panel and leaves non-comparable fields blank. It is a compact comparison aid, not a merged benchmark.
What these studies cannot establish
- Prefer’s questions all concern AI search, GEO or AEO. The results do not represent every query class.
- Prefer measured API responses on one day, not the consumer applications or long-term behavior.
- Eight questions named Profound or Peec. Prefer reports that 29 of the 103 answers citing
tryprofound.comcame from those questions. - Prefer and Epovest both sell AI-visibility products or services. Their disclosed methods and data improve inspectability but do not remove that commercial context.
- Neither study graded the factual correctness of every answer or proved why an engine selected a source.
- The two studies use different prompts, schedules and denominators. Their video results can be compared directionally, not pooled into one percentage.
The defensible conclusion is engine-specific
The same questions did not produce one shared citation market. Search activation, visible source capacity, repeated-run stability and platform preference all varied by engine. The strongest overlap appeared among frequently cited sites, while the long tail remained highly fragmented.
For publishers, that means one citation share should be treated as a summary after diagnosis, not as the diagnosis itself. Preserve the search gate and source-list denominator first. Then compare engines separately and repeat the test over time. Our YouTube citation comparison method applies the same discipline when the question is whether a cited video also appeared in a separately captured Google result set.
Study links
Keep learning
Continue this topic
Next in this topic
AI Answer Citations Can Fail at Four Separate Layers
Earlier in this topic
CiteShade Raised Wrong Answers From 1% to 68%
Research
Ask a question or join the discussion