Baidu and Google Gave Similar AI Answers From Different Source Webs

A 500-query bilingual study found high answer similarity but very low cited-host overlap between Baidu and Google AI overviews.

Sonar the Answer Whale compares two different source rivers that converge into similar AI answer bubbles

Baidu and Google produced semantically similar AI answers while drawing from sharply different source ecosystems. In a bilingual study of 500 information-seeking queries, the overlap between cited host sets was tiny even when answer similarity remained high.

For international AI visibility, translated prompt tracking is not enough. Measure which engine shows an answer, which hosts supply it, and how concentrated that source set is in each market.

The answers converged more than the sources

The study accepted at WAC@EMNLP 2026 compared Baidu and Google AI overviews using 500 English MS MARCO queries and their Simplified Chinese translations. Data collection produced 1,998 usable platform-query observations.

Median answer similarity was 0.813 for Google Chinese versus Baidu Chinese, 0.791 for Google Chinese versus Google English, and 0.701 for Google English versus Baidu Chinese. Yet the cited host overlap remained small.

Similar answer meaning did not require the same cited hosts
Comparison Shared hosts Union of hosts Jaccard overlap
Google Chinese vs Google English 179 5,341 0.034
Baidu Chinese vs Google Chinese 103 3,183 0.032
Baidu Chinese vs Google English 25 3,095 0.008

The smallest overlap, 0.008, means that fewer than one in one hundred hosts in the combined set appeared on both sides of that comparison. It does not mean that the answers shared less than one percent of their claims.

Google showed overviews much more consistently in this setup

Baidu returned a Chinese overview for 410 of 500 queries, an 82.0% overview appearance rate. It returned English overviews for only 20 of 500 queries, or 4.0%. Google returned overviews for 487 of 499 Chinese queries and 484 of 499 English queries, approximately 97.6% and 97.0%.

This is an experimental observation, not a global coverage promise. The tests ran from Milan between June 20 and July 10, 2026. Google was used in a logged-in English-interface state, there was no VPN, and the Chinese questions were translations rather than questions written originally by Chinese users.

Those conditions are useful because they are recorded. They also mean a mainland-China measurement program should not copy the percentages as its baseline.

Baidu’s Chinese source set was more concentrated

The reported host-level Gini coefficient was 0.895 for Baidu Chinese, compared with 0.500 for Google Chinese and 0.462 for Google English. A higher value means citation activity is concentrated among fewer hosts.

The leading host accounted for 27.1% of Baidu Chinese citations, versus 6.6% for Google Chinese and 5.5% for Google English. Mean distinct hosts per overview were 15.0, 18.6 and 16.7 respectively.

Audit four separate dimensions in every market

Overview availability
How often does the engine show an AI answer for the prompt set, locale, account state and date?
Answer similarity
Do the responses communicate similar claims, cautions and recommendations?
Source overlap
Which exact hosts and URLs appear on more than one engine or language surface?
Source concentration
How much of the citation share belongs to the leading domains?

Each measure can move independently. A translated page may support a similar answer without becoming a selected source. An engine may cite many links in each response while depending heavily on a small recurring domain group across the full sample.

A useful dashboard therefore needs both a query view and a domain view. The query view shows whether the engine answered and which sources appeared for that need. The domain view aggregates exposure across the complete sample, revealing whether a few hosts dominate the system. Preserve the exact-language prompt beside both views so translation changes are not mistaken for engine changes.

International visibility needs local source participation

Translating an English article is a distribution step, not proof of participation in another engine’s source ecosystem. A local strategy should identify the publishers, reference databases, communities and institutional sources that recur in that engine, then ask whether the brand can contribute useful evidence within those channels.

  • Build a native-language prompt set from local search, support and community evidence.
  • Run the same prompts on each engine with location, account and interface state recorded.
  • Separate host-level coverage from exact-URL coverage.
  • Inspect whether a translation preserves examples, units, regulations and source context.
  • Measure citations and referrals independently.

Our study of AI engines citing different parts of the web reaches a related operational conclusion: visibility on one answer surface cannot be assumed to transfer to another.

What this comparison cannot tell us

The paper studies source exposure and answer similarity. It does not determine whether every cited passage supports the adjacent claim, whether one engine’s sources are more trustworthy, or whether readers click them.

Host overlap also hides page-level differences. Two engines can cite the same domain but different articles, languages or timestamps. A claim-support audit therefore needs exact URLs and passages after the host-level ecosystem map is built.

Time is another missing dimension. The collection window captures one period in 2026. A platform can change source selection, overview eligibility or cited-page freshness after the study. Repeating a stable prompt panel is more informative than treating the published host list as permanent.

The useful conclusion is not that one engine is better. It is that similar answers can emerge from different source worlds, so a single-engine rank tracker is an incomplete international AI-search measurement system.

Keep learning

Continue this topic

Community discussion

Discuss: Baidu and Google Gave Similar AI Answers From Different Source Webs

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.