Which AI Search Assistant Cites the Most Credible Sources?

An EACL 2026 study compares GPT-4o, GPT-5, Perplexity and Qwen across 100 misinformation-prone claims—and reveals a tightly bounded result.

Sonar the Answer Whale weighs source cards from four assistant characters on a credibility balance inside a fenced test area.

Direct answer: in an EACL 2026 study of 100 claims across five misinformation-prone topics, Perplexity achieved the highest source credibility among the tested assistants, while GPT-4o cited more non-credible sources on sensitive topics. The tested systems were GPT-4o, GPT-5, Perplexity and Qwen Chat.

This is a scoped fact-checking result—not a universal ranking of assistants, a current-product guarantee or proof about every query type. Credibility and groundedness must be measured at the claim-and-source level.

What the study evaluated

Vykopal, Pikuliak, Ostermann and Simko introduce a methodology for evaluating web-search behavior with two linked questions: are cited sources credible, and is the response grounded in what those sources support? Their sample covers 100 claims across five topics selected for misinformation risk.

Source credibility alone is not enough. A system can cite a reputable page and still overstate it. Groundedness alone is not enough if the underlying source is unreliable. The evaluation therefore treats retrieval quality and answer use as separate components.

Overall source results reported in EACL 2026
AssistantCredibility rateNon-credibility rateTotal cited sourcesUnique domains
GPT-4o75.16% ± 1.622.27% ± 0.578,4161,863
GPT-571.37% ± 1.532.03% ± 0.5312,1032,425
Perplexity86.30% ± 1.670.69% ± 0.323,592754
Qwen Chat80.01% ± 2.061.07% ± 0.434,5871,130

Perplexity’s credibility rate was 11.14 percentage points above GPT-4o and 14.93 points above GPT-5 in this sample. The paper’s pairwise tests mark both differences statistically significant. Yet Perplexity also cited far fewer total sources, so the result describes the quality mix of its classified citations—not greater breadth.

The averages also hide meaningful topic spread. Perplexity’s reported credibility rate ranged from 78.95% for the Local topic to 92.28% for Climate Change. GPT-4o’s non-credibility rate reached 4.55% for Russia–Ukraine claims, compared with its 2.27% overall rate. A credible audit therefore needs topic rows, not only one product-wide average.

Do not turn the result into a permanent leaderboard

  • The sample contains misinformation-prone topics, not every informational or commercial task.
  • Products, models, search systems and policies change.
  • Location, account state and prompt wording can alter retrieval.
  • One average can hide topic-level failures.
  • Source reputation does not replace claim-level support.

A publisher should cite the venue, sample, systems and evaluation target whenever repeating the result. “Perplexity always has the best sources” would exceed the evidence. “Perplexity had the highest source credibility in this 100-claim study” stays inside it.

Run a publisher credibility audit

  1. Choose a bounded topic and a frozen set of factual claims.
  2. Run each assistant repeatedly under recorded location and account conditions.
  3. Save full answers, sources and page snapshots.
  4. Rate source credibility using a published rubric.
  5. Map each answer claim to supporting or contradicting passages.
  6. Report missing citations, unsupported synthesis and topic-level variation.

Use domain reputation as one input, not a substitute for inspecting the page. Primary documentation, transparent methods, named authorship and correction history make a source easier to evaluate, but none guarantees that an assistant will retrieve or represent it correctly.

The author identity-chain guide covers verifiable publication metadata. The evidence-led publishing guide covers sourcing and corrections.

Primary documentation

Community discussion

Discuss: Which AI Search Assistant Cites the Most Credible Sources?

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.