Which AI Search Assistant Cites the Most Credible Sources?
An EACL 2026 study compares GPT-4o, GPT-5, Perplexity and Qwen across 100 misinformation-prone claims—and reveals a tightly bounded result.
Direct answer: in an EACL 2026 study of 100 claims across five misinformation-prone topics, Perplexity achieved the highest source credibility among the tested assistants, while GPT-4o cited more non-credible sources on sensitive topics. The tested systems were GPT-4o, GPT-5, Perplexity and Qwen Chat.
This is a scoped fact-checking result—not a universal ranking of assistants, a current-product guarantee or proof about every query type. Credibility and groundedness must be measured at the claim-and-source level.
What the study evaluated
Vykopal, Pikuliak, Ostermann and Simko introduce a methodology for evaluating web-search behavior with two linked questions: are cited sources credible, and is the response grounded in what those sources support? Their sample covers 100 claims across five topics selected for misinformation risk.
Source credibility alone is not enough. A system can cite a reputable page and still overstate it. Groundedness alone is not enough if the underlying source is unreliable. The evaluation therefore treats retrieval quality and answer use as separate components.
| Assistant | Credibility rate | Non-credibility rate | Total cited sources | Unique domains |
|---|---|---|---|---|
| GPT-4o | 75.16% ± 1.62 | 2.27% ± 0.57 | 8,416 | 1,863 |
| GPT-5 | 71.37% ± 1.53 | 2.03% ± 0.53 | 12,103 | 2,425 |
| Perplexity | 86.30% ± 1.67 | 0.69% ± 0.32 | 3,592 | 754 |
| Qwen Chat | 80.01% ± 2.06 | 1.07% ± 0.43 | 4,587 | 1,130 |
Perplexity’s credibility rate was 11.14 percentage points above GPT-4o and 14.93 points above GPT-5 in this sample. The paper’s pairwise tests mark both differences statistically significant. Yet Perplexity also cited far fewer total sources, so the result describes the quality mix of its classified citations—not greater breadth.
The averages also hide meaningful topic spread. Perplexity’s reported credibility rate ranged from 78.95% for the Local topic to 92.28% for Climate Change. GPT-4o’s non-credibility rate reached 4.55% for Russia–Ukraine claims, compared with its 2.27% overall rate. A credible audit therefore needs topic rows, not only one product-wide average.
Do not turn the result into a permanent leaderboard
- The sample contains misinformation-prone topics, not every informational or commercial task.
- Products, models, search systems and policies change.
- Location, account state and prompt wording can alter retrieval.
- One average can hide topic-level failures.
- Source reputation does not replace claim-level support.
A publisher should cite the venue, sample, systems and evaluation target whenever repeating the result. “Perplexity always has the best sources” would exceed the evidence. “Perplexity had the highest source credibility in this 100-claim study” stays inside it.
Run a publisher credibility audit
- Choose a bounded topic and a frozen set of factual claims.
- Run each assistant repeatedly under recorded location and account conditions.
- Save full answers, sources and page snapshots.
- Rate source credibility using a published rubric.
- Map each answer claim to supporting or contradicting passages.
- Report missing citations, unsupported synthesis and topic-level variation.
Use domain reputation as one input, not a substitute for inspecting the page. Primary documentation, transparent methods, named authorship and correction history make a source easier to evaluate, but none guarantees that an assistant will retrieve or represent it correctly.
The author identity-chain guide covers verifiable publication metadata. The evidence-led publishing guide covers sourcing and corrections.
Ask a question or join the discussion