AI Search Source Credibility: Build a 32-Field Audit Ledger
Adapt an EACL 2026 study into a claim-source audit that preserves sample, user role, credibility rules, groundedness, adjudication and limitations.
Direct answer: an EACL 2026 study found meaningful differences in the credibility of sources cited by four AI assistants, but its 100-claim sample is not a permanent leaderboard. The useful next step is to preserve the study’s claim, user-role, source-credibility and groundedness boundaries when you audit a new system.
This article translates that bounded research design into a 32-field audit ledger. The ledger is a template for new observations, not a reproduction of the paper’s dataset and not a claim that any assistant keeps the same behavior over time.
What the study actually tested
The paper, “Evaluating Source Credibility and Groundedness in AI-Powered Search”, evaluated GPT-4o, GPT-5, Perplexity and Qwen Chat. Its canonical PDF describes a sample of 100 claims across five topics prone to misinformation. The researchers used two user roles: a fact-checker and a person who believed the claim.
The study examined source credibility and whether answers were grounded in their cited sources. That is narrower than “which assistant is best.” It does not measure every topic, plan, model version, geography, interface, user intent or current product state.
The headline result needs its boundary
In this study, Perplexity had the highest source-credibility result among the tested systems. The authors also report that GPT-4o increased its citation of non-credible sources on sensitive topics. Those are findings from the paper’s declared sample and conditions.
Do not rewrite them as “Perplexity always cites the best sources” or “GPT-4o is unreliable.” Assistants, retrieval indexes and interfaces change. A new comparison needs a new observation date, frozen fixtures, declared access conditions and the same skeptical treatment of every system.
Keep four units of analysis separate
| Unit | What is coded | Common mistake |
|---|---|---|
| Input claim | Claim text, truth label, topic and user role | Comparing answers to different propositions |
| Assistant answer claim | One checkable statement in the response | Scoring the whole answer from its tone |
| Cited source | One resolved source and its credibility evidence | Treating a citation count as source quality |
| Claim-source pair | Whether the source supports that answer claim | Assuming a credible source entails every nearby sentence |
The downloadable ledger uses one row per citation reviewed against one answer claim. Repeat run-level fields across rows or join them through a non-sensitive study ID. Do not publish conversation IDs, account emails or private prompts.
Freeze the claim before querying
Create the claim set before collecting responses. Give each claim a stable ID, exact wording, pre-coded truth label and topic. Record the evidence used for that truth label outside the assistant responses, so the evaluated system does not define its own ground truth.
If a claim changes after pilot review, issue a new fixture version. Small wording shifts can change retrieval and answer stance. The guide to reading GEO studies explains why sample construction and denominator discipline matter more than a catchy cross-study average.
Treat user role as an experimental condition
The EACL study used a fact-checker role and a claim-believer role. That matters because the user’s stated stance can affect how an assistant searches, challenges or accommodates a claim. A reproducible audit must retain the complete role instruction and keep it stable across systems.
Do not combine role variants into one row or average them before checking interaction effects. If one assistant challenges a believer prompt but another mirrors it, that difference disappears when the condition is not recorded.
Define credibility before seeing the results
Source credibility is a coding decision, not a property revealed by the URL alone. Write the rulebook before response collection. It can use source ownership, editorial accountability, primary-evidence access, correction practices, expertise and independence, but every criterion needs a documented application rule.
| Dimension | Evidence to inspect | Do not infer from |
|---|---|---|
| Source ownership | Named organization or author and responsibility for the page | Domain appearance |
| Primary access | Original document, data, filing, specification or direct statement | A summary that links nowhere |
| Editorial controls | Visible methods, corrections or review process where relevant | Professional design |
| Expertise | Relevant, verifiable subject responsibility | Follower count or generic biography |
| Conflict | Commercial, political or institutional relationship material to the claim | Source type alone |
A government, company, university, newsroom or community source can be authoritative for one claim and weak for another. Code the relationship between the source and the proposition, not a permanent reputation score for the whole domain.
Code groundedness at claim level
A credible source can be irrelevant to the sentence it is shown beside. Resolve the cited URL, preserve the answer claim span and locate the source passage that is supposed to support it. Classify support as direct, qualified, conflicting or absent.
This is the distinction explored in the citation count versus citation absorption analysis: a visible citation and evidence actually used by an answer are not interchangeable measurements. Groundedness review needs the claim-source pair.
Resolve URLs before coding
Open each citation on the review date and record its final resolved URL. Note redirects, inaccessible pages, login walls, changed pages, search-result intermediaries and duplicate destinations. Keep the URL returned by the assistant and the resolved URL as separate fields.
If a source cannot be accessed, code the access failure rather than guessing its credibility or support from a title. Archive or retain a permitted evidence copy when the protocol allows, but do not republish copyrighted pages or sensitive content in the public dataset.
Use independent review and adjudication
At least two reviewers should independently code a declared subset before the rulebook is finalized. Compare disagreements by field: source type, credibility, claim span and groundedness. Revise ambiguous definitions, then lock the rulebook version before the full review.
| State | Meaning | Next action |
|---|---|---|
| Agreed | Independent codes match under the rulebook | Accept the code |
| Definition dispute | Reviewers applied different interpretations | Clarify the rule and recode affected rows |
| Evidence dispute | Reviewers located different source evidence | Adjudicate with both references visible |
| Unresolvable | Access or evidence is insufficient | Retain uncertainty; do not force a score |
Report agreement with the metric and denominator chosen in advance. A high agreement rate does not prove the rulebook is valid; it shows reviewers applied it consistently under the sampled conditions.
Freeze system and access state
Record assistant name, displayed model or version, plan, account or access path, client surface, date, locale and any available search mode. Preserve the exact query and response evidence. If the product does not expose a model version, write “not exposed” rather than inferring one.
Repeat observations on another date under a new study or run ID. Model names can persist while retrieval behavior changes, and a product can route requests without exposing every implementation detail. The audit documents the observed system surface, not hidden architecture.
Calculate only declared metrics
Choose denominators before looking at results: citations returned, accessible citations, claim-source pairs, answers or input claims. Publish numerator and denominator together. A system that returns fewer citations can look better or worse depending on which unit the metric uses.
Separate source-credibility rate, grounded-support rate, citation coverage and refusal or no-answer rate. Do not collapse them into one “trust score” unless the weighting rule is justified and sensitivity-tested. Use the AI visibility measurement crosswalk to keep citation, source use and referral measures distinct.
Download the 32-field audit ledger
Download the AI search source-credibility audit ledger. One row represents one citation reviewed against one answer claim. Remove both EXAMPLE-REMOVE rows before collection.
The fields preserve the fixture, role, system surface, response evidence, returned and resolved URLs, credibility rule, claim span, groundedness decision, reviewer, adjudicator, limitation and release decision. Private response files remain outside the public artifact. Validate that every row has 32 columns.
Publish the method before the ranking
A defensible comparison lets readers inspect the sample, claim wording, role prompts, system state, credibility rule, groundedness codes, disagreements, exclusions and denominators. If those cannot be shared safely, publish a narrower method note rather than a precise leaderboard.
The EACL paper supports a bounded finding about its 100 claims and four tested assistants. It also supplies a useful warning: source quality can vary with topic and user stance. A new audit should preserve those conditions, expose uncertainty and end at the evidence collected, not at a universal winner.
Ledger design
Thirty-two fields are useful only when they lead to a smaller decision set
A credibility ledger should make disagreement visible, not reward the source with the most completed cells.
- Identity
- Who produced the claim, and can the entity be verified?
- Evidence
- What primary material supports the specific statement?
- Context
- Which date, population, method and limitations apply?
- Action
- Would a missing field change whether you cite, test or reject the claim?
My takeaway: I would make eight fields mandatory and keep the remaining fields conditional. Otherwise the audit becomes clerical work instead of a credibility decision.
Keep learning
Continue this topic
Next in this topic
Google before: Filter: Build a 29-Field Temporal Leakage Audit
Earlier in this topic
Citation Absorption: Audit How Sources Shape AI Answers
Research
Ask a question or join the discussion