AI Search Source Credibility: Build a 32-Field Audit Ledger

Adapt an EACL 2026 study into a claim-source audit that preserves sample, user role, credibility rules, groundedness, adjudication and limitations.

Sonar stops a trophy conveyor and weighs source ownership against claim support beside a bounded research sample.

Direct answer: an EACL 2026 study found meaningful differences in the credibility of sources cited by four AI assistants, but its 100-claim sample is not a permanent leaderboard. The useful next step is to preserve the study’s claim, user-role, source-credibility and groundedness boundaries when you audit a new system.

This article translates that bounded research design into a 32-field audit ledger. The ledger is a template for new observations, not a reproduction of the paper’s dataset and not a claim that any assistant keeps the same behavior over time.

What the study actually tested

The paper, “Evaluating Source Credibility and Groundedness in AI-Powered Search”, evaluated GPT-4o, GPT-5, Perplexity and Qwen Chat. Its canonical PDF describes a sample of 100 claims across five topics prone to misinformation. The researchers used two user roles: a fact-checker and a person who believed the claim.

The study examined source credibility and whether answers were grounded in their cited sources. That is narrower than “which assistant is best.” It does not measure every topic, plan, model version, geography, interface, user intent or current product state.

The headline result needs its boundary

In this study, Perplexity had the highest source-credibility result among the tested systems. The authors also report that GPT-4o increased its citation of non-credible sources on sensitive topics. Those are findings from the paper’s declared sample and conditions.

Do not rewrite them as “Perplexity always cites the best sources” or “GPT-4o is unreliable.” Assistants, retrieval indexes and interfaces change. A new comparison needs a new observation date, frozen fixtures, declared access conditions and the same skeptical treatment of every system.

Keep four units of analysis separate

One answer contains several different review units
UnitWhat is codedCommon mistake
Input claimClaim text, truth label, topic and user roleComparing answers to different propositions
Assistant answer claimOne checkable statement in the responseScoring the whole answer from its tone
Cited sourceOne resolved source and its credibility evidenceTreating a citation count as source quality
Claim-source pairWhether the source supports that answer claimAssuming a credible source entails every nearby sentence

The downloadable ledger uses one row per citation reviewed against one answer claim. Repeat run-level fields across rows or join them through a non-sensitive study ID. Do not publish conversation IDs, account emails or private prompts.

Freeze the claim before querying

Create the claim set before collecting responses. Give each claim a stable ID, exact wording, pre-coded truth label and topic. Record the evidence used for that truth label outside the assistant responses, so the evaluated system does not define its own ground truth.

If a claim changes after pilot review, issue a new fixture version. Small wording shifts can change retrieval and answer stance. The guide to reading GEO studies explains why sample construction and denominator discipline matter more than a catchy cross-study average.

Treat user role as an experimental condition

The EACL study used a fact-checker role and a claim-believer role. That matters because the user’s stated stance can affect how an assistant searches, challenges or accommodates a claim. A reproducible audit must retain the complete role instruction and keep it stable across systems.

Do not combine role variants into one row or average them before checking interaction effects. If one assistant challenges a believer prompt but another mirrors it, that difference disappears when the condition is not recorded.

Define credibility before seeing the results

Source credibility is a coding decision, not a property revealed by the URL alone. Write the rulebook before response collection. It can use source ownership, editorial accountability, primary-evidence access, correction practices, expertise and independence, but every criterion needs a documented application rule.

Example coding dimensions for a declared rulebook
DimensionEvidence to inspectDo not infer from
Source ownershipNamed organization or author and responsibility for the pageDomain appearance
Primary accessOriginal document, data, filing, specification or direct statementA summary that links nowhere
Editorial controlsVisible methods, corrections or review process where relevantProfessional design
ExpertiseRelevant, verifiable subject responsibilityFollower count or generic biography
ConflictCommercial, political or institutional relationship material to the claimSource type alone

A government, company, university, newsroom or community source can be authoritative for one claim and weak for another. Code the relationship between the source and the proposition, not a permanent reputation score for the whole domain.

Code groundedness at claim level

A credible source can be irrelevant to the sentence it is shown beside. Resolve the cited URL, preserve the answer claim span and locate the source passage that is supposed to support it. Classify support as direct, qualified, conflicting or absent.

This is the distinction explored in the citation count versus citation absorption analysis: a visible citation and evidence actually used by an answer are not interchangeable measurements. Groundedness review needs the claim-source pair.

Resolve URLs before coding

Open each citation on the review date and record its final resolved URL. Note redirects, inaccessible pages, login walls, changed pages, search-result intermediaries and duplicate destinations. Keep the URL returned by the assistant and the resolved URL as separate fields.

If a source cannot be accessed, code the access failure rather than guessing its credibility or support from a title. Archive or retain a permitted evidence copy when the protocol allows, but do not republish copyrighted pages or sensitive content in the public dataset.

Use independent review and adjudication

At least two reviewers should independently code a declared subset before the rulebook is finalized. Compare disagreements by field: source type, credibility, claim span and groundedness. Revise ambiguous definitions, then lock the rulebook version before the full review.

Keep disagreements visible until adjudication
StateMeaningNext action
AgreedIndependent codes match under the rulebookAccept the code
Definition disputeReviewers applied different interpretationsClarify the rule and recode affected rows
Evidence disputeReviewers located different source evidenceAdjudicate with both references visible
UnresolvableAccess or evidence is insufficientRetain uncertainty; do not force a score

Report agreement with the metric and denominator chosen in advance. A high agreement rate does not prove the rulebook is valid; it shows reviewers applied it consistently under the sampled conditions.

Freeze system and access state

Record assistant name, displayed model or version, plan, account or access path, client surface, date, locale and any available search mode. Preserve the exact query and response evidence. If the product does not expose a model version, write “not exposed” rather than inferring one.

Repeat observations on another date under a new study or run ID. Model names can persist while retrieval behavior changes, and a product can route requests without exposing every implementation detail. The audit documents the observed system surface, not hidden architecture.

Calculate only declared metrics

Choose denominators before looking at results: citations returned, accessible citations, claim-source pairs, answers or input claims. Publish numerator and denominator together. A system that returns fewer citations can look better or worse depending on which unit the metric uses.

Separate source-credibility rate, grounded-support rate, citation coverage and refusal or no-answer rate. Do not collapse them into one “trust score” unless the weighting rule is justified and sensitivity-tested. Use the AI visibility measurement crosswalk to keep citation, source use and referral measures distinct.

Download the 32-field audit ledger

Download the AI search source-credibility audit ledger. One row represents one citation reviewed against one answer claim. Remove both EXAMPLE-REMOVE rows before collection.

The fields preserve the fixture, role, system surface, response evidence, returned and resolved URLs, credibility rule, claim span, groundedness decision, reviewer, adjudicator, limitation and release decision. Private response files remain outside the public artifact. Validate that every row has 32 columns.

Publish the method before the ranking

A defensible comparison lets readers inspect the sample, claim wording, role prompts, system state, credibility rule, groundedness codes, disagreements, exclusions and denominators. If those cannot be shared safely, publish a narrower method note rather than a precise leaderboard.

The EACL paper supports a bounded finding about its 100 claims and four tested assistants. It also supplies a useful warning: source quality can vary with topic and user stance. A new audit should preserve those conditions, expose uncertainty and end at the evidence collected, not at a universal winner.

Ledger design

Thirty-two fields are useful only when they lead to a smaller decision set

A credibility ledger should make disagreement visible, not reward the source with the most completed cells.

Identity
Who produced the claim, and can the entity be verified?
Evidence
What primary material supports the specific statement?
Context
Which date, population, method and limitations apply?
Action
Would a missing field change whether you cite, test or reject the claim?

My takeaway: I would make eight fields mandatory and keep the remaining fields conditional. Otherwise the audit becomes clerical work instead of a credibility decision.

Keep learning

Continue this topic

Community discussion

Discuss: AI Search Source Credibility: Build a 32-Field Audit Ledger

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.