Perplexity Citation Audit: When a Link Does Not Support the Number

We recalculated two public Perplexity citation files, separated row and claim denominators, and built a four-state audit for numerical source support.

Sonar the Answer Whale checks whether a cited source supports a numerical claim

Direct answer: In a downloadable Haus Research file, 634 of 1,826 citation rows attached to claims containing a number did not pass its deterministic support check. That is 34.7%. When we grouped those rows into 872 numerical claims, 126 claims had no passing citation at all, or 14.4%.

Those rates answer different questions. The first asks whether each cited page contains at least one normalized number from the claim. The second asks whether any citation attached to a numerical claim passes. Neither rate proves that Perplexity’s answer was false, and neither represents every Perplexity product or query.

The research question is narrower than citation quality

Haus Research’s citation audit asks whether a page cited next to a numerical claim contains the cited number after basic normalization. That is a useful first check because a link can look authoritative while failing to expose the value a reader is trying to verify.

It is not a complete truth test. A source might express the number in a chart, image, script, rounded value, table the crawler did not extract, or a gated page. A page might contain the number but use it for a different entity, date or unit. “Number present” is necessary evidence in many cases, but it is not sufficient support by itself.

What SearchEngineAnswer recalculated

We downloaded the public claim-citation CSV linked by Haus and recalculated its row and claim denominators. The file contained 2,511 citation rows. Of those, 1,826 rows were attached to claims with at least one extracted number. The deterministic field recorded 1,192 passing rows and 634 non-passing rows.

A row-level failure rate and a claim-level unsupported rate are not interchangeable
UnitDenominatorNon-passing or unsupportedRate
Citation row attached to a numerical claim1,82663434.7%
Grouped numerical claim872126 with no passing citation14.4%

Our arithmetic uses the file’s supplied deterministic_pass field. We did not silently recode gated or unreadable pages as factual errors. Under the report’s method, a year alone can satisfy the normalized-number rule, which makes this a deliberately permissive support screen rather than a semantic validation of the whole claim.

The sample covered 310 questions and selected technology companies

The report says it tested 310 questions about 210 hand-picked technology companies using ten question templates. Responses came from the sonar and sonar-pro models through OpenRouter on September 2, 2026.

This was not the consumer Perplexity interface, a random sample of all user questions, or a longitudinal study. Hand-picked companies and numerical company questions create a valuable stress test, but they do not estimate a universal Perplexity citation error rate.

A second dataset measures domain popularity, not claim support

A Trellner report published a separate CSV of 7,534 citations across 2,055 domains for 380 software categories and 760 Perplexity API calls. We reproduced two calculations from that file: 1,766 cited domains were absent from the Tranco top million, or 23.4%, and 4,508 citation rows pointed to domains that were unranked or outside the top 100,000, or 59.8%.

Tranco rank is a popularity measure, not a quality score. A small specialist site can be the best source for a narrow category, and a highly visited domain can publish a weak page. The dataset raises a source-discovery question; it does not prove that an unranked domain is spam or that a cited recommendation is wrong.

The studies test different evidence layers
DatasetObservableCannot establish alone
Haus claim-citation fileWhether a readable page exposes a normalized number from the claimFull semantic support, answer truth or platform-wide error rate
Trellner citation-domain fileDomain popularity rank and first Wayback capturePage quality, common ownership, manipulation or causal recommendation impact

Do not treat the two reports as independent replication

The report pages name their teams, describe methods and publish CSVs under Creative Commons Attribution 4.0 terms. SearchEngineAnswer could reproduce the arithmetic above from those supplied files. We did not independently observe the API calls, authenticate the collection pipeline, or verify every cited page.

The two publishers also do not document an independent replication relationship. We therefore treat the files as two separate, inspectable datasets that ask different questions, not as independent confirmation of one shared claim about Perplexity.

Use four evidence states instead of pass or fail

A citation review needs more than a binary URL check
StateDefinitionEditorial action
SupportedThe cited page contains the number, entity, period, unit and relationship claimed.Keep the claim and cite the exact passage or table.
PartialThe page supports part of the claim but not every qualifier.Narrow the sentence or add another source.
UnverifiableThe page is gated, unavailable, script-only or otherwise unreadable in the audit.Do not call it false; find an inspectable source.
ContradictedThe accessible source states a materially different value or relationship.Correct or remove the claim and record the conflict.

A deterministic string match is a triage step. Human verification still needs to inspect the surrounding sentence, table header, footnote, publication date and entity. The same number can appear on a page for the wrong reason.

A seven-step citation audit for publishers

  1. Capture the complete answer, model or API surface, date, prompt and visible citation positions.
  2. Split the answer into atomic claims. Keep numbers, entities, periods and units together.
  3. Map every citation to the exact claim it appears to support.
  4. Open the canonical source and record access state: readable, gated, unavailable or script-only.
  5. Check the number, entity, date, unit and relationship in context.
  6. Classify the result as supported, partial, unverifiable or contradicted.
  7. Re-run a fixed sample after a model, retrieval or source change; never compare unmatched prompts as if they were the same test.

Our Perplexity SEO audit covers crawler, referral and source-label evidence. The AI-source credibility ledger adds owner, method and correction fields when a domain-level view is too coarse.

Publishers can make numerical evidence easier to verify

  • Keep the number, unit, period and entity in the same paragraph or table row.
  • Link to the primary dataset, filing, release or method rather than a summary that repeats the number.
  • Use descriptive table captions and headers so a value remains understandable outside the visual layout.
  • State whether a figure is measured, estimated, modeled or self-reported.
  • Preserve stable anchors or section headings for important evidence.
  • Publish a correction note when the number changes, and keep the original publication date distinct from the update date.

These steps improve human verification and machine extraction. They do not guarantee that an answer engine will retrieve, cite or rank the page.

Download the citation-support audit ledger

Download the CSV citation-support ledger. It records the answer surface, prompt, atomic claim, claimed number, cited URL, access state, exact supporting passage, entity-period-unit match, classification, reviewer and re-audit date.

All included rows are marked EXAMPLE-REMOVE. Replace or delete them before collecting real observations. The file contains no API keys, account identifiers, private prompts or copied source datasets.

Support check

A citation is useful only when the linked page supports the exact claim

The practical failure is not a missing link. It is a confident number placed beside a link that discusses the topic but does not establish that number.

Claim match
Can the cited page verify the same metric, population and period?
Location match
Does the evidence appear in the cited passage, table or downloadable data?
Qualification match
Are estimates, ranges and exclusions preserved in the answer?

My takeaway: My rule is simple: if a reader cannot reproduce the number from the cited source, I treat the statement as unsupported even when the domain is reputable.

Sources, method and limits

Sources: the Haus Research numerical citation audit and downloadable claim-citation CSV; the Trellner software-recommendation report and downloadable citation CSV. Each report page describes its model surface, sample and license.

SearchEngineAnswer contribution: We recalculated both files, separated citation-row and numerical-claim denominators, checked the domain-popularity arithmetic, mapped the two studies to different evidence questions, and created a four-state verification ledger that avoids converting unreadable sources into false claims.

Limits: The calculations reproduce fields supplied by the report publishers. We did not independently rerun their API collection, verify every source page, or test the consumer Perplexity interface. The findings should be read as dated audits of defined samples, not as a universal score for Perplexity.

Keep learning

Continue this topic

Community discussion

Discuss: Perplexity Citation Audit: When a Link Does Not Support the Number

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.