How to Measure AI Citations and Recommendations Separately

Separate citation, publisher naming, claim support, recommendation and referral outcomes, then calculate attribution yield only from matched saved answers.

Sonar separates a tangled AI answer into distinct citation and recommendation evidence tracks while an analyst records them independently.

A citation tells you that an answer visibly referenced a source. It does not tell you that the answer recommended the source owner, ranked its product first, sent a visit, or changed a purchase decision. Measure those outcomes separately, using the AI answer as the unit of observation and the entity as the unit of classification.

What changed in this guide: the earlier page preregistered a repeated-run experiment, but no completed dataset was published. This revision does not invent that missing dataset. It replaces the unfinished promise with a reusable coding method, a browser-local outcome coder, and a 48-field CSV ledger whose illustrative rows are marked EXAMPLE-REMOVE.

The short answer

For every entity in every saved answer, record at least four independent decisions: whether the entity is mentioned, whether one of its pages is cited, whether the cited page supports the nearby claim, and whether the answer recommends the entity for the reader’s task. Add comparison role, ordering, referral evidence, and exclusions as separate fields.

This prevents a common measurement error. A publisher can supply evidence for an answer while a competitor receives the recommendation. The reverse can also happen: a product may be recommended from model knowledge or an unlinked passage while its website receives no visible citation.

Why one AI visibility score hides the outcome

A composite score can combine observations that have different meanings and denominators. Citation presence is a source-use observation. Recommendation is an action-oriented statement about an entity. Referral is a recorded visit. Conversion is a later business event. Adding them together does not create a stronger metric; it removes the boundary between them.

Keep each evidence layer attached to the system that can observe it
LayerObservable evidenceWhat it cannot establish alone
CitationVisible source link or attribution in a saved answerEndorsement, recommendation, ranking, traffic, or conversion
Claim supportCited page and passage support, partially support, conflict with, or fail to verify the nearby claimWhy the system selected the source
RecommendationAction wording that proposes an entity for the stated task, with any conditionsWhether the user followed the advice
ReferralAnalytics or server evidence of a visit from the answer surfaceEvery citation or unclicked mention
ConversionPermissioned first-party outcome joined through a declared attribution methodUniversal causal impact of AI visibility

Bing’s current AI Performance documentation makes the boundary unusually explicit: its report measures visible citation activity, not ranking, authority, importance, clicks, or traffic. That documented metric is useful, but it cannot answer the recommendation question.

Do not turn cited pages and brand mentions into one ratio

A publisher reported on Reddit roughly 700 cited pages and five written brand mentions across several AI surfaces. The post raises a useful question, but the two numbers do not form a defensible 140-to-1 industry ratio.

“Cited pages” may be a deduplicated site-level inventory while “brand mentions” may count answers, prompts or monitoring events. The post does not supply the export, observation window, eligible answer count, prompt panel, repeated-run rule or cross-platform deduplication method. Its platform figures also add to more than the rounded cited-page total, which may be harmless rounding or overlap but prevents an exact calculation.

The useful signal is the mismatch: publisher pages can appear as source material without the answer naming the publisher in prose. Measure that relationship inside the same saved answers before testing possible explanations.

Calculate attribution yield from matched answers

Every numerator must come from the denominator’s saved answer set
MeasureNumeratorDenominatorReader question
Publisher naming yieldAnswers that name the publisher in proseEligible answers citing an owned pageWhen a page supplies evidence, how often does the outlet receive visible credit?
Entity recommendation yieldCited answers with a conditional or explicit recommendationEligible cited answers from recommendation-intent promptsHow often does citation accompany task-fit advice?
Support yieldReviewable citations whose page supports the linked claimReviewable citation-claim decisionsHow often is visible sourcing substantively useful?
Referral rateAttributed sessions under the declared analytics methodUse an analytics denominator, not cited-answer rowsDid any visible exposure become a measurable visit?

Report each count beside its rate. If publisher naming is not available from the same answer captures as citation presence, label the comparison unavailable instead of combining two vendor dashboard totals.

Use divergence as a diagnostic, not a score

Four common outcomes lead to different reviews
Observed answerWhat to inspect nextWhat not to conclude
Owned page cited, publisher unnamedSource-card label, byline, organization identity near the original evidence, and whether the answer names only the subject.That structured data caused the omission or that adding the brand name will force attribution.
Publisher named, no owned citationOther cited sources, unlinked answer wording and whether the name came from comparison context.That the publisher’s own page was retrieved.
Cited and named, not recommendedThe claim being supported and the prompt’s actual decision intent.That visible credit is a commercial endorsement.
Recommended, no visible citationRecommendation wording, conditions, source panel and repeat-run variability.That model memory, hidden retrieval or any specific source caused the choice.

A content or entity change can be tested against this matrix, but no markup, byline or writing pattern guarantees that an answer engine will name or recommend the publisher. The experiment should preserve a matched pre-change prompt set and change one identity or evidence element at a time.

Choose the unit before opening an answer

Use one row per entity-answer observation. If one answer discusses four products, create four rows linked by the same run ID. If the answer cites three pages for one entity, either add child citation rows or preserve the three citation IDs in a separate citation table. Do not count one answer as four independent runs.

Give every study, query, run, entity, response artifact, and citation a stable ID. Those keys let you revise a coding decision without losing the raw answer or silently duplicating a run.

Freeze the environment before coding

Record the platform, product surface, visible model label, plan, signed-in state, locale, geography, device, personalization or memory state, search state, and retrieval time. Save the exact prompt and full answer before opening source links.

OpenAI’s current ChatGPT Search documentation says web answers may include citations and that results can be incomplete, outdated, or incorrect. It also documents query rewriting and location-sensitive behavior. A later response can differ because the retrieval query, available web pages, location, memory, or product changed; the saved environment is part of the evidence.

Code the mention before the recommendation

Start with the lowest-level entity decision. Mark the canonical entity as absent, named, described without an exact name, or ambiguous. Preserve the exact sentence and any alias. A linked page owned by an entity is not automatically a mention of that entity in the answer text.

Entity mention states
StateRuleExample pattern
absentNo defensible reference to the entityThe answer and source panel omit it
namedCanonical name or registered alias appears“Tool A includes…”
describedIdentity is clear from context without the canonical name“The vendor’s free crawler…” after an unambiguous introduction
ambiguousSimilar names or missing context prevent a reliable matchA generic product name with no owner or URL

Record citation presence without adding meaning

Mark a citation only when the saved surface exposes a source link or explicit attribution. Preserve the displayed URL, final URL after redirects, source label, placement, and cited claim. Normalize fragments, tracking parameters, mobile variants, and redirects without deleting the original values.

Do not convert a source-card position into a quality grade. Perplexity’s source-label documentation says its labels describe a domain, not an individual page or claim, and are not endorsements. An unlabeled domain is not a negative judgment.

Review whether the cited page supports the claim

Citation presence and citation support need separate columns. Open the final page and compare the relevant passage with the claim immediately associated with the citation. Code support as supported, partial, conflict, inaccessible, or unverifiable. Save a short rationale and the source’s visible publication or update date when freshness matters.

Claim-support decision rules
DecisionUse whenDo not use when
supportedThe accessible passage supports the material claim in contextThe page merely discusses the same topic
partialThe source supports only part of a compound claim or needs a narrower qualifierThe unsupported portion is immaterial
conflictThe source contradicts the claim or states a materially different boundaryTwo sources simply use different terminology
inaccessibleLogin, block, removal, or technical failure prevents reviewThe reviewer chose not to open the source
unverifiableNo precise claim-passage relationship can be establishedA clear conflict is present

Require action language for a recommendation

Recommendation is a stronger label than mention or positive description. Require three elements: a stated reader task, an identifiable entity, and action-oriented wording that proposes the entity for that task. Preserve the wording and conditions instead of reducing it immediately to yes or no.

Recommendation strength is independent of citation status
StateOperational ruleTypical wording
noneNo action is proposedDefinition, neutral mention, or source attribution only
considerationEntity is placed in a candidate set without a task-fit judgment“Options include A, B, and C”
conditionalEntity is proposed only when declared conditions apply“Choose A if you need local processing”
explicitAnswer directly proposes the entity for the stated task“For this requirement, use A”
againstAnswer advises the reader not to choose the entity for the task“Avoid A when audit logs are required”

Positive adjectives are not enough. “Popular,” “well known,” or “feature-rich” may be descriptive or comparative. A recommendation needs an action or choice relationship.

Keep comparison and ordering separate

An entity can be included in a comparison without being recommended. Record its role as baseline, alternative, shortlist, winner, loser, or neutral comparator only when the answer supports that label. Preserve ordering as an observed presentation property, not a ranking claim.

Lists may be alphabetical, grouped by feature, or generated differently across runs. First position does not become “rank one” unless the answer explicitly defines and applies a ranked criterion.

Separate the source owner from the recommended entity

The cited publisher and the recommended product may be different organizations. Record the source owner, page subject, and recommended entity separately. A review site can support a product comparison; a documentation page can support a technical fact about a competitor; a government source can support a regulatory boundary without being a commercial option.

This is the row that exposes cited-but-not-recommended and recommended-but-not-cited outcomes instead of hiding them inside a single score.

Normalize repeated and redirected citations

Store every displayed citation, then resolve it to a final URL and normalized URL. Count repeated placements according to a declared rule. Page-level citation prevalence usually treats duplicate placements of one normalized URL in one answer as one cited page, while claim-support review may keep each placement because the associated claims differ.

Do not drop failed redirects, removed pages, paywalls, or login gates. They belong in the denominator for destination review.

Use a denominator that matches the question

Choose the denominator before calculating a rate
QuestionNumeratorDenominator
How often was the entity mentioned?Entity-answer rows coded named or describedEligible entity-answer rows
How often was the entity cited?Eligible rows with at least one entity-owned normalized citationEligible entity-answer rows
How often did citations support their claims?Supported citation-claim decisionsReviewable citation-claim decisions
How often was the entity recommended?Conditional or explicit recommendation rows under the declared ruleEligible entity-answer rows for recommendation-intent queries
How often did a citation lead to a visit?Attributed referral sessions under the declared methodDo not infer from answer rows; use the analytics contract

Always report counts beside percentages. A 50% recommendation rate means little if it represents one recommendation in two eligible rows, and it is invalid if informational prompts were mixed into a recommendation denominator without a declared reason.

Preserve refusals, errors, and excluded runs

Save blocked answers, tool failures, empty results, ambiguous entities, and incomplete captures. Give each row an exclusion reason and keep it out of the relevant rate only under a predeclared rule. Never replace a failed run with a more favorable prompt after seeing the answer.

Treat repeated runs as variability evidence

Repeat the same frozen prompt and environment on a declared schedule. A repeated answer is another observation, not proof of a stable rank. Report platform, surface, model label, time, location, and personalization alongside the distribution.

The prompt-family visibility study explains why the visible prompt may not equal the retrieval query. The small-experiment method helps predeclare the stopping rule and preserve failures.

Add adjudication where labels can change the conclusion

Have a second coder review every material disagreement and a sample of agreements. Preserve the original codes, reviewer IDs, adjudication outcome, and rationale. Calculate agreement only when both coders used the same frozen definitions and independent copies of the evidence.

Use the browser-local outcome coder

The AI citation and recommendation outcome coder stores rows only in the current browser tab. It makes no network requests and exports a CSV when you choose. It rejects contradictions such as a supported citation with no citation URL, or an explicit recommendation with no preserved wording.

Use the accompanying 48-field CSV ledger when you need the full environment, source-resolution, support, recommendation, adjudication, and recheck fields. Its six example rows are not observations; remove or replace every row marked EXAMPLE-REMOVE.

Apply platform documentation only to its own surface

Google’s generative AI Search guide describes links that support information in AI features and says indexing or serving is not guaranteed. OpenAI documents citations and relevant source links in ChatGPT Search. Bing documents aggregate citation activity across named Microsoft and partner experiences. Perplexity documents domain-level source labels. These are different products and metric contracts.

Do not merge their fields into a cross-platform leaderboard unless the study defines a compatible observation unit and preserves unavailable fields. A source link, a dashboard citation count, a source label, and an explicit recommendation are not interchangeable events.

Report a cross-tab before any composite

Start with a citation-by-recommendation cross-tab. Show cited and not cited on one axis; explicit or conditional recommendation, consideration, none, and against on the other. Add a separate support table and a list of excluded runs. This keeps the consequential cases visible.

Minimum reporting cross-tab
Citation stateRecommendedConsiderationNoneAgainst
CitedReport countReport countReport countReport count
Not citedReport countReport countReport countReport count

If stakeholders still require a composite, publish the formula, weights, missing-data rules, and component values. Never let the combined number replace the component table.

What this method does not prove

This coding method describes observed answers under recorded conditions. It does not reveal a platform’s hidden retrieval or ranking logic, establish that a content change caused a later citation, measure every unlinked mention, or prove that a recommendation caused a visit or sale. It also does not turn example rows into a benchmark.

Complete the audit only when the evidence can be reopened

A run is complete when the raw answer, exact prompt, environment, entity identity, displayed and final source URLs, citation-support decision, recommendation wording, exclusions, and coder decision can be reopened by a reviewer. If any material field is missing, report the gap instead of filling it from memory.

The practical conclusion is simple: citation is source evidence; recommendation is task-fit advice. Measure both, but never make one stand in for the other.

Two-axis measurement

Cited and recommended should be reported as separate events

A brand can be linked as evidence without being recommended, or recommended without receiving a visible citation.

Cited, not recommended
The site supports a claim but is not presented as the user’s choice.
Recommended, not cited
The brand is named, but the answer provides no inspectable source link.
Both
The brand is presented as a choice and linked as evidence.
Neither
No visible presence in the eligible answer.

My takeaway: This two-axis view makes the commercial interpretation much clearer than a single AI visibility score.

Keep learning

Continue this topic

Community discussion

Discuss: How to Measure AI Citations and Recommendations Separately

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.