How to Measure AI Citations and Recommendations Separately
Separate citation, publisher naming, claim support, recommendation and referral outcomes, then calculate attribution yield only from matched saved answers.
A citation tells you that an answer visibly referenced a source. It does not tell you that the answer recommended the source owner, ranked its product first, sent a visit, or changed a purchase decision. Measure those outcomes separately, using the AI answer as the unit of observation and the entity as the unit of classification.
What changed in this guide: the earlier page preregistered a repeated-run experiment, but no completed dataset was published. This revision does not invent that missing dataset. It replaces the unfinished promise with a reusable coding method, a browser-local outcome coder, and a 48-field CSV ledger whose illustrative rows are marked EXAMPLE-REMOVE.
The short answer
For every entity in every saved answer, record at least four independent decisions: whether the entity is mentioned, whether one of its pages is cited, whether the cited page supports the nearby claim, and whether the answer recommends the entity for the reader’s task. Add comparison role, ordering, referral evidence, and exclusions as separate fields.
This prevents a common measurement error. A publisher can supply evidence for an answer while a competitor receives the recommendation. The reverse can also happen: a product may be recommended from model knowledge or an unlinked passage while its website receives no visible citation.
Why one AI visibility score hides the outcome
A composite score can combine observations that have different meanings and denominators. Citation presence is a source-use observation. Recommendation is an action-oriented statement about an entity. Referral is a recorded visit. Conversion is a later business event. Adding them together does not create a stronger metric; it removes the boundary between them.
| Layer | Observable evidence | What it cannot establish alone |
|---|---|---|
| Citation | Visible source link or attribution in a saved answer | Endorsement, recommendation, ranking, traffic, or conversion |
| Claim support | Cited page and passage support, partially support, conflict with, or fail to verify the nearby claim | Why the system selected the source |
| Recommendation | Action wording that proposes an entity for the stated task, with any conditions | Whether the user followed the advice |
| Referral | Analytics or server evidence of a visit from the answer surface | Every citation or unclicked mention |
| Conversion | Permissioned first-party outcome joined through a declared attribution method | Universal causal impact of AI visibility |
Bing’s current AI Performance documentation makes the boundary unusually explicit: its report measures visible citation activity, not ranking, authority, importance, clicks, or traffic. That documented metric is useful, but it cannot answer the recommendation question.
Do not turn cited pages and brand mentions into one ratio
A publisher reported on Reddit roughly 700 cited pages and five written brand mentions across several AI surfaces. The post raises a useful question, but the two numbers do not form a defensible 140-to-1 industry ratio.
“Cited pages” may be a deduplicated site-level inventory while “brand mentions” may count answers, prompts or monitoring events. The post does not supply the export, observation window, eligible answer count, prompt panel, repeated-run rule or cross-platform deduplication method. Its platform figures also add to more than the rounded cited-page total, which may be harmless rounding or overlap but prevents an exact calculation.
The useful signal is the mismatch: publisher pages can appear as source material without the answer naming the publisher in prose. Measure that relationship inside the same saved answers before testing possible explanations.
Calculate attribution yield from matched answers
| Measure | Numerator | Denominator | Reader question |
|---|---|---|---|
| Publisher naming yield | Answers that name the publisher in prose | Eligible answers citing an owned page | When a page supplies evidence, how often does the outlet receive visible credit? |
| Entity recommendation yield | Cited answers with a conditional or explicit recommendation | Eligible cited answers from recommendation-intent prompts | How often does citation accompany task-fit advice? |
| Support yield | Reviewable citations whose page supports the linked claim | Reviewable citation-claim decisions | How often is visible sourcing substantively useful? |
| Referral rate | Attributed sessions under the declared analytics method | Use an analytics denominator, not cited-answer rows | Did any visible exposure become a measurable visit? |
Report each count beside its rate. If publisher naming is not available from the same answer captures as citation presence, label the comparison unavailable instead of combining two vendor dashboard totals.
Use divergence as a diagnostic, not a score
| Observed answer | What to inspect next | What not to conclude |
|---|---|---|
| Owned page cited, publisher unnamed | Source-card label, byline, organization identity near the original evidence, and whether the answer names only the subject. | That structured data caused the omission or that adding the brand name will force attribution. |
| Publisher named, no owned citation | Other cited sources, unlinked answer wording and whether the name came from comparison context. | That the publisher’s own page was retrieved. |
| Cited and named, not recommended | The claim being supported and the prompt’s actual decision intent. | That visible credit is a commercial endorsement. |
| Recommended, no visible citation | Recommendation wording, conditions, source panel and repeat-run variability. | That model memory, hidden retrieval or any specific source caused the choice. |
A content or entity change can be tested against this matrix, but no markup, byline or writing pattern guarantees that an answer engine will name or recommend the publisher. The experiment should preserve a matched pre-change prompt set and change one identity or evidence element at a time.
Choose the unit before opening an answer
Use one row per entity-answer observation. If one answer discusses four products, create four rows linked by the same run ID. If the answer cites three pages for one entity, either add child citation rows or preserve the three citation IDs in a separate citation table. Do not count one answer as four independent runs.
Give every study, query, run, entity, response artifact, and citation a stable ID. Those keys let you revise a coding decision without losing the raw answer or silently duplicating a run.
Freeze the environment before coding
Record the platform, product surface, visible model label, plan, signed-in state, locale, geography, device, personalization or memory state, search state, and retrieval time. Save the exact prompt and full answer before opening source links.
OpenAI’s current ChatGPT Search documentation says web answers may include citations and that results can be incomplete, outdated, or incorrect. It also documents query rewriting and location-sensitive behavior. A later response can differ because the retrieval query, available web pages, location, memory, or product changed; the saved environment is part of the evidence.
Code the mention before the recommendation
Start with the lowest-level entity decision. Mark the canonical entity as absent, named, described without an exact name, or ambiguous. Preserve the exact sentence and any alias. A linked page owned by an entity is not automatically a mention of that entity in the answer text.
| State | Rule | Example pattern |
|---|---|---|
absent | No defensible reference to the entity | The answer and source panel omit it |
named | Canonical name or registered alias appears | “Tool A includes…” |
described | Identity is clear from context without the canonical name | “The vendor’s free crawler…” after an unambiguous introduction |
ambiguous | Similar names or missing context prevent a reliable match | A generic product name with no owner or URL |
Record citation presence without adding meaning
Mark a citation only when the saved surface exposes a source link or explicit attribution. Preserve the displayed URL, final URL after redirects, source label, placement, and cited claim. Normalize fragments, tracking parameters, mobile variants, and redirects without deleting the original values.
Do not convert a source-card position into a quality grade. Perplexity’s source-label documentation says its labels describe a domain, not an individual page or claim, and are not endorsements. An unlabeled domain is not a negative judgment.
Review whether the cited page supports the claim
Citation presence and citation support need separate columns. Open the final page and compare the relevant passage with the claim immediately associated with the citation. Code support as supported, partial, conflict, inaccessible, or unverifiable. Save a short rationale and the source’s visible publication or update date when freshness matters.
| Decision | Use when | Do not use when |
|---|---|---|
supported | The accessible passage supports the material claim in context | The page merely discusses the same topic |
partial | The source supports only part of a compound claim or needs a narrower qualifier | The unsupported portion is immaterial |
conflict | The source contradicts the claim or states a materially different boundary | Two sources simply use different terminology |
inaccessible | Login, block, removal, or technical failure prevents review | The reviewer chose not to open the source |
unverifiable | No precise claim-passage relationship can be established | A clear conflict is present |
Require action language for a recommendation
Recommendation is a stronger label than mention or positive description. Require three elements: a stated reader task, an identifiable entity, and action-oriented wording that proposes the entity for that task. Preserve the wording and conditions instead of reducing it immediately to yes or no.
| State | Operational rule | Typical wording |
|---|---|---|
none | No action is proposed | Definition, neutral mention, or source attribution only |
consideration | Entity is placed in a candidate set without a task-fit judgment | “Options include A, B, and C” |
conditional | Entity is proposed only when declared conditions apply | “Choose A if you need local processing” |
explicit | Answer directly proposes the entity for the stated task | “For this requirement, use A” |
against | Answer advises the reader not to choose the entity for the task | “Avoid A when audit logs are required” |
Positive adjectives are not enough. “Popular,” “well known,” or “feature-rich” may be descriptive or comparative. A recommendation needs an action or choice relationship.
Keep comparison and ordering separate
An entity can be included in a comparison without being recommended. Record its role as baseline, alternative, shortlist, winner, loser, or neutral comparator only when the answer supports that label. Preserve ordering as an observed presentation property, not a ranking claim.
Lists may be alphabetical, grouped by feature, or generated differently across runs. First position does not become “rank one” unless the answer explicitly defines and applies a ranked criterion.
Separate the source owner from the recommended entity
The cited publisher and the recommended product may be different organizations. Record the source owner, page subject, and recommended entity separately. A review site can support a product comparison; a documentation page can support a technical fact about a competitor; a government source can support a regulatory boundary without being a commercial option.
This is the row that exposes cited-but-not-recommended and recommended-but-not-cited outcomes instead of hiding them inside a single score.
Normalize repeated and redirected citations
Store every displayed citation, then resolve it to a final URL and normalized URL. Count repeated placements according to a declared rule. Page-level citation prevalence usually treats duplicate placements of one normalized URL in one answer as one cited page, while claim-support review may keep each placement because the associated claims differ.
Do not drop failed redirects, removed pages, paywalls, or login gates. They belong in the denominator for destination review.
Use a denominator that matches the question
| Question | Numerator | Denominator |
|---|---|---|
| How often was the entity mentioned? | Entity-answer rows coded named or described | Eligible entity-answer rows |
| How often was the entity cited? | Eligible rows with at least one entity-owned normalized citation | Eligible entity-answer rows |
| How often did citations support their claims? | Supported citation-claim decisions | Reviewable citation-claim decisions |
| How often was the entity recommended? | Conditional or explicit recommendation rows under the declared rule | Eligible entity-answer rows for recommendation-intent queries |
| How often did a citation lead to a visit? | Attributed referral sessions under the declared method | Do not infer from answer rows; use the analytics contract |
Always report counts beside percentages. A 50% recommendation rate means little if it represents one recommendation in two eligible rows, and it is invalid if informational prompts were mixed into a recommendation denominator without a declared reason.
Preserve refusals, errors, and excluded runs
Save blocked answers, tool failures, empty results, ambiguous entities, and incomplete captures. Give each row an exclusion reason and keep it out of the relevant rate only under a predeclared rule. Never replace a failed run with a more favorable prompt after seeing the answer.
Treat repeated runs as variability evidence
Repeat the same frozen prompt and environment on a declared schedule. A repeated answer is another observation, not proof of a stable rank. Report platform, surface, model label, time, location, and personalization alongside the distribution.
The prompt-family visibility study explains why the visible prompt may not equal the retrieval query. The small-experiment method helps predeclare the stopping rule and preserve failures.
Add adjudication where labels can change the conclusion
Have a second coder review every material disagreement and a sample of agreements. Preserve the original codes, reviewer IDs, adjudication outcome, and rationale. Calculate agreement only when both coders used the same frozen definitions and independent copies of the evidence.
Use the browser-local outcome coder
The AI citation and recommendation outcome coder stores rows only in the current browser tab. It makes no network requests and exports a CSV when you choose. It rejects contradictions such as a supported citation with no citation URL, or an explicit recommendation with no preserved wording.
Use the accompanying 48-field CSV ledger when you need the full environment, source-resolution, support, recommendation, adjudication, and recheck fields. Its six example rows are not observations; remove or replace every row marked EXAMPLE-REMOVE.
Apply platform documentation only to its own surface
Google’s generative AI Search guide describes links that support information in AI features and says indexing or serving is not guaranteed. OpenAI documents citations and relevant source links in ChatGPT Search. Bing documents aggregate citation activity across named Microsoft and partner experiences. Perplexity documents domain-level source labels. These are different products and metric contracts.
Do not merge their fields into a cross-platform leaderboard unless the study defines a compatible observation unit and preserves unavailable fields. A source link, a dashboard citation count, a source label, and an explicit recommendation are not interchangeable events.
Report a cross-tab before any composite
Start with a citation-by-recommendation cross-tab. Show cited and not cited on one axis; explicit or conditional recommendation, consideration, none, and against on the other. Add a separate support table and a list of excluded runs. This keeps the consequential cases visible.
| Citation state | Recommended | Consideration | None | Against |
|---|---|---|---|---|
| Cited | Report count | Report count | Report count | Report count |
| Not cited | Report count | Report count | Report count | Report count |
If stakeholders still require a composite, publish the formula, weights, missing-data rules, and component values. Never let the combined number replace the component table.
What this method does not prove
This coding method describes observed answers under recorded conditions. It does not reveal a platform’s hidden retrieval or ranking logic, establish that a content change caused a later citation, measure every unlinked mention, or prove that a recommendation caused a visit or sale. It also does not turn example rows into a benchmark.
Complete the audit only when the evidence can be reopened
A run is complete when the raw answer, exact prompt, environment, entity identity, displayed and final source URLs, citation-support decision, recommendation wording, exclusions, and coder decision can be reopened by a reviewer. If any material field is missing, report the gap instead of filling it from memory.
The practical conclusion is simple: citation is source evidence; recommendation is task-fit advice. Measure both, but never make one stand in for the other.
Two-axis measurement
Cited and recommended should be reported as separate events
A brand can be linked as evidence without being recommended, or recommended without receiving a visible citation.
- Cited, not recommended
- The site supports a claim but is not presented as the user’s choice.
- Recommended, not cited
- The brand is named, but the answer provides no inspectable source link.
- Both
- The brand is presented as a choice and linked as evidence.
- Neither
- No visible presence in the eligible answer.
My takeaway: This two-axis view makes the commercial interpretation much clearer than a single AI visibility score.
Keep learning
Continue this topic
Next in this topic
How to Compare YouTube Citations With Google Search Results
Earlier in this topic
Google Subscription Labels: A Publisher Verification Method
AEO & AI Search
Ask a question or join the discussion