Citation Count vs Citation Absorption: The AEO Metric Most Reports Miss
A public cross-platform dataset shows why the number of citations and the depth of a source’s contribution to an answer must be measured separately.
Direct answer: citation count measures how many sources an answer exposes; citation absorption asks how much a cited page contributes to the answer’s language, evidence, structure or factual support. A 2026 preprint using a public cross-platform dataset finds that citation breadth and contribution depth diverge.
A page can be linked without materially shaping the answer. Another page can contribute substantial evidence while appearing among fewer citations. AEO reporting needs both events—and should label the paper’s cross-platform findings as descriptive preprint evidence.
What the dataset contains
Zhang Kai, He Xinyue and Yao Jingang analyze the public geo-citation-lab dataset: 602 controlled prompts across ChatGPT, Google AI Overview/Gemini and Perplexity; 21,143 valid search-layer citations; 23,745 citation-level feature records; 18,151 successfully fetched pages; and 72 extracted features.
The authors report that Perplexity and Google cite more sources on average, while ChatGPT cites fewer sources but shows higher average citation influence among fetched pages. High-influence pages tend to be longer, more structured, semantically aligned and richer in extractable definitions, numerical facts, comparisons and procedures.
Those are descriptive relationships in the collected dataset. They do not establish that increasing length or adding a template will cause absorption, and they do not guarantee the same pattern after platform changes.
| Dataset field | Published count | Calculated context |
|---|---|---|
| Controlled prompts | 602 | Denominator for per-prompt arithmetic |
| Valid search-layer citations | 21,143 | 35.12 citations per prompt |
| Citation-level feature records | 23,745 | 39.44 feature records per prompt |
| Successfully fetched pages | 18,151 | 85.85 pages per 100 valid citations, if compared descriptively |
| Extracted features | 72 | Feature breadth, not 72 causal factors |
The per-prompt values are SearchEngineAnswer calculations from the reported totals. The paper uses different record types, so the 85.85 comparison must not be interpreted as a one-to-one fetch success rate unless the underlying dataset keys are joined. That record-shape warning is itself useful: dashboard denominators need a data dictionary.
Measure selection and absorption separately
| Stage | Audit question | Failure example |
|---|---|---|
| Citation selection | Did the platform search and expose the page as a source? | Relevant page never enters the source set. |
| Citation absorption | Did the answer use the page’s evidence or structure faithfully? | Page is linked, but its claim is absent or distorted. |
For each saved answer, preserve the prompt, platform, date, answer text, cited URLs and page snapshots. Map answer claims to cited passages. Mark no contribution, weak contextual contribution, direct factual support, structural influence or contradiction. Use two reviewers for ambiguous cases.
Change the editorial scorecard
- Report source breadth and page-level contribution separately.
- Track unsupported synthesis and fidelity failures.
- Keep citations, brand mentions and recommendations distinct.
- Inspect which passages contributed, not just which domains appeared.
- Pair answer evidence with referral and on-site outcomes.
This framework does not make citation count useless. Breadth can reveal source diversity, retrieval changes and competitive coverage. It becomes misleading only when the count is presented as proof that a source shaped the answer or generated value.
Read Cited Is Not Recommended for another measurement boundary, and the citation-ready content guide for page-level evidence design.
Primary documentation
- From Citation Selection to Citation Absorption — preprint revised April 29, 2026.
- Public dataset and analysis pipeline linked by the paper.
Ask a question or join the discussion