What 7,473 AI Citations Can and Cannot Tell Us

BrandGhost's latest Observatory run offers useful signals about citation age and engine overlap, but its date and source-category coverage materially limit broader claims.

Sonar the Answer Whale compares 7,473 citations with the 2,600 dated and 8.1 percent classified subsets

Direct answer: BrandGhost’s latest Observatory run contains 7,473 AI citations and useful signals about source age and cross-engine overlap. It also has large coverage gaps: only 2,600 citations had a detectable publication date, and 91.9% were not assigned to an identifiable source category.

The dataset supports questions about its measured sample. It does not support a universal ranking recipe for ChatGPT, Claude, Gemini or Perplexity.

The headline findings

The BrandGhost AI Discovery Observatory reported 7,473 citations in its September 10, 2026 run. For citations with detectable dates, the median cited page age was 249 days, 20.9% were under 90 days old and 40.4% were more than one year old.

It also reported about 12.9% average source overlap across comparable engine pairs and 50.7% of comparable recommendation sets with no shared recommended brands.

Two denominators change the interpretation

Coverage audit for the latest Observatory run
Metric family Available records Coverage of 7,473 Safe interpretation
All analyzed citations 7,473 100% Run size
Citations with detectable dates 2,600 34.8% Age findings apply to the dated subset
Citations assigned a source category About 8.1% 8.1% Category shares describe a limited classified layer
Unclassified citations 91.9% 91.9% The dominant category is unknown, not “other”

The 34.8% figure is 2,600 divided by 7,473. The 8.1% classified share is 100% minus the reported 91.9% unclassified share. Both are SearchEngineAnswer calculations using the live page’s published values.

The age result rejects a simple freshness rule

Among citations with detectable dates, 5.7% were 0 to 30 days old, 15.2% were 31 to 90 days, 17.7% were 91 to 180 days and 20.9% were 181 to 365 days. The remaining reported buckets were one to two years at 13%, two to three years at 7% and more than three years at 20.4%.

That distribution shows that older pages can appear in the sample. It does not show that age caused citation selection. Topic, query, authority, availability and the ability to detect a publication date can all affect the observed set.

Low overlap is a portfolio signal, not proof of randomness

BrandGhost reports 15% source overlap for Perplexity and Claude, 12% for ChatGPT and Claude, and 11.8% for Perplexity and ChatGPT. Gemini is excluded because its citation representation is not directly comparable.

A low overlap measure means engine pairs often cited different domains for comparable prompts in this sample. It does not reveal whether differences came from retrieval indexes, query rewriting, answer policy, geography, account state or timing. Use it to justify multi-engine measurement, not to claim each engine has a fixed source preference.

Reddit’s share needs the unclassified denominator beside it

The Observatory lists Reddit at 245 citations, or 3.3% of the full run, and calls it the largest individually identifiable category. Because 91.9% of citations remain unclassified by source category, that statement should not become “Reddit is the largest source type in AI answers.” The page does not identify most source categories yet.

This also explains why the current 3.3% should not be joined casually to a different vendor’s earlier ChatGPT-only series. Our Reddit citation-share analysis covers another dataset with a different engine scope and time window.

How to use the dataset responsibly

  1. Record the run date and retrieval time because the Observatory is live.
  2. Copy the denominator beside every percentage.
  3. Label derived calculations separately from vendor-reported metrics.
  4. Keep dated-subset findings out of all-citation claims.
  5. Do not compare engine fields the publisher says are not comparable.
  6. Use the results as a benchmark sample, not a market-wide census.
  7. Repeat the audit when the classification or prompt mix changes.

Download the coverage audit

Download the AI citation dataset audit CSV. It records metric, numerator, denominator, coverage, engine scope, vendor statement, derived calculation and claim boundary.

Rows marked EXAMPLE-REMOVE are instructions. Replace them with values from the run you are analyzing.

The practical decision

The best use of the Observatory is to form sharper questions: how source age varies by engine, whether overlap changes across runs and where classification coverage improves. For a brand’s own decisions, pair that benchmark with a controlled prompt set and the separate citation-versus-recommendation protocol in our AI visibility measurement guide.

Evidence boundary

All vendor-reported values were retrieved from the live Observatory on September 12, 2026 and describe its September 10 run. BrandGhost operates a commercial visibility product. SearchEngineAnswer calculated coverage percentages but did not obtain the underlying citation rows.

Keep learning

Continue this topic

Community discussion

Discuss: What 7,473 AI Citations Can and Cannot Tell Us

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.