Can an AI Visibility Score Predict Referral Traffic? A 30-Day Test

A preregistered protocol for testing whether a vendor visibility score predicts page-matched AI referral traffic. Results pending.

Sonar the Answer Whale holds back an equals sign while dated observation cards connect an AI visibility gauge to a referral bucket.

Protocol status: testing pending. This 30-day protocol asks whether changes in a selected AI visibility score predict changes in observed AI referral traffic for the same pages. It does not assume that a score is accurate, that a citation creates a click, or that correlation establishes an AEO effect.

The output is a dated paired dataset: vendor observation, prompt cohort, cited URL, server arrival, analytics session and business outcome. Publish the rows and exclusions before publishing a conclusion.

Define the prediction

Choose one vendor metric and copy its current definition. Record its denominator, prompt source, platforms, countries, frequency and aggregation. The primary prediction should be directional and page-matched: when the score or cited-page share rises for the frozen cohort, do qualifying referral arrivals to those URLs rise in the following declared window?

Keep proxy and outcome events separate
LayerRequired field
Vendor observationPrompt, platform, country, run date, score, mention and citation
Page identityCanonical URL, template and treatment status
Server evidenceLanding request, referrer, timestamp and bot exclusion
Analytics evidenceSession source, landing page, consent and engagement rule
Business outcomeQualified lead, signup, sale or another predeclared event

Freeze the cohort and comparison

  1. Select at least 20 prompts and the pages expected to satisfy them.
  2. Keep a control set of comparable pages without an editorial change.
  3. Preserve zero-score, zero-citation and zero-traffic observations.
  4. Save daily vendor exports where available and weekly snapshots otherwise.
  5. Record platform releases, campaigns, tracking changes and news demand.
  6. Do not change prompt definitions or page membership mid-test.

A 30-day window is useful for detecting obvious alignment and measurement failures; it is usually too short for a durable causal claim. Referral traffic can be sparse, prompt execution can vary, and a vendor’s sample may not resemble real user demand.

Analyze without forcing a win

Plot the score and referral series, then inspect page-matched citation-to-arrival events. Report rank correlation only when enough nonzero observations exist, with uncertainty and lag choices declared. Also report false positives: high score with no observed arrival, and false negatives: arrivals to pages absent from the monitored sample.

A useful score may be directionally associated with citations but unrelated to visits. It may help prioritize research even when it cannot predict traffic. A null relationship does not prove the product is useless; it proves the tested score did not predict the declared outcome in this cohort and period.

Use the GA4 referral workflow and server logs together. Read Cited Is Not Recommended before selecting a brand outcome.

Predeclare the denominator

The minimum design above—20 prompts, three answer engines and 30 daily observations—creates 1,800 prompt-engine-day cells. The completed dataset must contain 1,800 rows or name every missing cell. A blank response is not the same event as an observed zero mention, zero citation or zero referral.

Proposed 30-day measurement budget
ComponentCalculationRequired output
Prompt observations20 × 3 engines × 30 days1,800 expected cells
Completenessobserved cells ÷ 1,800Report percentage and missing dates
Citation precisioncited cells with a matched arrival ÷ all cited cellsState the referral window and URL rule
Referral recallmatched arrivals represented in monitored cells ÷ all qualifying arrivalsExpose traffic outside the vendor sample

Precision and recall do not prove causality, but they reveal whether the score observes the same channel as the referral data. Publish a 2×2 table of cited/not cited by arrival/no arrival before reporting a correlation. If most cells are zero, use counts and exact intervals rather than a smooth trend line that suggests more information than the sample contains.

Publication template

  • Protocol version and registration date.
  • Vendor metric definition and retrieval date.
  • Prompt/page cohort and exclusions.
  • Raw daily or weekly rows, including zeros.
  • Correlation, lag and sensitivity results.
  • False positives, false negatives and tracking gaps.
  • What the result can and cannot support.

This page remains testing pending until a complete 30-day dataset is attached.

Community discussion

Discuss: Can an AI Visibility Score Predict Referral Traffic? A 30-Day Test

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.