Can an AI Visibility Score Predict Referral Traffic? A 30-Day Test
A preregistered protocol for testing whether a vendor visibility score predicts page-matched AI referral traffic. Results pending.
Protocol status: testing pending. This 30-day protocol asks whether changes in a selected AI visibility score predict changes in observed AI referral traffic for the same pages. It does not assume that a score is accurate, that a citation creates a click, or that correlation establishes an AEO effect.
The output is a dated paired dataset: vendor observation, prompt cohort, cited URL, server arrival, analytics session and business outcome. Publish the rows and exclusions before publishing a conclusion.
Define the prediction
Choose one vendor metric and copy its current definition. Record its denominator, prompt source, platforms, countries, frequency and aggregation. The primary prediction should be directional and page-matched: when the score or cited-page share rises for the frozen cohort, do qualifying referral arrivals to those URLs rise in the following declared window?
| Layer | Required field |
|---|---|
| Vendor observation | Prompt, platform, country, run date, score, mention and citation |
| Page identity | Canonical URL, template and treatment status |
| Server evidence | Landing request, referrer, timestamp and bot exclusion |
| Analytics evidence | Session source, landing page, consent and engagement rule |
| Business outcome | Qualified lead, signup, sale or another predeclared event |
Freeze the cohort and comparison
- Select at least 20 prompts and the pages expected to satisfy them.
- Keep a control set of comparable pages without an editorial change.
- Preserve zero-score, zero-citation and zero-traffic observations.
- Save daily vendor exports where available and weekly snapshots otherwise.
- Record platform releases, campaigns, tracking changes and news demand.
- Do not change prompt definitions or page membership mid-test.
A 30-day window is useful for detecting obvious alignment and measurement failures; it is usually too short for a durable causal claim. Referral traffic can be sparse, prompt execution can vary, and a vendor’s sample may not resemble real user demand.
Analyze without forcing a win
Plot the score and referral series, then inspect page-matched citation-to-arrival events. Report rank correlation only when enough nonzero observations exist, with uncertainty and lag choices declared. Also report false positives: high score with no observed arrival, and false negatives: arrivals to pages absent from the monitored sample.
A useful score may be directionally associated with citations but unrelated to visits. It may help prioritize research even when it cannot predict traffic. A null relationship does not prove the product is useless; it proves the tested score did not predict the declared outcome in this cohort and period.
Use the GA4 referral workflow and server logs together. Read Cited Is Not Recommended before selecting a brand outcome.
Predeclare the denominator
The minimum design above—20 prompts, three answer engines and 30 daily observations—creates 1,800 prompt-engine-day cells. The completed dataset must contain 1,800 rows or name every missing cell. A blank response is not the same event as an observed zero mention, zero citation or zero referral.
| Component | Calculation | Required output |
|---|---|---|
| Prompt observations | 20 × 3 engines × 30 days | 1,800 expected cells |
| Completeness | observed cells ÷ 1,800 | Report percentage and missing dates |
| Citation precision | cited cells with a matched arrival ÷ all cited cells | State the referral window and URL rule |
| Referral recall | matched arrivals represented in monitored cells ÷ all qualifying arrivals | Expose traffic outside the vendor sample |
Precision and recall do not prove causality, but they reveal whether the score observes the same channel as the referral data. Publish a 2×2 table of cited/not cited by arrival/no arrival before reporting a correlation. If most cells are zero, use counts and exact intervals rather than a smooth trend line that suggests more information than the sample contains.
Publication template
- Protocol version and registration date.
- Vendor metric definition and retrieval date.
- Prompt/page cohort and exclusions.
- Raw daily or weekly rows, including zeros.
- Correlation, lag and sensitivity results.
- False positives, false negatives and tracking gaps.
- What the result can and cannot support.
This page remains testing pending until a complete 30-day dataset is attached.
Ask a question or join the discussion