How to Measure the Noise Floor of an AI Visibility Score
A 1,200-observation research protocol for measuring how AI mentions, recommendations, and citations move when the monitored site does not change.
If an AI visibility score rises while the site, prompt list, and market stay unchanged, the movement is not automatically progress. The answer engine may have produced a different answer, the monitoring product may have sampled another run, or the score may have crossed an internal classification boundary.
This research protocol measures that natural wobble before any optimization experiment begins. It publishes the denominator and collection rules now. It publishes no result because the 1,200 planned observations have not been collected.
The decision this test should unlock
Teams need a threshold for deciding whether a reported change deserves investigation. A one-point move may be meaningful in one instrument and ordinary variation in another. The threshold should come from repeated no-change observations on the same surface, not from a vendor’s color change or a community anecdote.
Two waves create 1,200 answer observations
| Control | Planned value | Why it is fixed |
|---|---|---|
| Prompt panel | 30 buyer and informational prompts | Preserves the questions being measured |
| Repeated runs | 5 per prompt and surface | Exposes within-wave answer variation |
| Consumer surfaces | ChatGPT web, Gemini web, Perplexity web, and Microsoft Copilot web | Keeps API and consumer behavior separate |
| Wave 1 | 600 eligible observations | Creates the first no-change distribution |
| Wave 2 | 600 observations after 48 hours | Measures short-interval drift without a site release |
| Market | One country and language | Limits location and language mixing |
The denominator is the planned number of eligible answer attempts, including refusals and failed runs. A silent rerun would remove inconvenient zeros and make the instrument look more stable than it was.
Freeze the site before the first wave
Create a release log for the monitored domain. Record the deployment identifier, robots rules, canonical URLs, structured-data state, sitemap timestamp, and content hashes for the pages likely to appear. Avoid publishing, updating, redirecting, or changing those pages until the second wave completes.
If an unavoidable change occurs, preserve it and mark the affected prompt families. Do not describe the interval as “no change.” A clean null condition is more valuable than a larger sample with an unknown intervention.
Record events before calculating scores
- Mention
- The answer names the target entity under a written matching rule.
- Recommendation
- The answer presents the entity as suitable for the prompt’s task or decision, with conditions preserved.
- Citation
- The answer exposes a source reference that resolves to the target domain or page.
- Source URL
- The displayed and resolved destination are both retained.
- Failure
- A refusal, timeout, blocked run, or missing answer remains in the eligible attempt ledger.
Also record surface, visible model label, login state, country, language, timestamp, and run order. Record position only when the interface exposes an order that can be defined consistently. Do not infer a ranking from citation layout.
Download the 1,200-row collection ledger. Every starter row is labeled EXAMPLE-REMOVE and contains no result.
Three numbers will describe the noise floor
- Within-wave agreement: for each prompt and surface, how often do the five runs agree on mention, recommendation, and citation?
- Between-wave movement: how much does each rate change after 48 hours with the site frozen?
- Prompt-level stability: how many prompts keep the same majority outcome across both waves?
Report counts beside percentages and keep the four surfaces separate. The final study can then define a practical investigation band: movement inside the observed no-change range is a monitoring signal, while movement outside it deserves a page, prompt, or platform review. It will still not prove the cause.
Do not compare two different instruments as one series
A vendor dashboard can use a stored prompt corpus, a generated prompt set, scheduled consumer-interface runs, APIs, or a mixture. Adobe, for example, documents an instant corpus-based AI Visibility view powered by Semrush Enterprise AIO separately from prompts an organization chooses to track. Those are useful products, but they are not automatically the same observation instrument.
This protocol uses named consumer surfaces and preserves each answer. A future comparison with a commercial score must disclose prompt source, engine set, market, refresh schedule, denominator, missing-run treatment, and whether the underlying answers can be exported.
A 960-answer study shows why one run is not a score
Prefer ran 80 questions three times across four API-based answer engines on September 13, producing 960 answers in a 16-minute window. Its aggregate repeat-run source overlap was 96% for Perplexity, 68% for Claude, 37% for ChatGPT and 35% for Gemini.
The gap is large enough to change a reporting rule. A single ChatGPT or Gemini run in that sample would have captured only one realization of a relatively unstable source set. Repetition is not a cosmetic confidence feature; it changes which domains can enter the numerator.
| Engine | Overlap | Protocol response |
|---|---|---|
| Perplexity | 96% | Retain repetitions, but do not assume the same stability on another prompt set. |
| Claude | 68% | Report prompt-level agreement beside the aggregate rate. |
| ChatGPT | 37% | Do not label one captured source set as the engine’s stable answer. |
| Gemini | 35% | Preserve failures and repeat order so volatility remains visible. |
The study used US-English questions biased toward AI search, GEO and AEO; API behavior is not the same instrument as consumer apps. Prefer sells AI-visibility software and released aggregate files rather than answer-level rows. Use the numbers as a design warning, not as a universal platform benchmark. Review the study and downloads.
Audit the prompt set before trusting the score
A September 2026 preprint tested 1,851 synthetic prompts against 322 authentic prompts, but only 165 authentic prompts had static source URLs that could be evaluated directly. The remaining 157, or 49%, were not directly scorable under that source-based setup. That missing half is a warning: a clean benchmark can become easier by excluding the queries that are hardest to verify.
| Check | Reported evidence | Protocol response |
|---|---|---|
| Query length | Synthetic prompts averaged 15.7 words; authentic prompts averaged 6.8. Cliff’s delta was .917. | Publish the full length distribution and weight common real queries appropriately. |
| Source concentration | The top source held 5.5% of synthetic references versus 25.3% of authentic references; top-three share was 15.1% versus 44.4%. | Report top-source share, top-three share and category concentration. |
| Retrieval performance | Hybrid Hit@5 fell from .896 on synthetic prompts to .527 on authentic prompts. | Keep authentic and synthetic results separate instead of blending them. |
| Latency tradeoff | Dense retrieval reached .545 Hit@5 on authentic prompts in 1.57 seconds; hybrid reached .527 in 2.70 seconds. | Measure quality and latency together; the more complex pipeline was not automatically better. |
| Evaluability | Only 165 of 322 authentic prompts had static source URLs. | Publish the excluded share and explain how non-URL answers are handled. |
The paper reports a Jensen-Shannon divergence of .203 bits between the source-category distributions. It also found the hybrid configuration was eight times slower than the fastest tested configuration. These results do not define a universal score. They show why prompt realism, source concentration, missing labels and latency must sit beside the headline visibility number.
Source: preprint on synthetic and authentic query evaluation for retrieval systems.
Current status
The protocol is frozen for collection. No SearchEngineAnswer result, stability percentage, or vendor comparison is reported on this page yet. The page will be materially updated only after both 600-observation waves, retained failures, quality checks, and the derived calculations are complete.
The question was prompted by practitioner discussions about disagreement among AI visibility tools and sensitivity to prompt design. Those discussions are leads, not benchmarks: reporting ChatGPT visibility and tracking AI citations.
Keep learning
Continue this topic
Next in this topic
Q2D-Web Shows Why AI Citations Depend on First-Stage Retrieval
Earlier in this topic
The 16% Synthetic-Source Finding Depends on an AI Detector
Research
Ask a question or join the discussion