How to Measure the Noise Floor of an AI Visibility Score

A 1,200-observation research protocol for measuring how AI mentions, recommendations, and citations move when the monitored site does not change.

Sonar measures a noise-floor band between two waves of AI answer cards while the monitored website remains locked and unchanged.

If an AI visibility score rises while the site, prompt list, and market stay unchanged, the movement is not automatically progress. The answer engine may have produced a different answer, the monitoring product may have sampled another run, or the score may have crossed an internal classification boundary.

This research protocol measures that natural wobble before any optimization experiment begins. It publishes the denominator and collection rules now. It publishes no result because the 1,200 planned observations have not been collected.

The decision this test should unlock

Teams need a threshold for deciding whether a reported change deserves investigation. A one-point move may be meaningful in one instrument and ordinary variation in another. The threshold should come from repeated no-change observations on the same surface, not from a vendor’s color change or a community anecdote.

Two waves create 1,200 answer observations

Planned collection design
ControlPlanned valueWhy it is fixed
Prompt panel30 buyer and informational promptsPreserves the questions being measured
Repeated runs5 per prompt and surfaceExposes within-wave answer variation
Consumer surfacesChatGPT web, Gemini web, Perplexity web, and Microsoft Copilot webKeeps API and consumer behavior separate
Wave 1600 eligible observationsCreates the first no-change distribution
Wave 2600 observations after 48 hoursMeasures short-interval drift without a site release
MarketOne country and languageLimits location and language mixing

The denominator is the planned number of eligible answer attempts, including refusals and failed runs. A silent rerun would remove inconvenient zeros and make the instrument look more stable than it was.

Freeze the site before the first wave

Create a release log for the monitored domain. Record the deployment identifier, robots rules, canonical URLs, structured-data state, sitemap timestamp, and content hashes for the pages likely to appear. Avoid publishing, updating, redirecting, or changing those pages until the second wave completes.

If an unavoidable change occurs, preserve it and mark the affected prompt families. Do not describe the interval as “no change.” A clean null condition is more valuable than a larger sample with an unknown intervention.

Record events before calculating scores

Mention
The answer names the target entity under a written matching rule.
Recommendation
The answer presents the entity as suitable for the prompt’s task or decision, with conditions preserved.
Citation
The answer exposes a source reference that resolves to the target domain or page.
Source URL
The displayed and resolved destination are both retained.
Failure
A refusal, timeout, blocked run, or missing answer remains in the eligible attempt ledger.

Also record surface, visible model label, login state, country, language, timestamp, and run order. Record position only when the interface exposes an order that can be defined consistently. Do not infer a ranking from citation layout.

Download the 1,200-row collection ledger. Every starter row is labeled EXAMPLE-REMOVE and contains no result.

Three numbers will describe the noise floor

  1. Within-wave agreement: for each prompt and surface, how often do the five runs agree on mention, recommendation, and citation?
  2. Between-wave movement: how much does each rate change after 48 hours with the site frozen?
  3. Prompt-level stability: how many prompts keep the same majority outcome across both waves?

Report counts beside percentages and keep the four surfaces separate. The final study can then define a practical investigation band: movement inside the observed no-change range is a monitoring signal, while movement outside it deserves a page, prompt, or platform review. It will still not prove the cause.

Do not compare two different instruments as one series

A vendor dashboard can use a stored prompt corpus, a generated prompt set, scheduled consumer-interface runs, APIs, or a mixture. Adobe, for example, documents an instant corpus-based AI Visibility view powered by Semrush Enterprise AIO separately from prompts an organization chooses to track. Those are useful products, but they are not automatically the same observation instrument.

This protocol uses named consumer surfaces and preserves each answer. A future comparison with a commercial score must disclose prompt source, engine set, market, refresh schedule, denominator, missing-run treatment, and whether the underlying answers can be exported.

A 960-answer study shows why one run is not a score

Prefer ran 80 questions three times across four API-based answer engines on September 13, producing 960 answers in a 16-minute window. Its aggregate repeat-run source overlap was 96% for Perplexity, 68% for Claude, 37% for ChatGPT and 35% for Gemini.

The gap is large enough to change a reporting rule. A single ChatGPT or Gemini run in that sample would have captured only one realization of a relatively unstable source set. Repetition is not a cosmetic confidence feature; it changes which domains can enter the numerator.

Repeat-run source overlap reported by Prefer
EngineOverlapProtocol response
Perplexity96%Retain repetitions, but do not assume the same stability on another prompt set.
Claude68%Report prompt-level agreement beside the aggregate rate.
ChatGPT37%Do not label one captured source set as the engine’s stable answer.
Gemini35%Preserve failures and repeat order so volatility remains visible.

The study used US-English questions biased toward AI search, GEO and AEO; API behavior is not the same instrument as consumer apps. Prefer sells AI-visibility software and released aggregate files rather than answer-level rows. Use the numbers as a design warning, not as a universal platform benchmark. Review the study and downloads.

Audit the prompt set before trusting the score

A September 2026 preprint tested 1,851 synthetic prompts against 322 authentic prompts, but only 165 authentic prompts had static source URLs that could be evaluated directly. The remaining 157, or 49%, were not directly scorable under that source-based setup. That missing half is a warning: a clean benchmark can become easier by excluding the queries that are hardest to verify.

Prompt-set checks that belong beside every AI visibility score
CheckReported evidenceProtocol response
Query lengthSynthetic prompts averaged 15.7 words; authentic prompts averaged 6.8. Cliff’s delta was .917.Publish the full length distribution and weight common real queries appropriately.
Source concentrationThe top source held 5.5% of synthetic references versus 25.3% of authentic references; top-three share was 15.1% versus 44.4%.Report top-source share, top-three share and category concentration.
Retrieval performanceHybrid Hit@5 fell from .896 on synthetic prompts to .527 on authentic prompts.Keep authentic and synthetic results separate instead of blending them.
Latency tradeoffDense retrieval reached .545 Hit@5 on authentic prompts in 1.57 seconds; hybrid reached .527 in 2.70 seconds.Measure quality and latency together; the more complex pipeline was not automatically better.
EvaluabilityOnly 165 of 322 authentic prompts had static source URLs.Publish the excluded share and explain how non-URL answers are handled.

The paper reports a Jensen-Shannon divergence of .203 bits between the source-category distributions. It also found the hybrid configuration was eight times slower than the fastest tested configuration. These results do not define a universal score. They show why prompt realism, source concentration, missing labels and latency must sit beside the headline visibility number.

Source: preprint on synthetic and authentic query evaluation for retrieval systems.

Current status

The protocol is frozen for collection. No SearchEngineAnswer result, stability percentage, or vendor comparison is reported on this page yet. The page will be materially updated only after both 600-observation waves, retained failures, quality checks, and the derived calculations are complete.

The question was prompted by practitioner discussions about disagreement among AI visibility tools and sensitivity to prompt design. Those discussions are leads, not benchmarks: reporting ChatGPT visibility and tracking AI citations.

Keep learning

Continue this topic

Community discussion

Discuss: How to Measure the Noise Floor of an AI Visibility Score

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.