What 45 GEO Studies Actually Prove—and What They Do Not

A critical 2026 survey separates retrieval, citation, absorption and downstream behavior—and finds the evidence narrower than most GEO claims.

Sonar the Answer Whale sorts green evidence cards from violet uncertainty cards while puncturing an inflated claim bubble.

Direct answer: a July 2026 critical survey of 45 GEO studies finds credible evidence that already-retrieved content can change how often it is cited or used inside an answer. It does not find a reviewed technique with a stable, longitudinal, cross-platform causal effect on organic discovery or downstream behavior.

The survey is itself a preprint, but it supplies a useful correction to “GEO tactic” lists: AI visibility is a partially observable pipeline, and a result at one stage does not prove movement at the next.

Treat GEO as a pipeline

A result must name the stage it measured
StageQuestion
ActivationDid the system use web search or another retrieval path?
Crawl and indexCould the source be accessed and represented?
Retrieval and rerankingDid the source enter the candidate context?
Citation and prominenceWas it linked, and where?
Absorption and fidelityDid the answer use the source, and accurately?
BehaviorDid a person visit, trust, subscribe or buy?

A citation test that starts with a fixed context says nothing about organic retrieval. A prompt-tracker score does not prove referral traffic. A referral session does not prove that the cited passage drove the visit. Preserve these boundaries in headlines and reports.

What appears most reproducible

The survey identifies topical relevance and context position as the most reproducible levers within reviewed settings. It also reports that generic heuristics transfer poorly, competition can erode an individual publisher’s gain, and citation-oriented rewrites can impair retrieval. The foundational GEO gains remain valid inside their experimental setting, where a source was already present in a fixed context; they should not be retold as proof of durable discovery or traffic.

The right editorial conclusion is not “GEO does not work.” It is “define the stage, control the comparison and keep the inference local.” Clear definitions, direct evidence and useful structure still have reader value even when a citation effect is null.

Raise the evidence standard

  • Repeat measurements across time, prompt paraphrases and clean sessions.
  • Keep an untreated control or a credible baseline.
  • Record platform, location, account state, model or surface, and search activation.
  • Inspect answer text and citations manually instead of trusting only a score.
  • Test multi-actor interference: competitors can adopt the same change.
  • Report discovery, citation, absorption, referral and conversion separately.

The authors provide a literature matrix and search protocol with the preprint. Use those artifacts to inspect inclusion rules instead of treating the number 45 as authority by itself. For an operational measurement model, see Google, Bing and ChatGPT report different signals. For experimental discipline, use the small SEO experiment guide.

Inspect the 45-study corpus

We downloaded the authors’ ancillary literature matrix and counted its rows rather than relying only on the abstract. Of the 45 included studies, 27 are dated 2026 (60%), 12 are dated 2025, five are dated 2024 and one is dated 2023. Twenty-two rows (48.9%) have a status containing the word “preprint.”

SearchEngineAnswer calculation from the published 45-row literature matrix
Corpus sliceStudiesShare
Dated 202627 of 4560.0%
Dated 202512 of 4526.7%
Dated 20245 of 4511.1%
Dated 20231 of 452.2%
Status contains “preprint”22 of 4548.9%

The arithmetic explains why broad certainty is premature: the corpus is recent, heterogeneous and nearly half preprint-labelled. It does not invalidate those studies. It means a buyer or publisher should weight peer review, system realism, outcome stage and replication instead of counting every row as equivalent evidence. The survey also notes that heterogeneous effect sizes cannot support a meaningful pooled average.

Primary documentation

Community discussion

Discuss: What 45 GEO Studies Actually Prove—and What They Do Not

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.