Citation Wars: Can GEO Optimization Make Content Worse?
A new simulation shows how citation-seeking rewrites can degrade quality—and how to keep AEO experiments inside factual and reader-value guardrails.
Direct answer: optimizing a page for AI citations can make the page worse when citation probability becomes the target instead of a constraint. An August 2026 preprint models that failure as a “citation war”: publishers repeatedly rewrite documents to win citations, defenses adapt, and the rewrites can introduce unsupported claims or reduce document quality.
This is a warning from a simulation, not proof that a named search product currently rewards bad content. The practical lesson is narrower: treat citation visibility as one measured outcome, while factual support, reader usefulness, retrieval, and downstream behavior remain release gates.
What the paper tested
Xu, Guo and Xiong formulate the publisher–platform interaction as a repeated Stackelberg game with partial monitoring. In their simulations, content providers use citation-seeking GEO attacks, while the generative engine applies defenses. The authors report that adaptive rewrites can bypass conventional defenses by trading away document quality and adding unsupported claims.
They propose a mechanism called VCR, based on verifiable-content rewards. Instead of only penalizing suspicious rewrites, the platform also rewards checkable factual substance. Across three benchmarks, the paper reports that VCR produced the strongest net defense-utility result under its evaluation design.
| Supported | Not supported |
|---|---|
| Citation incentives can create harmful strategic behavior in a repeated simulation. | A claim that Google, ChatGPT, Gemini or Perplexity currently uses this mechanism. |
| Checkable factual substance can be included in an incentive design. | A universal content formula that guarantees citations. |
| Platform and publisher objectives may diverge. | A finding that all GEO rewrites reduce quality. |
Recognize citation over-optimization
A page is drifting toward the wrong objective when a rewrite adds assertive language without stronger evidence, converts uncertainty into certainty, repeats quotable claims that do not help the reader, or removes conditions that make a statement accurate. Another warning sign is a test that records only whether a URL was cited and ignores whether the answer used the source faithfully.
The problem is not concise writing. Direct definitions, inspectable methods, original data and clear sourcing help people and retrieval systems. The problem begins when the page is optimized to trigger a visible event while the meaning becomes less reliable.
- Keep a claim-to-source ledger for material factual statements.
- Label experience, inference and verified fact separately.
- Preserve dates, sample definitions, exclusions and limitations.
- Measure citation, absorption, referral and reader outcome as different events.
- Reject a winning variant if support or usefulness declines.
Use a quality-preserving AEO test
- Freeze the original page, evidence set and target reader job.
- Write one hypothesized change, such as moving a supported definition closer to the question.
- Predeclare quality guardrails: unsupported-claim count, source proximity, factual review and task completion.
- Run repeated prompt observations with dates, platform state and citations preserved.
- Compare visibility with the guardrails and downstream evidence.
- Keep, revise or roll back the change. Preserve null results.
The citation-ready content guide provides an editorial baseline. The small-experiment method shows how to change one important variable without turning an observation into a ranking claim.
Put a number on the quality gate
The paper evaluates VCR across three benchmarks and reports an average 12.1 percentage-point improvement in net defense utility over the strongest baseline. That is a result for the authors’ simulated defense objective—not a 12.1% traffic, ranking or citation lift for publishers.
| Number | Type | What it means |
|---|---|---|
| 3 benchmarks | Reported | Evaluation breadth in the preprint. |
| +12.1 percentage points | Reported | Average net defense-utility advantage of VCR over the strongest baseline in that setup. |
| 0 new unsupported claims | Recommended gate | A SearchEngineAnswer revision fails review if it gains unsupported certainty. |
| 100% material claims mapped | Recommended gate | Every consequential factual claim must point to inspectable evidence. |
A practical quality-adjusted result should therefore be reported as a vector, not one score: citation change, unsupported-claim change, factual-fidelity failures and reader-task completion. If citations rise while any material claim becomes unsupported, the variant does not ship. This release rule is our editorial method; it is not a metric proposed by the paper.
Primary documentation
- Xu, Guo and Xiong: Mechanism Design for Generative Engines — preprint submitted August 11, 2026.
- Google Search Central: creating helpful, reliable, people-first content.
Ask a question or join the discussion