Citation Wars: Can GEO Optimization Make Content Worse?
A 2026 simulation shows how citation-seeking rewrites can degrade quality—and how to test AEO changes without sacrificing evidence or reader value.
Direct answer: yes;optimizing for citation probability can make content worse when the score replaces the reader’s job. A 2026 preprint models publishers repeatedly rewriting documents to win citations from an AI answer engine. In the simulation, competitive citation-seeking can reduce content quality and increase unsupported claims. The result is a warning about incentives, not proof that a named live platform behaves the same way.
A safe GEO or AEO workflow therefore needs two gates: did source selection improve, and did the page remain accurate, useful and verifiable? A citation gain fails the release if evidence quality falls.
What the citation-wars paper actually tested
Chen Xu, Zitian Guo and Chenyan Xiong model GEO as a repeated Stackelberg game with partial monitoring. Publishers adapt documents after observing citation outcomes, while an answer engine chooses sources. The authors test a defense called Verifiable Content Reward, or VCR, that rewards factual substance that can be checked rather than surface tactics alone.
Across three benchmarks, the paper reports a 12.1 percentage-point improvement in Net Defense Utility over its strongest baseline. That is a result inside the paper’s experimental setup. It is not a measured uplift for an ordinary publisher and should not be converted into a promise about traffic, rankings or revenue.
Use the public repository without overstating it
The authors link a public reference implementation. Its README describes E-commerce, GEO-Bench and Researchy-GEO datasets, a five-round configuration, Gemini 2.5 Flash-Lite as the default answer engine, and GPT-4o mini for extraction or judging roles. The example commands describe 1,000 test examples and 3,000 training examples.
Those are repository implementation details, not universal properties of every result or production answer engine. The repository also offers a deterministic offline demo, which is useful for checking workflow plumbing without an API key. Reproduction still requires recording the commit, configuration, datasets, provider state and deviations from the authors’ setup.
| Statement | Evidence class | Safe use |
|---|---|---|
| 12.1 percentage-point Net Defense Utility improvement | Reported paper result | Describe with benchmark and simulation scope |
| Five rounds and named model defaults | Repository configuration | Attach commit and local changes |
| Every publisher should use one quality score | Unsupported generalization | Reject; choose job-specific guardrails |
Recognize citation over-optimization in an edit
- Certainty rises while evidence stays flat: “may help” becomes “will rank” without a new source.
- Caveats disappear: scope, date, population or failure conditions are removed to make a sentence easier to quote.
- Source density becomes decoration: links multiply, but the cited pages do not support the nearby claims.
- The user job fragments: one useful page becomes several thin URLs because each sentence is treated as a citation target.
- Measurement collapses: a mention, citation, click and conversion are reported as one visibility score.
The danger is not concise writing. The danger is optimizing the visible signal while discarding the substance that made the answer worth citing.
A worked repair: from quotable certainty to bounded evidence
Illustrative unsafe rewrite: “Adding statistics and schema makes any article more likely to be cited by AI engines.” It sounds decisive, but it invents a universal effect, merges two treatments and supplies no denominator or platform scope.
Evidence-first repair: “For this 20-page test, we added one source-backed statistic to ten comparable pages and left ten unchanged. We will record citation presence for a frozen prompt set over four weekly observations. The test cannot isolate effects on platforms outside the recorded set.”
The second version is less dramatic but more useful: it names the intervention, comparison, sample, observation window and limit. This example is an editorial demonstration, not a result from the paper or from a SearchEngineAnswer experiment.
Run a quality-preserving AEO test
- Define the user job: write the decision or task the page must support before changing it.
- Freeze the treatment: change one meaningful element, such as claim-level sourcing or a worked example;not length, schema, headings and links together.
- Declare the quality floor: require zero unsupported material claims, intact caveats, accessible references and a subject-aware reviewer.
- Freeze the prompt set: preserve wording, platform, mode, locale, account context and observation schedule.
- Record the full denominator: keep uncited, unavailable and contradictory responses instead of saving only favorable examples.
- Use a rollback rule: revert or repair when usefulness or factual support declines, even if citation presence rises.
Pair the claim-level citation workflow with the small SEO experiment method. The first protects the passage; the second protects the inference.
Score the release gate without pretending it is universal
A practical editorial gate can require 100% of material factual claims to map to evidence, zero unsupported material claims, preserved limitations and a named reviewer. Those thresholds are publishing controls, not ranking factors. Add job-specific checks for legal, medical, financial or safety-sensitive content.
Download the GEO rewrite quality-gate worksheet. Its rows are marked EXAMPLE-REMOVE so examples cannot silently become reported evidence.
Google’s people-first content guidance asks whether a page provides original information, substantial value, clear sourcing and a satisfying experience. It also says Google has no preferred word count. That is why “make it longer” is not a recovery method; stronger evidence, examples, decisions and limits are.
The decision rule
Keep an AEO edit only when it improves or preserves the reader’s ability to verify and act. If a variant wins more citations but loses a condition, weakens a source, invents certainty or fragments the task, it has failed. Our AI search ranking evidence register helps separate documented requirements, local observations and unknown selection mechanics.
Verification note: we rechecked the preprint and its linked repository on September 3, 2026. The paper was submitted on August 11, 2026 and had not been treated here as peer-reviewed production-platform evidence.
Quality guardrail
Optimization becomes harmful when the answer unit loses necessary context
A concise passage can be easy to retrieve and still be misleading if it removes the conditions that make the claim true.
- Keep the direct answer near the top.
- Attach the population, date or condition to the claim itself.
- Link to the strongest evidence and retain material uncertainty.
My takeaway: My GEO test is reader-first: if a standalone excerpt would create a worse decision than the complete paragraph, the passage is not ready to optimize.
Primary documentation and reproducibility
September 4 update: defenses need a separate evaluation
Two papers submitted September 2 examine a different part of this problem: defending the answer system. Counter-GEO-Bench tests information-distorting rewrites, while GEO Defender also addresses manipulation of source selection with factually consistent material. Our comparison of the two GEO defense papers separates their threat models, reported results, and limits. It adds a claim-support and source-selection review sheet; it does not establish that either defense is deployed by a commercial search engine.
Keep learning
Continue this topic
Next in this topic
How to Validate AI Visibility Scores Against Referral Traffic
Earlier in this topic
How to Test Gemini 3.7 Flash Before Switching Grounded Answers
AEO & AI Search
Ask a question or join the discussion