CITECHOICE Study: More Citation Credit, Uncertain First-Citation Gains
A controlled AI-search replay found more citation markers for structured source text, but no clear gain in whether a source was cited at all. Here is what publishers should measure.
Making a retrieved source easier for an AI answer system to use may earn it more citation markers without making it substantially more likely to receive its first citation. That is the practical distinction in CITECHOICE, a September 2026 preprint by Selvam and Ghosh. It matters because a report that counts links in answers can mistake repeated credit for broader visibility.
The researchers did not edit live websites and watch their search traffic change. They replayed already-retrieved source material through an experimental answer agent, changing how one source was presented while holding the surrounding transcript fixed. The result is unusually useful for understanding citation allocation after retrieval, but it is not a recipe for getting a page discovered or indexed.
Two questions that sound the same, but are not
Suppose an answer cites your article three times instead of once. Your citation-marker count triples. Yet the answer still contains just one cited source from your site, and the proportion of answers that cite you at all may be unchanged. CITECHOICE measured these outcomes separately:
- First-citation incidence: did the target source receive at least one citation in the answer?
- Citation count: how many citation markers pointed to that target?
- Sentence coverage: how many distinct answer sentences cited it?
That separation should also guide publisher AI-visibility reporting. A total citation count cannot tell you by itself whether more prompts included your brand, whether a newly published page became a source, or whether users clicked through.
What the controlled replay actually found
The researchers began with 129 completed everyday-query transcripts and selected 113 matched document pairs that supported the same pre-specified fact. Their main replay contained 89 independent answer-target families and 452 trials. A two-by-two design changed the order of a target and rival source and whether the target appeared as structured material or prose. The answer agent used a GPT-5.4-based setup with Exa retrieval; it did not browse the live web during replay. The paper’s methods and results give the denominator and uncertainty intervals.
The clearest result was more credit for the target after restructuring: an estimated 0.50 additional target citations per answer, with a 95% confidence interval of +0.20 to +0.84 and a Holm-adjusted p-value of .033. Distinct sentences citing the target also increased by an estimated 0.50. The study did not establish a comparable increase in total citations across the whole answer.
Its primary yes-or-no outcome was less certain: the estimated increase in answers citing the target at least once was 4.5 percentage points, but the 95% confidence interval ran from -1.4 to +10.4 points and p was .168. That interval includes no gain. It is inaccurate to report this experiment as proof that structured pages are more likely to enter the cited-source set.
The paper uses the phrase “source admission” for that binary outcome. In this design it means admission to the answer’s citation list after the documents were already retrieved. It does not mean admission to a search index, the retrieval set, or a consumer product’s source-selection pipeline.
A single case shows why the metrics diverge
One worked case in the paper’s appendix paired two U.S. State Department pages relevant to a passport-validity question. The replay crossed source order with prose versus structured rendering. The target’s structured version received three or four citation markers across the two order conditions; its prose version received none. The rival was cited in every cell.
That case makes the mechanism tangible: a source can supply the same supporting fact yet receive different credit in the final answer. It is one illustrative pair, not an estimate for all travel queries or a finding about the State Department’s live traffic. The aggregate first-citation test above remains inconclusive even though this example looks decisive.
Position looked powerful until it was manipulated
In the original observed material, the difference in citation incidence between first- and fifth-ranked sources was 42.3 percentage points. That is a correlation: sources placed early can differ from late sources in relevance, quality, or other unmeasured ways. In the controlled order swap, the main estimate was +7.9 points with a Holm-adjusted p-value of .350; in the held-out subset of 56 pairs it was 0.0 points. The study therefore cannot turn the observational gradient into a causal “move up and get cited” claim.
There is another caution for anyone selling a one-run citation audit. When the team regenerated 120 cells, 15% of the binary cited-or-not decisions changed. Exact target-citation counts agreed in 61.7% of those repeats. A small difference in one answer is not a stable trend without repeated runs and a defined prompt set.
What I would change in an editorial test
If an article is already being retrieved and cited, restructuring a passage may be worth testing for clearer attribution. Keep the factual claim, source links, and answerable scope intact. Create a before/after version with genuinely scannable headings, concise definitions, and evidence next to the claim. The CITECHOICE treatment also involved lexical differences between versions, so this study does not isolate “headings alone” as the winning ingredient.
For a publisher-run test, I would freeze a small set of prompts, engine settings, dates, and source URLs; record the full answers and citations; repeat each prompt; then report three columns separately: share of answers citing the page at least once, citation markers per answer, and whether the cited passage actually supports the answer’s claim. I would also record the retrieval or access evidence if available. A shift in marker count without a shift in first-citation incidence would be a useful finding, not a failed test.
Do not generalize this replay to ChatGPT Search, Google AI Mode, Gemini, or Perplexity without platform-specific tests. The experiment edited serialized, model-visible text after retrieval, not a site’s HTML, CSS, indexability, or live ranking. Its code and derived trial tables are linked from the preprint, while licensed page text and raw traces are not fully redistributed. For broader measurement, see our breakdown of citation count versus citation absorption and the different failure layers between access and attribution.
Bottom line: this controlled study gives publishers a reason to separate citation frequency from citation incidence. It does not establish that formatting a live page will get it discovered, selected, or cited for the first time.
Keep learning
Continue this topic
Next in this topic
Perplexity Photon Speeds Retrieval but Does Not Explain Citations
Earlier in this topic
ChatGPT and Gemini Shared No Source Domains in 76.7% of Product Checks
AEO & AI Search
Ask a question or join the discussion