What 30,000 Reddit Answers Reveal About Which Voices Survive AI Search
A 30,000-output study finds retrieval drives most variation while Reddit Answers favors early, top-level and more formal comments in its selected evidence.
A study of 30,000 Reddit Answers outputs suggests that the system does not merely summarize a community. It narrows a large, messy discussion through retrieval and comment selection, then rewrites the surviving evidence in a less personal voice. The important finding is not that one writing style always wins. It is that each stage can remove a different kind of evidence.
The research in one sentence: across 10,000 queries drawn from 20 advice and support communities, retrieval determined most answer variation, while the final selection disproportionately favored early, top-level and more formal comments.
The voice-selection funnel has three separate losses
1. The corpus becomes a retrieved source set
The researchers started with 14.68 million comments and generated Reddit Answers three times for each of 10,000 queries. In large communities, 41.8% of output pairs used an identical source set and 27.2% of queries used the same sources in all three runs. In smaller communities, those figures fell to 24.6% and 11.6%.
When the retrieved sources were identical, the outputs were usually nearly identical: 89.6% of large-community pairs exceeded .990 cosine similarity. When retrieval changed, only one of 8,657 pairs reached that level. This makes retrieval the main source of variation in the tested system.
2. Retrieved threads become selected comments
The system heavily favored comments that were structurally easy to lift. In the selected set, 92% were top-level comments, compared with 53% in the broader corpus. Selected comments arrived after a median 1.2 hours, while the corpus median was 5.9 hours.
| Feature | Reported association | What it means |
|---|---|---|
| Score percentile, +1 SD | Odds ratio 2.88 | More highly ranked comments were much more likely to survive selection. |
| Top-level comment | Odds ratio 1.97 | Direct replies to the post were favored over deeper discussion. |
| Reply depth | Odds ratio 0.50 | Each move deeper into a thread was associated with lower selection odds. |
| Formality | Odds ratio 1.488 unadjusted; 1.213 adjusted | More formal language retained a positive association after controls. |
| Experiential voice | Odds ratio .789 unadjusted; .860 adjusted | First-hand framing was less likely to survive selection in this sample. |
These are associations in one system and dataset. They should not be converted into a universal instruction to make community writing formal or impersonal.
3. Selected comments become a synthesized answer
The strongest shift appears in pronoun use. First-person singular language represented 3.3% of quoted community comments but .06% of Reddit Answers text. Personal pronouns fell from 9.5% to 3.8%, and past-focused language fell from 2.7% to .9%.
A support community can contain useful information precisely because someone describes what happened to them. When synthesis removes that voice, the factual advice may remain while the conditions, uncertainty and emotional context become harder to see.
The answer often leaves the community where the question began
A typical answer drew from about seven communities. Only 17.7% of large-community outputs and 15.7% of small-community outputs retrieved posts from the question’s home community. The system was therefore assembling cross-community evidence rather than simply summarizing the local thread.
Original posts were retrieved for 35.7% of large-community queries and 42.6% of small-community queries. Higher-scoring posts were more likely to be retrieved: each e-fold score increase raised the odds by 24.1% in large communities and 16.4% in small communities.
What publishers and community operators can do without gaming the system
- Preserve provenance. Keep author, community, thread and timestamp attached when quoting or summarizing lived experience.
- Expose minority evidence. A useful synthesis should show material disagreement, not only the most highly ranked comment.
- Separate experience from general advice. Label what happened to one person and what the evidence supports more broadly.
- Audit depth loss. Compare selected top-level comments with informative replies buried deeper in a thread.
- Test paraphrase drift. Check whether certainty, chronology or risk changed when first-person testimony became neutral prose.
For AI visibility measurement, this is another reason to distinguish a platform-level citation trend from a system’s internal source selection. A domain can be present while particular voices inside it remain absent.
What this study does not prove
The paper is a preprint focused on Reddit Answers, 20 selected communities and a defined query sample. It does not establish that ChatGPT Search, Google AI Overviews, Gemini or Perplexity use the same selection process. It does not prove that formality causes selection. It also does not show that rewriting comments to imitate the favored features will improve retrieval.
The strongest practical lesson is methodological: audit the full path from corpus to retrieval, selection and synthesis. Counting final links alone cannot tell you which people and contexts disappeared before the answer reached the screen.
Research source
The measurements and odds ratios in this analysis come from the September 2026 preprint “Whose Voices Get Heard? Selection and Representation in Reddit Answers”. SearchEngineAnswer independently checked the paper’s methods, tables and stated limitations before preparing this interpretation.
Keep learning
Continue this topic
Next in this topic
Chrome Now Measures Ad Count, Density and Weight in CrUX
Earlier in this topic
Google Research Cut Retrieval Fan-Out Time 12× to 20×
Research
Ask a question or join the discussion