The 16% Synthetic-Source Finding Depends on an AI Detector

A 712-query audit found detector-classified synthetic pages among citations, but 27.1% of cited URLs were not analyzed and page accuracy was not tested.

Sonar routes cited pages past an unread tray and through a detector so excluded pages remain separate from analyzed results.

A 712-query audit reported that about 16% of the cited pages it could analyze were classified as AI-generated. That is a meaningful warning signal, but it is not a finding that 16% of all AI-search citations are fake or low quality.

The percentage depends on an AI detector, applies to 19,154 extractable pages rather than all 26,266 unique cited URLs, and does not measure whether a page’s claims were accurate. Those boundaries change how the result should be used.

What the research actually asked

Northwestern University researchers Mowafak Allaham and Nicholas Diakopoulos audited citations shown by the consumer interfaces of ChatGPT with web search, Copilot, Gemini, and Perplexity. Their May 2026 paper used 712 English queries relevant to the United States: 175 about politics, 257 about health, and 280 about the environment.

The study asked whether cited source pages showed evidence of AI-generated text. It did not test the factual accuracy of every page, whether the answer engine relied on the page for a particular claim, or whether AI authorship caused the citation.

The 712-query sample
TopicQueriesUnique cited URLsAverage URLs per query
Environment28012,25210.94
Health2578,6408.40
Politics1755,3757.68

The topic URL counts overlap across engines and queries. They should not be added to reconstruct the number of unique pages in the final dataset.

The 16% result sits at the end of a funnel

Which pages reached the production classifier
StagePages or domainsShareWhat it means
Unique cited URLs collected26,266 URLs100%A citation appeared in at least one audited answer.
Main text extracted19,154 URLs72.9%Text was available for detector analysis.
Not analyzed by the detector7,112 URLs27.1%Extraction failed or the source used another format, including PDFs and images.
Classified likely or highly likely AI3,056 URLs16.0% of extracted URLsPangram assigned one of the two included AI-likelihood categories.

The 16% denominator is 19,154, not 26,266. Expressed against the complete cited-URL set, the 3,056 classified pages equal 11.6%, but that calculation does not resolve the 7,112 unclassified URLs. Their composition is unknown, so neither 11.6% nor 16% is a ground-truth prevalence rate for every citation.

Download the evidence-funnel CSV with the counts, denominators, and claim boundaries.

The detector boundary is not a footnote

The researchers compared Pangram and GPTZero before choosing Pangram for the production run. Both detectors achieved the reported perfect result on a curated validation set containing 200 human-written and 200 AI-generated samples. A clean benchmark is useful, but it cannot reproduce every edited, mixed, translated, or domain-specific page encountered on the live web.

The paper added a second stress test using 105 recent articles from seven domains previously reported as AI-slop sources. Pangram classified 58 articles, or 55.2%, as highly likely AI; 17, or 16.2%, as mixed; and 30, or 28.6%, as unlikely AI. The original paper rounds the last figure to 28.5%.

The study also excluded 283 pages categorized as “Possibly AI” from the final AI-generated count because that label could reflect mixed authorship or classifier uncertainty. That conservative decision reduces one kind of overstatement but does not remove detector dependence.

Three claims that should not be collapsed

The page was cited.
The study captured the page in the visible cited-source list for an audited answer. This is an interface observation within the collection period.
The page was classified as synthetic.
Pangram assigned a likelihood category to extracted text. This is a model output under the study’s selected threshold and category rules.
The page contained poor information.
This would require a separate assessment of claims, sources, accuracy, omissions, and context. The study did not conduct that page-level quality audit.

AI-generated text can contain accurate, well-sourced information. Human-written text can be false or misleading. Authorship is a useful provenance signal, but it is not a replacement for checking whether a cited page supports the answer.

The citation market had a concentrated head and a long tail

The 26,266 unique cited URLs came from 7,675 domains. The aggregate domain-frequency distribution had a Gini index of 0.68, while the top 25 domains accounted for 23.8% of citations. At the same time, 59.1% of cited domains appeared once and 16.5% appeared twice.

Those figures are not contradictory. A small group can collect many citations while thousands of domains appear at the edge of the distribution. Our separate 16,069-citation recomputation found a similar need to inspect the head and long tail separately, although it used a different dataset and should not be merged with this paper’s percentages.

For risk review, repeated domains and one-off domains deserve different sampling plans. Repeated sources may create scale risk if a problem recurs. One-off sources may reveal how broadly a system reaches into pages that have little prior visibility.

A better audit starts with claims, not authorship labels

  1. Save the answer, query, interface, date, account state, and complete visible citation list.
  2. Choose material claims before looking at detector labels.
  3. Open each cited page and record the exact passage that supports or contradicts the claim.
  4. Record page provenance, author identity, dates, primary sources, and AI-use disclosure when available.
  5. Use detector output as a review signal, not as a verdict.
  6. Report inaccessible and non-text sources in the denominator rather than silently removing them.
  7. Repeat the query panel before generalizing beyond one collection.

The paper establishes that all four audited systems visibly cited pages Pangram classified as likely or highly likely AI-generated. It does not establish that those pages were inaccurate, that the current interfaces behave the same way, or that synthetic text caused a citation. The next useful study is a claim-level accuracy and provenance audit, especially in health, politics, and environmental decisions where a source error can matter most.

For a policy boundary on scaled AI content, see when utility becomes scaled content abuse. For methods that separate citations from recommendations and referrals, use the AI visibility measurement protocol.

Keep learning

Continue this topic

Community discussion

Discuss: The 16% Synthetic-Source Finding Depends on an AI Detector

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.