GPT-6 Astra’s BrowseComp Score Is Not a Citation Quality Score

OpenAI reports a 91.5% BrowseComp score for GPT-6 Astra. A new Perplexity customer story adds engineering context, not a consumer citation benchmark.

Sonar the Answer Whale tests citation support separately from a 91.5 percent BrowseComp score

Direct answer: OpenAI reports a 91.5% BrowseComp score for GPT-6 Astra, compared with 90.4% for GPT-5.6 Sol. That benchmark can support a claim about performance in OpenAI’s evaluation setup. It is not a citation-quality score for live ChatGPT Search.

Publishers need a separate test for source correctness, passage support, freshness, diversity and reproducibility in the product they actually use.

What OpenAI reported

OpenAI introduced GPT-6 Astra on September 9, 2026. Its published professional benchmark table reports BrowseComp at 91.5% for Astra and 90.4% for GPT-5.6 Sol.

On the same page, OpenAI says its evaluations use the maximum score at any effort and may produce outputs that differ from production ChatGPT because system prompts, tools and other settings can differ.

Why BrowseComp is not a citation-quality score

A browsing benchmark can test whether an agent finds and uses information under a defined harness. A publisher citation audit asks different questions: Did the answer cite the correct URL? Does the cited passage support the exact claim? Was a more current primary source available? Did the source set remain stable across repeated runs?

A one-point benchmark difference cannot answer those questions. It also cannot establish how often a particular publisher will appear in ChatGPT Search.

Use a five-part citation contract

A live citation audit needs fields the benchmark headline does not provide
Dimension Pass condition Common failure
Resolution Citation opens the intended canonical source Redirect, inaccessible page or wrong destination
Support The cited passage supports the answer claim Topical link with no numeric or factual support
Freshness Source date fits the claim’s change rate Old policy or product documentation
Provenance Primary and secondary evidence are labeled A summary replaces the source that owns the fact
Reproducibility Repeated runs preserve the core supported answer Material citation drift across equivalent runs

A 40-prompt live test

Build eight prompt families with five prompts each: current product facts, policy changes, numerical comparisons, definitions, local recommendations, multi-source synthesis, disputed claims and time-sensitive news. Record the exact product surface, model, reasoning effort, account plan, location and time.

Run each prompt at least twice if budget allows. Score answer correctness and citation support independently. The downloadable fixture contains 40 starter rows, but no results. It is a test design, not a SearchEngineAnswer benchmark.

Score the answer and the citation separately

  1. Mark whether the answer addresses the question.
  2. Break the answer into material claims.
  3. Open every citation and identify the supporting passage.
  4. Mark full support, partial support, contradiction or no support.
  5. Record source type, date and canonical status.
  6. Compare repeated runs for answer and source drift.
  7. Calculate rates only after preserving the fixed denominator.

Our number-support audit gives a four-state classification for numerical claims. The same discipline applies here.

Keep the comparison environment matched

If Astra is compared with another model, hold the product surface, prompt set, date window, account, location and tool availability as constant as possible. Record timeouts, refusals and missing citations rather than silently rerunning only the failures.

ChatGPT web, an API browsing harness and a third-party agent can expose different search tools and system instructions. A model label alone does not make the runs equivalent. The web and app citation-drift protocol covers surface-specific comparisons.

Perplexity’s customer story changes the engineering context

On September 14, OpenAI published a Perplexity customer story describing Astra use for communications, software changes, production monitoring, and end-to-end testing. Perplexity co-founder Johnny Ho also connected stronger code-writing models with programs that search web or internal information and summarize results.

That supports a narrower conclusion than “Astra powers Perplexity answers.” The story does not identify the model serving Perplexity’s consumer search responses, publish a citation benchmark, or report a latency, traffic, or publisher-visibility effect. It makes the engineering use case more concrete without changing the benchmark boundary in this article.

What publishers can learn

A matched test can reveal which claim types fail, which sources resolve poorly and where a page lacks a citable passage. It can also show that the model’s answer changes even when the source page does not.

It cannot isolate a hidden ranking factor or promise a future citation. Use failures to improve source clarity and evidence, then rerun the same fixture. Keep changes that help readers even when the citation outcome remains unstable.

Download the 40-prompt fixture

Download the GPT-6 Astra citation-quality test CSV. It contains 40 clearly marked starter prompts plus fields for environment, answer score, claim support, source freshness and repeated-run drift.

Every row is labeled EXAMPLE-REMOVE. Replace the prompts or remove the marker before treating the sheet as an executed study.

The practical decision

Use BrowseComp to understand a vendor-reported browsing evaluation. Use a matched live audit to judge citation quality in the surface your readers use. Do not convert one benchmark score into a publisher visibility forecast.

Evidence boundary

Model availability and benchmark figures come from OpenAI’s launch page. The Perplexity engineering details come from OpenAI’s customer story. Both were checked September 14, 2026. SearchEngineAnswer has not run the 40-prompt fixture, independently reproduced BrowseComp, or verified which model serves Perplexity consumer answers.

Keep learning

Continue this topic

Community discussion

Discuss: GPT-6 Astra’s BrowseComp Score Is Not a Citation Quality Score

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.