GPT-6 Astra’s BrowseComp Score Is Not a Citation Quality Score
OpenAI reports a 91.5% BrowseComp score for GPT-6 Astra. A new Perplexity customer story adds engineering context, not a consumer citation benchmark.
Direct answer: OpenAI reports a 91.5% BrowseComp score for GPT-6 Astra, compared with 90.4% for GPT-5.6 Sol. That benchmark can support a claim about performance in OpenAI’s evaluation setup. It is not a citation-quality score for live ChatGPT Search.
Publishers need a separate test for source correctness, passage support, freshness, diversity and reproducibility in the product they actually use.
What OpenAI reported
OpenAI introduced GPT-6 Astra on September 9, 2026. Its published professional benchmark table reports BrowseComp at 91.5% for Astra and 90.4% for GPT-5.6 Sol.
On the same page, OpenAI says its evaluations use the maximum score at any effort and may produce outputs that differ from production ChatGPT because system prompts, tools and other settings can differ.
Why BrowseComp is not a citation-quality score
A browsing benchmark can test whether an agent finds and uses information under a defined harness. A publisher citation audit asks different questions: Did the answer cite the correct URL? Does the cited passage support the exact claim? Was a more current primary source available? Did the source set remain stable across repeated runs?
A one-point benchmark difference cannot answer those questions. It also cannot establish how often a particular publisher will appear in ChatGPT Search.
Use a five-part citation contract
| Dimension | Pass condition | Common failure |
|---|---|---|
| Resolution | Citation opens the intended canonical source | Redirect, inaccessible page or wrong destination |
| Support | The cited passage supports the answer claim | Topical link with no numeric or factual support |
| Freshness | Source date fits the claim’s change rate | Old policy or product documentation |
| Provenance | Primary and secondary evidence are labeled | A summary replaces the source that owns the fact |
| Reproducibility | Repeated runs preserve the core supported answer | Material citation drift across equivalent runs |
A 40-prompt live test
Build eight prompt families with five prompts each: current product facts, policy changes, numerical comparisons, definitions, local recommendations, multi-source synthesis, disputed claims and time-sensitive news. Record the exact product surface, model, reasoning effort, account plan, location and time.
Run each prompt at least twice if budget allows. Score answer correctness and citation support independently. The downloadable fixture contains 40 starter rows, but no results. It is a test design, not a SearchEngineAnswer benchmark.
Score the answer and the citation separately
- Mark whether the answer addresses the question.
- Break the answer into material claims.
- Open every citation and identify the supporting passage.
- Mark full support, partial support, contradiction or no support.
- Record source type, date and canonical status.
- Compare repeated runs for answer and source drift.
- Calculate rates only after preserving the fixed denominator.
Our number-support audit gives a four-state classification for numerical claims. The same discipline applies here.
Keep the comparison environment matched
If Astra is compared with another model, hold the product surface, prompt set, date window, account, location and tool availability as constant as possible. Record timeouts, refusals and missing citations rather than silently rerunning only the failures.
ChatGPT web, an API browsing harness and a third-party agent can expose different search tools and system instructions. A model label alone does not make the runs equivalent. The web and app citation-drift protocol covers surface-specific comparisons.
Perplexity’s customer story changes the engineering context
On September 14, OpenAI published a Perplexity customer story describing Astra use for communications, software changes, production monitoring, and end-to-end testing. Perplexity co-founder Johnny Ho also connected stronger code-writing models with programs that search web or internal information and summarize results.
That supports a narrower conclusion than “Astra powers Perplexity answers.” The story does not identify the model serving Perplexity’s consumer search responses, publish a citation benchmark, or report a latency, traffic, or publisher-visibility effect. It makes the engineering use case more concrete without changing the benchmark boundary in this article.
What publishers can learn
A matched test can reveal which claim types fail, which sources resolve poorly and where a page lacks a citable passage. It can also show that the model’s answer changes even when the source page does not.
It cannot isolate a hidden ranking factor or promise a future citation. Use failures to improve source clarity and evidence, then rerun the same fixture. Keep changes that help readers even when the citation outcome remains unstable.
Download the 40-prompt fixture
Download the GPT-6 Astra citation-quality test CSV. It contains 40 clearly marked starter prompts plus fields for environment, answer score, claim support, source freshness and repeated-run drift.
Every row is labeled EXAMPLE-REMOVE. Replace the prompts or remove the marker before treating the sheet as an executed study.
The practical decision
Use BrowseComp to understand a vendor-reported browsing evaluation. Use a matched live audit to judge citation quality in the surface your readers use. Do not convert one benchmark score into a publisher visibility forecast.
Evidence boundary
Model availability and benchmark figures come from OpenAI’s launch page. The Perplexity engineering details come from OpenAI’s customer story. Both were checked September 14, 2026. SearchEngineAnswer has not run the 40-prompt fixture, independently reproduced BrowseComp, or verified which model serves Perplexity consumer answers.
Keep learning
Continue this topic
Next in this topic
Search and AI Changes: September 7–13, 2026
Earlier in this topic
ChatGPT Library Citations Are Not Public Web Visibility
AEO & AI Search
Ask a question or join the discussion