ChatGPT Voice Can Switch Models for Search: What to Audit
OpenAI says Voice can use GPT-5.6 or GPT-6 Astra for search and harder reasoning. This paired protocol separates model controls, transcripts, search and citations.
Direct answer: OpenAI says ChatGPT Voice can now use GPT-5.6 or GPT-6 Astra when a spoken request needs web search or harder reasoning. The model and reasoning controls used in text chat now apply to Voice, but the plan, Voice mode, workspace, app version and usage limit can all change the experience.
The practical consequence is simple: “asked in Voice” is not enough to reproduce an answer. A useful audit records the spoken words, displayed transcript, Voice mode, selected model, search event, citations and spoken attribution as separate fields.
What OpenAI changed
OpenAI’s September 9 release notes say Voice can use GPT-5.6 or GPT-6 Astra when it needs to search or reason through a harder question. OpenAI also says the model and reasoning controls available in text are now respected in Voice.
The release retires the separate Instant, Medium and High Voice intelligence labels. The useful test variables are now the account plan, selected model, reasoning control, Voice mode, input transcript and whether a search actually ran.
Voice mode can change the input before search begins
OpenAI’s current Voice documentation distinguishes Live, Advanced and Standard experiences. Standard transcribes a turn before generating a response. Live is a real-time conversation, while Advanced retains supported mobile capabilities such as video or screen sharing.
That distinction matters in a search audit. Standard gives the reviewer a discrete transcript to compare with the typed control. A real-time conversation can carry interruptions, repairs and prior-turn context into the request. Record the mode instead of assuming every microphone input follows the same path.
| Voice mode | Useful observation | Main comparison risk |
|---|---|---|
| Live | Conversation state, interruptions and spoken response | The turn may depend on prior dialogue. |
| Advanced | Mobile capability and any visual context supplied | Video or screen context can make the input unlike a text prompt. |
| Standard | Displayed transcript before answer generation | A transcription error can change the query. |
Usage limits are part of the test environment
At the September 9 check, OpenAI lists three hours of GPT-Live-1 for Plus, 15 hours for the $100 Pro plan and unlimited GPT-Live-1 for the $200 Pro plan. Go receives three hours with GPT-Live-1 mini. Business and institutional plans use different allowances or credit rules.
A limit can change the available experience during a long audit. Record the plan and whether a warning or fallback appeared. Do not mix results collected before and after a usage-limit event in one comparison group.
What the announcement does not prove
- Voice does not always route a search request to GPT-6 Astra.
- A model switch does not establish better source ranking, citation quality or answer accuracy.
- Spoken and typed versions may not carry the same transcript or conversational context.
- The selected model does not reveal every internal step used during a turn.
A selector is a test condition, not a complete routing log. Report the controls and outputs that are visible, then leave the hidden route unresolved.
The minimum audit contract
| Layer | Record | Reason |
|---|---|---|
| Environment | Date, locale, device, app version and account plan | Access and interfaces can vary. |
| Controls | Voice mode, selected model and reasoning setting | Each can alter the test conditions. |
| Input | Exact speech, displayed transcript and typed counterpart | Transcription changes can alter the task. |
| Search | Visible search indicator, latency and answer timestamp | A factual answer does not prove the web was searched. |
| Sources | Spoken attribution, visible citations and destination URLs | Voice presentation and on-screen evidence are different outputs. |
| Repeat | Run number, warnings and changed source set | One answer is not a stable ranking. |
A paired 24-query protocol
Use six prompts in each of four families: fresh facts, local information, shopping comparisons and multi-step research. For each prompt, run one spoken version and one pasted text version with the same plan, selected model, reasoning setting, locale and observation window. Randomize which mode goes first.
Repeat the pair on a second day. That produces 48 observations without pretending that 24 questions represent the product. Preserve the Voice transcript because a changed place name, date or product can explain a changed source set more directly than model switching.
- Write the reference prompt before opening either mode.
- Speak naturally and save the exact transcript shown by ChatGPT.
- Paste the reference prompt in text using the same controls.
- Record search visibility, citations and spoken attribution separately.
- Classify a difference as transcript, search, answer, citation or presentation.
Diagnose the first observable difference
| Observed difference | First explanation to test | Safe conclusion |
|---|---|---|
| Transcript differs | Speech recognition, punctuation or named entity | The inputs were not matched. |
| Only one mode searched | Query interpretation or tool decision | Search behavior differed in this pair. |
| Same answer, different spoken attribution | Presentation constraints | Voice exposed fewer or different source cues. |
| Different cited URLs | Retrieval variation, timing or input difference | The source sets differed, but the cause is unknown. |
| Repeat changes | Normal variation or fresh results | The first result was not stable. |
Stop the diagnosis at the first unmatched layer. If the transcript changed “Portland, Maine” to “Portland,” a later source difference cannot be assigned to the selected model. Fix the input pair and run it again.
Score source exposure without inventing a Voice ranking
For each matched pair, record five observations: search indicator visible, on-screen citation count, spoken source count, citation destination status and claim support. This produces a source-exposure profile, not an accuracy score.
| Field | Value | What it establishes |
|---|---|---|
| Search visible | Yes or no | The interface displayed a search event. |
| Citation visible | Count | How many linked sources the screen exposed. |
| Source spoken | Count | How many sources Voice named aloud. |
| Destination opens | Pass, fail or restricted | Whether a reviewer can reach the cited page. |
| Claim supported | Full, partial, none or unclear | Whether the destination supports the specific answer claim. |
Compare each field separately. Adding them into one percentage would hide the difference between source presentation, access and factual support.
What this changes for publishers
Voice creates two source surfaces. A page can appear as a visible citation while receiving no spoken attribution, or be named aloud while the user never opens the citation tray. Spoken mention, visible link, supporting passage and referral visit are separate events.
Do not optimize for a speculative Voice ranking factor. Improve the evidence that helps every surface: stable author and publication identity, clear dates, primary-source links, descriptive headings and passages that state a verifiable answer before qualification. Our AI citation guide explains that page-level work.
Download the paired source-audit ledger
Download the CSV template. The example rows are labeled EXAMPLE-REMOVE; delete them before collecting observations. Do not store private conversations, account tokens or recordings containing personal information in a public worksheet.
GPT-Live-1 separates the voice layer from the search and reasoning layer
OpenAI released GPT-Live-1 in the API on September 10, 2026 at $0.05 per minute for the front-end voice layer. The developer chooses the backend model, tools and agent harness. OpenAI says GPT-Live-1 can delegate reasoning and tool calls to GPT-6 Astra or a third-party model.
This architecture sharpens the audit in this article. A voice result can fail at speech handling, backend reasoning, tool selection, web retrieval, citation support or spoken rendering. “The voice model searched” is too broad to locate the cause.
Record each handoff in the voice evidence map
| Layer | Evidence to retain | What it cannot prove alone |
|---|---|---|
| Audio interaction | Interruption, pause, noise and turn state | Which source supported the answer |
| Transcript | Recognized words and corrections | That the backend interpreted them correctly |
| Backend delegation | Selected model, effort and request ID | That web search was invoked |
| Tool activity | Search call, query, results and errors | That a cited passage supports the answer |
| Spoken response | Final audio and response text | That every material claim was cited |
GPT-Live-1 also provides ASR transcripts and response text, supports keyword biasing and turn detection, and can run through WebRTC, WebSockets or telephony integrations. Preserve the client and transport because latency and failure modes can differ.
The 30-point interaction gain is not a search-quality result
OpenAI reports a 30 percentage-point improvement on Full Duplex Bench over GPT-Realtime-2.1. That evaluation covers turn taking, interruptions, backchannels and related interactive behavior. It does not measure public-web citation accuracy.
Score spoken interaction and evidence support separately. A natural answer can cite the wrong page, while a well-supported answer can still arrive with poor timing. Use our Astra citation-quality test for the evidence layer and the existing protocol on this page for voice behavior.
Download the voice-search handoff audit
Download the GPT-Live voice-search audit CSV. It records audio conditions, transcript, backend model, tool calls, citations, spoken answer and failure owner.
Rows marked EXAMPLE-REMOVE are examples and not observed product results.
Official references
Current product details come from OpenAI’s ChatGPT release notes and Voice documentation, checked September 9, 2026.
Keep learning
Continue this topic
Next in this topic
Google Election Answers Across Search and Gemini: Source Audit
Earlier in this topic
Search and AI Changes: August 31–September 6, 2026
AEO & AI Search
Ask a question or join the discussion