ChatGPT Voice Can Switch Models for Search: What to Audit

OpenAI says Voice can use GPT-5.6 or GPT-6 Astra for search and harder reasoning. This paired protocol separates model controls, transcripts, search and citations.

Sonar routes voice and text questions through search and compares the returning source cards

Direct answer: OpenAI says ChatGPT Voice can now use GPT-5.6 or GPT-6 Astra when a spoken request needs web search or harder reasoning. The model and reasoning controls used in text chat now apply to Voice, but the plan, Voice mode, workspace, app version and usage limit can all change the experience.

The practical consequence is simple: “asked in Voice” is not enough to reproduce an answer. A useful audit records the spoken words, displayed transcript, Voice mode, selected model, search event, citations and spoken attribution as separate fields.

What OpenAI changed

OpenAI’s September 9 release notes say Voice can use GPT-5.6 or GPT-6 Astra when it needs to search or reason through a harder question. OpenAI also says the model and reasoning controls available in text are now respected in Voice.

The release retires the separate Instant, Medium and High Voice intelligence labels. The useful test variables are now the account plan, selected model, reasoning control, Voice mode, input transcript and whether a search actually ran.

Voice mode can change the input before search begins

OpenAI’s current Voice documentation distinguishes Live, Advanced and Standard experiences. Standard transcribes a turn before generating a response. Live is a real-time conversation, while Advanced retains supported mobile capabilities such as video or screen sharing.

That distinction matters in a search audit. Standard gives the reviewer a discrete transcript to compare with the typed control. A real-time conversation can carry interruptions, repairs and prior-turn context into the request. Record the mode instead of assuming every microphone input follows the same path.

What each Voice mode changes in a source audit
Voice mode Useful observation Main comparison risk
Live Conversation state, interruptions and spoken response The turn may depend on prior dialogue.
Advanced Mobile capability and any visual context supplied Video or screen context can make the input unlike a text prompt.
Standard Displayed transcript before answer generation A transcription error can change the query.

Usage limits are part of the test environment

At the September 9 check, OpenAI lists three hours of GPT-Live-1 for Plus, 15 hours for the $100 Pro plan and unlimited GPT-Live-1 for the $200 Pro plan. Go receives three hours with GPT-Live-1 mini. Business and institutional plans use different allowances or credit rules.

A limit can change the available experience during a long audit. Record the plan and whether a warning or fallback appeared. Do not mix results collected before and after a usage-limit event in one comparison group.

What the announcement does not prove

  • Voice does not always route a search request to GPT-6 Astra.
  • A model switch does not establish better source ranking, citation quality or answer accuracy.
  • Spoken and typed versions may not carry the same transcript or conversational context.
  • The selected model does not reveal every internal step used during a turn.

A selector is a test condition, not a complete routing log. Report the controls and outputs that are visible, then leave the hidden route unresolved.

The minimum audit contract

Fields needed to compare a voice answer with a typed answer
Layer Record Reason
Environment Date, locale, device, app version and account plan Access and interfaces can vary.
Controls Voice mode, selected model and reasoning setting Each can alter the test conditions.
Input Exact speech, displayed transcript and typed counterpart Transcription changes can alter the task.
Search Visible search indicator, latency and answer timestamp A factual answer does not prove the web was searched.
Sources Spoken attribution, visible citations and destination URLs Voice presentation and on-screen evidence are different outputs.
Repeat Run number, warnings and changed source set One answer is not a stable ranking.

A paired 24-query protocol

Use six prompts in each of four families: fresh facts, local information, shopping comparisons and multi-step research. For each prompt, run one spoken version and one pasted text version with the same plan, selected model, reasoning setting, locale and observation window. Randomize which mode goes first.

Repeat the pair on a second day. That produces 48 observations without pretending that 24 questions represent the product. Preserve the Voice transcript because a changed place name, date or product can explain a changed source set more directly than model switching.

  1. Write the reference prompt before opening either mode.
  2. Speak naturally and save the exact transcript shown by ChatGPT.
  3. Paste the reference prompt in text using the same controls.
  4. Record search visibility, citations and spoken attribution separately.
  5. Classify a difference as transcript, search, answer, citation or presentation.

Diagnose the first observable difference

A difference ladder for cautious diagnosis
Observed difference First explanation to test Safe conclusion
Transcript differs Speech recognition, punctuation or named entity The inputs were not matched.
Only one mode searched Query interpretation or tool decision Search behavior differed in this pair.
Same answer, different spoken attribution Presentation constraints Voice exposed fewer or different source cues.
Different cited URLs Retrieval variation, timing or input difference The source sets differed, but the cause is unknown.
Repeat changes Normal variation or fresh results The first result was not stable.

Stop the diagnosis at the first unmatched layer. If the transcript changed “Portland, Maine” to “Portland,” a later source difference cannot be assigned to the selected model. Fix the input pair and run it again.

Score source exposure without inventing a Voice ranking

For each matched pair, record five observations: search indicator visible, on-screen citation count, spoken source count, citation destination status and claim support. This produces a source-exposure profile, not an accuracy score.

Five observations that can be counted consistently
Field Value What it establishes
Search visible Yes or no The interface displayed a search event.
Citation visible Count How many linked sources the screen exposed.
Source spoken Count How many sources Voice named aloud.
Destination opens Pass, fail or restricted Whether a reviewer can reach the cited page.
Claim supported Full, partial, none or unclear Whether the destination supports the specific answer claim.

Compare each field separately. Adding them into one percentage would hide the difference between source presentation, access and factual support.

What this changes for publishers

Voice creates two source surfaces. A page can appear as a visible citation while receiving no spoken attribution, or be named aloud while the user never opens the citation tray. Spoken mention, visible link, supporting passage and referral visit are separate events.

Do not optimize for a speculative Voice ranking factor. Improve the evidence that helps every surface: stable author and publication identity, clear dates, primary-source links, descriptive headings and passages that state a verifiable answer before qualification. Our AI citation guide explains that page-level work.

Download the paired source-audit ledger

Download the CSV template. The example rows are labeled EXAMPLE-REMOVE; delete them before collecting observations. Do not store private conversations, account tokens or recordings containing personal information in a public worksheet.

GPT-Live-1 separates the voice layer from the search and reasoning layer

OpenAI released GPT-Live-1 in the API on September 10, 2026 at $0.05 per minute for the front-end voice layer. The developer chooses the backend model, tools and agent harness. OpenAI says GPT-Live-1 can delegate reasoning and tool calls to GPT-6 Astra or a third-party model.

This architecture sharpens the audit in this article. A voice result can fail at speech handling, backend reasoning, tool selection, web retrieval, citation support or spoken rendering. “The voice model searched” is too broad to locate the cause.

Record each handoff in the voice evidence map

A voice answer crosses several observable layers
Layer Evidence to retain What it cannot prove alone
Audio interaction Interruption, pause, noise and turn state Which source supported the answer
Transcript Recognized words and corrections That the backend interpreted them correctly
Backend delegation Selected model, effort and request ID That web search was invoked
Tool activity Search call, query, results and errors That a cited passage supports the answer
Spoken response Final audio and response text That every material claim was cited

GPT-Live-1 also provides ASR transcripts and response text, supports keyword biasing and turn detection, and can run through WebRTC, WebSockets or telephony integrations. Preserve the client and transport because latency and failure modes can differ.

The 30-point interaction gain is not a search-quality result

OpenAI reports a 30 percentage-point improvement on Full Duplex Bench over GPT-Realtime-2.1. That evaluation covers turn taking, interruptions, backchannels and related interactive behavior. It does not measure public-web citation accuracy.

Score spoken interaction and evidence support separately. A natural answer can cite the wrong page, while a well-supported answer can still arrive with poor timing. Use our Astra citation-quality test for the evidence layer and the existing protocol on this page for voice behavior.

Download the voice-search handoff audit

Download the GPT-Live voice-search audit CSV. It records audio conditions, transcript, backend model, tool calls, citations, spoken answer and failure owner.

Rows marked EXAMPLE-REMOVE are examples and not observed product results.

Official references

Current product details come from OpenAI’s ChatGPT release notes and Voice documentation, checked September 9, 2026.

Keep learning

Continue this topic

Community discussion

Discuss: ChatGPT Voice Can Switch Models for Search: What to Audit

Have a question, a useful example, or a different perspective? Join the discussion, share evidence, and help other readers reach a better answer.

0 replies Moderated
No replies yet.

Be the first to ask a focused question, share a practical example, or add useful evidence.

Ask a question or join the discussion

Share evidence, a useful example, or a clear question. Be specific, stay on topic, and challenge ideas without attacking people. First-time replies may be held for moderation.