Robots Atlas>ROBOTS ATLAS
GPT-Live-1

GPT-Live-1

1 · Family: GPT
OpenAI's conversational voice model released on 8 July 2026, built on a full-duplex architecture that listens and speaks simultaneously. It replaces Advanced Voice Mode in ChatGPT for the paid Go, Plus and Pro tiers, and delegates queries requiring search or reasoning to GPT-5.5 running in the background.
✓ Active✓ Public accessAudioMultimodal📁 GPT
Release date
8 July 2026
Access:HostedDeployment:☁ Cloud

Overview

GPT-Live-1 is OpenAI's new generation of conversational voice models, launched on 8 July 2026 to replace Advanced Voice Mode in ChatGPT. The model introduces a full-duplex architecture: it processes audio input and generates output simultaneously, deciding whether to speak, listen, pause or interrupt many times per second. Users can interject and trail off without the rigid turn-taking rhythm of earlier voice interfaces, and the model signals active listening with brief acknowledgements such as "mhmm" or "got it".

A key architectural element is the separation between the voice model and the reasoning model: when a question requires web search or deeper reasoning, GPT-Live-1 delegates it to GPT-5.5 running in the background while keeping the conversation going. Users can choose an Instant, Medium or High reasoning effort. GPT-Live-1 is the default ChatGPT Voice model for the paid Go, Plus and Pro tiers (GPT-Live-1 mini is the default for the Free tier). ChatGPT ships with nine remastered, predefined voices, and the system includes dedicated voice safeguards described in its System Card.

At launch GPT-Live-1 does not support video or screen sharing and is not yet available through the API — OpenAI has said developer access will follow in the coming months. Some languages (e.g. Hindi in the demo) may sound artificial.

Classification
AudioMultimodal
Family: GPT
Access & deployment
Hosted
Cloud
Weights: Closed
Key parameters
Tools
📥 Input: text, audio

Technical specification

License
Proprietary
Features:Tool use
Modalities
⬇ Input
textaudio
⬆ Output
textaudio

Capabilities and applications

Native model capabilities
Voice Conversation
Ability to conduct multi-turn real-time voice conversations with context retention and natural speech pacing.
Category: speech
Speech to text
Category: speech
Text to speech
Category: speech
Real-time inference
The model's ability to generate responses with very low latency (>1000 tokens/sec) on specialized inference hardware (e.g. Cerebras WSE), enabling interactive, turn-by-turn collaboration with a human.
Category: coding
Streaming Speech-to-Text
Real-time conversion of speech to text with immediate output as the speaker is talking.
Category: speech
Live Translation
Real-time speech translation between multiple languages without interrupting the audio stream.
Category: speech
Multilingual
Competence in many natural languages (from a few to over a hundred): understanding, generation, translation, and code-switching within a single conversation. Frontier models support a wide range of languages with comparable quality.
Category: language

Benchmark results

3 benchmarks
GPQA
accuracy · GPT-Live-1 substantially outperforms Advanced Voice Mode; exact score not publicly disclosed by OpenAI.
📄 OpenAI blog — Introducing GPT-Live
Expert-level scientific reasoning benchmark. OpenAI published only a comparison chart, without the raw score.
BrowseComp
accuracy · Better than Advanced Voice Mode; exact number not published.
📄 OpenAI blog — Introducing GPT-Live
Tests agentic web search and the ability to find hard-to-locate information.
τ³-Voice Telecom (internal variant)
task success · GPT-Live-1 outperforms Advanced Voice Mode; numeric score not disclosed.
📄 OpenAI blog — Introducing GPT-Live
Realistic, multi-turn voice telecom support scenarios. Customized user model variant.

Pricing

Technical architecture

Deployment and security

🔒 Security / Enterprise
✓ Verified enterprise information
Updated: 8 Jul 2026↗ Security documentation