For the better part of three years, talking to ChatGPT has meant taking turns. You spoke, you waited, and a beat of silence later the assistant spoke back. On July 8, OpenAI announced it is tearing up that script.

The company launched GPT-Live, a new generation of voice models built on what it calls a full-duplex architecture — one that can listen and speak at the same time, rather than trading discrete turns. Two versions, GPT-Live-1 and GPT-Live-1 mini, began rolling out to ChatGPT users worldwide across iOS, Android, and the web on launch day. It is the first ground-up rebuild of the ChatGPT Voice experience since OpenAI shipped Advanced Voice Mode in 2024.

The difference is architectural, not cosmetic. Where older systems processed conversation as a sequence of separate messages, GPT-Live "continuously processes input while generating output," OpenAI wrote in its announcement. That lets the model "make interaction decisions many times per second: whether to speak, continue listening, pause, interrupt, or invoke a tool." In practice, that means the assistant can backchannel with a quick "mhmm" or "yeah" to show it is following along, hold its tongue while you gather your thoughts, or jump in when you trail off — the small timing cues that make human conversation feel alive.

From Cascades to Continuous Listening

To appreciate the leap, it helps to see what came before. The original 2023 ChatGPT Voice chained three separate models together: speech-to-text to transcribe you, a language model to think, and text-to-speech to answer. That "cascaded" pipeline was slow and stilted, and information leaked between the handoffs. Advanced Voice Mode, introduced in 2024, collapsed that into a single model to cut latency — but it still operated in rigid turns. Because it detected the end of your turn by listening for silence, a thoughtful pause or a bit of passing traffic could trip it into interrupting at the wrong moment.

GPT-Live attacks that problem with a second design change: it decouples the fast, conversational layer from heavier cognition. When a question needs web search, deeper reasoning, or agentic work, GPT-Live delegates to a frontier model in the background — GPT-5.5 at launch — and folds the result back into the conversation when it is ready, all while keeping the chat flowing. OpenAI says it will swap in newer frontier models over time without changing the voice layer. In head-to-head human evaluations, the company reported, GPT-Live-1 and its mini sibling were "strongly preferred" over Advanced Voice Mode on turn-taking, interruptions, and overall naturalness, and posted gains on reasoning and web-search benchmarks like GPQA and BrowseComp.

The rollout maps neatly onto OpenAI's subscription tiers. GPT-Live-1 becomes the default for paying Go, Plus, and Pro users; GPT-Live-1 mini becomes the default for free accounts. OpenAI says it plans to bring the models to its API "soon" and has opened a signup form for developers and enterprises. There is no separate consumer price — GPT-Live is folded into existing ChatGPT plans. At launch it does not support video or screen sharing, and legacy Standard and Advanced Voice modes remain available for users who want those features.

OpenAI is framing this as more than a nicer chatbot. "Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work," ChatGPT Voice product lead Atty Eleti said in a press briefing reported by TechCrunch. "The kind of amazing use cases that we see people using Codex and ChatGPT to accomplish, we think voice can be the future interface to all kinds of work." Eleti added that he has held 30- to 40-minute conversations with the feature on walks.

Why It Matters

More than 150 million people already talk to ChatGPT each week through Voice and Dictation, and OpenAI is betting that voice — not text boxes — is how a large share of AI use eventually happens. Reports have suggested the company is developing AI earbuds this year, hardware for which a full-duplex, always-listening model would be a natural fit, though OpenAI declined to discuss devices.

The competitive stakes are high. Google has pushed conversational Gemini Live across Android; ElevenLabs has built a business on lifelike synthetic voices and agents; and startups like Sesame, from Oculus co-founder Brendan Iribe, are chasing the same feeling of talking to something that actually listens. Apple and Amazon have both retooled Siri and Alexa to sound more natural. Whoever nails the ambient, hands-free interface stands to own a new front door to computing — one that does not require a screen at all.

Not every reaction has been glowing. Early users have complained that GPT-Live can be over-enthusiastic, with its constant "mhmm" and "yeah" backchanneling reading as verbose or even irritating rather than reassuring. TechCrunch also found the launch's live-translation demo rough, noting the assistant spoke Hindi with a heavy American accent and a stilted, "bookish" tone. OpenAI acknowledged the model is optimized for only its most popular languages.

What to Watch

Three things will tell us whether GPT-Live is a genuine inflection point. First, whether OpenAI can dial back the over-eager backchanneling without making the model feel dead again — a tuning problem that is harder than it sounds. Second, the API launch, which will determine how quickly developers build full-duplex voice into their own products. And third, hardware: if OpenAI ships earbuds this year, GPT-Live is almost certainly the engine inside them, and the phone-screen assistant will start to look like a transitional step toward something you simply talk to.

“Over time, we think this will also unlock the ability to use voice as a kind of primary interface to computing, and to manage increasingly complex long-running agentic work.”
— Atty Eleti, ChatGPT Voice product lead, OpenAI
150M+
Weekly voice users
GPT-5.5
Background model