OpenAI reveals build process behind GPT‑Live

Justin Uberti and Zahan Malkani have shared details of GPT‑Live.

OpenAI has revealed how it built a real-time system for responsive voice AI in just six months.

Justin Uberti and Zahan Malkani at OpenAI have shared how the company managed to build GPT‑Live⁠, their third-generation voice system.

In a post on the OpenAI website, Uberti and Malkani explained: “For voice AI, knowing when to speak is harder than it sounds. Human speakers effortlessly hand off to each other in a fraction of a second, but previous voice AI systems couldn’t keep up with this rhythm.

“Their turn-based architecture relied on tiny models known as turn detectors, which faced an unenviable task: guess too soon, and the user gets cut off; guess too late, and the response feels sluggish. Only after the detector made its decision could the much larger LLM get to work.

“GPT‑Live⁠, our third-generation voice system, removes the turn detector from the audio path. Its voice model is full-duplex, which means it can listen and speak at the same time. That eliminates the need for a separate detector and makes conversation feel more immediate and natural.”

OpenAI feels the latest model is a significant improvement on previous iterations.

Uberti and Malkani said: “Earlier voice architectures inherited the turn-based nature of text LLMs, but with each turn represented as a discrete audio blob rather than text. In cascaded systems, speech-to-text, the LLM, and text-to-speech each ran in series. This sequencing added latency and ignored cues such as tone and pacing.

“Speech-to-speech models improved on this approach by processing audio directly. Training the model to natively understand and generate speech allowed it to preserve details lost in transcription and respond more quickly. But the system still relied on the turn detector to decide when inference could begin. The model handled more of the interaction, but the interaction remained turn-based.

“GPT‑Live puts the voice model in control of the conversation: audio flows in and out of the model, while deeper reasoning and tool use happen asynchronously.”

Close Bitnami banner
Bitnami