Natural conversation includes hesitation, interruption and small signs of attention. GPT-Live tries to reproduce that rhythm with a full-duplex architecture: the model receives audio while generating output and decides several times per second whether to speak, wait, interrupt or call a tool. OpenAI also separates quick conversation from deeper reasoning performed by a frontier model in the background.

The challenge is therefore larger than producing a pleasant voice. The system must maintain low-latency audio, preserve state through a long session, move that session between accelerators and return research without an awkward silence.

The short answer

QuestionAnswer
What does full duplex mean?The system listens and speaks simultaneously instead of enforcing strict turns.
Why use two models?GPT-Live manages interaction while a stronger model searches or reasons.
Which transport carries audio?WebRTC provides real-time media and handles packet loss.
How does a long conversation survive?Context is compacted and transferred to an instance prepared in parallel.
Is the API available?OpenAI says an API is coming; the system already powers ChatGPT Voice.

From a three-stage pipeline to continuous audio

Early voice assistants chained speech recognition, a language model and speech synthesis. Every stage added delay and could lose tone or timing information. Later audio models reduced latency but still waited for an end of turn inferred from silence.

GPT-Live processes a continuous stream. A pause no longer automatically means the user has finished. The model can offer a brief acknowledgement without treating it as a complete reply, or remain quiet when asked to listen.

That freedom complicates records. The interface, safety systems and analytics still expect discrete messages. The server builds a provisional view from partial transcripts and timing, then finalizes history once speaker attribution is reliable enough.

WebRTC keeps audio on a priority path

OpenAI uses WebRTC as the transport foundation. It handles packet loss, clock drift and connection changes. It can subtly stretch late audio and briefly speed playback to catch up to real time.

The target is sub-second responsiveness. That requires minimizing buffers and blocking work across the entire path, not just accelerating inference. A saturated media service or distant regional capacity can erase the benefit of a fast model.

For startup, OpenAI worked on WARP, proposals that reduce multiple WebRTC negotiations from six network round trips to one. Instant Connect also prepares parameters before the session is materialized. In the best case, the first UDP packet is enough for the server to start the media path.

Speaking quickly and thinking deeply become separate jobs

GPT-Live handles immediate interaction. When a question requires search, tools or deeper reasoning, it delegates to GPT-5.5 in the background. The voice model can keep the exchange moving while the result is prepared.

The server creates a reasoning session in advance and prefills initial context. Session affinity and prompt caching prevent repeated setup on every delegation. Tool latency, reasoning effort and output size still count against the responsiveness budget.

The design offers a broader agent lesson: a real-time interface should not block on its slowest task. It needs a responsive channel, asynchronous work and a clear way to insert results without misrepresenting progress.

Context must survive model-instance changes

A long call grows its context while model instances appear and disappear with demand. OpenAI warms a replacement in parallel, transfers state, briefly runs both and cuts over when the new instance is ready.

The same mechanism handles compaction. Summarizing history invalidates the attention key-value cache. Instead of blocking voice while rebuilding it, the platform compacts and prefills elsewhere, then transitions without interrupting media.

Capacity is measured in sessions, not just requests

Shadow tests on production traffic showed that GPU throughput was an incomplete benchmark. A voice session remains open and continuously sends frames, consuming CPU, network queues and state storage. A supporting component can saturate before inference and amplify latency everywhere.

OpenAI therefore measured how many concurrent sessions could keep every frame on schedule. Long calls, reconnects and ordinary disconnects exposed failures that short load tests did not reproduce.

What developers should retain

A future GPT-Live API will not remove application responsibilities. Products must handle recording consent, interruption, allowed tools, long-session costs and fallback under poor connectivity. SynthID provenance added to supported audio helps identify some generated output without solving every abuse case.

GPT-Live's fluidity comes less from one model than from a system protecting the audio path. Streaming, delegation, state transfer and faster transport work together so conversation continues while infrastructure does the heavy work.