Future TechnologyFuture Technology
AI

OpenAI GPT-Live Drops the Transcript, and That Is the Whole Point

· 2 min read · By Future Technology

Key takeaways

  • GPT-Live is a native voice model, so audio stays as audio instead of being flattened into a transcript
  • OpenAI claims sub 300 millisecond latency, roughly the gap between turns in natural human conversation
  • These are launch figures, and real latency over a mobile network will be worse

Three hundred milliseconds is roughly the gap between turns in ordinary human conversation. That is the number OpenAI is claiming for GPT-Live, the native voice model now running behind ChatGPT Voice.

The claim rests on removing a step rather than adding one. Voice assistants have generally worked in three stages: speech to text, then the language model, then text to speech. Every stage adds delay, and every stage throws information away. Tone, hesitation, emphasis, the half second where somebody changes their mind mid-sentence, all of it dies in the transcript.

What native voice changes

A native voice model keeps the audio as audio the whole way through. There is no transcript in the middle to reconstruct meaning from, so the model can hear that a question was tentative rather than just reading the words of it, and can answer with matching emphasis.

The latency figure matters more than the emotional nuance, at least in the short term. Below roughly 300 milliseconds, interruption starts working properly. You can cut in, the model stops, and the exchange feels like talking rather than taking turns with a kiosk. Above it, everyone falls back into the walkie-talkie rhythm that has made voice assistants tiring since 2011.

The caveat worth stating plainly

These are launch claims measured under launch conditions. Real latency over a mobile network, on a congested evening, with a cheap Bluetooth headset in the chain, will be higher. Bluetooth alone can add well over 100 milliseconds before the model has heard anything.

Where the processing happens matters too, and that trade is the same one covered in our comparison of local AI agents against cloud agents. Native voice is heavy enough that it stays in the cloud for now, which means the network is part of the latency budget whether OpenAI likes it or not.

What to watch

The interesting question is what happens to everything built on the assumption that a transcript exists. Logging, moderation, analytics and most agent tooling all read text. If the audio never becomes text, that plumbing has to be rebuilt or run as a separate pass, which is a cost nobody mentions in a launch post. Agent interoperability work such as A2A and MCP is written around text payloads as well.

Watch for independent latency measurements over the next few weeks. That number, taken on a normal connection rather than a demo stage, is the one that decides whether this changes how people use voice or just how it sounds.

Get the briefing, free

The biggest tech story, explained in 3 minutes every weekday. Choose your briefings →

Free. No spam. Unsubscribe in one click.