On 10 September OpenAI brought the voice model GPT-Live-1, so far in ChatGPT, to the API. It listens and speaks at the same moment, can leave the thinking and the tools to another model behind it, and costs $0.05 per minute for the voice layer.
- One model handles incoming and outgoing audio together, instead of the speech-to-text, language model, text-to-speech chain.
- Reasoning and tools can be delegated to a backend model, for example GPT-6 Astra or a third-party model.
- Per OpenAI: +30 percentage points on Full Duplex Bench over GPT-Realtime-2.1; Speak reports, in early evaluations, almost 80% fewer interruptions during thinking pauses with its language tutor.
Almost 80% fewer interruptions during thinking pauses, in early evaluations. That is Speak's number, a language-learning app, and it isn't about speed. The voice tutor talks over you far less often while you are still thinking.
Anyone who has learned a language by voice with a machine knows the problem. You pause to find the word, and it is already answering. For the machine, silence is the end of a sentence.
Full duplex, in plain words
Until now voice agents worked like a two-way radio: you talk, release the button, it talks. Full duplex is the telephone. Both of you can speak at once, interrupt each other, say "uh-huh" in the middle of the other person's sentence, without the conversation falling apart.
The clever part is the split. The voice model keeps the conversation alive while the heavy thinking can go to another model at the back. So you can put a lighter model on booking an appointment and a heavier one on a complex problem.
The benchmark numbers are OpenAI's. The price of $0.05, that is 5 cents a minute, is for the voice only; the model at the back is paid separately and its bill depends on how much it thinks.
If you are building a voice assistant for customers in your own language, first check whether that language is supported. Then test not the speed but how patient it is when the person on the other end is searching for a word.