GPT-Live-1 Lands in the API: An Engineering Guide to Full-Duplex Voice Agents
Voice agents have been stuck in an awkward architecture for two years: transcribe the user with speech-to-text, hand that to an LLM to think, then synthesize the reply with text-to-speech. Chaining three stages stacks latency on latency, and interruption handling is nearly impossible: the user starts a new sentence while the system is still reading out the previous one. On September 10, 2026, OpenAI put GPT-Live-1 in the API at $0.05 per minute, and it takes the full-duplex route instead. It listens and speaks at the same moment, and it reacts to pauses, interruptions, and short acknowledgements. This is not a faster voice model. It replaces the architecture of voice agents.
"Full duplex: listening and speaking at once, not taking turns"
1. What Actually Shipped on September 10
GPT-Live-1 moved from ChatGPT Voice into the API. A few words in the announcement deserve to be read slowly: full duplex, stronger instruction following, custom voices, and telephony support. It ships with built-in transcripts for both speech and response text, which means you no longer have to assemble your own STT component just to get logs and compliance records. On safety, audio produced through ChatGPT Voice and the API is watermarked with SynthID; OpenAI's public verification tool detects that provenance signal, and there is now API access for verification so teams can fold "is this audio AI-generated" into their own pipelines.
// 1) Open a full-duplex session and pick a voice. The model listens
// and speaks concurrently; you never wait for a "turn" to end.
const session = await fetch("https://api.openai.com/v1/realtime/sessions", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.OPENAI_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "gpt-live-1",
voice: "alloy-custom",
instructions: "You are a support agent. Keep replies under 20 seconds.",
input_audio_transcription: { model: "whisper-1" }, // built-in transcripts
}),
}).then((r) => r.json());
console.log("session", session.id, "expires", session.expires_at);2. Why Full Duplex Is More Than "A Bit Faster"
In a three-stage pipeline every step waits for the previous one, so latency adds up. Full duplex makes the work concurrent: the model consumes an audio stream and produces one at the same time, with no hard "turn" boundary in between. Three capabilities actually get unlocked. First, natural interruption: the model can tell the difference between a user cutting in and a user just acknowledging. Second, waiting strategy: as the user's pause grows, the system does not have to rush to fill the silence. Third, parallelism: the conversation stays alive in the voice model while the slow work is handed to an agent behind it, so the user still has someone to talk to while they wait. If you have shipped a turn-based bot before, you know the hidden tax it carried: a state machine tracking whose turn it is, plus a queue of pending utterances. Full duplex deletes that machine, and it also deletes your safety net. In a turn-based design the model never speaks while the user is speaking, so a bad filler phrase cannot talk over them. Here you enforce that yourself, with playback gates and an explicit barge-in policy.
// 2) Interruptions are events, not exceptions. React to the stream.
function wire(ws, { onUserSpeech, onModelSpeech }) {
ws.on("message", (raw) => {
const evt = JSON.parse(raw);
switch (evt.type) {
case "input_audio_buffer.speech_started":
// user cut in: stop local playback immediately, do not finish the sentence
stopPlayback();
onUserSpeech();
break;
case "response.audio.delta":
queueAudio(evt.delta);
onModelSpeech();
break;
case "response.done":
logTurn(evt.response_id, evt.usage);
break;
}
});
}3. Delegate the Reasoning Instead of Stuffing the Voice Model
A tempting mistake is treating the voice model as the whole brain: have it query databases, compute prices, run long chains of reasoning. The better design keeps GPT-Live-1 on conversation only, and pushes heavy reasoning and actions to a backend agent. The documentation is explicit that you choose the backend model, the tools, and the agent framework; GPT-Live-1 keeps the conversation going and delegates deeper reasoning and actions to that backend. Once that boundary is clean, the voice layer stays low-latency while hard tasks run in your own agent at their own pace and report back into the conversation.
// 3) Keep the conversation alive while the backend agent does the work.
async function handleToolCall(call, voice) {
// fire the slow work behind the voice layer
const job = await backendAgent.start({
task: call.name,
args: call.arguments,
traceId: call.id,
});
// never let the caller sit in silence
voice.say("Let me pull that up for you.");
const result = await job.done();
voice.say(result.spokenSummary); // voice layer only receives the summary
return result.payload; // structured data stays server-side
}4. Doing the Unit Economics: What $0.05 a Minute Means
Per-minute pricing is a meaningful signal for voice products: cost grows linearly with call duration instead of jittering with token counts. A six-minute call costs about $0.30; a thousand of them cost about $300. That makes peak concurrency and average call duration your two most important metrics. Conversely, if your backend agent burns tokens on top, the real cost is the voice fee plus backend inference. Do the two sums separately before launch, and set an idle timeout so a user who leaves the microphone open does not silently run up a bill. Watch the distribution rather than the average, too. A handful of very long calls usually dominates the bill, so alert on p95 call duration and on calls still open past your cap.
# 4) Per-minute economics: two separate bill lines, not one.
VOICE_PER_MIN = 0.05 # USD, GPT-Live-1
BACKEND_IN = 3.0 / 1_000_000 # USD per input token
BACKEND_OUT = 15.0 / 1_000_000
def call_cost(minutes, backend_in, backend_out):
voice = minutes * VOICE_PER_MIN
agent = backend_in * BACKEND_IN + backend_out * BACKEND_OUT
return round(voice + agent, 4), round(voice, 4), round(agent, 4)
total, voice, agent = call_cost(6.0, 40_000, 3_000)
print("total", total, "| voice", voice, "| backend", agent)
# Guard the silent-mic case: cap duration, do not bill by surprise.
MAX_MINUTES = 305. Audio Provenance and Compliance: Do Not Bolt It On Later
Voice is one of the easiest media to fake and abuse. Because GPT-Live-1 includes transcripts, your call records come with a text version by default, which helps both audits and quality review. SynthID watermarking plus the verification API give you a tool for judging provenance. In practice: store audio and transcript together, record the watermark check result, and carry an "is this caller human or AI" field through your compliance path. For customer support, healthcare, and financial scenarios this is not a nice-to-have; it is the entry requirement.
// 5) Provenance: store audio + transcript + watermark verdict together.
async function recordTurn(turn) {
const verdict = await verifyProvenance(turn.audioUrl); // SynthID check
await db.calls.insert({
callId: turn.callId,
audioUrl: turn.audioUrl, // never store audio without its text
transcript: turn.transcript, // GPT-Live-1 gives this to you
watermark: verdict.status, // "openai-synthid" | "none" | "unknown"
generatedBy: "gpt-live-1",
checkedAt: new Date().toISOString(),
});
}6. A Go-Live Checklist
Five rules. First, draw the boundary between the voice layer and the backend agent in your architecture diagram, and never let the voice layer hold a production database connection. Second, set a maximum call length and an idle threshold so cost and safety are managed together. Third, route every tool call through the backend agent so the voice layer receives results, not credentials. Fourth, persist audio and transcript as a pair, with the watermark check attached. Fifth, run an interruption test with real user recordings first; "mm-hm", "right", and "hold on" show up far more often than synthetic demos suggest, and handling them badly makes the whole thing feel dumb. Full duplex is not a new model; it is a new set of constraints. Write them down and it will finally sound natural. And track one metric you would not expect: interruption rate. A sudden rise almost always means your prompts got longer, the model started monologuing, and users are cutting in because they already heard what they needed.
"An interruption is an event, not an exception"
"Store audio and transcript as a pair, with provenance attached"
📌 Frequently Asked Questions
What is actually different about GPT-Live-1 versus older voice models?
The architecture. Older systems chained speech-to-text, an LLM, and text-to-speech, so latency added up and interruptions were awkward. GPT-Live-1 is full duplex: it listens and speaks at the same time, reacts to pauses and interruptions, and delegates deeper reasoning to a backend agent you choose.
Is $0.05 per minute expensive?
It depends on call duration. A six-minute call is about $0.30 and a thousand calls about $300, which is predictable linear pricing. If your backend agent also burns tokens, add that inference cost, and keep the two lines separate.
Do I still need my own speech-to-text layer?
No. GPT-Live-1 ships with built-in transcripts for both speech and response text, which removes a self-hosted STT component and gives you logs, quality review, and compliance records for free.
Can generated audio be traced back to its origin?
Yes. Audio generated through ChatGPT Voice and the API carries a SynthID watermark, and OpenAI offers a public verification tool plus API access for verification, so you can fold provenance checks into your own pipeline.
What is the first trap to avoid?
Treating the voice model as the whole brain and giving it direct production database access. Keep the voice layer on conversation only, and hold tools and credentials in the backend agent.