GPT-Live Voice Sessions
GPT-Live (gpt-live-1) is a full-duplex realtime voice model: it listens while it speaks and can be interrupted at any time. When a request needs a lookup or a tool, it delegates that work to your application or to a Responses model while the conversation continues. Turing offers two transports: browsers and apps use WebRTC, where your server creates the session and audio flows directly to the service; servers and devices can also stream audio through a Turing WebSocket.
This page covers the integration flow, session control, and billing. For request and response fields, see API Reference → Create a GPT-Live WebRTC session and Stop a running GPT-Live session.
GPT-Live uses its own session protocol; its endpoints and events are not interchangeable with the turing/gpt-realtime series. For GPT Realtime, see the GPT Realtime Usage Guide.
Overview
| Item | Description |
|---|---|
| Create a session (WebRTC) | POST /api/v1/live/sessions |
| WebSocket session | wss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions; see WebSocket transport |
| Stop a session | POST /api/v1/live/sessions/{session_id}/close |
| Model | gpt-live-1 |
| Transport | WebRTC: audio on the media track, session events on a data channel named oai-events; WebSocket: audio and events on one connection |
| Delegation | client (default) or responses; see Delegation |
| Authentication | Authorization: Bearer $TURING_API_KEY, server side only |
| Billing | Voice by session duration at $0.05 / minute, billed per second; Responses delegation is billed separately by the model it uses |
| Supported regions | China region |
Choosing a transport
| Transport | Use it for | Notes |
|---|---|---|
| WebRTC | Browsers and apps | Lowest latency, with the browser's echo cancellation and codecs; media connects directly to the service on port 3478 (UDP / TCP), which the network must allow |
| WebSocket | Your servers, devices, and networks that block UDP | Runs over TCP 443, so it works wherever HTTPS does; the caller captures, encodes, and plays 24 kHz PCM16 audio |
Session events, delegation, the stop endpoint, and billing are the same for both transports.
WebRTC integration flow
- The browser creates an
RTCPeerConnection, adds the microphone track and a data channel namedoai-events, and creates an SDP offer. - The browser sends the offer to your server.
- Your server calls the create-session endpoint with your Turing API Key and receives
session.idand the SDP answer. - Once the browser applies the answer, audio flows directly between the browser and GPT-Live, and session events travel over the data channel.
GPT-Live issues no ephemeral client keys, so sessions must be created on a server that holds the API Key. The browser only exchanges SDP with your server.
Browser
const audioElement = document.querySelector("audio");
const pc = new RTCPeerConnection();
pc.ontrack = (event) => {
audioElement.srcObject = event.streams[0];
};
const microphone = await navigator.mediaDevices.getUserMedia({ audio: true });
pc.addTrack(microphone.getAudioTracks()[0], microphone);
const events = pc.createDataChannel("oai-events");
events.addEventListener("message", ({ data }) => {
const event = JSON.parse(data);
console.log(event.type, event);
});
const offer = await pc.createOffer();
await pc.setLocalDescription(offer);
// Your own server endpoint, which calls POST /live/sessions (see the next section)
const { session_id, answer_sdp } = await fetch("/your-server/live-session", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ sdp: offer.sdp }),
}).then((response) => response.json());
await pc.setRemoteDescription({ type: "answer", sdp: answer_sdp });
Server: create the session
Your server sends the browser's SDP offer together with the session configuration, then returns the SDP answer to the browser. Requests and responses follow the GPT-Live WebRTC protocol and are passed through unchanged.
- cURL
- Python
- Node.js
# OFFER_SDP holds the SDP offer the browser sent to your server
curl "https://live-turing.cn.llm.tcljd.com/api/v1/live/sessions" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d "$(jq -n --arg sdp "$OFFER_SDP" '{
session: {model: "gpt-live-1", instructions: "Be concise.", delegation: {type: "client"}},
transport: {type: "webrtc", sdp: $sdp}
}')"
import os
import requests
def create_live_session(offer_sdp: str) -> dict:
response = requests.post(
"https://live-turing.cn.llm.tcljd.com/api/v1/live/sessions",
headers={"Authorization": f"Bearer {os.environ['TURING_API_KEY']}"},
json={
"session": {
"model": "gpt-live-1",
"instructions": "Be concise.",
"delegation": {"type": "client"},
},
"transport": {"type": "webrtc", "sdp": offer_sdp},
},
timeout=60,
)
response.raise_for_status()
created = response.json()
return {
"session_id": created["session"]["id"],
"answer_sdp": created["transport"]["sdp"],
"trace_id": response.headers.get("x-turing-trace-id"),
}
export async function createLiveSession(offerSdp) {
const response = await fetch("https://live-turing.cn.llm.tcljd.com/api/v1/live/sessions", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.TURING_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({
session: { model: "gpt-live-1", instructions: "Be concise.", delegation: { type: "client" } },
transport: { type: "webrtc", sdp: offerSdp },
}),
});
if (!response.ok) {
throw new Error(`GPT-Live session failed: ${response.status} ${await response.text()}`);
}
const created = await response.json();
return {
sessionId: created.session.id,
answerSdp: created.transport.sdp,
traceId: response.headers.get("x-turing-trace-id"),
};
}
- A successful call returns
201. Usesession.idto stop the session and when troubleshooting. - The
x-turing-trace-idresponse header is the session's trace ID; use it to look up the bill after settlement (see Billing and limits). - The upstream validates the
sessionobject strictly and rejects unknown fields. Available fields:model(required,gpt-live-1),instructions,audio.output.voice, anddelegation(see Delegation).model,instructions, and the voice cannot change once the session starts. - Optional query parameter
idle_timeout_seconds: Turing closes the session once nobody has spoken in it (no transcript events) for this many seconds, for examplePOST /api/v1/live/sessions?idle_timeout_seconds=300. Without it the platform's one hour applies, and a larger value is treated as one hour. It must be a positive integer, or the call returns400(error code1004).
WebSocket transport
Connect to wss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions (or /api/openai/v1/live/sessions) and authenticate with the Authorization: Bearer $TURING_API_KEY header, from a server or device only. Then exchange events following GPT-Live's WebSocket protocol:
- Send
session.startwith the samesessionobject as a WebRTC session; the session begins whensession.startedarrives. - Send base64-encoded 24 kHz mono PCM16 audio with
session.input_audio.append, and read the model's speech fromsession.output_audio.delta. - Send
session.close; the connection closes aftersession.closedarrives.
The x-turing-trace-id handshake response header is the session's trace ID. The connection URL can carry idle_timeout_seconds too (for example /ws/v1/live/sessions?idle_timeout_seconds=300), with the same rules as WebRTC. The connection is checked against your budget, rate limits, concurrent-session limit and idle_timeout_seconds too; if one is not met, you receive an error message and the connection closes with code 1008. A command Turing rejects (such as switching the delegation to an unsupported model) is answered with an error event carrying its client_event_id, is not sent upstream, and the session continues.
- Python
- Node.js
import asyncio
import base64
import json
import os
import websockets
async def speak_then_close(ws, audio: str) -> None:
await ws.send(json.dumps({"type": "session.input_audio.append", "audio": audio}))
await asyncio.sleep(10) # leave time for the model to answer
await ws.send(json.dumps({"type": "session.close"}))
async def main() -> None:
async with websockets.connect(
"wss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions",
extra_headers={"Authorization": f"Bearer {os.environ['TURING_API_KEY']}"},
) as ws:
print("trace id:", ws.response_headers.get("x-turing-trace-id"))
await ws.send(
json.dumps({"type": "session.start", "session": {"model": "gpt-live-1", "instructions": "Be concise."}})
)
# hello_24khz_mono.pcm: headerless 24 kHz mono PCM16 audio
audio = base64.b64encode(open("hello_24khz_mono.pcm", "rb").read()).decode()
async for message in ws:
event = json.loads(message)
if event["type"] == "session.started":
asyncio.ensure_future(speak_then_close(ws, audio))
elif event["type"] == "session.output_audio.delta":
pcm16 = base64.b64decode(event["delta"]) # hand to your player
elif event["type"] == "session.closed":
print("usage:", event["usage"])
asyncio.run(main())
import fs from "node:fs";
import WebSocket from "ws";
const ws = new WebSocket("wss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions", {
headers: { Authorization: `Bearer ${process.env.TURING_API_KEY}` },
});
ws.on("upgrade", (response) => console.log("trace id:", response.headers["x-turing-trace-id"]));
ws.on("open", () => {
ws.send(JSON.stringify({ type: "session.start", session: { model: "gpt-live-1", instructions: "Be concise." } }));
});
ws.on("message", (data) => {
const event = JSON.parse(data.toString());
if (event.type === "session.started") {
// hello_24khz_mono.pcm: headerless 24 kHz mono PCM16 audio
const audio = fs.readFileSync("hello_24khz_mono.pcm").toString("base64");
ws.send(JSON.stringify({ type: "session.input_audio.append", audio }));
setTimeout(() => ws.send(JSON.stringify({ type: "session.close" })), 10_000);
} else if (event.type === "session.output_audio.delta") {
const pcm16 = Buffer.from(event.delta, "base64"); // hand to your player
} else if (event.type === "session.closed") {
console.log("usage:", event.usage);
}
});
Session events
Common events on both transports (the data channel for WebRTC):
| Event | Description |
|---|---|
session.started | The session has started |
session.input_transcript.delta / session.output_transcript.delta | Transcript fragments of the user's and the model's speech, split by audio cadence rather than complete turns |
session.delegation.created | The model handed a unit of work to your application; when it is done, return the result with session.thinking.append or session.commentary.append and its delegation_id |
response.event | Envelope for Responses delegation events; dispatch on the nested event.type, such as response.output_item.done and response.completed |
session.usage.updated | Cumulative voice duration (usage.seconds), about once a minute; it is a running total, so do not add snapshots together |
session.closed | The session has ended, with the final usage.seconds and a reason |
error | Startup, validation, or command error |
For the full event definitions, see the Azure GPT-Live event reference.
Delegation
When the conversation needs a lookup, a tool, or deeper reasoning, GPT-Live delegates that work while the voice conversation continues. Choose the delegation mode with session.delegation when you create the session.
Client delegation
The default. When the data channel delivers session.delegation.created with target set to client, your application does the work and returns the result with session.thinking.append (quiet context) or session.commentary.append (for the model to say aloud), carrying the delegation_id. Only the voice duration is billed.
Responses delegation
A Responses model does the work on the server side:
{
"delegation": {
"type": "responses",
"responses": {
"model": "gpt-6-sol",
"instructions": "Use tools when current information is required.",
"tools": [{ "type": "web_search" }]
}
}
}
- Available models:
gpt-6-sol,gpt-6-luna,gpt-6-astra,gpt-5.6-sol,gpt-5.6-terra, andgpt-5.6-luna, also accepted with theturing/prefix. See the model list for prices. - Available tools:
functionandweb_search. Other tool types are rejected (400, error code1002). - Each delegated run is billed as the chosen model, the same way as a
/responsescall: cached tokens and per-callweb_searchincluded. Aservice_tierofpriorityis billed at Priority prices, andflexat standard prices. - Delegation events arrive inside
response.eventenvelopes. For afunctioncall, read thecall_idand arguments from the nestedresponse.output_item.done, run it in your application, submit afunction_call_outputwithresponse.item.create, then sendresponse.createto continue the delegated run. - If a
session.updateswitches the delegation to an unsupported model or tool mid-session, Turing closes the session immediately. A delegated run on a model Turing cannot recognize is billed at the price of the most expensive available model.
Ending a session
Either way below ends the session, and Turing then settles it with the upstream's final usage:
- The browser sends
{"type": "session.close"}on the data channel. The upstream closes the data channel right aftersession.closed, and that event does not always reach the browser first, so treat eithersession.closedor the data channel closing as the end of the session, then close theRTCPeerConnection. - Your server calls the stop endpoint. The session's creator or an admin may call it; Turing usually closes the session within 5 seconds. The browser then receives no
session.closed; it sees the data channel and the connection close.
curl -X POST "https://live-turing.cn.llm.tcljd.com/api/v1/live/sessions/$SESSION_ID/close" \
-H "Authorization: Bearer $TURING_API_KEY"
The stop endpoint returns 202. data.status is closing (the session was told to close and the upstream has not confirmed yet) or closed (the session has already ended).
Billing and limits
- Voice sessions are billed by duration:
$0.05/ minute, billed per second. The billed duration covers the whole session, including silence and the time the model spends on delegated work. - Sessions are settled asynchronously after they end. Call
/trace/{trace_id}/billingwith thex-turing-trace-idreturned at session creation to find thegpt-live-1billing record. - While a session runs, each cumulative usage report from the upstream (about once a minute) and each finished delegation run is charged to your monthly spend right away; the rest is settled after the session ends. The billing record still covers the whole session.
- When the upstream returns no final usage, Turing settles the session by its duration from creation to close.
- Each Responses delegation run is billed separately as the chosen model, in the same trace as the voice session;
/trace/{trace_id}/billingshows a billing record for that model. - Turing checks your budget and rate limits, including those of the delegation model, before creating a session. If they are exceeded, the call returns
429(error codes1110/1104/1115) and no session is created. While a session runs, Turing checks the budget against the user's current quota after each charge, and closes the session once the monthly budget has run out, including when the quota is lowered during the session. - Each user can run at most 3 sessions at a time. Beyond that, creating a session returns
429(error code1115), and a WebSocket connection receives an error with the same code before it closes; the slot is freed when a session ends. - Silence is billed too, so when nobody has spoken in a session (no transcript events) for an hour, Turing closes it; pass
idle_timeout_secondswhen creating the session for a shorter idle timeout. - A single session lasts at most about 3 hours 10 minutes (11,400 seconds, about
$9.50). It costs at most$10, counting voice and delegation together. When either limit is reached, Turing closes the session. A delegated run counts toward the limit only when it finishes, so the last run can take a session slightly past$10. - When Turing closes a session, a WebSocket caller receives
session.closed, while a WebRTC browser sees the data channel and the connection close. The trace records why inextra_info.gpt_live_terminated_by:max_duration/max_cost(the per-session limits),over_budget(the budget ran out),idle(the idle timeout),manual(the stop endpoint), ordelegation_policy(the delegation left the allowed range).
For billing rules and queries, see Billing and Usage.
Limitations
- WebRTC media connects directly to the service on port 3478 (UDP / TCP). If the network does not allow it, the data channel never opens; use the WebSocket transport instead.
- Responses delegation supports only the models listed above and the
function/web_searchtools. - The upstream limits concurrent GPT-Live sessions per subscription, so creating a session at peak times may return an upstream rate-limit error.
Error handling
401: missing or invalid Turing API Key.400: the delegation model or tool is not supported (error code1002), the delegation is malformed (error code1050), oridle_timeout_secondsis not a positive integer (error code1004); no session is created.429: budget, rate limit or concurrent-session limit exceeded; no session is created.- Other
4xx: the upstream rejected the request (for example an unparseable SDP or unknown session fields); the status code and error body are passed through unchanged. - Stop endpoint:
403means you can only stop sessions you created, and404means the session does not exist or has already been settled.
Use the trace_id in an error response for request tracing; see the error code reference for code meanings.