Skip to main content

GPT-Live Voice Sessions

GPT-Live (gpt-live-1) is a full-duplex realtime voice model: it listens while it speaks and can be interrupted at any time. When a request needs a lookup or a tool, it delegates that work to your application or to a Responses model while the conversation continues. Turing offers two transports: browsers and apps use WebRTC, where your server creates the session and audio flows directly to the service; servers and devices can also stream audio through a Turing WebSocket.

Full schema and Try-It

This page covers the integration flow, session control, and billing. For request and response fields, see API Reference → Create a GPT-Live WebRTC session and Stop a running GPT-Live session.

How it differs from GPT Realtime

GPT-Live uses its own session protocol; its endpoints and events are not interchangeable with the turing/gpt-realtime series. For GPT Realtime, see the GPT Realtime Usage Guide.

Overview​

ItemDescription
Create a session (WebRTC)POST /api/v1/live/sessions
WebSocket sessionwss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions; see WebSocket transport
Stop a sessionPOST /api/v1/live/sessions/{session_id}/close
Modelgpt-live-1
TransportWebRTC: audio on the media track, session events on a data channel named oai-events; WebSocket: audio and events on one connection
Delegationclient (default) or responses; see Delegation
AuthenticationAuthorization: Bearer $TURING_API_KEY, server side only
BillingVoice by session duration at $0.05 / minute, billed per second; Responses delegation is billed separately by the model it uses
Supported regionsChina region

Choosing a transport​

TransportUse it forNotes
WebRTCBrowsers and appsLowest latency, with the browser's echo cancellation and codecs; media connects directly to the service on port 3478 (UDP / TCP), which the network must allow
WebSocketYour servers, devices, and networks that block UDPRuns over TCP 443, so it works wherever HTTPS does; the caller captures, encodes, and plays 24 kHz PCM16 audio

Session events, delegation, the stop endpoint, and billing are the same for both transports.

WebRTC integration flow​

  1. The browser creates an RTCPeerConnection, adds the microphone track and a data channel named oai-events, and creates an SDP offer.
  2. The browser sends the offer to your server.
  3. Your server calls the create-session endpoint with your Turing API Key and receives session.id and the SDP answer.
  4. Once the browser applies the answer, audio flows directly between the browser and GPT-Live, and session events travel over the data channel.
Never use your Turing API Key in the browser

GPT-Live issues no ephemeral client keys, so sessions must be created on a server that holds the API Key. The browser only exchanges SDP with your server.

Browser​

const audioElement = document.querySelector("audio");
const pc = new RTCPeerConnection();
pc.ontrack = (event) => {
audioElement.srcObject = event.streams[0];
};

const microphone = await navigator.mediaDevices.getUserMedia({ audio: true });
pc.addTrack(microphone.getAudioTracks()[0], microphone);

const events = pc.createDataChannel("oai-events");
events.addEventListener("message", ({ data }) => {
const event = JSON.parse(data);
console.log(event.type, event);
});

const offer = await pc.createOffer();
await pc.setLocalDescription(offer);

// Your own server endpoint, which calls POST /live/sessions (see the next section)
const { session_id, answer_sdp } = await fetch("/your-server/live-session", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ sdp: offer.sdp }),
}).then((response) => response.json());

await pc.setRemoteDescription({ type: "answer", sdp: answer_sdp });

Server: create the session​

Your server sends the browser's SDP offer together with the session configuration, then returns the SDP answer to the browser. Requests and responses follow the GPT-Live WebRTC protocol and are passed through unchanged.

# OFFER_SDP holds the SDP offer the browser sent to your server
curl "https://live-turing.cn.llm.tcljd.com/api/v1/live/sessions" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d "$(jq -n --arg sdp "$OFFER_SDP" '{
session: {model: "gpt-live-1", instructions: "Be concise.", delegation: {type: "client"}},
transport: {type: "webrtc", sdp: $sdp}
}')"
  • A successful call returns 201. Use session.id to stop the session and when troubleshooting.
  • The x-turing-trace-id response header is the session's trace ID; use it to look up the bill after settlement (see Billing and limits).
  • The upstream validates the session object strictly and rejects unknown fields. Available fields: model (required, gpt-live-1), instructions, audio.output.voice, and delegation (see Delegation). model, instructions, and the voice cannot change once the session starts.
  • Optional query parameter idle_timeout_seconds: Turing closes the session once nobody has spoken in it (no transcript events) for this many seconds, for example POST /api/v1/live/sessions?idle_timeout_seconds=300. Without it the platform's one hour applies, and a larger value is treated as one hour. It must be a positive integer, or the call returns 400 (error code 1004).

WebSocket transport​

Connect to wss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions (or /api/openai/v1/live/sessions) and authenticate with the Authorization: Bearer $TURING_API_KEY header, from a server or device only. Then exchange events following GPT-Live's WebSocket protocol:

  1. Send session.start with the same session object as a WebRTC session; the session begins when session.started arrives.
  2. Send base64-encoded 24 kHz mono PCM16 audio with session.input_audio.append, and read the model's speech from session.output_audio.delta.
  3. Send session.close; the connection closes after session.closed arrives.

The x-turing-trace-id handshake response header is the session's trace ID. The connection URL can carry idle_timeout_seconds too (for example /ws/v1/live/sessions?idle_timeout_seconds=300), with the same rules as WebRTC. The connection is checked against your budget, rate limits, concurrent-session limit and idle_timeout_seconds too; if one is not met, you receive an error message and the connection closes with code 1008. A command Turing rejects (such as switching the delegation to an unsupported model) is answered with an error event carrying its client_event_id, is not sent upstream, and the session continues.

import asyncio
import base64
import json
import os

import websockets


async def speak_then_close(ws, audio: str) -> None:
await ws.send(json.dumps({"type": "session.input_audio.append", "audio": audio}))
await asyncio.sleep(10) # leave time for the model to answer
await ws.send(json.dumps({"type": "session.close"}))


async def main() -> None:
async with websockets.connect(
"wss://live-turing.cn.llm.tcljd.com/ws/v1/live/sessions",
extra_headers={"Authorization": f"Bearer {os.environ['TURING_API_KEY']}"},
) as ws:
print("trace id:", ws.response_headers.get("x-turing-trace-id"))
await ws.send(
json.dumps({"type": "session.start", "session": {"model": "gpt-live-1", "instructions": "Be concise."}})
)
# hello_24khz_mono.pcm: headerless 24 kHz mono PCM16 audio
audio = base64.b64encode(open("hello_24khz_mono.pcm", "rb").read()).decode()
async for message in ws:
event = json.loads(message)
if event["type"] == "session.started":
asyncio.ensure_future(speak_then_close(ws, audio))
elif event["type"] == "session.output_audio.delta":
pcm16 = base64.b64decode(event["delta"]) # hand to your player
elif event["type"] == "session.closed":
print("usage:", event["usage"])


asyncio.run(main())

Session events​

Common events on both transports (the data channel for WebRTC):

EventDescription
session.startedThe session has started
session.input_transcript.delta / session.output_transcript.deltaTranscript fragments of the user's and the model's speech, split by audio cadence rather than complete turns
session.delegation.createdThe model handed a unit of work to your application; when it is done, return the result with session.thinking.append or session.commentary.append and its delegation_id
response.eventEnvelope for Responses delegation events; dispatch on the nested event.type, such as response.output_item.done and response.completed
session.usage.updatedCumulative voice duration (usage.seconds), about once a minute; it is a running total, so do not add snapshots together
session.closedThe session has ended, with the final usage.seconds and a reason
errorStartup, validation, or command error

For the full event definitions, see the Azure GPT-Live event reference.

Delegation​

When the conversation needs a lookup, a tool, or deeper reasoning, GPT-Live delegates that work while the voice conversation continues. Choose the delegation mode with session.delegation when you create the session.

Client delegation​

The default. When the data channel delivers session.delegation.created with target set to client, your application does the work and returns the result with session.thinking.append (quiet context) or session.commentary.append (for the model to say aloud), carrying the delegation_id. Only the voice duration is billed.

Responses delegation​

A Responses model does the work on the server side:

{
"delegation": {
"type": "responses",
"responses": {
"model": "gpt-6-sol",
"instructions": "Use tools when current information is required.",
"tools": [{ "type": "web_search" }]
}
}
}
  • Available models: gpt-6-sol, gpt-6-luna, gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna, also accepted with the turing/ prefix. See the model list for prices.
  • Available tools: function and web_search. Other tool types are rejected (400, error code 1002).
  • Each delegated run is billed as the chosen model, the same way as a /responses call: cached tokens and per-call web_search included. A service_tier of priority is billed at Priority prices, and flex at standard prices.
  • Delegation events arrive inside response.event envelopes. For a function call, read the call_id and arguments from the nested response.output_item.done, run it in your application, submit a function_call_output with response.item.create, then send response.create to continue the delegated run.
  • If a session.update switches the delegation to an unsupported model or tool mid-session, Turing closes the session immediately. A delegated run on a model Turing cannot recognize is billed at the price of the most expensive available model.

Ending a session​

Either way below ends the session, and Turing then settles it with the upstream's final usage:

  • The browser sends {"type": "session.close"} on the data channel. The upstream closes the data channel right after session.closed, and that event does not always reach the browser first, so treat either session.closed or the data channel closing as the end of the session, then close the RTCPeerConnection.
  • Your server calls the stop endpoint. The session's creator or an admin may call it; Turing usually closes the session within 5 seconds. The browser then receives no session.closed; it sees the data channel and the connection close.
curl -X POST "https://live-turing.cn.llm.tcljd.com/api/v1/live/sessions/$SESSION_ID/close" \
-H "Authorization: Bearer $TURING_API_KEY"

The stop endpoint returns 202. data.status is closing (the session was told to close and the upstream has not confirmed yet) or closed (the session has already ended).

Billing and limits​

  • Voice sessions are billed by duration: $0.05 / minute, billed per second. The billed duration covers the whole session, including silence and the time the model spends on delegated work.
  • Sessions are settled asynchronously after they end. Call /trace/{trace_id}/billing with the x-turing-trace-id returned at session creation to find the gpt-live-1 billing record.
  • While a session runs, each cumulative usage report from the upstream (about once a minute) and each finished delegation run is charged to your monthly spend right away; the rest is settled after the session ends. The billing record still covers the whole session.
  • When the upstream returns no final usage, Turing settles the session by its duration from creation to close.
  • Each Responses delegation run is billed separately as the chosen model, in the same trace as the voice session; /trace/{trace_id}/billing shows a billing record for that model.
  • Turing checks your budget and rate limits, including those of the delegation model, before creating a session. If they are exceeded, the call returns 429 (error codes 1110 / 1104 / 1115) and no session is created. While a session runs, Turing checks the budget against the user's current quota after each charge, and closes the session once the monthly budget has run out, including when the quota is lowered during the session.
  • Each user can run at most 3 sessions at a time. Beyond that, creating a session returns 429 (error code 1115), and a WebSocket connection receives an error with the same code before it closes; the slot is freed when a session ends.
  • Silence is billed too, so when nobody has spoken in a session (no transcript events) for an hour, Turing closes it; pass idle_timeout_seconds when creating the session for a shorter idle timeout.
  • A single session lasts at most about 3 hours 10 minutes (11,400 seconds, about $9.50). It costs at most $10, counting voice and delegation together. When either limit is reached, Turing closes the session. A delegated run counts toward the limit only when it finishes, so the last run can take a session slightly past $10.
  • When Turing closes a session, a WebSocket caller receives session.closed, while a WebRTC browser sees the data channel and the connection close. The trace records why in extra_info.gpt_live_terminated_by: max_duration / max_cost (the per-session limits), over_budget (the budget ran out), idle (the idle timeout), manual (the stop endpoint), or delegation_policy (the delegation left the allowed range).

For billing rules and queries, see Billing and Usage.

Limitations​

  • WebRTC media connects directly to the service on port 3478 (UDP / TCP). If the network does not allow it, the data channel never opens; use the WebSocket transport instead.
  • Responses delegation supports only the models listed above and the function / web_search tools.
  • The upstream limits concurrent GPT-Live sessions per subscription, so creating a session at peak times may return an upstream rate-limit error.

Error handling​

  • 401: missing or invalid Turing API Key.
  • 400: the delegation model or tool is not supported (error code 1002), the delegation is malformed (error code 1050), or idle_timeout_seconds is not a positive integer (error code 1004); no session is created.
  • 429: budget, rate limit or concurrent-session limit exceeded; no session is created.
  • Other 4xx: the upstream rejected the request (for example an unparseable SDP or unknown session fields); the status code and error body are passed through unchanged.
  • Stop endpoint: 403 means you can only stop sessions you created, and 404 means the session does not exist or has already been settled.

Use the trace_id in an error response for request tracing; see the error code reference for code meanings.