Skip to main content

Audio Models

Icon Legend

Input: Text · Audio · Output: Text · Audio

Text-to-Speech / TTS​

Azure OpenAI​

VendorModel IDModel CapabilitiesEndpointPrice (per million tokens)Launch DatePlanned DeprecationCapacitySupported RegionsNotes
Azureturing/tts-1Input:
Output:
Tools: Not supported
v1/audio/speech$152024-12-16-GlobalChina region
Europe region
North America region
-
Azureturing/tts-1-hdInput:
Output:
Tools: Not supported
v1/audio/speech$302024-12-16-GlobalChina region
Europe region
North America region
-

Speech-to-Text / ASR​

For meeting minutes, audio transcription, speaker diarization, and proper noun recognition. The API uses an asynchronous task model: first create a transcription job, then poll for results.

VendorService NameModel CapabilitiesEndpointBillingLaunch DatePlanned DeprecationCapacitySupported RegionsNotes
Alibaba Cloudaliyun/tingwuInput:
Output:
Speaker diarization / Proper noun recognition / Audio & video format conversion
v1/audio/transcriptions/runsBilled by audio duration---China regionMeeting minutes ASR; see Speech-to-Text / STT for usage instructions

Realtime Voice / GPT-Live​

Full-duplex realtime voice conversation: the model listens while it speaks, can be interrupted at any time, and delegates lookups and tool calls to your application or to a Responses model. Browsers and apps connect over WebRTC with audio flowing directly to the service; servers and devices can stream audio through a Turing WebSocket.

VendorModel IDModel CapabilitiesEndpointPriceLaunch DatePlanned DeprecationCapacitySupported RegionsNotes
Azuregpt-live-1Input:
Output:
Tools: Client / Responses delegation
v1/live/sessions (WebRTC)
ws/v1/live/sessions (WebSocket)
$0.05 / minute, billed per second--GlobalChina regionAt most $10 per session; see GPT-Live Voice Sessions for usage instructions