Audio Models
Icon Legend
Input: Text · Audio · Output: Text · Audio
Text-to-Speech / TTS
Azure OpenAI
| Vendor | Model ID | Model Capabilities | Endpoint | Price (per million tokens) | Launch Date | Planned Deprecation | Capacity | Supported Regions | Notes |
|---|---|---|---|---|---|---|---|---|---|
| Azure | turing/tts-1 | Input: Output: Tools: Not supported | v1/audio/speech | $15 | 2024-12-16 | - | Global | China region Europe region North America region | - |
| Azure | turing/tts-1-hd | Input: Output: Tools: Not supported | v1/audio/speech | $30 | 2024-12-16 | - | Global | China region Europe region North America region | - |
Speech-to-Text / ASR
For meeting minutes, audio transcription, speaker diarization, and proper noun recognition. The API uses an asynchronous task model: first create a transcription job, then poll for results.
| Vendor | Service Name | Model Capabilities | Endpoint | Billing | Launch Date | Planned Deprecation | Capacity | Supported Regions | Notes |
|---|---|---|---|---|---|---|---|---|---|
| Alibaba Cloud | aliyun/tingwu | Input: Output: Speaker diarization / Proper noun recognition / Audio & video format conversion | v1/audio/transcriptions/runs | Billed by audio duration | - | - | - | China region | Meeting minutes ASR; see Speech-to-Text / STT for usage instructions |
Realtime Voice / GPT-Live
Full-duplex realtime voice conversation: the model listens while it speaks, can be interrupted at any time, and delegates lookups and tool calls to your application or to a Responses model. Browsers and apps connect over WebRTC with audio flowing directly to the service; servers and devices can stream audio through a Turing WebSocket.
| Vendor | Model ID | Model Capabilities | Endpoint | Price | Launch Date | Planned Deprecation | Capacity | Supported Regions | Notes |
|---|---|---|---|---|---|---|---|---|---|
| Azure | gpt-live-1 | Input: Output: Tools: Client / Responses delegation | v1/live/sessions (WebRTC)ws/v1/live/sessions (WebSocket) | $0.05 / minute, billed per second | - | - | Global | China region | At most $10 per session; see GPT-Live Voice Sessions for usage instructions |