Skip to main content

Audio Models

Icon Legend

Input: Text · Audio · Output: Text · Audio

Text-to-Speech / TTS

Azure OpenAI

VendorModel IDModel CapabilitiesEndpointPrice (per million tokens)Launch DatePlanned DeprecationCapacitySupported RegionsNotes
Azureturing/tts-1Input:
Output:
Tools: Not supported
v1/audio/speech$152024-12-16-GlobalChina region
Europe region
North America region
-
Azureturing/tts-1-hdInput:
Output:
Tools: Not supported
v1/audio/speech$302024-12-16-GlobalChina region
Europe region
North America region
-

Speech-to-Text / ASR

For meeting minutes, audio transcription, speaker diarization, and proper noun recognition. The API uses an asynchronous task model: first create a transcription job, then poll for results.

VendorService NameModel CapabilitiesEndpointBillingLaunch DatePlanned DeprecationCapacitySupported RegionsNotes
Alibaba Cloudaliyun/tingwuInput:
Output:
Speaker diarization / Proper noun recognition / Audio & video format conversion
v1/audio/transcriptions/runsBilled by audio duration---China regionMeeting minutes ASR; see Speech-to-Text / STT for usage instructions