Skip to main content

Chat Model List

New Model: deepseek-v4-flash-0731

deepseek-v4-flash-0731 now calls DeepSeek's own API; the Aliyun Model Studio copy of the same snapshot is published separately as aliyun/deepseek-v4-flash-0731. Context limits, request parameters and rate-limit quota are identical, and so are the input/output rates -- they differ only on the cache-hit rate (official ¥0.1 vs Model Studio ¥0.3) and on which hours count as off-peak. See Deepseek for details.

OpenAI

Price reduction: turing/gpt-5.6-terra / turing/gpt-5.6-luna

Token prices for short-context usage on both models have been reduced. Call methods, context limits, and rate-limit quotas are unchanged — no code changes required:

  • turing/gpt-5.6-terra: Input $2.5 → $2, Output $15 → $12, Cache read $0.25 → $0.2
  • turing/gpt-5.6-luna: Input $1 → $0.2, Output $6 → $1.2, Cache read $0.1 → $0.02

The price of turing/gpt-5.6-sol is unchanged. Long-context pricing above 272K tokens is not included in this reduction; refer to the notes in the table for those rates.

Filters25/33
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Azureturing/gpt-5.6-sol
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $30
Cache read: $0.5
At >272K input, the entire request is priced at 2× input / 1.5× output; cache writes are billed at 1.25× the input price
Azure Global
Launched: 2026-07-10
China
Europe
North America
Asia Pacific
GPT-5.6 flagship model, suited for complex reasoning tasks; supports max reasoning effort
Azureturing/gpt-5.6-terra
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2.5 $2
Output: $15 $12
Cache read: $0.25 $0.2
At >272K input, the entire request is billed at long-context pricing: input $5 / cache read $0.5 / output $22.5; cache writes are billed at 1.25× the input price
Azure Global
Launched: 2026-07-10
China
Europe
North America
Asia Pacific
GPT-5.6 balanced model, weighing capability against cost; supports max reasoning effort
Azureturing/gpt-5.6-luna
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1 $0.2
Output: $6 $1.2
Cache read: $0.1 $0.02
At >272K input, the entire request is billed at long-context pricing: input $2 / cache read $0.2 / output $9; cache writes are billed at 1.25× the input price
Azure Global
Launched: 2026-07-10
China
Europe
North America
Asia Pacific
GPT-5.6's fastest and lowest-cost model, suited for high-throughput tasks; supports max reasoning effort
Azureturing/gpt-5.5
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $30
Cache read: $0.5
Azure Global
China
Europe
North America
Asia Pacific
Next-generation flagship model; Response API recommended for the thinking feature
Azureturing/gpt-5.4
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2.5
Output: $15
Cache read: $0.25
Azure Global
China
Europe
North America
Asia Pacific
Upgraded version of the previous generation; Response API recommended for the thinking feature
Azureturing/gpt-5.4-mini
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.75
Output: $4.5
Cache read: $0.07
Azure Global
China
Europe
North America
Asia Pacific
Cost-effective mini model; Response API recommended for the thinking feature
Azureturing/gpt-5.4-nano
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.2
Output: $1.25
Cache read: $0.02
Azure Global
China
Europe
North America
Asia Pacific
Ultra-low-cost nano model, suited for classification and sub-agent tasks
Azureturing/gpt-5.3-codex
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/responses
SDK: OpenAI SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2026-02-26
China
Europe
North America
Asia Pacific
Supports Response API only; supports reasoning tokens
Azureturing/gpt-5.3-chat
max_input_tokens: 128,000
max_output_tokens: 16,384
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2026-03-16
China
Europe
North America
Asia Pacific
Azureturing/gpt-5.2
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2025-12-11
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.2-chat
max_input_tokens: 1,047,576
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2025-12-11
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.2-codex
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/responses
SDK: OpenAI SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2025-12-11
China
Europe
North America
Asia Pacific
Supports Response API only
Azureturing/gpt-5.1
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-11-13
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.1-chat
max_input_tokens: 1,047,576
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-11-13
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.1-codex
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5.1-codex-mini
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
Input: $0.25
Output: $2
Cache read: $0.03
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5-mini
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.25
Output: $2
Cache read: $0.03
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5-nano
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.05
Output: $0.4
Cache read: $0.01
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5-chat
max_input_tokens: 1,047,576
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-08-10
Expected retirement: 2026-04-15
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/o3
max_input_tokens: 200,000
max_output_tokens: 100,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2
Output: $8
Cache read: $0.5
Azure Global
Launched: 2025-04-17
China
Europe
North America
Asia Pacific
Thinking model; Response API recommended
Azureturing/gpt-4.1
max_input_tokens: 1,024,000
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2
Output: $8
Cache read: $0.50
Azure Global
Launched: 2025-04-16
China
Europe
North America
Asia Pacific
Azureturing/o4-mini
max_input_tokens: 200,000
max_output_tokens: 100,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.1
Output: $4.4
Cache read: $0.275
Azure Global
Launched: 2025-04-16
China
Europe
North America
Asia Pacific
Thinking model; Response API recommended
Azureturing/gpt-4.1-mini
max_input_tokens: 1,024,000
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.4
Output: $1.6
Cache read: $0.10
Azure Global
Launched: 2025-04-14
China
Europe
North America
Asia Pacific
Azureturing/gpt-4.1-nano
max_input_tokens: 1,024,000
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.1
Output: $0.4
Cache read: $0.03
Azure Global
Launched: 2025-04-14
China
Europe
North America
Asia Pacific

Gemini

Gemini Model Load Notice

Gemini models run on Google Cloud Vertex AI under a shared quota, so throughput cannot be guaranteed. Under high-concurrency workloads, 429 rate-limit errors may occur frequently. Direct use in production environments with strict stability requirements is not recommended. The Turing Platform provides Retry and Fallback as engineering mitigations, but these cannot fundamentally resolve the shared-quota issue, and switching models via Fallback may result in inconsistent outputs. To resolve this at the root, additional Provisioned Throughput must be purchased. See: Gemini 429 Rate Limiting and Provisioned Throughput

Filters9/22
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Vertex AIturing/gemini-3.7-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.50 $0.75
Output: $7.50 $3.75
Cache read: $0.15 $0.075
Web search: $14 / 1K
Introductory pricing is valid through 2026-12-31 (provided as 50% credits back on net spend); from 2027-01-01, standard pricing is input $1.50 / output $7.50 / cache read $0.15
Vertex AI shared quota
Launched: 2026-08-17
China
Europe
North America
Asia Pacific
Latest Flash model; supports PDF input, thinking mode, and built-in web search
Vertex AIturing/gemini-3.6-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.5
Output: $7.5
Cache read: $0.15
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2026-07-23
China
Europe
North America
Asia Pacific
Optimized for multi-step orchestration, full-stack code refactoring, agent execution, and spatial reasoning; supports PDF input, thinking mode, and built-in web search
Vertex AIturing/gemini-3.5-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,535
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.5
Output: $9
Cache read: $0.15
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2026-05-19
China
Europe
North America
Asia Pacific
Supports thinking mode & built-in web search
Vertex AIturing/gemini-3.1-pro-latest
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $2
Output: $12
Cache read: $0.4
Web search: $14 / 1K
Tiered pricing: above 200K input, input: $4 / output: $18
Vertex AI shared quota
Launched: 2026-02-19
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode (including MEDIUM level) & built-in web search
Vertex AIturing/gemini-3.1-flash-lite-latest
max_input_tokens: 1,048,576
max_output_tokens: 65,535
Input:
Output:
Tools: Supported
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.25
Output: $1.5
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2025-12-18
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode & built-in web search
Vertex AIturing/gemini-3.1-flash-image
max_input_tokens: 131,072
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.50
Output: $3(Text)/$60(Image)
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2025-12-18
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode & image output
Vertex AIturing/gemini-3.1-flash-lite-image
max_input_tokens: 65,536
max_output_tokens: 4,096
Input:
Output:
Tools: Not supported
Reasoning: Supported
Content moderation: Supported content moderation
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.25
Output: $1.5(Text)/$30(Image)
Vertex AI shared quota
Launched: 2026-07-01
China
Europe
North America
Asia Pacific
Nano Banana 2 Lite, lightweight image generation, output resolution up to 1K
Vertex AIturing/gemini-3-flash-latest
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.5
Output: $3
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2025-12-18
Expected retirement: 2026-06-15
China
Europe
North America
Asia Pacific
Supports thinking mode & built-in web search
Vertex AIturing/gemini-3-pro-image
max_input_tokens: 65,000
max_output_tokens: 32,000
Input:
Output:
Tools: -
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $2
Output: $12(Text)/$120(Image)
Web search: $14 / 1K
Tiered pricing
Vertex AI shared quota
Launched: 2025-11-19
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode

Claude

caution

Due to Anthropic policy restrictions, Claude models may become unavailable at any time. They are recommended for personal projects and experimentation only.

Filters8/20
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Anthropicturing/claude-sonnet-5
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $3 $2
Output: $15 $10
Cache read: $0.3 $0.2
Cache write 5m: $3.75 $2.5
Cache write 1h: $6 $4
Web search: $10 / 1K
Promotional pricing valid through 2026-08-31; from 2026-09-01 the standard price is input $3 / output $15 / cache read $0.3 / cache write $3.75 (5m), $6 (1h)
Vertex AI (global deployment)
Launched: 2026-07-01
China
Europe
North America
Asia Pacific
Latest Sonnet model, adaptive thinking on by default; does not support enabled/budget_tokens; supports tools, explicit caching, and web search
Anthropicturing/claude-fable-5
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $10
Output: $50
Cache read: $1
Cache write 5m: $12.5
Cache write 1h: $20
Web search: $10 / 1K
Vertex AI (global deployment)
China
Europe
North America
Asia Pacific
Next-generation Claude model, adaptive thinking always on and cannot be disabled; supports tools, explicit caching, and web search; use for model distillation is strictly prohibited, and the Turing platform reserves the right to pursue liability against violators
Anthropicturing/claude-opus-5
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
China
Europe
North America
Asia Pacific
New: latest Opus flagship model; supports adaptive thinking mode only; supports tools, explicit caching, and web search
Anthropicturing/claude-opus-4.8
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-06-01
China
Europe
North America
Asia Pacific
Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive)
Anthropicturing/claude-opus-4.7
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-04-15
China
Europe
North America
Asia Pacific
Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive)
Anthropicturing/claude-opus-4.6
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-02-25
China
Europe
North America
Asia Pacific
Supports tools; adaptive thinking mode recommended
Anthropicturing/claude-sonnet-4.6
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $3
Output: $15
Cache read: $0.3
Cache write 5m: $3.75
Cache write 1h: $6
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-02-25
China
Europe
North America
Asia Pacific
Latest version; supports tools
Anthropicturing/claude-haiku-4.5
max_input_tokens: 200,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $1
Output: $5
Cache read: $0.1
Cache write 5m: $1.25
Cache write 1h: $2
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-01-19
China
Europe
North America
Asia Pacific
Supports tools

ByteDance Volcano

Filters8/22
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
ByteDance Volcanodoubao-seed-2.1-pro
max_input_tokens: 262,144
max_output_tokens: 262,144
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥6
Output: ¥30
Cache read: ¥1.2
ByteDance Volcano
Launched: 2026-06-24
ChinaSeed 2.1 flagship edition, supports video understanding
ByteDance Volcanodoubao-seed-2.1-turbo
max_input_tokens: 262,144
max_output_tokens: 262,144
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥3
Output: ¥15
Cache read: ¥0.6
ByteDance Volcano
Launched: 2026-06-24
ChinaSeed 2.1 balanced edition, supports video understanding + GUI agent
ByteDance Volcanodoubao-seed-character
max_input_tokens: 131,072
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥0.8
Output: ¥2
Cache read: ¥0.16
ByteDance Volcano
Launched: 2026-06-28
ChinaRole-play/virtual companionship, supports image-text understanding, deep thinking, and tool calling
ByteDance Volcanodoubao-seed-2-0-pro-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaLatest version Seed 2.0 flagship edition, supports video understanding
ByteDance Volcanodoubao-seed-2-0-lite-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaLatest version Seed 2.0 mid-range edition, supports video understanding
ByteDance Volcanodoubao-seed-2-0-lite-260428
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-04-28
ChinaSeed 2.0 omni-modal edition, supports unified understanding of text/image/video/audio + deep thinking
ByteDance Volcanodoubao-seed-2-0-mini-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaLatest version Seed 2.0 lightweight edition, supports video understanding
ByteDance Volcanodoubao-seed-2-0-code-preview-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaCode preview edition

Alibaba DashScope

Filters9/41
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Alibaba DashScopeqwen3.8-max
max_input_tokens: 991,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥12
Output: ¥36
Cache read: ¥1.5
Alibaba Cloud
Launched: 2026-08-03
ChinaFlagship edition 3.8, vision-language (image/video input)
Alibaba DashScopeqwen3.7-plus
max_input_tokens: 991,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-06-08
ChinaCost-effective edition 3.7, vision-language (image input)
Alibaba DashScopeqwen3.7-flash
max_input_tokens: 991,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥0.2
Output: ¥0.8
Cache read: ¥0.04
Alibaba Cloud
Launched: 2026-07-28
ChinaLightweight edition 3.7, vision-language (image input)
Alibaba DashScopeqwen3.7-max
max_input_tokens: 991,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥6
Output: ¥18
Cache read: ¥1.2
Alibaba Cloud
Launched: 2026-05-28
ChinaFlagship edition 3.7
Alibaba DashScopeqwen3.6-plus
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-04-03
ChinaStable edition, currently identical in capability to qwen3.6-plus-2026-04-02
Alibaba DashScopeqwen3.6-flash
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-04-23
ChinaLightweight edition 3.6
Alibaba DashScopeqwen3.6-max
max_input_tokens: 229,376
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-04-23
Expected retirement: 2026-09-08
ChinaFlagship edition 3.6
Alibaba DashScopeqwen3.5-plus
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-02-25
ChinaSnapshot version
Alibaba DashScopeqwen3.5-flash
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-02-25
ChinaSnapshot version

Deepseek

Dated stable releases

deepseek-v4-pro-0813 (2026-08-13) and deepseek-v4-flash-0731 (2026-07-31) are the stable releases of the two DeepSeek V4 tiers, available under fixed version IDs. The undated deepseek-v4-pro / deepseek-v4-flash remain available with unchanged behavior. The calling method is the same, so switching requires only a model name change.

Note the billing difference: both dated versions bill on a peak/off-peak schedule while the two undated versions keep a single flat rate. Rate-limit quota is shared within a tier (pro with pro-0813, flash with flash-0731 -- aliyun/deepseek-v4-flash-0731 draws on that same flash quota). The window follows the upstream: deepseek-v4-flash-0731 goes to DeepSeek's own API, where peak is only UTC+8 09:00-12:00 and 14:00-18:00; deepseek-v4-pro-0813 and aliyun/deepseek-v4-flash-0731 sit on Aliyun Model Studio, where off-peak is UTC+8 22:00 to 08:00 the next day. Both are settled by bill time -- see Usage & Billing.

New Model: bytedance/deepseek-v4-flash

A Volcano Engine deployment of DeepSeek V4 Flash, with the same per-token pricing as deepseek-v4-flash on Alibaba Cloud. It additionally supports Ark's built-in web search, but that capability is only available on /v1/responses. See Web Search → ByteDance Ark for details.

Filters7/14
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Deepseekdeepseek-v4-flash-0731
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥3
Output: ¥9
Cache read: ¥0.3 ¥0.1
峰谷定价:高峰时段为东八区 09:00-12:00、14:00-18:00(以账单时间为准),其余时间均为闲时,输入 ¥1.5 / 输出 ¥4.5 / 缓存命中 ¥0.05
DeepSeek 官方
Launched: 2026-07-31
ChinaDirect to DeepSeek's own API; a measured 99% cache-hit rate on long sessions. See the DeepSeek Harness guide for agent setup
Alibaba DashScopealiyun/deepseek-v4-flash-0731
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥3
Output: ¥9
Cache read: ¥0.3
Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥1.5 / output ¥4.5 / cache read ¥0.15
Alibaba Cloud
Launched: 2026-07-31
ChinaThe Model Studio-hosted V4 Flash stable release -- the same snapshot as the officially routed deepseek-v4-flash-0731, sharing one rate-limit quota; they differ on the cache-hit rate (¥0.3 vs ¥0.1) and on which hours are off-peak (overnight vs outside business peak)
Deepseekdeepseek-v4-pro
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥12
Output: ¥24
Cache read: ¥1
Alibaba Cloud
Launched: 2026-04-25
China
Deepseekdeepseek-v4-pro-0813
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥9
Output: ¥27
Cache read: ¥0.9
Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥4.5 / output ¥13.5 / cache read ¥0.45
Alibaba Cloud
Launched: 2026-08-13
ChinaNew: V4 Pro stable release (fixed version id), available alongside deepseek-v4-pro and sharing one rate-limit quota; the two are priced differently and this one bills on a peak/off-peak schedule
Deepseekdeepseek-v4-flash
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥1
Output: ¥2
Cache read: ¥0.2
Alibaba Cloud
Launched: 2026-04-25
China
ByteDance Volcanobytedance/deepseek-v4-flash
max_input_tokens: 1,048,576
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥1
Output: ¥2
ByteDance Volcano
Launched: 2026-07-28
ChinaVolcano deployment of DeepSeek V4 Flash; built-in web search is available only on v1/responses
Deepseekturing/deepseek-v3-2
max_input_tokens: 64,000
max_output_tokens: 8,192
Input:
Output:
Tools: Not supported
Reasoning: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2
Output: ¥3
[0, 32] input: ¥2 output: ¥3 / (32, 128] input: ¥4 output: ¥6
Volcano Engine (China)
Launched: 2025-12-11
Chinav1/responses supports streaming only

Zhipu

Filters6/17
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Zhipuglm-5.3
max_input_tokens: 1,048,576
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥8
Output: ¥28
Cache read: ¥2
Zhipu
Launched: 2026-08-19
China开源旗舰,1M 上下文
Zhipuglm-5.2
max_input_tokens: 1,048,576
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥8
Output: ¥28
Cache read: ¥2
Zhipu
Launched: 2026-06-17
ChinaLatest flagship, 1M context
Zhipuglm-5.1
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2026-04-09
China
Zhipuglm-5
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2026-02-24
China
Zhipuglm-4.7
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2025-12-23
China
Zhipuglm-4.6
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2025-11-11
China

MiniMax

Filters5/11
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
MiniMaxminimax-m3
max_input_tokens: 1,048,576
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: ¥4.2 ¥2.1
Output: ¥16.8 ¥8.4
Cache read: ¥0.84 ¥0.42
Limited-time 50% off; context 512K-1M input ¥4.2 / output ¥16.8 / cache read ¥0.84
MiniMax
Launched: 2026-06-17
ChinaMiniMax M3; v1/responses supports streaming only
MiniMaxminimax-m2.7
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2.1
Output: ¥8.4
Alibaba Cloud
Launched: 2026-03-23
ChinaLatest version M2.7
MiniMaxminimax-m2.5-highspeed
max_input_tokens: 204,800
max_output_tokens: 204,800
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: ¥4.2
Output: ¥16.8
MiniMax
Launched: 2026-02-25
China
MiniMaxminimax-m2.5
max_input_tokens: 204,800
max_output_tokens: 204,800
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2.1
Output: ¥8.4
MiniMax
Launched: 2026-02-25
China
MiniMaxminimax-text-01
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: ¥1
Output: ¥8
MiniMax
Launched: 2025-04-02
China

KIMI

Filters1/6
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
KIMIkimi-k3
max_input_tokens: 1,048,576
max_output_tokens: 1,048,576
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions
SDK: OpenAI SDK / Anthropic SDK
Input: ¥20
Output: ¥100
Cache read: ¥2
Alibaba Cloud
Launched: 2026-07-21
ChinaKimi flagship model, supports 1M context, image understanding, deep thinking, and tool calling; v1/responses supports streaming only

Xiaomi MiMo

Filters1/1
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Xiaomimimo-v2.5-pro
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥7
Output: ¥21
Cache read: ¥1.4
Alibaba Cloud
Launched: 2026-05-28
ChinaXiaomi open-source reasoning model

Bedrock

Filters4/4
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Bedrockturing/nova-lite-v2
max_input_tokens: 1,000,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Cache: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.30
Output: $2.50
Cache read: $0.075
AWS Bedrock (US)
Launched: 2026-07-22
China
Europe
North America
Asia Pacific
Nova 2 Lite
Bedrockturing/nova-pro
max_input_tokens: 300,000
max_output_tokens: 10,000
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.80
Output: $3.20
AWS Bedrock (US)
Launched: 2026-07-22
China
Europe
North America
Asia Pacific
Nova Pro
Bedrocknova-lite-v1
max_input_tokens: 290,000
max_output_tokens: 10,000
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.06
Output: $0.24
AWS Bedrock (US)
Launched: 2025-08-21
China
Europe
North America
Asia Pacific
Bedrockturing/titan-nove-lite-v1
max_input_tokens: 290,000
max_output_tokens: 10,000
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.06
Output: $0.24
AWS Bedrock (US)
Launched: 2025-08-21
China
Europe
North America
Asia Pacific

Grok

Filters4/4
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
xAIturing/grok-4.3
max_input_tokens: 1,000,000
max_output_tokens: 1,000,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $2.5
Cache read: $0.2
xAI API
Launched: 2026-05-28
China
Europe
North America
Asia Pacific
xAIturing/grok-4-0709
max_input_tokens: 256,000
max_output_tokens: 256,000
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $3
Output: $15
xAI API
Launched: 2025-10-13
China
Europe
North America
Asia Pacific
xAIturing/grok-4-fast-non-reasoning
max_input_tokens: 2,000,000
max_output_tokens: 2,000,000
Input:
Output:
Tools: Not supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.2
Output: $0.5
xAI API
Launched: 2025-10-13
China
Europe
North America
Asia Pacific
xAIturing/grok-4-fast-reasoning
max_input_tokens: 2,000,000
max_output_tokens: 2,000,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.2
Output: $0.5
xAI API
Launched: 2025-10-13
China
Europe
North America
Asia Pacific