Chat Model List
deepseek-v4-flash-0731deepseek-v4-flash-0731 now calls DeepSeek's own API; the Aliyun Model Studio copy of the same snapshot is published separately as aliyun/deepseek-v4-flash-0731. Context limits, request parameters and rate-limit quota are identical, and so are the input/output rates -- they differ only on the cache-hit rate (official ¥0.1 vs Model Studio ¥0.3) and on which hours count as off-peak. See Deepseek for details.
OpenAI
turing/gpt-5.6-terra / turing/gpt-5.6-lunaToken prices for short-context usage on both models have been reduced. Call methods, context limits, and rate-limit quotas are unchanged — no code changes required:
turing/gpt-5.6-terra: Input $2.5 → $2, Output $15 → $12, Cache read $0.25 → $0.2turing/gpt-5.6-luna: Input $1 → $0.2, Output $6 → $1.2, Cache read $0.1 → $0.02
The price of turing/gpt-5.6-sol is unchanged. Long-context pricing above 272K tokens is not included in this reduction; refer to the notes in the table for those rates.
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Azure | turing/gpt-5.6-sol | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $30Cache read: $0.5At >272K input, the entire request is priced at 2× input / 1.5× output; cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-07-10 | China Europe North America Asia Pacific | GPT-5.6 flagship model, suited for complex reasoning tasks; supports max reasoning effort | |
| Azure | turing/gpt-5.6-terra | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: At >272K input, the entire request is billed at long-context pricing: input $5 / cache read $0.5 / output $22.5; cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-07-10 | China Europe North America Asia Pacific | GPT-5.6 balanced model, weighing capability against cost; supports max reasoning effort | |
| Azure | turing/gpt-5.6-luna | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: At >272K input, the entire request is billed at long-context pricing: input $2 / cache read $0.2 / output $9; cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-07-10 | China Europe North America Asia Pacific | GPT-5.6's fastest and lowest-cost model, suited for high-throughput tasks; supports max reasoning effort | |
| Azure | turing/gpt-5.5 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $30Cache read: $0.5 | Azure Global | China Europe North America Asia Pacific | Next-generation flagship model; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.4 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $2.5Output: $15Cache read: $0.25 | Azure Global | China Europe North America Asia Pacific | Upgraded version of the previous generation; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.4-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.75Output: $4.5Cache read: $0.07 | Azure Global | China Europe North America Asia Pacific | Cost-effective mini model; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.4-nano | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.2Output: $1.25Cache read: $0.02 | Azure Global | China Europe North America Asia Pacific | Ultra-low-cost nano model, suited for classification and sub-agent tasks | |
| Azure | turing/gpt-5.3-codex | API: v1/responses SDK: OpenAI SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2026-02-26 | China Europe North America Asia Pacific | Supports Response API only; supports reasoning tokens | |
| Azure | turing/gpt-5.3-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2026-03-16 | China Europe North America Asia Pacific | ||
| Azure | turing/gpt-5.2 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2025-12-11 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.2-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2025-12-11 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.2-codex | API: v1/responses SDK: OpenAI SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2025-12-11 | China Europe North America Asia Pacific | Supports Response API only | |
| Azure | turing/gpt-5.1 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-11-13 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.1-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-11-13 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.1-codex | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.1-codex-mini | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | Input: $0.25Output: $2Cache read: $0.03 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.25Output: $2Cache read: $0.03 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5-nano | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.05Output: $0.4Cache read: $0.01 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-08-10 Expected retirement: 2026-04-15 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/o3 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Azure Global Launched: 2025-04-17 | China Europe North America Asia Pacific | Thinking model; Response API recommended | ||
| Azure | turing/gpt-4.1 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $2Output: $8Cache read: $0.50 | Azure Global Launched: 2025-04-16 | China Europe North America Asia Pacific | ||
| Azure | turing/o4-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Azure Global Launched: 2025-04-16 | China Europe North America Asia Pacific | Thinking model; Response API recommended | ||
| Azure | turing/gpt-4.1-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.4Output: $1.6Cache read: $0.10 | Azure Global Launched: 2025-04-14 | China Europe North America Asia Pacific | ||
| Azure | turing/gpt-4.1-nano | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.1Output: $0.4Cache read: $0.03 | Azure Global Launched: 2025-04-14 | China Europe North America Asia Pacific |
Gemini
Gemini models run on Google Cloud Vertex AI under a shared quota, so throughput cannot be guaranteed. Under high-concurrency workloads, 429 rate-limit errors may occur frequently. Direct use in production environments with strict stability requirements is not recommended. The Turing Platform provides Retry and Fallback as engineering mitigations, but these cannot fundamentally resolve the shared-quota issue, and switching models via Fallback may result in inconsistent outputs. To resolve this at the root, additional Provisioned Throughput must be purchased. See: Gemini 429 Rate Limiting and Provisioned Throughput
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Vertex AI | turing/gemini-3.7-flash | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Input: Output: Cache read: Web search: $14 / 1KIntroductory pricing is valid through 2026-12-31 (provided as 50% credits back on net spend); from 2027-01-01, standard pricing is input $1.50 / output $7.50 / cache read $0.15 → | Vertex AI shared quota Launched: 2026-08-17 | China Europe North America Asia Pacific | Latest Flash model; supports PDF input, thinking mode, and built-in web search |
| Vertex AI | turing/gemini-3.6-flash | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2026-07-23 | China Europe North America Asia Pacific | Optimized for multi-step orchestration, full-stack code refactoring, agent execution, and spatial reasoning; supports PDF input, thinking mode, and built-in web search | |
| Vertex AI | turing/gemini-3.5-flash | max_input_tokens: 1,048,576max_output_tokens: 65,535Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2026-05-19 | China Europe North America Asia Pacific | Supports thinking mode & built-in web search | |
| Vertex AI | turing/gemini-3.1-pro-latest | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Input: $2Output: $12Cache read: $0.4Web search: $14 / 1KTiered pricing: above 200K input, input: $4 / output: $18 → | Vertex AI shared quota Launched: 2026-02-19 | China Europe North America Asia Pacific | Latest version; supports thinking mode (including MEDIUM level) & built-in web search |
| Vertex AI | turing/gemini-3.1-flash-lite-latest | max_input_tokens: 1,048,576max_output_tokens: 65,535Output: Tools: Supported Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-12-18 | China Europe North America Asia Pacific | Latest version; supports thinking mode & built-in web search | |
| Vertex AI | turing/gemini-3.1-flash-image | max_input_tokens: 131,072max_output_tokens: 32,768Output: Tools: Supported Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-12-18 | China Europe North America Asia Pacific | Latest version; supports thinking mode & image output | |
| Vertex AI | turing/gemini-3.1-flash-lite-image | max_input_tokens: 65,536max_output_tokens: 4,096Output: Tools: Not supported Reasoning: Supported Content moderation: Supported content moderation | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2026-07-01 | China Europe North America Asia Pacific | Nano Banana 2 Lite, lightweight image generation, output resolution up to 1K | |
| Vertex AI | turing/gemini-3-flash-latest | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-12-18 Expected retirement: 2026-06-15 | China Europe North America Asia Pacific | Supports thinking mode & built-in web search | |
| Vertex AI | turing/gemini-3-pro-image | max_input_tokens: 65,000max_output_tokens: 32,000Output: Tools: - Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-11-19 | China Europe North America Asia Pacific | Latest version; supports thinking mode |
Claude
Due to Anthropic policy restrictions, Claude models may become unavailable at any time. They are recommended for personal projects and experimentation only.
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Anthropic | turing/claude-sonnet-5 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: Cache write 5m: Cache write 1h: Web search: $10 / 1KPromotional pricing valid through 2026-08-31; from 2026-09-01 the standard price is input $3 / output $15 / cache read $0.3 / cache write $3.75 (5m), $6 (1h) → | Vertex AI (global deployment) Launched: 2026-07-01 | China Europe North America Asia Pacific | Latest Sonnet model, adaptive thinking on by default; does not support enabled/budget_tokens; supports tools, explicit caching, and web search |
| Anthropic | turing/claude-fable-5 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $10Output: $50Cache read: $1Cache write 5m: $12.5Cache write 1h: $20Web search: $10 / 1K | Vertex AI (global deployment) | China Europe North America Asia Pacific | Next-generation Claude model, adaptive thinking always on and cannot be disabled; supports tools, explicit caching, and web search; use for model distillation is strictly prohibited, and the Turing platform reserves the right to pursue liability against violators |
| Anthropic | turing/claude-opus-5 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) | China Europe North America Asia Pacific | New: latest Opus flagship model; supports adaptive thinking mode only; supports tools, explicit caching, and web search |
| Anthropic | turing/claude-opus-4.8 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-06-01 | China Europe North America Asia Pacific | Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive) |
| Anthropic | turing/claude-opus-4.7 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-04-15 | China Europe North America Asia Pacific | Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive) |
| Anthropic | turing/claude-opus-4.6 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-02-25 | China Europe North America Asia Pacific | Supports tools; adaptive thinking mode recommended |
| Anthropic | turing/claude-sonnet-4.6 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $3Output: $15Cache read: $0.3Cache write 5m: $3.75Cache write 1h: $6Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-02-25 | China Europe North America Asia Pacific | Latest version; supports tools |
| Anthropic | turing/claude-haiku-4.5 | max_input_tokens: 200,000max_output_tokens: 64,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $1Output: $5Cache read: $0.1Cache write 5m: $1.25Cache write 1h: $2Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-01-19 | China Europe North America Asia Pacific | Supports tools |
ByteDance Volcano
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| ByteDance Volcano | doubao-seed-2.1-pro | max_input_tokens: 262,144max_output_tokens: 262,144Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-06-24 | China | Seed 2.1 flagship edition, supports video understanding | |
| ByteDance Volcano | doubao-seed-2.1-turbo | max_input_tokens: 262,144max_output_tokens: 262,144Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-06-24 | China | Seed 2.1 balanced edition, supports video understanding + GUI agent | |
| ByteDance Volcano | doubao-seed-character | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-06-28 | China | Role-play/virtual companionship, supports image-text understanding, deep thinking, and tool calling | ||
| ByteDance Volcano | doubao-seed-2-0-pro-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Latest version Seed 2.0 flagship edition, supports video understanding | ||
| ByteDance Volcano | doubao-seed-2-0-lite-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Latest version Seed 2.0 mid-range edition, supports video understanding | ||
| ByteDance Volcano | doubao-seed-2-0-lite-260428 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-04-28 | China | Seed 2.0 omni-modal edition, supports unified understanding of text/image/video/audio + deep thinking | ||
| ByteDance Volcano | doubao-seed-2-0-mini-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Latest version Seed 2.0 lightweight edition, supports video understanding | ||
| ByteDance Volcano | doubao-seed-2-0-code-preview-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Code preview edition |
Alibaba DashScope
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Alibaba DashScope | qwen3.8-max | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-08-03 | China | Flagship edition 3.8, vision-language (image/video input) | ||
| Alibaba DashScope | qwen3.7-plus | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-06-08 | China | Cost-effective edition 3.7, vision-language (image input) | ||
| Alibaba DashScope | qwen3.7-flash | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-07-28 | China | Lightweight edition 3.7, vision-language (image input) | ||
| Alibaba DashScope | qwen3.7-max | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-05-28 | China | Flagship edition 3.7 | ||
| Alibaba DashScope | qwen3.6-plus | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-04-03 | China | Stable edition, currently identical in capability to qwen3.6-plus-2026-04-02 | |
| Alibaba DashScope | qwen3.6-flash | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-04-23 | China | Lightweight edition 3.6 | |
| Alibaba DashScope | qwen3.6-max | max_input_tokens: 229,376max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-04-23 Expected retirement: 2026-09-08 | China | Flagship edition 3.6 | |
| Alibaba DashScope | qwen3.5-plus | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-02-25 | China | Snapshot version | |
| Alibaba DashScope | qwen3.5-flash | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-02-25 | China | Snapshot version |
Deepseek
deepseek-v4-pro-0813 (2026-08-13) and deepseek-v4-flash-0731 (2026-07-31) are the stable releases of the two DeepSeek V4 tiers, available under fixed version IDs. The undated deepseek-v4-pro / deepseek-v4-flash remain available with unchanged behavior. The calling method is the same, so switching requires only a model name change.
Note the billing difference: both dated versions bill on a peak/off-peak schedule while the two undated versions keep a single flat rate. Rate-limit quota is shared within a tier (pro with pro-0813, flash with flash-0731 -- aliyun/deepseek-v4-flash-0731 draws on that same flash quota). The window follows the upstream: deepseek-v4-flash-0731 goes to DeepSeek's own API, where peak is only UTC+8 09:00-12:00 and 14:00-18:00; deepseek-v4-pro-0813 and aliyun/deepseek-v4-flash-0731 sit on Aliyun Model Studio, where off-peak is UTC+8 22:00 to 08:00 the next day. Both are settled by bill time -- see Usage & Billing.
bytedance/deepseek-v4-flashA Volcano Engine deployment of DeepSeek V4 Flash, with the same per-token pricing as deepseek-v4-flash on Alibaba Cloud. It additionally supports Ark's built-in web search, but that capability is only available on /v1/responses. See Web Search → ByteDance Ark for details.
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Deepseek | deepseek-v4-flash-0731 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥3Output: ¥9Cache read: 峰谷定价:高峰时段为东八区 09:00-12:00、14:00-18:00(以账单时间为准),其余时间均为闲时,输入 ¥1.5 / 输出 ¥4.5 / 缓存命中 ¥0.05 → | DeepSeek 官方 Launched: 2026-07-31 | China | Direct to DeepSeek's own API; a measured 99% cache-hit rate on long sessions. See the DeepSeek Harness guide for agent setup → | |
| Alibaba DashScope | aliyun/deepseek-v4-flash-0731 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥3Output: ¥9Cache read: ¥0.3Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥1.5 / output ¥4.5 / cache read ¥0.15 → | Alibaba Cloud Launched: 2026-07-31 | China | The Model Studio-hosted V4 Flash stable release -- the same snapshot as the officially routed deepseek-v4-flash-0731, sharing one rate-limit quota; they differ on the cache-hit rate (¥0.3 vs ¥0.1) and on which hours are off-peak (overnight vs outside business peak) | |
| Deepseek | deepseek-v4-pro | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥12Output: ¥24Cache read: ¥1 | Alibaba Cloud Launched: 2026-04-25 | China | ||
| Deepseek | deepseek-v4-pro-0813 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥9Output: ¥27Cache read: ¥0.9Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥4.5 / output ¥13.5 / cache read ¥0.45 | Alibaba Cloud Launched: 2026-08-13 | China | New: V4 Pro stable release (fixed version id), available alongside deepseek-v4-pro and sharing one rate-limit quota; the two are priced differently and this one bills on a peak/off-peak schedule | |
| Deepseek | deepseek-v4-flash | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥1Output: ¥2Cache read: ¥0.2 | Alibaba Cloud Launched: 2026-04-25 | China | ||
| ByteDance Volcano | bytedance/deepseek-v4-flash | max_input_tokens: 1,048,576max_output_tokens: 393,216Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-07-28 | China | Volcano deployment of DeepSeek V4 Flash; built-in web search is available only on v1/responses | |
| Deepseek | turing/deepseek-v3-2 | max_input_tokens: 64,000max_output_tokens: 8,192Input: Output: Tools: Not supported Reasoning: Supported | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: ¥2Output: ¥3[0, 32] input: ¥2 output: ¥3 / (32, 128] input: ¥4 output: ¥6 | Volcano Engine (China) Launched: 2025-12-11 | China | v1/responses supports streaming only |
Zhipu
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Zhipu | glm-5.3 | max_input_tokens: 1,048,576max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-08-19 | China | 开源旗舰,1M 上下文 | |
| Zhipu | glm-5.2 | max_input_tokens: 1,048,576max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-06-17 | China | Latest flagship, 1M context | |
| Zhipu | glm-5.1 | max_input_tokens: 204,800max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-04-09 | China | ||
| Zhipu | glm-5 | max_input_tokens: 204,800max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-02-24 | China | ||
| Zhipu | glm-4.7 | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2025-12-23 | China | |||
| Zhipu | glm-4.6 | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2025-11-11 | China |
MiniMax
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| MiniMax | minimax-m3 | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: Limited-time 50% off; context 512K-1M input ¥4.2 / output ¥16.8 / cache read ¥0.84 → | MiniMax Launched: 2026-06-17 | China | MiniMax M3; v1/responses supports streaming only | |
| MiniMax | minimax-m2.7 | max_input_tokens: 204,800max_output_tokens: 131,072Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-03-23 | China | Latest version M2.7 | |
| MiniMax | minimax-m2.5-highspeed | max_input_tokens: 204,800max_output_tokens: 204,800Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | MiniMax Launched: 2026-02-25 | China | ||
| MiniMax | minimax-m2.5 | max_input_tokens: 204,800max_output_tokens: 204,800Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK / Anthropic SDK | MiniMax Launched: 2026-02-25 | China | ||
| MiniMax | minimax-text-01 | Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | MiniMax Launched: 2025-04-02 | China |
KIMI
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| KIMI | kimi-k3 | API: v1/chat/completions SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-07-21 | China | Kimi flagship model, supports 1M context, image understanding, deep thinking, and tool calling; v1/responses supports streaming only |
Xiaomi MiMo
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Xiaomi | mimo-v2.5-pro | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-05-28 | China | Xiaomi open-source reasoning model |
Bedrock
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Bedrock | turing/nova-lite-v2 | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.30Output: $2.50Cache read: $0.075 | AWS Bedrock (US) Launched: 2026-07-22 | China Europe North America Asia Pacific | Nova 2 Lite | |
| Bedrock | turing/nova-pro | max_input_tokens: 300,000max_output_tokens: 10,000Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.80Output: $3.20 | AWS Bedrock (US) Launched: 2026-07-22 | China Europe North America Asia Pacific | Nova Pro |
| Bedrock | nova-lite-v1 | max_input_tokens: 290,000max_output_tokens: 10,000Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.06Output: $0.24 | AWS Bedrock (US) Launched: 2025-08-21 | China Europe North America Asia Pacific | |
| Bedrock | turing/titan-nove-lite-v1 | max_input_tokens: 290,000max_output_tokens: 10,000Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.06Output: $0.24 | AWS Bedrock (US) Launched: 2025-08-21 | China Europe North America Asia Pacific |
Grok
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| xAI | turing/grok-4.3 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $2.5Cache read: $0.2 | xAI API Launched: 2026-05-28 | China Europe North America Asia Pacific | ||
| xAI | turing/grok-4-0709 | max_input_tokens: 256,000max_output_tokens: 256,000Input: Output: Tools: Supported | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $3Output: $15 | xAI API Launched: 2025-10-13 | China Europe North America Asia Pacific | |
| xAI | turing/grok-4-fast-non-reasoning | max_input_tokens: 2,000,000max_output_tokens: 2,000,000Input: Output: Tools: Not supported | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.2Output: $0.5 | xAI API Launched: 2025-10-13 | China Europe North America Asia Pacific | |
| xAI | turing/grok-4-fast-reasoning | max_input_tokens: 2,000,000max_output_tokens: 2,000,000Input: Output: Tools: Supported Reasoning: Supported | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.2Output: $0.5 | xAI API Launched: 2025-10-13 | China Europe North America Asia Pacific |