Chat Model List
OpenAI
turing/gpt-6-sol / turing/gpt-6-lunaGPT-6 Sol balances intelligence and speed for reasoning and coding; GPT-6 Luna targets cost-sensitive, high-volume workloads. Both share GPT-6 Astra's request contract: v1/responses is recommended; v1/chat/completions has limited support for text requests, and tool calling requires the Responses API; reasoning effort supports low / medium / high / xhigh / max, but not none or minimal. Short-context pricing: Sol $2 input / $10 output / $0.2 cached input, Luna $0.1 input / $0.5 output / $0.01 cached input; requests over 272K input tokens use long-context rates for the full request.
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Azure | turing/gpt-6-astra | max_input_tokens: 922,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions (Limited support), v1/responses SDK: OpenAI SDK | Input: $10Output: $50Cache read: $1Cache write 5m: $12.5Web search: $10 / 1KFor requests with >272K input tokens, the entire request uses long-context rates (input $20 / cache read $2 / cache write $25 / output $75); cache writes are billed at 1.25× the input rate. Azure's pricing page is still being updated; use the Azure bill as the source of truth → | Azure Global Launched: 2026-09-03 | China Europe North America Asia Pacific | GPT-6 Astra for complex reasoning, coding, research, and document creation; v1/chat/completions has limited support (text only); reasoning effort supports low/medium/high/xhigh/max; use the Responses API for tool calling |
| Azure | turing/gpt-6-sol | API: v1/chat/completions (Limited support), v1/responses SDK: OpenAI SDK | Input: $2Output: $10Cache read: $0.2Cache write 5m: $2.5For requests with >272K input tokens, the entire request uses long-context rates (input $4 / cache read $0.4 / cache write $5 / output $15); cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-09-23 | China Europe North America Asia Pacific | GPT-6 Sol balances intelligence and speed for reasoning and coding; v1/chat/completions has limited support (text only); reasoning effort supports low/medium/high/xhigh/max; use the Responses API for tool calling | |
| Azure | turing/gpt-6-luna | API: v1/chat/completions (Limited support), v1/responses SDK: OpenAI SDK | Input: $0.1Output: $0.5Cache read: $0.01Cache write 5m: $0.125For requests with >272K input tokens, the entire request uses long-context rates (input $0.2 / cache read $0.02 / cache write $0.25 / output $0.75); cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-09-23 | China Europe North America Asia Pacific | GPT-6 Luna is the cost-efficient model for cost-sensitive, high-volume workloads; v1/chat/completions has limited support (text only); reasoning effort supports low/medium/high/xhigh/max; use the Responses API for tool calling | |
| Azure | turing/gpt-5.6-sol | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: $0.5The short-context promotional price takes effect 2026-09-01 and runs through at least 2026-11-30; it reduces input and output only -- cache read stays $0.5 and cache write stays $6.25. At >272K input, the entire request is billed at long-context pricing: input $10 / cache read $1 / cache write $12.5 / output $45 → | Azure Global Launched: 2026-07-10 | China Europe North America Asia Pacific | GPT-5.6 flagship model, suited for complex reasoning tasks; supports max reasoning effort | |
| Azure | turing/gpt-5.6-terra | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: At >272K input, the entire request is billed at long-context pricing: input $5 / cache read $0.5 / output $22.5; cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-07-10 | China Europe North America Asia Pacific | GPT-5.6 balanced model, weighing capability against cost; supports max reasoning effort | |
| Azure | turing/gpt-5.6-luna | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: At >272K input, the entire request is billed at long-context pricing: input $2 / cache read $0.2 / output $9; cache writes are billed at 1.25× the input price → | Azure Global Launched: 2026-07-10 | China Europe North America Asia Pacific | GPT-5.6's fastest and lowest-cost model, suited for high-throughput tasks; supports max reasoning effort | |
| Azure | turing/gpt-5.5 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $30Cache read: $0.5 | Azure Global | China Europe North America Asia Pacific | Next-generation flagship model; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.4 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $2.5Output: $15Cache read: $0.25 | Azure Global | China Europe North America Asia Pacific | Upgraded version of the previous generation; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.4-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.75Output: $4.5Cache read: $0.07 | Azure Global | China Europe North America Asia Pacific | Cost-effective mini model; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.4-nano | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.2Output: $1.25Cache read: $0.02 | Azure Global | China Europe North America Asia Pacific | Ultra-low-cost nano model, suited for classification and sub-agent tasks | |
| Azure | turing/gpt-5.3-codex | API: v1/responses SDK: OpenAI SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2026-02-26 | China Europe North America Asia Pacific | Supports Response API only; supports reasoning tokens | |
| Azure | turing/gpt-5.3-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2026-03-16 | China Europe North America Asia Pacific | ||
| Azure | turing/gpt-5.2 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2025-12-11 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.2-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2025-12-11 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.2-codex | API: v1/responses SDK: OpenAI SDK | Input: $1.75Output: $14Cache read: $0.17 | Azure Global Launched: 2025-12-11 | China Europe North America Asia Pacific | Supports Response API only | |
| Azure | turing/gpt-5.1 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-11-13 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.1-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-11-13 | China Europe North America Asia Pacific | Upgraded version; Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.1-codex | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5.1-codex-mini | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | Input: $0.25Output: $2Cache read: $0.03 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.25Output: $2Cache read: $0.03 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5-nano | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.05Output: $0.4Cache read: $0.01 | Azure Global Launched: 2025-08-10 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/gpt-5-chat | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $1.25Output: $10Cache read: $0.12 | Azure Global Launched: 2025-08-10 Expected retirement: 2026-04-15 | China Europe North America Asia Pacific | Response API recommended for the thinking feature | |
| Azure | turing/o3 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Azure Global Launched: 2025-04-17 | China Europe North America Asia Pacific | Thinking model; Response API recommended | ||
| Azure | turing/gpt-4.1 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $2Output: $8Cache read: $0.50 | Azure Global Launched: 2025-04-16 | China Europe North America Asia Pacific | ||
| Azure | turing/o4-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Azure Global Launched: 2025-04-16 | China Europe North America Asia Pacific | Thinking model; Response API recommended | ||
| Azure | turing/gpt-4.1-mini | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.4Output: $1.6Cache read: $0.10 | Azure Global Launched: 2025-04-14 | China Europe North America Asia Pacific | ||
| Azure | turing/gpt-4.1-nano | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $0.1Output: $0.4Cache read: $0.03 | Azure Global Launched: 2025-04-14 | China Europe North America Asia Pacific |
Gemini
Gemini models run on Google Cloud Vertex AI under a shared quota, so throughput cannot be guaranteed. Under high-concurrency workloads, 429 rate-limit errors may occur frequently. Direct use in production environments with strict stability requirements is not recommended. The Turing Platform provides Retry and Fallback as engineering mitigations, but these cannot fundamentally resolve the shared-quota issue, and switching models via Fallback may result in inconsistent outputs. To resolve this at the root, additional Provisioned Throughput must be purchased. See: Gemini 429 Rate Limiting and Provisioned Throughput
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Vertex AI | turing/gemini-3.8-flash | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Input: Output: Cache read: Web search: $14 / 1KIntroductory pricing is valid through 2026-12-31 (provided as 50% credits back on net spend); from 2027-01-01, standard pricing is input $1.50 / output $7.50 / cache read $0.15 → | Vertex AI shared quota Launched: 2026-09-03 | China Europe North America Asia Pacific | Latest Flash model, designed for long-horizon software engineering, autonomous agents, and complex enterprise workflows; supports PDF input, thinking mode, and built-in web search |
| Vertex AI | turing/gemini-3.7-flash | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Input: Output: Cache read: Web search: $14 / 1KIntroductory pricing is valid through 2026-12-31 (provided as 50% credits back on net spend); from 2027-01-01, standard pricing is input $1.50 / output $7.50 / cache read $0.15 → | Vertex AI shared quota Launched: 2026-08-17 | China Europe North America Asia Pacific | Designed for general agentic workflows, multi-step orchestration, and coding tasks; supports PDF input, thinking mode, and built-in web search |
| Vertex AI | turing/gemini-3.6-flash | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2026-07-23 | China Europe North America Asia Pacific | Optimized for multi-step orchestration, full-stack code refactoring, agent execution, and spatial reasoning; supports PDF input, thinking mode, and built-in web search | |
| Vertex AI | turing/gemini-3.5-flash | max_input_tokens: 1,048,576max_output_tokens: 65,535Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2026-05-19 | China Europe North America Asia Pacific | Supports thinking mode & built-in web search | |
| Vertex AI | turing/gemini-3.1-pro-latest | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Cache: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Input: $2Output: $12Cache read: $0.4Web search: $14 / 1KTiered pricing: above 200K input, input: $4 / output: $18 → | Vertex AI shared quota Launched: 2026-02-19 | China Europe North America Asia Pacific | Latest version; supports thinking mode (including MEDIUM level) & built-in web search |
| Vertex AI | turing/gemini-3.1-flash-lite-latest | max_input_tokens: 1,048,576max_output_tokens: 65,535Output: Tools: Supported Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-12-18 | China Europe North America Asia Pacific | Latest version; supports thinking mode & built-in web search | |
| Vertex AI | turing/gemini-3.1-flash-image | max_input_tokens: 131,072max_output_tokens: 32,768Output: Tools: Supported Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-12-18 | China Europe North America Asia Pacific | Latest version; supports thinking mode & image output | |
| Vertex AI | turing/gemini-3.1-flash-lite-image | max_input_tokens: 65,536max_output_tokens: 4,096Output: Tools: Not supported Reasoning: Supported Content moderation: Supported content moderation | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2026-07-01 | China Europe North America Asia Pacific | Nano Banana 2 Lite, lightweight image generation, output resolution up to 1K | |
| Vertex AI | turing/gemini-3-flash-latest | max_input_tokens: 1,048,576max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-12-18 Expected retirement: 2026-06-15 | China Europe North America Asia Pacific | Supports thinking mode & built-in web search | |
| Vertex AI | turing/gemini-3-pro-image | max_input_tokens: 65,000max_output_tokens: 32,000Output: Tools: - Reasoning: Supported Content moderation: Supported content moderation Built-in tools: 🔍 Web search | API: v1/chat/completions SDK: OpenAI SDK | Vertex AI shared quota Launched: 2025-11-19 | China Europe North America Asia Pacific | Latest version; supports thinking mode |
Claude
Due to Anthropic policy restrictions, Claude models may become unavailable at any time. They are recommended for personal projects and experimentation only.
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Anthropic | turing/claude-sonnet-5 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $2Output: $10Cache read: $0.2Cache write 5m: $2.5Cache write 1h: $4Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-07-01 | China Europe North America Asia Pacific | Latest Sonnet model, adaptive thinking on by default; does not support enabled/budget_tokens; supports tools, explicit caching, and web search |
| Anthropic | turing/claude-opus-4.8 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-06-01 | China Europe North America Asia Pacific | Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive) |
| Anthropic | turing/claude-opus-4.7 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-04-15 | China Europe North America Asia Pacific | Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive) |
| Anthropic | turing/claude-opus-4.6 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $5Output: $25Cache read: $0.5Cache write 5m: $6.25Cache write 1h: $10Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-02-25 | China Europe North America Asia Pacific | Supports tools; adaptive thinking mode recommended |
| Anthropic | turing/claude-sonnet-4.6 | max_input_tokens: 1,000,000max_output_tokens: 128,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $3Output: $15Cache read: $0.3Cache write 5m: $3.75Cache write 1h: $6Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-02-25 | China Europe North America Asia Pacific | Latest version; supports tools |
| Anthropic | turing/claude-haiku-4.5 | max_input_tokens: 200,000max_output_tokens: 64,000Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: $1Output: $5Cache read: $0.1Cache write 5m: $1.25Cache write 1h: $2Web search: $10 / 1K | Vertex AI (global deployment) Launched: 2026-01-19 | China Europe North America Asia Pacific | Supports tools |
ByteDance Volcano
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| ByteDance Volcano | doubao-seed-2.1-pro | max_input_tokens: 262,144max_output_tokens: 262,144Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-06-24 | China | Seed 2.1 flagship edition, supports video understanding | |
| ByteDance Volcano | doubao-seed-2.1-turbo | max_input_tokens: 262,144max_output_tokens: 262,144Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-06-24 | China | Seed 2.1 balanced edition, supports video understanding + GUI agent | |
| ByteDance Volcano | doubao-seed-character | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-06-28 | China | Role-play/virtual companionship, supports image-text understanding, deep thinking, and tool calling | ||
| ByteDance Volcano | doubao-seed-2-0-pro-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Latest version Seed 2.0 flagship edition, supports video understanding | ||
| ByteDance Volcano | doubao-seed-2-0-lite-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Latest version Seed 2.0 mid-range edition, supports video understanding | ||
| ByteDance Volcano | doubao-seed-2-0-lite-260428 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-04-28 | China | Seed 2.0 omni-modal edition, supports unified understanding of text/image/video/audio + deep thinking | ||
| ByteDance Volcano | doubao-seed-2-0-mini-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Latest version Seed 2.0 lightweight edition, supports video understanding | ||
| ByteDance Volcano | doubao-seed-2-0-code-preview-260215 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK | ByteDance Volcano Launched: 2026-02-25 | China | Code preview edition |
Alibaba DashScope
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Alibaba DashScope | qwen3.8-max | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-08-03 | China | Flagship edition 3.8, vision-language (image/video input) | ||
| Alibaba DashScope | qwen3.8-flash | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-08-31 | China | Lightweight edition 3.8, vision-language (image/video input) | ||
| Alibaba DashScope | qwen3.8-27b | max_input_tokens: 991,808max_output_tokens: 131,072Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-09-10 | China | 3.8 27B edition, vision-language (image/video input) | |
| Alibaba DashScope | qwen3.7-plus | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-06-08 | China | Cost-effective edition 3.7, vision-language (image input) | ||
| Alibaba DashScope | qwen3.7-flash | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-07-28 | China | Lightweight edition 3.7, vision-language (image input) | ||
| Alibaba DashScope | qwen3.7-max | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-05-28 | China | Flagship edition 3.7 | ||
| Alibaba DashScope | qwen3.6-plus | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-04-03 | China | Stable edition, currently identical in capability to qwen3.6-plus-2026-04-02 | |
| Alibaba DashScope | qwen3.6-flash | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-04-23 | China | Lightweight edition 3.6 | |
| Alibaba DashScope | qwen3.6-max | max_input_tokens: 229,376max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-04-23 Expected retirement: 2026-09-08 | China | Flagship edition 3.6 | |
| Alibaba DashScope | qwen3.5-plus | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-02-25 | China | Snapshot version | |
| Alibaba DashScope | qwen3.5-flash | max_input_tokens: 983,616max_output_tokens: 65,536Output: Tools: Supported Reasoning: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-02-25 | China | Snapshot version |
Deepseek
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens, billed = list price × discount) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Deepseek | deepseek-v4.1-flash | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥2Output: ¥8Cache read: ¥0.04Peak/off-peak pricing: peak hours are UTC+8 09:00-12:00 and 14:00-18:00, Monday to Friday only (by bill time); the whole weekend and every other weekday hour are off-peak at input ¥1 / output ¥4 / cache read ¥0.02 → | DeepSeek 官方 Launched: 2026-09-11 | China | New: V4.1 Flash general release, direct to DeepSeek's own API; natively supports image input, shares one rate-limit quota with deepseek-v4-flash-0731, and replaces the discontinued older V4 Flash | |
| ByteDance Volcano | bytedance/deepseek-v4.1-flash | max_input_tokens: 1,048,576max_output_tokens: 393,216Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | 60% off Input: ¥2Output: ¥8Cache read: ¥0.0460% off effective 2026-09-24 09:52:17 Beijing time (applies to both peak and off-peak rates); peak/off-peak pricing: peak hours are UTC+8 09:00-12:00 and 14:00-18:00, Monday to Friday only (by bill time); the whole weekend and every other weekday hour are off-peak at input ¥1 / output ¥4 / cache read ¥0.02 → | ByteDance Volcano Launched: 2026-09-17 | China | New: Volcano Ark deployment of DeepSeek V4.1 Flash (Ark snapshot deepseek-v4-1-flash-260910); natively supports image input, shares one rate-limit quota with deepseek-v4.1-flash but is priced separately; built-in web search is available only on v1/responses |
| Deepseek | deepseek-v4-pro-0813 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥9Output: ¥27Cache read: ¥0.9Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥4.5 / output ¥13.5 / cache read ¥0.45 | Alibaba Cloud Launched: 2026-08-13 | China | New: V4 Pro stable release (fixed version id), available alongside deepseek-v4-pro and sharing one rate-limit quota; the two are priced differently and this one bills on a peak/off-peak schedule | |
| Alibaba DashScope | aliyun/deepseek-v4-flash-0731 | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥3Output: ¥9Cache read: ¥0.3Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥1.5 / output ¥4.5 / cache read ¥0.15 → | Alibaba Cloud Launched: 2026-07-31 | China | The Aliyun Model Studio route for the old V4 Flash 0731 snapshot; the official direct route is discontinued. The cache-read rate is ¥0.3, and off-peak runs 22:00 to 08:00 UTC+8 daily | |
| Deepseek | deepseek-v4-pro | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: ¥12Output: ¥24Cache read: ¥1 | Alibaba Cloud Launched: 2026-04-25 | China | ||
| ByteDance Volcano | bytedance/deepseek-v4-flash | max_input_tokens: 1,048,576max_output_tokens: 393,216Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | ByteDance Volcano Launched: 2026-07-28 | China | Volcano deployment of DeepSeek V4 Flash; built-in web search is available only on v1/responses | |
| Deepseek | turing/deepseek-v3-2 | max_input_tokens: 64,000max_output_tokens: 8,192Input: Output: Tools: Not supported Reasoning: Supported | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: ¥2Output: ¥3[0, 32] input: ¥2 output: ¥3 / (32, 128] input: ¥4 output: ¥6 | Volcano Engine (China) Launched: 2025-12-11 | China | v1/responses supports streaming only |
Zhipu
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens, billed = list price × discount) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Zhipu | glm-5.3-flash | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-08-26 | China | Native multimodal model with 1M context; thinking is always enabled, with low, high, and max reasoning effort → | ||
| Zhipu | glm-5.3 | max_input_tokens: 1,048,576max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-08-19 | China | 开源旗舰,1M 上下文 | |
| Alibaba DashScope | dashscope/glm-5.3 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-09-20 | China | New: Aliyun Model Studio route for GLM-5.3 (Model Studio's own glm-5.3 listing); shares one rate-limit quota with glm-5.3 but is priced separately; thinking mode only, 1M context, implicit cache hits billed at 25% of the input price | ||
| Zhipu | glm-5.2 | max_input_tokens: 1,048,576max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-06-17 | China | Latest flagship, 1M context | |
| Alibaba DashScope | dashscope/glm-5.2 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-09-20 | China | New: Aliyun Model Studio route for GLM-5.2 (Model Studio's own glm-5.2 listing); shares one rate-limit quota with glm-5.2 but is priced separately; supports thinking and non-thinking modes, 1M context, implicit cache hits billed at 25% of the input price | ||
| Zhipu | glm-5.1 | max_input_tokens: 204,800max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-04-09 | China | ||
| Zhipu | glm-5 | max_input_tokens: 204,800max_output_tokens: 131,072Input: Output: Tools: Supported Reasoning: Supported Cache: Supported Built-in tools: 🔍 Web search | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2026-02-24 | China | ||
| Zhipu | glm-4.7 | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2025-12-23 | China | |||
| Zhipu | glm-4.6 | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Zhipu Launched: 2025-11-11 | China |
MiniMax
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| MiniMax | minimax-m3 | API: v1/chat/completions, v1/messages SDK: OpenAI SDK / Anthropic SDK | Input: Output: Cache read: Limited-time 50% off; context 512K-1M input ¥4.2 / output ¥16.8 / cache read ¥0.84 → | MiniMax Launched: 2026-06-17 | China | MiniMax M3; v1/responses supports streaming only | |
| MiniMax | minimax-m2.7 | max_input_tokens: 204,800max_output_tokens: 131,072Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-03-23 | China | Latest version M2.7 | |
| MiniMax | minimax-m2.5-highspeed | max_input_tokens: 204,800max_output_tokens: 204,800Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | MiniMax Launched: 2026-02-25 | China | ||
| MiniMax | minimax-m2.5 | max_input_tokens: 204,800max_output_tokens: 204,800Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK / Anthropic SDK | MiniMax Launched: 2026-02-25 | China | ||
| MiniMax | minimax-text-01 | Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | MiniMax Launched: 2025-04-02 | China |
KIMI
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| KIMI | kimi-k3 | API: v1/chat/completions SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-07-21 | China | Kimi flagship model, supports 1M context, image understanding, deep thinking, and tool calling; v1/responses supports streaming only |
Xiaomi MiMo
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Xiaomi | mimo-v2.5-pro | API: v1/chat/completions, v1/responses SDK: OpenAI SDK / Anthropic SDK | Alibaba Cloud Launched: 2026-05-28 | China | Xiaomi open-source reasoning model |
Bedrock
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Bedrock | turing/nova-lite-v2 | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.30Output: $2.50Cache read: $0.075 | AWS Bedrock (US) Launched: 2026-07-22 | China Europe North America Asia Pacific | Nova 2 Lite | |
| Bedrock | turing/nova-pro | max_input_tokens: 300,000max_output_tokens: 10,000Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.80Output: $3.20 | AWS Bedrock (US) Launched: 2026-07-22 | China Europe North America Asia Pacific | Nova Pro |
| Bedrock | nova-lite-v1 | max_input_tokens: 290,000max_output_tokens: 10,000Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.06Output: $0.24 | AWS Bedrock (US) Launched: 2025-08-21 | China Europe North America Asia Pacific | |
| Bedrock | turing/titan-nove-lite-v1 | max_input_tokens: 290,000max_output_tokens: 10,000Input: Output: Tools: Supported | API: v1/chat/completions SDK: OpenAI SDK | Input: $0.06Output: $0.24 | AWS Bedrock (US) Launched: 2025-08-21 | China Europe North America Asia Pacific |
Grok
| Provider | Model ID | Capabilities | endpoint | Price (per million Tokens) | Deployment & lifecycle | Supported regions | Notes |
|---|---|---|---|---|---|---|---|
| Bedrock | turing/grok-4.6 | API: v1/chat/completions, v1/messages, v1/responses SDK: OpenAI SDK / Anthropic SDK | Input: $2Output: $6Cache read: $0.5≥200K 输入时整次请求按 2× 输入 / 2× 输出计价,缓存输入同样按 2× 计费 | AWS Bedrock (US) Launched: 2026-08-21 | China Europe North America Asia Pacific | Supports Grok Build; see the integration guide for configuration → |