Skip to main content

Chat Model List

OpenAI​

New Models: turing/gpt-6-sol / turing/gpt-6-luna

GPT-6 Sol balances intelligence and speed for reasoning and coding; GPT-6 Luna targets cost-sensitive, high-volume workloads. Both share GPT-6 Astra's request contract: v1/responses is recommended; v1/chat/completions has limited support for text requests, and tool calling requires the Responses API; reasoning effort supports low / medium / high / xhigh / max, but not none or minimal. Short-context pricing: Sol $2 input / $10 output / $0.2 cached input, Luna $0.1 input / $0.5 output / $0.01 cached input; requests over 272K input tokens use long-context rates for the full request.

Filters28/36
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Azureturing/gpt-6-astra
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions (Limited support), v1/responses
SDK: OpenAI SDK
Input: $10
Output: $50
Cache read: $1
Cache write 5m: $12.5
Web search: $10 / 1K
For requests with >272K input tokens, the entire request uses long-context rates (input $20 / cache read $2 / cache write $25 / output $75); cache writes are billed at 1.25× the input rate. Azure's pricing page is still being updated; use the Azure bill as the source of truth →
Azure Global
Launched: 2026-09-03
China
Europe
North America
Asia Pacific
GPT-6 Astra for complex reasoning, coding, research, and document creation; v1/chat/completions has limited support (text only); reasoning effort supports low/medium/high/xhigh/max; use the Responses API for tool calling
Azureturing/gpt-6-sol
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions (Limited support), v1/responses
SDK: OpenAI SDK
Input: $2
Output: $10
Cache read: $0.2
Cache write 5m: $2.5
For requests with >272K input tokens, the entire request uses long-context rates (input $4 / cache read $0.4 / cache write $5 / output $15); cache writes are billed at 1.25× the input price →
Azure Global
Launched: 2026-09-23
China
Europe
North America
Asia Pacific
GPT-6 Sol balances intelligence and speed for reasoning and coding; v1/chat/completions has limited support (text only); reasoning effort supports low/medium/high/xhigh/max; use the Responses API for tool calling
Azureturing/gpt-6-luna
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions (Limited support), v1/responses
SDK: OpenAI SDK
Input: $0.1
Output: $0.5
Cache read: $0.01
Cache write 5m: $0.125
For requests with >272K input tokens, the entire request uses long-context rates (input $0.2 / cache read $0.02 / cache write $0.25 / output $0.75); cache writes are billed at 1.25× the input price →
Azure Global
Launched: 2026-09-23
China
Europe
North America
Asia Pacific
GPT-6 Luna is the cost-efficient model for cost-sensitive, high-volume workloads; v1/chat/completions has limited support (text only); reasoning effort supports low/medium/high/xhigh/max; use the Responses API for tool calling
Azureturing/gpt-5.6-sol
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $5 $4
Output: $30 $20
Cache read: $0.5
The short-context promotional price takes effect 2026-09-01 and runs through at least 2026-11-30; it reduces input and output only -- cache read stays $0.5 and cache write stays $6.25. At >272K input, the entire request is billed at long-context pricing: input $10 / cache read $1 / cache write $12.5 / output $45 →
Azure Global
Launched: 2026-07-10
China
Europe
North America
Asia Pacific
GPT-5.6 flagship model, suited for complex reasoning tasks; supports max reasoning effort
Azureturing/gpt-5.6-terra
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2.5 $2
Output: $15 $12
Cache read: $0.25 $0.2
At >272K input, the entire request is billed at long-context pricing: input $5 / cache read $0.5 / output $22.5; cache writes are billed at 1.25× the input price →
Azure Global
Launched: 2026-07-10
China
Europe
North America
Asia Pacific
GPT-5.6 balanced model, weighing capability against cost; supports max reasoning effort
Azureturing/gpt-5.6-luna
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1 $0.2
Output: $6 $1.2
Cache read: $0.1 $0.02
At >272K input, the entire request is billed at long-context pricing: input $2 / cache read $0.2 / output $9; cache writes are billed at 1.25× the input price →
Azure Global
Launched: 2026-07-10
China
Europe
North America
Asia Pacific
GPT-5.6's fastest and lowest-cost model, suited for high-throughput tasks; supports max reasoning effort
Azureturing/gpt-5.5
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $30
Cache read: $0.5
Azure Global
China
Europe
North America
Asia Pacific
Next-generation flagship model; Response API recommended for the thinking feature
Azureturing/gpt-5.4
max_input_tokens: 922,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2.5
Output: $15
Cache read: $0.25
Azure Global
China
Europe
North America
Asia Pacific
Upgraded version of the previous generation; Response API recommended for the thinking feature
Azureturing/gpt-5.4-mini
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.75
Output: $4.5
Cache read: $0.07
Azure Global
China
Europe
North America
Asia Pacific
Cost-effective mini model; Response API recommended for the thinking feature
Azureturing/gpt-5.4-nano
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.2
Output: $1.25
Cache read: $0.02
Azure Global
China
Europe
North America
Asia Pacific
Ultra-low-cost nano model, suited for classification and sub-agent tasks
Azureturing/gpt-5.3-codex
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/responses
SDK: OpenAI SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2026-02-26
China
Europe
North America
Asia Pacific
Supports Response API only; supports reasoning tokens
Azureturing/gpt-5.3-chat
max_input_tokens: 128,000
max_output_tokens: 16,384
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2026-03-16
China
Europe
North America
Asia Pacific
Azureturing/gpt-5.2
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2025-12-11
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.2-chat
max_input_tokens: 1,047,576
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2025-12-11
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.2-codex
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/responses
SDK: OpenAI SDK
Input: $1.75
Output: $14
Cache read: $0.17
Azure Global
Launched: 2025-12-11
China
Europe
North America
Asia Pacific
Supports Response API only
Azureturing/gpt-5.1
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-11-13
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.1-chat
max_input_tokens: 1,047,576
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-11-13
China
Europe
North America
Asia Pacific
Upgraded version; Response API recommended for the thinking feature
Azureturing/gpt-5.1-codex
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5.1-codex-mini
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
Input: $0.25
Output: $2
Cache read: $0.03
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5-mini
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.25
Output: $2
Cache read: $0.03
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5-nano
max_input_tokens: 400,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.05
Output: $0.4
Cache read: $0.01
Azure Global
Launched: 2025-08-10
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/gpt-5-chat
max_input_tokens: 1,047,576
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.25
Output: $10
Cache read: $0.12
Azure Global
Launched: 2025-08-10
Expected retirement: 2026-04-15
China
Europe
North America
Asia Pacific
Response API recommended for the thinking feature
Azureturing/o3
max_input_tokens: 200,000
max_output_tokens: 100,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2
Output: $8
Cache read: $0.5
Azure Global
Launched: 2025-04-17
China
Europe
North America
Asia Pacific
Thinking model; Response API recommended
Azureturing/gpt-4.1
max_input_tokens: 1,024,000
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2
Output: $8
Cache read: $0.50
Azure Global
Launched: 2025-04-16
China
Europe
North America
Asia Pacific
Azureturing/o4-mini
max_input_tokens: 200,000
max_output_tokens: 100,000
Input:
Output:
Tools: Supported
Reasoning: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $1.1
Output: $4.4
Cache read: $0.275
Azure Global
Launched: 2025-04-16
China
Europe
North America
Asia Pacific
Thinking model; Response API recommended
Azureturing/gpt-4.1-mini
max_input_tokens: 1,024,000
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.4
Output: $1.6
Cache read: $0.10
Azure Global
Launched: 2025-04-14
China
Europe
North America
Asia Pacific
Azureturing/gpt-4.1-nano
max_input_tokens: 1,024,000
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $0.1
Output: $0.4
Cache read: $0.03
Azure Global
Launched: 2025-04-14
China
Europe
North America
Asia Pacific

Gemini​

Gemini Model Load Notice

Gemini models run on Google Cloud Vertex AI under a shared quota, so throughput cannot be guaranteed. Under high-concurrency workloads, 429 rate-limit errors may occur frequently. Direct use in production environments with strict stability requirements is not recommended. The Turing Platform provides Retry and Fallback as engineering mitigations, but these cannot fundamentally resolve the shared-quota issue, and switching models via Fallback may result in inconsistent outputs. To resolve this at the root, additional Provisioned Throughput must be purchased. See: Gemini 429 Rate Limiting and Provisioned Throughput

Filters10/23
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Vertex AIturing/gemini-3.8-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.50 $0.75
Output: $7.50 $3.75
Cache read: $0.15 $0.075
Web search: $14 / 1K
Introductory pricing is valid through 2026-12-31 (provided as 50% credits back on net spend); from 2027-01-01, standard pricing is input $1.50 / output $7.50 / cache read $0.15 →
Vertex AI shared quota
Launched: 2026-09-03
China
Europe
North America
Asia Pacific
Latest Flash model, designed for long-horizon software engineering, autonomous agents, and complex enterprise workflows; supports PDF input, thinking mode, and built-in web search
Vertex AIturing/gemini-3.7-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.50 $0.75
Output: $7.50 $3.75
Cache read: $0.15 $0.075
Web search: $14 / 1K
Introductory pricing is valid through 2026-12-31 (provided as 50% credits back on net spend); from 2027-01-01, standard pricing is input $1.50 / output $7.50 / cache read $0.15 →
Vertex AI shared quota
Launched: 2026-08-17
China
Europe
North America
Asia Pacific
Designed for general agentic workflows, multi-step orchestration, and coding tasks; supports PDF input, thinking mode, and built-in web search
Vertex AIturing/gemini-3.6-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.5
Output: $7.5
Cache read: $0.15
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2026-07-23
China
Europe
North America
Asia Pacific
Optimized for multi-step orchestration, full-stack code refactoring, agent execution, and spatial reasoning; supports PDF input, thinking mode, and built-in web search
Vertex AIturing/gemini-3.5-flash
max_input_tokens: 1,048,576
max_output_tokens: 65,535
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $1.5
Output: $9
Cache read: $0.15
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2026-05-19
China
Europe
North America
Asia Pacific
Supports thinking mode & built-in web search
Vertex AIturing/gemini-3.1-pro-latest
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $2
Output: $12
Cache read: $0.4
Web search: $14 / 1K
Tiered pricing: above 200K input, input: $4 / output: $18 →
Vertex AI shared quota
Launched: 2026-02-19
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode (including MEDIUM level) & built-in web search
Vertex AIturing/gemini-3.1-flash-lite-latest
max_input_tokens: 1,048,576
max_output_tokens: 65,535
Input:
Output:
Tools: Supported
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.25
Output: $1.5
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2025-12-18
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode & built-in web search
Vertex AIturing/gemini-3.1-flash-image
max_input_tokens: 131,072
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.50
Output: $3(Text)/$60(Image)
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2025-12-18
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode & image output
Vertex AIturing/gemini-3.1-flash-lite-image
max_input_tokens: 65,536
max_output_tokens: 4,096
Input:
Output:
Tools: Not supported
Reasoning: Supported
Content moderation: Supported content moderation
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.25
Output: $1.5(Text)/$30(Image)
Vertex AI shared quota
Launched: 2026-07-01
China
Europe
North America
Asia Pacific
Nano Banana 2 Lite, lightweight image generation, output resolution up to 1K
Vertex AIturing/gemini-3-flash-latest
max_input_tokens: 1,048,576
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.5
Output: $3
Web search: $14 / 1K
Vertex AI shared quota
Launched: 2025-12-18
Expected retirement: 2026-06-15
China
Europe
North America
Asia Pacific
Supports thinking mode & built-in web search
Vertex AIturing/gemini-3-pro-image
max_input_tokens: 65,000
max_output_tokens: 32,000
Input:
Output:
Tools: -
Reasoning: Supported
Content moderation: Supported content moderation
Built-in tools: 🔍 Web search
API: v1/chat/completions
SDK: OpenAI SDK
Input: $2
Output: $12(Text)/$120(Image)
Web search: $14 / 1K
Tiered pricing →
Vertex AI shared quota
Launched: 2025-11-19
China
Europe
North America
Asia Pacific
Latest version; supports thinking mode

Claude​

caution

Due to Anthropic policy restrictions, Claude models may become unavailable at any time. They are recommended for personal projects and experimentation only.

Filters6/20
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Anthropicturing/claude-sonnet-5
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $2
Output: $10
Cache read: $0.2
Cache write 5m: $2.5
Cache write 1h: $4
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-07-01
China
Europe
North America
Asia Pacific
Latest Sonnet model, adaptive thinking on by default; does not support enabled/budget_tokens; supports tools, explicit caching, and web search
Anthropicturing/claude-opus-4.8
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-06-01
China
Europe
North America
Asia Pacific
Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive)
Anthropicturing/claude-opus-4.7
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-04-15
China
Europe
North America
Asia Pacific
Supports tools; supports adaptive thinking mode only (off by default, requires explicitly passing thinking.type=adaptive)
Anthropicturing/claude-opus-4.6
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $5
Output: $25
Cache read: $0.5
Cache write 5m: $6.25
Cache write 1h: $10
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-02-25
China
Europe
North America
Asia Pacific
Supports tools; adaptive thinking mode recommended
Anthropicturing/claude-sonnet-4.6
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $3
Output: $15
Cache read: $0.3
Cache write 5m: $3.75
Cache write 1h: $6
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-02-25
China
Europe
North America
Asia Pacific
Latest version; supports tools
Anthropicturing/claude-haiku-4.5
max_input_tokens: 200,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: $1
Output: $5
Cache read: $0.1
Cache write 5m: $1.25
Cache write 1h: $2
Web search: $10 / 1K
Vertex AI (global deployment)
Launched: 2026-01-19
China
Europe
North America
Asia Pacific
Supports tools

ByteDance Volcano​

Filters8/22
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
ByteDance Volcanodoubao-seed-2.1-pro
max_input_tokens: 262,144
max_output_tokens: 262,144
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥6
Output: ¥30
Cache read: ¥1.2
ByteDance Volcano
Launched: 2026-06-24
ChinaSeed 2.1 flagship edition, supports video understanding
ByteDance Volcanodoubao-seed-2.1-turbo
max_input_tokens: 262,144
max_output_tokens: 262,144
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥3
Output: ¥15
Cache read: ¥0.6
ByteDance Volcano
Launched: 2026-06-24
ChinaSeed 2.1 balanced edition, supports video understanding + GUI agent
ByteDance Volcanodoubao-seed-character
max_input_tokens: 131,072
max_output_tokens: 32,768
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥0.8
Output: ¥2
Cache read: ¥0.16
ByteDance Volcano
Launched: 2026-06-28
ChinaRole-play/virtual companionship, supports image-text understanding, deep thinking, and tool calling
ByteDance Volcanodoubao-seed-2-0-pro-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaLatest version Seed 2.0 flagship edition, supports video understanding
ByteDance Volcanodoubao-seed-2-0-lite-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaLatest version Seed 2.0 mid-range edition, supports video understanding
ByteDance Volcanodoubao-seed-2-0-lite-260428
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-04-28
ChinaSeed 2.0 omni-modal edition, supports unified understanding of text/image/video/audio + deep thinking
ByteDance Volcanodoubao-seed-2-0-mini-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaLatest version Seed 2.0 lightweight edition, supports video understanding
ByteDance Volcanodoubao-seed-2-0-code-preview-260215
max_input_tokens: 262,144
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK
ByteDance Volcano
Launched: 2026-02-25
ChinaCode preview edition

Alibaba DashScope​

Filters11/43
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Alibaba DashScopeqwen3.8-max
max_input_tokens: 991,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥12
Output: ¥36
Cache read: ¥1.5
Alibaba Cloud
Launched: 2026-08-03
ChinaFlagship edition 3.8, vision-language (image/video input)
Alibaba DashScopeqwen3.8-flash
max_input_tokens: 991,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥0.8
Output: ¥2.7
Cache read: ¥0.1
Alibaba Cloud
Launched: 2026-08-31
ChinaLightweight edition 3.8, vision-language (image/video input)
Alibaba DashScopeqwen3.8-27b
max_input_tokens: 991,808
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥3
Output: ¥12
Cache read: ¥0.6
Explicit cache creation ¥3.75 / explicit cache hit ¥0.3 →
Alibaba Cloud
Launched: 2026-09-10
China3.8 27B edition, vision-language (image/video input)
Alibaba DashScopeqwen3.7-plus
max_input_tokens: 991,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-06-08
ChinaCost-effective edition 3.7, vision-language (image input)
Alibaba DashScopeqwen3.7-flash
max_input_tokens: 991,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥0.2
Output: ¥0.8
Cache read: ¥0.04
Alibaba Cloud
Launched: 2026-07-28
ChinaLightweight edition 3.7, vision-language (image input)
Alibaba DashScopeqwen3.7-max
max_input_tokens: 991,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥6
Output: ¥18
Cache read: ¥1.2
Alibaba Cloud
Launched: 2026-05-28
ChinaFlagship edition 3.7
Alibaba DashScopeqwen3.6-plus
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-04-03
ChinaStable edition, currently identical in capability to qwen3.6-plus-2026-04-02
Alibaba DashScopeqwen3.6-flash
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-04-23
ChinaLightweight edition 3.6
Alibaba DashScopeqwen3.6-max
max_input_tokens: 229,376
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-04-23
Expected retirement: 2026-09-08
ChinaFlagship edition 3.6
Alibaba DashScopeqwen3.5-plus
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-02-25
ChinaSnapshot version
Alibaba DashScopeqwen3.5-flash
max_input_tokens: 983,616
max_output_tokens: 65,536
Input:
Output:
Tools: Supported
Reasoning: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Alibaba Cloud
Launched: 2026-02-25
ChinaSnapshot version

Deepseek​

Filters7/17
ProviderModel IDCapabilitiesendpointPrice (per million Tokens, billed = list price × discount)Deployment & lifecycleSupported regionsNotes
Deepseekdeepseek-v4.1-flash
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2
Output: ¥8
Cache read: ¥0.04
Peak/off-peak pricing: peak hours are UTC+8 09:00-12:00 and 14:00-18:00, Monday to Friday only (by bill time); the whole weekend and every other weekday hour are off-peak at input ¥1 / output ¥4 / cache read ¥0.02 →
DeepSeek 官方
Launched: 2026-09-11
ChinaNew: V4.1 Flash general release, direct to DeepSeek's own API; natively supports image input, shares one rate-limit quota with deepseek-v4-flash-0731, and replaces the discontinued older V4 Flash
ByteDance Volcanobytedance/deepseek-v4.1-flash
max_input_tokens: 1,048,576
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
60% off
Input: ¥2
Output: ¥8
Cache read: ¥0.04
60% off effective 2026-09-24 09:52:17 Beijing time (applies to both peak and off-peak rates); peak/off-peak pricing: peak hours are UTC+8 09:00-12:00 and 14:00-18:00, Monday to Friday only (by bill time); the whole weekend and every other weekday hour are off-peak at input ¥1 / output ¥4 / cache read ¥0.02 →
ByteDance Volcano
Launched: 2026-09-17
ChinaNew: Volcano Ark deployment of DeepSeek V4.1 Flash (Ark snapshot deepseek-v4-1-flash-260910); natively supports image input, shares one rate-limit quota with deepseek-v4.1-flash but is priced separately; built-in web search is available only on v1/responses
Deepseekdeepseek-v4-pro-0813
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥9
Output: ¥27
Cache read: ¥0.9
Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥4.5 / output ¥13.5 / cache read ¥0.45
Alibaba Cloud
Launched: 2026-08-13
ChinaNew: V4 Pro stable release (fixed version id), available alongside deepseek-v4-pro and sharing one rate-limit quota; the two are priced differently and this one bills on a peak/off-peak schedule
Alibaba DashScopealiyun/deepseek-v4-flash-0731
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥3
Output: ¥9
Cache read: ¥0.3
Peak/off-peak pricing: off-peak (UTC+8 22:00 to 08:00 next day, by bill time) input ¥1.5 / output ¥4.5 / cache read ¥0.15 →
Alibaba Cloud
Launched: 2026-07-31
ChinaThe Aliyun Model Studio route for the old V4 Flash 0731 snapshot; the official direct route is discontinued. The cache-read rate is ¥0.3, and off-peak runs 22:00 to 08:00 UTC+8 daily
Deepseekdeepseek-v4-pro
max_input_tokens: 1,000,000
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥12
Output: ¥24
Cache read: ¥1
Alibaba Cloud
Launched: 2026-04-25
China
ByteDance Volcanobytedance/deepseek-v4-flash
max_input_tokens: 1,048,576
max_output_tokens: 393,216
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥1
Output: ¥2
ByteDance Volcano
Launched: 2026-07-28
ChinaVolcano deployment of DeepSeek V4 Flash; built-in web search is available only on v1/responses
Deepseekturing/deepseek-v3-2
max_input_tokens: 64,000
max_output_tokens: 8,192
Input:
Output:
Tools: Not supported
Reasoning: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2
Output: ¥3
[0, 32] input: ¥2 output: ¥3 / (32, 128] input: ¥4 output: ¥6
Volcano Engine (China)
Launched: 2025-12-11
Chinav1/responses supports streaming only

Zhipu​

Filters9/20
ProviderModel IDCapabilitiesendpointPrice (per million Tokens, billed = list price × discount)Deployment & lifecycleSupported regionsNotes
Zhipuglm-5.3-flash
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥0.8
Output: ¥2.8
Cache read: ¥0.23
Zhipu
Launched: 2026-08-26
ChinaNative multimodal model with 1M context; thinking is always enabled, with low, high, and max reasoning effort →
Zhipuglm-5.3
max_input_tokens: 1,048,576
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥8
Output: ¥28
Cache read: ¥2
Zhipu
Launched: 2026-08-19
China开源旗舰,1M 上下文
Alibaba DashScopedashscope/glm-5.3
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
20% off
Input: ¥8
Output: ¥28
Cache read: ¥2
20% off, effective 2026-09-24 09:52:17 Beijing time →
Alibaba Cloud
Launched: 2026-09-20
ChinaNew: Aliyun Model Studio route for GLM-5.3 (Model Studio's own glm-5.3 listing); shares one rate-limit quota with glm-5.3 but is priced separately; thinking mode only, 1M context, implicit cache hits billed at 25% of the input price
Zhipuglm-5.2
max_input_tokens: 1,048,576
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥8
Output: ¥28
Cache read: ¥2
Zhipu
Launched: 2026-06-17
ChinaLatest flagship, 1M context
Alibaba DashScopedashscope/glm-5.2
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
50% off
Input: ¥8
Output: ¥28
Cache read: ¥2
50% off, effective 2026-09-24 09:52:17 Beijing time →
Alibaba Cloud
Launched: 2026-09-20
ChinaNew: Aliyun Model Studio route for GLM-5.2 (Model Studio's own glm-5.2 listing); shares one rate-limit quota with glm-5.2 but is priced separately; supports thinking and non-thinking modes, 1M context, implicit cache hits billed at 25% of the input price
Zhipuglm-5.1
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2026-04-09
China
Zhipuglm-5
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
Built-in tools: 🔍 Web search
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2026-02-24
China
Zhipuglm-4.7
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2025-12-23
China
Zhipuglm-4.6
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Zhipu
Launched: 2025-11-11
China

MiniMax​

Filters5/11
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
MiniMaxminimax-m3
max_input_tokens: 1,048,576
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages
SDK: OpenAI SDK / Anthropic SDK
Input: ¥4.2 ¥2.1
Output: ¥16.8 ¥8.4
Cache read: ¥0.84 ¥0.42
Limited-time 50% off; context 512K-1M input ¥4.2 / output ¥16.8 / cache read ¥0.84 →
MiniMax
Launched: 2026-06-17
ChinaMiniMax M3; v1/responses supports streaming only
MiniMaxminimax-m2.7
max_input_tokens: 204,800
max_output_tokens: 131,072
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2.1
Output: ¥8.4
Alibaba Cloud
Launched: 2026-03-23
ChinaLatest version M2.7
MiniMaxminimax-m2.5-highspeed
max_input_tokens: 204,800
max_output_tokens: 204,800
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: ¥4.2
Output: ¥16.8
MiniMax
Launched: 2026-02-25
China
MiniMaxminimax-m2.5
max_input_tokens: 204,800
max_output_tokens: 204,800
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK / Anthropic SDK
Input: ¥2.1
Output: ¥8.4
MiniMax
Launched: 2026-02-25
China
MiniMaxminimax-text-01
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: ¥1
Output: ¥8
MiniMax
Launched: 2025-04-02
China

KIMI​

Filters1/6
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
KIMIkimi-k3
max_input_tokens: 1,048,576
max_output_tokens: 1,048,576
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions
SDK: OpenAI SDK / Anthropic SDK
Input: ¥20
Output: ¥100
Cache read: ¥2
Alibaba Cloud
Launched: 2026-07-21
ChinaKimi flagship model, supports 1M context, image understanding, deep thinking, and tool calling; v1/responses supports streaming only

Xiaomi MiMo​

Filters1/1
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Xiaomimimo-v2.5-pro
max_input_tokens: 1,000,000
max_output_tokens: 128,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: ¥7
Output: ¥21
Cache read: ¥1.4
Alibaba Cloud
Launched: 2026-05-28
ChinaXiaomi open-source reasoning model

Bedrock​

Filters4/4
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Bedrockturing/nova-lite-v2
max_input_tokens: 1,000,000
max_output_tokens: 64,000
Input:
Output:
Tools: Supported
Cache: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.30
Output: $2.50
Cache read: $0.075
AWS Bedrock (US)
Launched: 2026-07-22
China
Europe
North America
Asia Pacific
Nova 2 Lite
Bedrockturing/nova-pro
max_input_tokens: 300,000
max_output_tokens: 10,000
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.80
Output: $3.20
AWS Bedrock (US)
Launched: 2026-07-22
China
Europe
North America
Asia Pacific
Nova Pro
Bedrocknova-lite-v1
max_input_tokens: 290,000
max_output_tokens: 10,000
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.06
Output: $0.24
AWS Bedrock (US)
Launched: 2025-08-21
China
Europe
North America
Asia Pacific
Bedrockturing/titan-nove-lite-v1
max_input_tokens: 290,000
max_output_tokens: 10,000
Input:
Output:
Tools: Supported
API: v1/chat/completions
SDK: OpenAI SDK
Input: $0.06
Output: $0.24
AWS Bedrock (US)
Launched: 2025-08-21
China
Europe
North America
Asia Pacific

Grok​

Filters1/5
ProviderModel IDCapabilitiesendpointPrice (per million Tokens)Deployment & lifecycleSupported regionsNotes
Bedrockturing/grok-4.6
max_input_tokens: 500,000
max_output_tokens: 500,000
Input:
Output:
Tools: Supported
Reasoning: Supported
Cache: Supported
API: v1/chat/completions, v1/messages, v1/responses
SDK: OpenAI SDK / Anthropic SDK
Input: $2
Output: $6
Cache read: $0.5
≥200K 输入时整次请求按 2× 输入 / 2× 输出计价,缓存输入同样按 2× 计费
AWS Bedrock (US)
Launched: 2026-08-21
China
Europe
North America
Asia Pacific
Supports Grok Build; see the integration guide for configuration →