Model Built-in Search
Qwen / Gemini / Claude use the standard /v1/chat/completions endpoint, with provider differences expressed in tools / extra_body; GPT-6 Astra and ByteDance Volcano (Ark) built-in search are only available on /v1/responses. For the full request body schema see /api/create-chat-completion; for the Claude native protocol see Messages API; for the Responses protocol see Responses API.
Supported Models
The full list of supported models is determined by the provider sections in the model list and the Web Search / built-in tool markers:
OpenAI / Azure GPT-6 Astra
GPT-6 Astra executes its native web_search through the Responses API. Put the tool declaration in tools; the model decides whether to search and returns citations with the answer. Chat Completions does not accept this native tool. Each search is billed at $10 / 1K requests; use the Azure bill as the source of truth.
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.responses.create(
model="turing/gpt-6-astra",
input="Check today's weather in Shanghai and provide source links.",
tools=[{"type": "web_search"}],
)
print(response.output_text)
curl $TURING_BASE_URL/responses \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "turing/gpt-6-astra",
"input": "Check today's weather in Shanghai and provide source links.",
"tools": [{"type": "web_search"}]
}'
Qwen
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.chat.completions.create(
model="qwen-plus-latest",
messages=[
{"role": "user", "content": "What will the weather be like in Hangzhou tomorrow?"}
],
# enable_search is not a standard OpenAI parameter; pass it via extra_body in the Python SDK
extra_body={
"enable_search": True
}
)
print(response.choices[0].message.content)
curl $TURING_BASE_URL/chat/completions \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-plus-latest",
"messages": [
{"role": "user", "content": "What will the weather be like in Hangzhou tomorrow?"}
],
"enable_search": true
}'
Gemini
The Gemini family configures Google Search via the tools parameter. Two search modes are available:
- googleSearch: Standard Google Search
- enterpriseWebSearch: Enterprise-grade search providing safer, more compliant results
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.chat.completions.create(
model="turing/gemini-3-pro-latest",
messages=[
{"role": "user", "content": "What is the weather like in Shanghai today?"}
],
extra_body={
"tools": [{"googleSearch": {}}]
}
)
print(response.choices[0].message.content)
To switch to enterprise mode, simply replace googleSearch with enterpriseWebSearch:
extra_body={
"tools": [{"enterpriseWebSearch": {}}]
}
Claude
The Claude family accepts Anthropic's native web_search_20250305 tool type via the tools parameter. The model automatically decides when to search and integrates search results with citations into its response.
Official documentation: Web search tool (Anthropic)
Optional fields:
max_uses: Maximum number of search calls allowed per conversation turnallowed_domains/blocked_domains: Domain allowlist / blocklist (mutually exclusive)user_location: Geographic context to guide the model's retrieval
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.chat.completions.create(
model="turing/claude-sonnet-5",
max_tokens=4096,
messages=[
{"role": "user", "content": "What was the closing price of the Shanghai Composite Index today?"}
],
extra_body={
"tools": [
{
"type": "web_search_20250305",
"name": "web_search",
"max_uses": 3
}
]
}
)
print(response.choices[0].message.content)
Claude Response Structure
The Claude response includes search call counts and citation metadata; the exact field locations depend on which API you use. Regardless of the API, usage.server_tool_use.web_search_requests always returns the number of searches actually triggered by the request, which you can use for billing reconciliation.
Via /v1/messages (Anthropic native API) — the content array contains the following blocks in order:
server_tool_use— the search request issued by the model (includesinput.query)web_search_tool_result— list of search results (title/url/encrypted_content/page_age)text— the final answer; citations are inserted as separate text blocks, each carrying acitations[]array pointing to the correspondingweb_search_result_location
Example (excerpt):
{
"model": "claude-sonnet-5",
"content": [
{
"type": "server_tool_use",
"id": "srvtoolu_vrtx_01FZ...",
"name": "web_search",
"input": { "query": "Shanghai Composite Index closing price today" }
},
{
"type": "web_search_tool_result",
"tool_use_id": "srvtoolu_vrtx_01FZ...",
"content": [
{
"type": "web_search_result",
"title": "Shanghai Composite (SSEC) Real-Time Quote…",
"url": "https://cn.investing.com/indices/shanghai-composite",
"page_age": "5 days ago"
}
]
},
{
"type": "text",
"text": "The Shanghai Composite Index latest quote is 4,051.43 points",
"citations": [
{
"type": "web_search_result_location",
"cited_text": "Shanghai Composite (SSEC) latest index quote is 4,051.43…",
"url": "https://cn.investing.com/indices/shanghai-composite",
"title": "Shanghai Composite (SSEC) Real-Time Quote…"
}
]
}
],
"usage": {
"input_tokens": 11505,
"output_tokens": 235,
"server_tool_use": {
"web_search_requests": 1
}
}
}
Via /chat/completions (OpenAI-compatible API) — the response shape is closer to OpenAI:
choices[0].message.content— the model's final text responsechoices[0].message.tool_calls— eachweb_searchcall actually made by the modelchoices[0].message.provider_specific_fields.citations— citation array containingcited_text/url/title/supported_textusage.server_tool_use.web_search_requests— number of searches triggered in this request
The encrypted_content / encrypted_index fields are opaque data used by Anthropic for multi-turn conversation signature verification. Pass them through to the next request as-is; do not modify them.
ByteDance Volcano Ark
ByteDance Volcano's web search is executed server-side by Ark and is enabled via tools: [{"type": "web_search"}]. The model decides when to search and writes citations back into the response.
/v1/responsesThis is the key difference from the other three providers. Passing the same tools to /chat/completions will be rejected outright by Ark:
{"error": {"message": "VolcengineException - The request failed because it is missing `tools.function` parameter", "code": "400"}}
tools on /chat/completions only accepts type: function. Use /v1/responses for web search.
Official documentation: Web Content Plugin (Volcano Ark) ↗
Supported models are indicated by the built-in tool markers in Model List → ByteDance Volcano.
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.responses.create(
model="bytedance/deepseek-v4-flash",
input="Use web search to check today's weather in Shenzhen and provide source links",
tools=[{"type": "web_search"}],
)
print(response.output_text)
curl $TURING_BASE_URL/responses \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bytedance/deepseek-v4-flash",
"input": "Use web search to check today'\''s weather in Shenzhen and provide source links",
"tools": [{"type": "web_search"}]
}'
Ark Response Structure
In addition to reasoning and message, the output array includes web_search_call items that record each retrieval actually initiated by the model:
{
"output": [
{ "type": "reasoning", "summary": [], "status": "completed" },
{
"type": "web_search_call",
"id": "ws_0217858...",
"status": "completed",
"action": {
"type": "search",
"query": "Shenzhen weather August 4, 2026; Shenzhen today's weather"
}
},
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "..." }]
}
]
}
Citations are annotated as url_citation attached to output_text. In streaming mode (stream: true), three additional events are emitted — response.web_search_call.in_progress / .searching / .completed — and citations are delivered incrementally via response.output_text.annotation.added.
The usage field includes the actual number of searches performed. The field names differ from Claude / Gemini: tool_usage is the total count and tool_usage_details breaks it down by content source.
"usage": {
"input_tokens": 3996,
"output_tokens": 1050,
"tool_usage": { "web_search": 3 },
"tool_usage_details": { "web_search": { "search_engine": 3 } }
}
The schema is consistent between streaming and non-streaming; tool_usage is returned with the final response.completed event.