Skip to main content

Model Built-in Search

Qwen / Gemini / Claude use the standard /v1/chat/completions endpoint, with provider differences expressed in tools / extra_body; GPT-6 Astra and ByteDance Volcano (Ark) built-in search are only available on /v1/responses. For the full request body schema see /api/create-chat-completion; for the Claude native protocol see Messages API; for the Responses protocol see Responses API.

Supported Models​

The full list of supported models is determined by the provider sections in the model list and the Web Search / built-in tool markers:

OpenAI / Azure GPT-6 Astra​

GPT-6 Astra executes its native web_search through the Responses API. Put the tool declaration in tools; the model decides whether to search and returns citations with the answer. Chat Completions does not accept this native tool. Each search is billed at $10 / 1K requests; use the Azure bill as the source of truth.

from openai import OpenAI

client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)

response = client.responses.create(
model="turing/gpt-6-astra",
input="Check today's weather in Shanghai and provide source links.",
tools=[{"type": "web_search"}],
)

print(response.output_text)
curl $TURING_BASE_URL/responses \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "turing/gpt-6-astra",
"input": "Check today's weather in Shanghai and provide source links.",
"tools": [{"type": "web_search"}]
}'

Qwen​

from openai import OpenAI

client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)

response = client.chat.completions.create(
model="qwen-plus-latest",
messages=[
{"role": "user", "content": "What will the weather be like in Hangzhou tomorrow?"}
],
# enable_search is not a standard OpenAI parameter; pass it via extra_body in the Python SDK
extra_body={
"enable_search": True
}
)

print(response.choices[0].message.content)
curl $TURING_BASE_URL/chat/completions \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-plus-latest",
"messages": [
{"role": "user", "content": "What will the weather be like in Hangzhou tomorrow?"}
],
"enable_search": true
}'

Gemini​

The Gemini family configures Google Search via the tools parameter. Two search modes are available:

  • googleSearch: Standard Google Search
  • enterpriseWebSearch: Enterprise-grade search providing safer, more compliant results
from openai import OpenAI

client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)

response = client.chat.completions.create(
model="turing/gemini-3-pro-latest",
messages=[
{"role": "user", "content": "What is the weather like in Shanghai today?"}
],
extra_body={
"tools": [{"googleSearch": {}}]
}
)

print(response.choices[0].message.content)

To switch to enterprise mode, simply replace googleSearch with enterpriseWebSearch:

extra_body={
"tools": [{"enterpriseWebSearch": {}}]
}

Claude​

The Claude family accepts Anthropic's native web_search_20250305 tool type via the tools parameter. The model automatically decides when to search and integrates search results with citations into its response.

Official documentation: Web search tool (Anthropic)

Optional fields:

  • max_uses: Maximum number of search calls allowed per conversation turn
  • allowed_domains / blocked_domains: Domain allowlist / blocklist (mutually exclusive)
  • user_location: Geographic context to guide the model's retrieval
from openai import OpenAI

client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)

response = client.chat.completions.create(
model="turing/claude-sonnet-5",
max_tokens=4096,
messages=[
{"role": "user", "content": "What was the closing price of the Shanghai Composite Index today?"}
],
extra_body={
"tools": [
{
"type": "web_search_20250305",
"name": "web_search",
"max_uses": 3
}
]
}
)

print(response.choices[0].message.content)

Claude Response Structure​

The Claude response includes search call counts and citation metadata; the exact field locations depend on which API you use. Regardless of the API, usage.server_tool_use.web_search_requests always returns the number of searches actually triggered by the request, which you can use for billing reconciliation.

Via /v1/messages (Anthropic native API) — the content array contains the following blocks in order:

  1. server_tool_use — the search request issued by the model (includes input.query)
  2. web_search_tool_result — list of search results (title / url / encrypted_content / page_age)
  3. text — the final answer; citations are inserted as separate text blocks, each carrying a citations[] array pointing to the corresponding web_search_result_location

Example (excerpt):

{
"model": "claude-sonnet-5",
"content": [
{
"type": "server_tool_use",
"id": "srvtoolu_vrtx_01FZ...",
"name": "web_search",
"input": { "query": "Shanghai Composite Index closing price today" }
},
{
"type": "web_search_tool_result",
"tool_use_id": "srvtoolu_vrtx_01FZ...",
"content": [
{
"type": "web_search_result",
"title": "Shanghai Composite (SSEC) Real-Time Quote…",
"url": "https://cn.investing.com/indices/shanghai-composite",
"page_age": "5 days ago"
}
]
},
{
"type": "text",
"text": "The Shanghai Composite Index latest quote is 4,051.43 points",
"citations": [
{
"type": "web_search_result_location",
"cited_text": "Shanghai Composite (SSEC) latest index quote is 4,051.43…",
"url": "https://cn.investing.com/indices/shanghai-composite",
"title": "Shanghai Composite (SSEC) Real-Time Quote…"
}
]
}
],
"usage": {
"input_tokens": 11505,
"output_tokens": 235,
"server_tool_use": {
"web_search_requests": 1
}
}
}

Via /chat/completions (OpenAI-compatible API) — the response shape is closer to OpenAI:

  • choices[0].message.content — the model's final text response
  • choices[0].message.tool_calls — each web_search call actually made by the model
  • choices[0].message.provider_specific_fields.citations — citation array containing cited_text / url / title / supported_text
  • usage.server_tool_use.web_search_requests — number of searches triggered in this request
tip

The encrypted_content / encrypted_index fields are opaque data used by Anthropic for multi-turn conversation signature verification. Pass them through to the next request as-is; do not modify them.

ByteDance Volcano Ark​

ByteDance Volcano's web search is executed server-side by Ark and is enabled via tools: [{"type": "web_search"}]. The model decides when to search and writes citations back into the response.

Only available on /v1/responses

This is the key difference from the other three providers. Passing the same tools to /chat/completions will be rejected outright by Ark:

{"error": {"message": "VolcengineException - The request failed because it is missing `tools.function` parameter", "code": "400"}}

tools on /chat/completions only accepts type: function. Use /v1/responses for web search.

Official documentation: Web Content Plugin (Volcano Ark) ↗

Supported models are indicated by the built-in tool markers in Model List → ByteDance Volcano.

from openai import OpenAI

client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)

response = client.responses.create(
model="bytedance/deepseek-v4-flash",
input="Use web search to check today's weather in Shenzhen and provide source links",
tools=[{"type": "web_search"}],
)

print(response.output_text)
curl $TURING_BASE_URL/responses \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bytedance/deepseek-v4-flash",
"input": "Use web search to check today'\''s weather in Shenzhen and provide source links",
"tools": [{"type": "web_search"}]
}'

Ark Response Structure​

In addition to reasoning and message, the output array includes web_search_call items that record each retrieval actually initiated by the model:

{
"output": [
{ "type": "reasoning", "summary": [], "status": "completed" },
{
"type": "web_search_call",
"id": "ws_0217858...",
"status": "completed",
"action": {
"type": "search",
"query": "Shenzhen weather August 4, 2026; Shenzhen today's weather"
}
},
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "..." }]
}
]
}

Citations are annotated as url_citation attached to output_text. In streaming mode (stream: true), three additional events are emitted — response.web_search_call.in_progress / .searching / .completed — and citations are delivered incrementally via response.output_text.annotation.added.

The usage field includes the actual number of searches performed. The field names differ from Claude / Gemini: tool_usage is the total count and tool_usage_details breaks it down by content source.

"usage": {
"input_tokens": 3996,
"output_tokens": 1050,
"tool_usage": { "web_search": 3 },
"tool_usage_details": { "web_search": { "search_engine": 3 } }
}

The schema is consistent between streaming and non-streaming; tool_usage is returned with the final response.completed event.