模型内置搜索
Qwen / Gemini / Claude 走标准 /v1/chat/completions endpoint,参数差异体现在 tools / extra_body 里;GPT-6 Astra 和字节火山(Ark)的内置搜索只在 /v1/responses 上可用。完整请求体 schema 见 /api/create-chat-completion;Claude 原生协议见 Messages 接口;Responses 协议见 Response 接口。
支持的模型
完整支持范围以模型列表中的厂商段落和 Web Search / 内置工具标记为准:
OpenAI / Azure GPT-6 Astra
GPT-6 Astra 的原生 web_search 由 Responses API 执行。把工具声明放在 tools 中,模型会决定是否搜索并在回答中返回引用;Chat Completions 不接受该原生工具。每次搜索按 $10 / 1K 次请求计费,具体账单以 Azure 为准。
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.responses.create(
model="turing/gpt-6-astra",
input="查一下今天上海的天气,并给出信息来源链接",
tools=[{"type": "web_search"}],
)
print(response.output_text)
curl $TURING_BASE_URL/responses \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "turing/gpt-6-astra",
"input": "查一下今天上海的天气,并给出信息来源链接",
"tools": [{"type": "web_search"}]
}'
Qwen
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.chat.completions.create(
model="qwen-plus-latest",
messages=[
{"role": "user", "content": "杭州明天天气如何"}
],
# enable_search 非 OpenAI 标准参数,Python SDK 通过 extra_body 传入
extra_body={
"enable_search": True
}
)
print(response.choices[0].message.content)
curl $TURING_BASE_URL/chat/completions \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen-plus-latest",
"messages": [
{"role": "user", "content": "杭州明天天气如何"}
],
"enable_search": true
}'
Gemini
Gemini 系列用 tools 参数配置 Google Search,两种搜索模式:
- googleSearch:标准 Google 搜索
- enterpriseWebSearch:企业级搜索,提供更安全、合规的搜索结果
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.chat.completions.create(
model="turing/gemini-3-pro-latest",
messages=[
{"role": "user", "content": "上海今天天气如何"}
],
extra_body={
"tools": [{"googleSearch": {}}]
}
)
print(response.choices[0].message.content)
切换到企业级模式只需把 googleSearch 换成 enterpriseWebSearch:
extra_body={
"tools": [{"enterpriseWebSearch": {}}]
}
Claude
Claude 系列通过 tools 参数传入 Anthropic 原生的 web_search_20250305 工具类型,模型会自动判断何时检索,并把搜索结果与引用整合到回答中。
可选字段:
max_uses:单次对话最多调用搜索的次数上限allowed_domains/blocked_domains:域名白/黑名单(二选一)user_location:提示模型检索时结合的地理位置信息
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.chat.completions.create(
model="turing/claude-sonnet-5",
max_tokens=4096,
messages=[
{"role": "user", "content": "今日上证指数收盘价是多少?"}
],
extra_body={
"tools": [
{
"type": "web_search_20250305",
"name": "web_search",
"max_uses": 3
}
]
}
)
print(response.choices[0].message.content)
Claude 响应结构
Claude 的响应里会包含搜索调用次数和引用元数据,具体字段位置取决于所用接口。无论走哪条接口,usage.server_tool_use.web_search_requests 都会返回本次请求实际触发的搜索次数,用于计费对账。
通过 /v1/messages(Anthropic 原生接口) —— content 数组按顺序出现以下 block:
server_tool_use— 模型发起的搜索请求(含input.query)web_search_tool_result— 搜索结果列表(title/url/encrypted_content/page_age)text— 最终回答;引用以独立 text block 形式插入,block 上的citations[]数组指向对应的web_search_result_location
示例(节选):
{
"model": "claude-sonnet-5",
"content": [
{
"type": "server_tool_use",
"id": "srvtoolu_vrtx_01FZ...",
"name": "web_search",
"input": { "query": "上证指数今日收盘价" }
},
{
"type": "web_search_tool_result",
"tool_use_id": "srvtoolu_vrtx_01FZ...",
"content": [
{
"type": "web_search_result",
"title": "上证指数 (SSEC) 实时行情…",
"url": "https://cn.investing.com/indices/shanghai-composite",
"page_age": "5 days ago"
}
]
},
{
"type": "text",
"text": "上证指数最新报价为4,051.43点",
"citations": [
{
"type": "web_search_result_location",
"cited_text": "上证指数(SSEC)最新指数报价为4,051.43…",
"url": "https://cn.investing.com/indices/shanghai-composite",
"title": "上证指数 (SSEC) 实时行情…"
}
]
}
],
"usage": {
"input_tokens": 11505,
"output_tokens": 235,
"server_tool_use": {
"web_search_requests": 1
}
}
}
通过 /chat/completions(OpenAI 兼容接口) —— 响应形态贴近 OpenAI:
choices[0].message.content— 模型的最终文本回答choices[0].message.tool_calls— 模型实际发起的每一次web_search调用choices[0].message.provider_specific_fields.citations— 引用数组,含cited_text/url/title/supported_textusage.server_tool_use.web_search_requests— 本次触发的搜索次数
encrypted_content / encrypted_index 字段是 Anthropic 侧用于多轮对话签名校验的不透明数据,透传回下一轮请求即可,不要篡改。
字节火山 Ark
字节火山的联网搜索由 Ark 在服务端执行,通过 tools: [{"type": "web_search"}] 启用,模型自己决定检索时机并把引用写回回答。
/v1/responses 上可用这是和上面三家最大的差别。往 /chat/completions 传同样的 tools 会被 Ark 直接拒绝:
{"error": {"message": "VolcengineException - The request failed because it is missing `tools.function` parameter", "code": "400"}}
/chat/completions 上的 tools 只接受 type: function。需要联网就走 /v1/responses。
官方文档:联网内容插件(火山方舟)↗
适用模型以 模型列表 → 字节火山 里的内置工具标记为准。
from openai import OpenAI
client = OpenAI(
api_key="your-api-key",
base_url="https://live-turing.cn.llm.tcljd.com/api/v1"
)
response = client.responses.create(
model="bytedance/deepseek-v4-flash",
input="用联网搜索查一下今天深圳的天气,给出信息来源链接",
tools=[{"type": "web_search"}],
)
print(response.output_text)
curl $TURING_BASE_URL/responses \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "bytedance/deepseek-v4-flash",
"input": "用联网搜索查一下今天深圳的天气,给出信息来源链接",
"tools": [{"type": "web_search"}]
}'
Ark 响应结构
output 数组里除了 reasoning 和 message,还会多出 web_search_call 项,记录模型实际发起的每一次检索:
{
"output": [
{ "type": "reasoning", "summary": [], "status": "completed" },
{
"type": "web_search_call",
"id": "ws_0217858...",
"status": "completed",
"action": {
"type": "search",
"query": "深圳天气 2026年8月4日;深圳今日天气"
}
},
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "..." }]
}
]
}
引用以 url_citation 标注挂在 output_text 上。流式模式(stream: true)会额外推 response.web_search_call.in_progress / .searching / .completed 三个事件,引用则通过 response.output_text.annotation.added 增量下发。
usage 里会带上本次实际发起的检索次数,字段和 Claude / Gemini 都不一样:tool_usage 是总数,tool_usage_details 是按内容来源拆开的明细。
"usage": {
"input_tokens": 3996,
"output_tokens": 1050,
"tool_usage": { "web_search": 3 },
"tool_usage_details": { "web_search": { "search_engine": 3 } }
}
流式与非流式口径一致,tool_usage 随最终的 response.completed 事件返回。