Skip to main content

Turing Headers Protocol

In addition to standard HTTP headers, the Turing Platform defines a set of X-Turing-* / x-turing-* headers: response headers return the trace ID, actual routing result, and processing time for each call; request headers are used for proactive debugging, tagging, and trace chaining. This page is the authoritative reference for these headers — other pages (Request Tracing, Timeouts, Retries & Fallback) cover usage in their respective contexts, but definitions here take precedence.

Response Headers (Returned by the Platform)​

HeaderDescriptionScope
X-Turing-Trace-IdThe Turing Platform request trace ID, used for log lookup and billing inquiriesEvery response
X-Turing-Process-TimePlatform-side processing time in millisecondsEvery response
x-turing-model-idThe model ID that actually served the request (especially useful with fallback / auto routing)Non-streaming LLM responses
x-turing-retriesThe actual number of retries that occurred for the requestNon-streaming LLM responses
x-turing-fallbacksThe actual number of fallback switches that occurred for the requestNon-streaming LLM responses
X-Turing-Memory-*The Managed Memory read and capture outcome for the turn; six headers, see Managed Memory HeadersChat Completions responses with turing_options.memory enabled, including streaming

A typical non-streaming response:

curl -i $TURING_BASE_URL/chat/completions \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "turing/gpt-4.1",
"messages": [{"role": "user", "content": "Hello!"}]
}'

# Response Headers:
# x-turing-trace-id: xxxxxxxx-xxxx-xxxx-xxxx-xxxxxxxxxxxx
# x-turing-process-time: 812.34
# x-turing-model-id: gpt-4.1
# x-turing-retries: 0
# x-turing-fallbacks: 0

For the complete workflow of using X-Turing-Trace-Id to debug issues and look up billing records, see Request Tracing.


Response Observability Headers: Actual Model / Retries / Fallbacks​

When fallback / retry is configured, x-turing-model-id / x-turing-retries / x-turing-fallbacks let you confirm on the client side which model actually served the request and how many retries and fallbacks occurred:

HeaderDescription
x-turing-model-idThe model ID that actually served the request. When fallbacks (especially auto) is configured, use this to confirm which model was ultimately used.
x-turing-retriesThe actual number of retries that occurred for the request (corresponding to turing_options.max_retries)
x-turing-fallbacksThe actual number of fallback switches that occurred for the request
from openai import OpenAI

client = OpenAI()

response = client.chat.completions.with_raw_response.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Hello!"}],
turing_options={
"max_retries": 2,
"fallbacks": "auto",
},
)

print("served by:", response.headers.get("x-turing-model-id"))
print("retries: ", response.headers.get("x-turing-retries"))
print("fallbacks:", response.headers.get("x-turing-fallbacks"))

completion = response.parse()
info
  • Non-streaming only: Streaming responses (stream=True) are delivered as SSE chunks and do not carry these three headers.
  • Also present on error responses: When all fallbacks are exhausted and an error is returned, these three headers are still included — which is precisely when you need them most to determine how far the chain progressed.
  • Browser-accessible: All three headers are included in the CORS expose_headers list, so frontend code can read them directly via response.headers.get(...).

For configuration options (max_retries / fallbacks), see Timeouts, Retries & Fallback.


Managed Memory Headers​

When a Chat Completions request enables turing_options.memory, the response carries the headers below to show whether memory was used and captured for the turn. Response bodies and SSE events keep their format, so clients that ignore these headers are unaffected. To enable Managed Memory, see LTM integrations.

HeaderDescription
X-Turing-Memory-Read-Statusused: memory was used for this turn; empty: no relevant memory exists, which is normal; degraded: memory was not used, and the answer is still returned
X-Turing-Memory-Capture-Statusaccepted: the turn was written as an Event and memory is derived asynchronously; skipped: the turn is not written; failed: the write did not succeed; conflict: the same turn was already written with different content and was not overwritten; pending: a streaming response whose write happens after the output ends
X-Turing-Memory-Event-IdThe written Event ID, returned only with accepted
X-Turing-Memory-Turn-IdThe memory turn ID, returned when a memory read ran
X-Turing-Memory-Evidence-SourceThe memory evidence source, returned when a memory read ran; currently always computed
X-Turing-Memory-ReasonA reason code when memory was not used or the turn was not written, for example fallback_enabled (the request configures fallbacks, so memory is bypassed for the whole request), ltm_unavailable (the memory read or write could not complete), or caller_tool_call (the model returned a tool call, so the turn is not written). Other values are for troubleshooting; record them verbatim
info
  • Only three kinds of error fail the request: an invalid turing_options.memory (code 1004), a credential that may not use this client's memory (code 1308), and a Space that does not exist or does not belong to the current client and environment (code 1751). Other memory failures do not interrupt the answer and only appear in the headers above; monitor degraded, failed, and X-Turing-Memory-Reason.
  • Streaming responses carry them too: the headers are sent before the output starts, before anything is written, so X-Turing-Memory-Capture-Status is pending (skipped when memory is bypassed).
  • Browser-accessible: all of these headers are in the CORS expose_headers list.

For what each error code means, see Error Code Reference.


Request Headers (Set by the Client)​

HeaderDescription
AuthorizationBearer <API_KEY> — required for all requests
X-Client-Request-IdA request ID generated by the client. Can still be used for debugging when a request fails or receives no response (see Request Tracing)
X-Turing-TagsRequest tags as a JSON string. Used for usage categorization and source identification
X-Turing-Trace-Id(Optional) Pass this proactively to reuse a trace ID and link multiple calls within the same business flow under a single ID. Only valid within a 10-minute freshness window; expired IDs are ignored and a new one is generated automatically
import json
from openai import OpenAI

client = OpenAI()

completion = client.chat.completions.create(
model="turing/gpt-4.1",
messages=[{"role": "user", "content": "Hello!"}],
extra_headers={
"X-Client-Request-Id": "req-20260715-abcdef",
"X-Turing-Tags": json.dumps({"team": "growth", "scene": "summarizer"}),
},
)

See also​