Long-term Memory / LTM
Turing LTM asynchronously derives Profile, Facts, and Summary memories from accepted conversation Events. Browser, Search, Context, and Reflection then expose those memories at different levels. The Long-term Memory workspace in Developer Platform is where you configure, test, and integrate project-scoped memory.
Think of LTM as an event → derivation → read pipeline. The application does not need to maintain its own vector store or summary jobs, but it still chooses which events are worth storing, which Space owns them, and when results enter the model context.
Every Event, read, and Managed Chat request must bind to one Space. Data, models, strategies, and Owner scope never cross the selected Space. A USER or CLIENT_ADMIN in the same project can use the workspace; destructive controls such as deleting a Space still require CLIENT_ADMIN.
Choose an Integration Surface
| Scenario | Integration surface | Data boundary |
|---|---|---|
| An application or project needs Actor or shared memory | The project LTM workspace in Developer Platform; use the Public LTM API or Managed Chat | Current project, Environment, and Space |
| Local agents such as Codex, Claude Code, and OpenClaw share personal memory | Agent Memory + Turing CLI Memory | One independent Personal Space per Portal user |
Agent Memory is an experimental personal integration, not a Host Adapter for project LTM. Both surfaces reuse the same underlying Public LTM capabilities, but they read and write the same memory only when explicitly bound to the same space_id.
Complete the First Loop
- In Developer Platform, open Long-term Memory for the target project and create or select a Space.
- Configure Embedding, the default LLM, Profile / Facts / Summary strategies, and optionally Reflection.
- Submit an Actor or Space Event in Event Playground.
Acceptedonly means the Event entered the asynchronous processing queue; it does not mean derived memory is already readable. - In Memory Explorer, use Browser to inspect records, Search to verify semantic recall, Context to verify bounded model-ready context, and Reflections to read or generate reports when needed.
- Based on who controls the application loop, either call Public LTM directly or let Chat Completions manage Recall, Evidence injection, and terminal Capture.
Run the API Loop in 5 Minutes
The sequence below mirrors Spaces → Event Playground → Memory Explorer in Developer Platform. Replace the environment variables with the current project's base URL and API key; TURING_BASE_URL must include /api/v1.
1. Discover models and create a Space
Read the LLM families published in the current environment, then put the stable returned family alias into model_profile.default. The embedding model is immutable after creation.
export TURING_BASE_URL="https://live-turing.cn.llm.tcljd.com/api/v1"
export TURING_API_KEY="<your-api-key>"
curl "$TURING_BASE_URL/ltm/families" \
-H "Authorization: Bearer $TURING_API_KEY"
curl "$TURING_BASE_URL/ltm/spaces" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "customer-support",
"description": "Support conversations for the customer-facing assistant",
"embedding_model": "<embedding-model-from-your-environment>",
"model_profile": {"default": "<llm-family-from-ltm-families>"}
}'
The create response's data.space_id is the unique boundary for every later request. Do not hard-code a Space ID outside your client configuration or reuse it across projects or environments.
2. Write one Event
Event is the only public V1 write entry point. Send only newly appended messages; never resend the cumulative transcript. Reusing session_id identifies one conversation.
export SPACE_ID="<space-id>"
export ACTOR_ID="user-123"
curl "$TURING_BASE_URL/ltm/spaces/$SPACE_ID/events" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: support-session-20260912-01" \
-d '{
"owner_type": "actor",
"actor_id": "user-123",
"session_id": "support-session-20260912",
"messages": [
{"role": "user", "content": "I prefer Chinese, and I want a project summary every Monday."},
{"role": "assistant", "content": "Got it. I will remember that."}
],
"metadata": {"channel": "web"}
}'
A successful response is HTTP 202 with data.status = "accepted". This only means that the Event entered the asynchronous derivation queue. Do not treat accepted as “memory is ready”; verify it later with Browser or Context/Search.
3. Read and inject memory
Choose the read primitive by the job you need to do:
| Problem to solve | API | Best time to call | Calls an LLM |
|---|---|---|---|
| Inspect exactly which records were derived | List memories | Debugging, admin UI, pagination | No |
| Find memories relevant to the current question | Search memories | Before each answer when using semantic retrieval | No; it creates one query embedding |
| Assemble bounded, model-ready context | Get context | Session initialization or prompt assembly | No |
| Let the platform own Recall and Capture | Chat Completions with turing_options.memory | When you do not want to maintain the Agent loop | The main Chat request makes one main-model call |
# Session briefing: preload at the start of a new session
curl "$TURING_BASE_URL/ltm/spaces/$SPACE_ID/context?actor_id=$ACTOR_ID&include_shared=true&token_budget=4000" \
-H "Authorization: Bearer $TURING_API_KEY"
# Semantic retrieval: call before answering the current question
curl "$TURING_BASE_URL/ltm/spaces/$SPACE_ID/memories/search" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "What notification cadence does this user prefer?",
"actor_id": "user-123",
"include_shared": true,
"top_k": 5
}'
Every public LTM response uses Turing's code / message / data envelope. Context.data.rendered is Markdown reference material with explicit data boundaries. Pass it to the model as untrusted context, never as system or developer instructions.
4. Do not want to maintain the Agent loop? Use Managed Chat
curl "$TURING_BASE_URL/chat/completions" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "turing/gpt-5.4-mini",
"messages": [{"role": "user", "content": "Summarize this project's notification preferences"}],
"turing_options": {
"memory": {
"enabled": true,
"space_id": "<space-id>",
"owner_type": "actor",
"actor_id": "user-123",
"include_shared": true,
"session_id": "support-session-20260912"
}
}
}'
Managed Chat automatically Captures a qualifying terminal answer; it does not write during an intermediate tool call. If the model falls back, Memory is bypassed for the whole request, producing no LTM, embedding, or Capture cost.
Workspace Capabilities
| Workspace | Purpose | Key boundary |
|---|---|---|
| Spaces | Create, select, and configure Spaces, models, strategies, and Reflection | The Embedding model is immutable after creation; the first Actor Event permanently freezes the Profile Schema; deleting a Space is destructive and requires CLIENT_ADMIN |
| Event Playground | Submit Actor or Space Events with a session_id and message sequence, with an exact request preview | Optional idempotency key, timestamp, and scalar metadata; processing is asynchronous |
| Memory Explorer | Browser, Search, Context, and Reflections | Every read uses an explicit Owner scope; Reflection must first be enabled on the Space |
| Managed Chat | Build and execute a real Chat Completions request with a Memory Binding | Manages request-level memory only; it does not change the caller's chat protocol or return Evidence |
Space settings control extraction and retrieval. Manage derived memory records through the Event and Browser flows rather than editing them directly in Space settings.
Owner and Read Scope
- Actor-only reads: Browser, Search, and Context send
actor_idand either omitinclude_sharedor set it tofalseto read that Actor's Profile, Facts, and Summary. - Actor + shared reads: send both
actor_idandinclude_shared: trueto include shared Space Facts with the Actor's memory. - Shared-only reads: omit
actor_identirely and sendinclude_shared: true; shared scope supports Facts only. - Reflections: Actor Portrait and Activity Insights are Actor-scoped and do not support shared-only scope.
owner_type is not part of the Browser, Search, or Context read scope. Event and Managed Chat bindings explicitly use owner_type: "actor" or owner_type: "space"; Space Owner requests must omit actor_id.
Browser paginates through derived records. Search performs semantic retrieval. Context assembles bounded, model-ready context within a token budget. Reflections generates or reads Actor reports on demand.
Managed Chat Boundaries
When an application wants to keep calling /chat/completions without assembling the Prompt or maintaining Capture itself, add a Memory Binding to the request:
{
"turing_options": {
"memory": {
"enabled": true,
"space_id": "ltmspace_...",
"owner_type": "actor",
"actor_id": "user-123",
"include_shared": true,
"session_id": "session-2026-08-24"
}
}
}
- Terminal answer: one Recall, one main LLM call, then one Event Capture.
- Caller tool call: one Recall and one main LLM call; Capture is skipped until a terminal continuation.
- Any fallback topology: Memory is bypassed completely—0 LTM, 0 Embedding, and 0 Capture.
- Authentication: only an authenticated Turing Bearer JWT or Turing API Key is accepted. Other schemes are rejected before any LTM, Embedding, or main LLM cost.
Evidence is not returned to the caller. Developer Platform shows the public Chat response, Memory outcome headers, and Trace so you can diagnose Recall and Capture. If the application owns Prompt assembly or the Agent loop, use Context, Search, and Event directly instead of Managed Chat.
The core Public LTM endpoints and turing_options.memory are now included in this API documentation's OpenAPI reference; typed SDK schemas are still planned for a later version. Developer Platform's Managed Chat Builder continues to provide exact JSON, cURL, and real-request diagnostics; do not depend on unpublished SDK-generated types.