Skip to main content

Long-term Memory / LTM

Turing LTM asynchronously derives Profile, Facts, and Summary memories from accepted conversation Events. Browser, Search, Context, and Reflection then expose those memories at different levels. The Long-term Memory workspace in Developer Platform is where you configure, test, and integrate project-scoped memory.

Think of LTM as an event → derivation → read pipeline. The application does not need to maintain its own vector store or summary jobs, but it still chooses which events are worth storing, which Space owns them, and when results enter the model context.

Space is a hard boundary

Every Event, read, and Managed Chat request must bind to one Space. Data, models, strategies, and Owner scope never cross the selected Space. A USER or CLIENT_ADMIN in the same project can use the workspace; destructive controls such as deleting a Space still require CLIENT_ADMIN.

Choose an Integration Surface​

ScenarioIntegration surfaceData boundary
An application or project needs Actor or shared memoryThe project LTM workspace in Developer Platform; use the Public LTM API or Managed ChatCurrent project, Environment, and Space
Local agents such as Codex, Claude Code, and OpenClaw share personal memoryAgent Memory + Turing CLI MemoryOne independent Personal Space per Portal user

Agent Memory is an experimental personal integration, not a Host Adapter for project LTM. Both surfaces reuse the same underlying Public LTM capabilities, but they read and write the same memory only when explicitly bound to the same space_id.

Complete the First Loop​

  1. In Developer Platform, open Long-term Memory for the target project and create or select a Space.
  2. Configure Embedding, the default LLM, Profile / Facts / Summary strategies, and optionally Reflection.
  3. Submit an Actor or Space Event in Event Playground. Accepted only means the Event entered the asynchronous processing queue; it does not mean derived memory is already readable.
  4. In Memory Explorer, use Browser to inspect records, Search to verify semantic recall, Context to verify bounded model-ready context, and Reflections to read or generate reports when needed.
  5. Based on who controls the application loop, either call Public LTM directly or let Chat Completions manage Recall, Evidence injection, and terminal Capture.

Run the API Loop in 5 Minutes​

The sequence below mirrors Spaces → Event Playground → Memory Explorer in Developer Platform. Replace the environment variables with the current project's base URL and API key; TURING_BASE_URL must include /api/v1.

1. Discover models and create a Space​

Read the LLM families published in the current environment, then put the stable returned family alias into model_profile.default. The embedding model is immutable after creation.

export TURING_BASE_URL="https://live-turing.cn.llm.tcljd.com/api/v1"
export TURING_API_KEY="<your-api-key>"

curl "$TURING_BASE_URL/ltm/families" \
-H "Authorization: Bearer $TURING_API_KEY"

curl "$TURING_BASE_URL/ltm/spaces" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "customer-support",
"description": "Support conversations for the customer-facing assistant",
"embedding_model": "<embedding-model-from-your-environment>",
"model_profile": {"default": "<llm-family-from-ltm-families>"}
}'

The create response's data.space_id is the unique boundary for every later request. Do not hard-code a Space ID outside your client configuration or reuse it across projects or environments.

2. Write one Event​

Event is the only public V1 write entry point. Send only newly appended messages; never resend the cumulative transcript. Reusing session_id identifies one conversation.

export SPACE_ID="<space-id>"
export ACTOR_ID="user-123"

curl "$TURING_BASE_URL/ltm/spaces/$SPACE_ID/events" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: support-session-20260912-01" \
-d '{
"owner_type": "actor",
"actor_id": "user-123",
"session_id": "support-session-20260912",
"messages": [
{"role": "user", "content": "I prefer Chinese, and I want a project summary every Monday."},
{"role": "assistant", "content": "Got it. I will remember that."}
],
"metadata": {"channel": "web"}
}'

A successful response is HTTP 202 with data.status = "accepted". This only means that the Event entered the asynchronous derivation queue. Do not treat accepted as “memory is ready”; verify it later with Browser or Context/Search.

3. Read and inject memory​

Choose the read primitive by the job you need to do:

Problem to solveAPIBest time to callCalls an LLM
Inspect exactly which records were derivedList memoriesDebugging, admin UI, paginationNo
Find memories relevant to the current questionSearch memoriesBefore each answer when using semantic retrievalNo; it creates one query embedding
Assemble bounded, model-ready contextGet contextSession initialization or prompt assemblyNo
Let the platform own Recall and CaptureChat Completions with turing_options.memoryWhen you do not want to maintain the Agent loopThe main Chat request makes one main-model call
# Session briefing: preload at the start of a new session
curl "$TURING_BASE_URL/ltm/spaces/$SPACE_ID/context?actor_id=$ACTOR_ID&include_shared=true&token_budget=4000" \
-H "Authorization: Bearer $TURING_API_KEY"

# Semantic retrieval: call before answering the current question
curl "$TURING_BASE_URL/ltm/spaces/$SPACE_ID/memories/search" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"query": "What notification cadence does this user prefer?",
"actor_id": "user-123",
"include_shared": true,
"top_k": 5
}'

Every public LTM response uses Turing's code / message / data envelope. Context.data.rendered is Markdown reference material with explicit data boundaries. Pass it to the model as untrusted context, never as system or developer instructions.

4. Do not want to maintain the Agent loop? Use Managed Chat​

curl "$TURING_BASE_URL/chat/completions" \
-H "Authorization: Bearer $TURING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "turing/gpt-5.4-mini",
"messages": [{"role": "user", "content": "Summarize this project's notification preferences"}],
"turing_options": {
"memory": {
"enabled": true,
"space_id": "<space-id>",
"owner_type": "actor",
"actor_id": "user-123",
"include_shared": true,
"session_id": "support-session-20260912"
}
}
}'

Managed Chat automatically Captures a qualifying terminal answer; it does not write during an intermediate tool call. If the model falls back, Memory is bypassed for the whole request, producing no LTM, embedding, or Capture cost.

Workspace Capabilities​

WorkspacePurposeKey boundary
SpacesCreate, select, and configure Spaces, models, strategies, and ReflectionThe Embedding model is immutable after creation; the first Actor Event permanently freezes the Profile Schema; deleting a Space is destructive and requires CLIENT_ADMIN
Event PlaygroundSubmit Actor or Space Events with a session_id and message sequence, with an exact request previewOptional idempotency key, timestamp, and scalar metadata; processing is asynchronous
Memory ExplorerBrowser, Search, Context, and ReflectionsEvery read uses an explicit Owner scope; Reflection must first be enabled on the Space
Managed ChatBuild and execute a real Chat Completions request with a Memory BindingManages request-level memory only; it does not change the caller's chat protocol or return Evidence

Space settings control extraction and retrieval. Manage derived memory records through the Event and Browser flows rather than editing them directly in Space settings.

Owner and Read Scope​

  • Actor-only reads: Browser, Search, and Context send actor_id and either omit include_shared or set it to false to read that Actor's Profile, Facts, and Summary.
  • Actor + shared reads: send both actor_id and include_shared: true to include shared Space Facts with the Actor's memory.
  • Shared-only reads: omit actor_id entirely and send include_shared: true; shared scope supports Facts only.
  • Reflections: Actor Portrait and Activity Insights are Actor-scoped and do not support shared-only scope.

owner_type is not part of the Browser, Search, or Context read scope. Event and Managed Chat bindings explicitly use owner_type: "actor" or owner_type: "space"; Space Owner requests must omit actor_id.

Browser paginates through derived records. Search performs semantic retrieval. Context assembles bounded, model-ready context within a token budget. Reflections generates or reads Actor reports on demand.

Managed Chat Boundaries​

When an application wants to keep calling /chat/completions without assembling the Prompt or maintaining Capture itself, add a Memory Binding to the request:

{
"turing_options": {
"memory": {
"enabled": true,
"space_id": "ltmspace_...",
"owner_type": "actor",
"actor_id": "user-123",
"include_shared": true,
"session_id": "session-2026-08-24"
}
}
}
  • Terminal answer: one Recall, one main LLM call, then one Event Capture.
  • Caller tool call: one Recall and one main LLM call; Capture is skipped until a terminal continuation.
  • Any fallback topology: Memory is bypassed completely—0 LTM, 0 Embedding, and 0 Capture.
  • Authentication: only an authenticated Turing Bearer JWT or Turing API Key is accepted. Other schemes are rejected before any LTM, Embedding, or main LLM cost.

Evidence is not returned to the caller. Developer Platform shows the public Chat response, Memory outcome headers, and Trace so you can diagnose Recall and Capture. If the application owns Prompt assembly or the Agent loop, use Context, Search, and Event directly instead of Managed Chat.

V1 interface scope

The core Public LTM endpoints and turing_options.memory are now included in this API documentation's OpenAPI reference; typed SDK schemas are still planned for a later version. Developer Platform's Managed Chat Builder continues to provide exact JSON, cURL, and real-request diagnostics; do not depend on unpublished SDK-generated types.

See also​