Managed Deep Agents is in public beta and available on LangSmith Cloud in the US region only.
Memory layers require
managed-deepagents>=0.8.0. Earlier versions declare a single deployment-shared scope instead.Choose memory layers
A layer names whose memory the agent reads and writes. Each enabled layer mounts its own read/write tree in the agent filesystem.
Omitting a layer disables it. The runtime never copies content between layers.
Use durable memory for knowledge the agent should learn while it runs and reuse later. For always-on behavior, use instructions instead. For task-specific procedures, use skills instead.
User memory is stored in an opaque Context Hub repository keyed on the caller’s authenticated principal, so one person cannot reach another person’s memory.
User memory prerequisite: user memory mounts only for a caller the deployment authenticates as a person. A LangSmith API key authenticates a service, so the default identity provider never mounts user memory. To authenticate signed-in end users, configure Supabase. Runs that arrive through a connected Slack workspace resolve to the linked person and satisfy this requirement.Agent memory has no identity requirement. Skip this prerequisite if you enable only the agent layer.
Enable memory
1
Add the memory declaration
Export a named Each layer uses its default access policy unless you supply your own.
memory declaration that enables the layers you want:memory.py
2
Guide what to remember (Optional)
The agent decides what to remember based on prompting. To make the policy explicit, add guidance like the following to
instructions.md and adapt it to your application:instructions.md is always read-only. The agent never updates it. Deploys sync project-owned instructions and skills, but do not overwrite durable content already stored in Context Hub.Control access to a layer
Declaring a layer makes it available. Whether the agent actually reaches it on a given run is a second decision, made once at the start of that run:- The runtime resolves the caller. User memory stops here unless the caller is an authenticated person.
- The layer’s
allow(context)policy runs, or its default applies when you declare no policy. - Layers that pass mount at their paths, and their hot memory loads.
allow, Managed Deep Agents applies these defaults:
Agent memory is deployment-shared, so it is available by default. User memory is personal, so it mounts by default only where the conversation is already private to one person. A Slack channel, a group DM, and a direct API run all fall outside that, and the runtime cannot tell from the outside whether such a run is private. Denying by default means personal memory never reaches a shared conversation unless you opt in with a policy of your own.
The runtime evaluates a policy once per run, before it mounts storage. Returning
false removes only that layer’s mount and its automatic memory content for that run. Stored memory is not deleted.
Read the run context
allow receives the run context. You only need to read it when you want user memory somewhere the defaults deny, such as your own application calling the API.
Managed channel runs carry the current delivery under context.channel, with the original provider event alongside it. The default user-memory policy reads the Slack channel type, where "im" means a one-to-one DM:
Replace a default policy
Supplyallow to replace a layer’s default policy. Policies can be synchronous or asynchronous. This configuration keeps agent memory available to everyone and enables user memory when the caller supplies remember: true:
memory.py
Identify the person behind user memory
User memory is keyed on the authenticated principal for the run, not on anything the caller passes in. Where that principal comes from depends on how the run started:
Each person’s memory lives in a separate Context Hub repository derived from the deployment and the principal. Two deployments therefore never share a person’s memory, even for the same person in the same Slack workspace. Agent memory is the layer to use for knowledge that should reach everyone on one deployment.
How the agent uses memory
Each mount holds hot memory and cold memory:
Keep hot memory compact because it consumes context on every run. Put detailed material, such as procedures, decision logs, and research notes, in cold files, and link to them from hot memory when useful.
The agent reads and updates memory with the built-in
read_file, edit_file, and write_file tools. A hot file that does not exist yet stays empty until the first write.
Writes to other locations, including elsewhere under /memories/, are not durable.
Test memory
Studio is the way to exercise user memory without a Slack workspace. A verified Studio user receives every declared memory layer, and the runtime does not call the user layer’sallow policy for them. A declared layer that appears to do nothing in Studio is usually not declared at all, so check memory.py or memory.ts first.
Where Studio writes depends on what you are running:
- A deployment: memory goes to Context Hub, keyed on your logged-in LangSmith user.
mda dev: memory stays on disk under.mda/__contexthub__, keyed on thelangsmith-devAgent Auth principal that the CLI resolves from your personal LangSmith API key.
mda dev does not appear after you deploy.
The runtime evaluates a policy only while a run executes. Inspecting a graph does not call it. Native Node additionally requires Agent Server to pass run context to the graph factory.
For more information, see Local development.
Disable memory
Omit a layer to disable it. Remove the memory declaration to turn durable memory off entirely. Disabling a layer stops the agent from reaching it, but does not delete stored memory.Deployment
When you runmda deploy, Managed Deep Agents enables the layers declared in the project and backs them with Context Hub. Deploys do not overwrite durable content already stored in Context Hub.
For how memory relates to deploy-owned instructions and skills in Context Hub, see Context Hub.
When to use memory
See also
Connect these docs to your agent of choice via MCP for real-time answers.

