Skip to Content
DeerFlow

Memory

💾

Memory lets DeerFlow carry useful information across sessions. The agent remembers user preferences, project context, and recurring facts so it can give better responses without starting from zero every time.

Memory is a runtime feature of the DeerFlow Harness. It is not a simple conversation log — it is a structured store of facts and context summaries that persist across separate sessions and inform the agent’s behavior in future conversations.

What memory stores

The memory store holds several categories of information:

  • Work context: summaries of ongoing projects, goals, and recurring topics the user works on.
  • Personal context: preferences, communication style, and other user-specific details the agent has learned.
  • Top of mind: the most recent focus areas and active tasks.
  • History: recent months’ context, earlier background, and long-term facts.
  • Facts: discrete, specific facts the agent has extracted from conversations (e.g., preferred tools, team names, project constraints).

Each category is updated over time as the agent learns from ongoing conversations.

How it works

Memory has two operation modes:

  1. Injection in both modes: at the start of each conversation, the agent’s current memory is injected into the system prompt at a controlled token budget (max_injection_tokens).
  2. Middleware mode (default): MemoryMiddleware runs after each Lead Agent turn. It filters the conversation, queues a background update, and extracts new facts automatically. Updates are debounced by debounce_seconds to batch rapid changes.
  3. Tool mode (experimental): DeerFlow registers memory_search, memory_add, memory_update, and memory_delete as agent tools and skips MemoryMiddleware. The model decides when to search or write memory, so effectiveness depends on model tool-use behavior. These explicit CRUD tools are not the passive staleness-review path; operators who enable tool mode are opting into model-directed updates/deletes instead of the middleware staleness guardrails.
  4. Per-agent memory: when a custom agent is active, its memory is stored separately from the global memory. This keeps different agents’ knowledge isolated.

In tool mode, the four memory_* tool names are reserved. If an MCP server or custom tool already uses memory_search, memory_add, memory_update, or memory_delete, the existing tool keeps that name and the colliding memory tool is skipped with a warning.

Memory Manager: the pluggable memory manager

At the core of the memory feature is a backend-neutral MemoryManager (defined in backend/packages/harness/deerflow/agents/memory/manager.py). It abstracts what memory can do into a stable contract, while where memory lives and how it is extracted and retrieved is left to swappable backends.

Design philosophy

The design goal is a single sentence: swapping the memory backend requires zero changes anywhere else in DeerFlow.

  • The contract says what, not how. get_context() only has to return injection-ready text — whether that comes from a local file load or a remote retrieval call is the backend’s own business. The contract does not even assume memory is a set of “facts”.
  • Tiered by implementation cost. Only three members are mandatory: the from_config classmethod (assembly), add (write) and get_context (read-and-inject). Search, clear, import and fact CRUD ship with a default “unsupported” implementation, and the buffered-write helpers (add_nowait, shutdown_flush, cancel_by_agent) default to simple pass-throughs, so a backend exposes exactly as much capability as it implements.
  • Fail loud, never silently. Memory is persistent data. A misconfigured backend or a corrupted file raises directly (e.g. ValueError at build time, HTTP 409 on write conflicts) instead of quietly falling back to another store — which would silently write your data to the wrong place.
  • Host and backend decoupled. Host capabilities (Langfuse spans, the default LLM, extraction metrics) are injected as hooks that backends consume as needed, so the backend package itself never depends on a DeerFlow-specific concept and can be shipped independently.

The MemoryManager contract at a glance:

TierMethodsPurpose
Mandatoryfrom_config / add / get_contextBuild the instance from backend_config and host hooks / queue conversations for write (debounced) / return injection-ready memory text
Management opssearch, get_memory, clear_memory, import_memory, create_fact / update_fact / delete_fact, add_nowait, cancel_by_agent, shutdown_flushSearch & inspect, clear & import, fact CRUD for tool mode, emergency flush before summarization, cancel pending extraction, bounded drain on graceful shutdown
Optional hookswarm, reload_memory, on_pre_compress / on_turn_start, plus async variantsStartup warm-up, cache reload, future extension points

How to use it

Select a backend via memory.manager_class in config.yaml; backend-private settings all live under memory.backend_config:

memory: enabled: true mode: middleware # middleware (default) | tool (model calls memory_* tools directly) manager_class: deermem # backend selector: deermem | mem0 | honcho | openviking | noop backend_config: {} # this backend's own knobs, interpreted by the backend

Built-in backends compared:

manager_classTypePositioning
deermemLocal (default)The bundled full memory implementation: Markdown fact storage, summaries, full-text retrieval, and fact lifecycle management
openvikingRemoteConnects to an independent OpenViking memory server, recalling through the official langchain-openviking adapter
mem0 / honchoRemoteAdapters for the Mem0 / Honcho memory products
noopLocalEmpty implementation: memory call sites stay, nothing happens

Two things to know:

  • Tool mode requires search support. mode: tool needs the memory_search tool, so the backend must implement search(); a backend without it fails at startup rather than silently returning empty results at runtime.
  • Custom backends are drop-in. Create a package under deerflow/agents/memory/backends/<name>/ exposing MANAGER_CLASS = <your MemoryManager subclass>, set manager_class: <name>, done. Dotted import paths (pkg.mod:Cls) also work.

How deermem is integrated

deermem is the default backend and implements the full contract:

  • The factory get_memory_manager() reads manager_class: deermem and passes backend_config plus a set of host hooks (Langfuse tracing callbacks, the default extraction LLM, hidden-message filtering, extraction metrics) to DeerMem.from_config(), which assembles the instance.
  • Storage is bucketed per (user, agent): memory.json holds only shared summaries and revision metadata, every fact is its own Markdown file (with YAML front matter), and a SQLite FTS5 full-text index backs memory_search.
  • Writes are queued by MemoryMiddleware calling add() after each turn (debounced and batched); summarization triggers an add_nowait() emergency flush so compacted content is never lost.
  • On top of that sit fact-governance capabilities, all set under memory.backend_config: staleness review (staleness_review_enabled, on by default), plus opt-in write-side near-duplicate merging (fact_dedup_enabled), consolidation (consolidation_enabled), a capacity eviction policy (fact_eviction_policy: hybrid-v1; the default is confidence), and relevance-aware retrieval (retrieval_relevance_enabled).

How openviking is integrated

OpenViking is a remote memory backend with a crisp responsibility split: DeerFlow owns capture timing, the recall query, and the transcript cursor; the langchain-openviking package owns transport, message conversion, batching, and Session commits.

memory: enabled: true injection_enabled: true manager_class: openviking mode: middleware # required: openviking supports middleware mode only backend_config: base_url: http://openviking:1933 owner_user_id: default # use default when auth is disabled api_key_env: OPENVIKING_API_KEY failure_policy: read: fail_open # read failure returns empty results (raise aborts the turn) write: log_and_drop retrieval: top_k: 8 score_threshold: 0.25 max_injection_chars: 12000

Integration notes:

  • One API key is bound to one DeerFlow owner; accessing another owner is rejected. Keep the USER key in the server environment (api_key_env), never in config.yaml.
  • One DeerFlow thread maps to one stable OpenViking Session; bounded hash-only cursors live under {storage_path}/openviking/sessions/.
  • This version supports one user with one key in middleware mode; multi-user key provisioning is out of scope. See docs/OPENVIKING.md for the full boundary and startup guide.

Implementing your own memory backend

Hooking up your own memory system (a database, an internal service, or a new product) takes four steps — nothing else in DeerFlow changes.

Step 1: Create the backend package

Create a new directory named after your backend under backend/packages/harness/deerflow/agents/memory/backends/. As long as its __init__.py exposes a MANAGER_CLASS attribute (a MemoryManager subclass), the factory discovers and registers it automatically — the folder name is the manager_class value in config.

Step 2: Implement the contract

Create mybackend/manager.py:

from deerflow.agents.memory.manager import MemoryManager class MyMemoryManager(MemoryManager): # Declare search support (True only if you actually override search(); # a mismatch between the flag and the implementation fails at # instantiation, so the two can never drift apart) supports_search = True @classmethod def from_config(cls, backend_config, *, mode="middleware", **host_hooks): # backend_config: the dict from memory.backend_config, yours to interpret # host_hooks: optional host-provided capabilities (see below); ignore if unused return cls(backend_config=backend_config, mode=mode) def add(self, thread_id, messages, *, agent_name=None, user_id=None, trace_id=None): """Mandatory: queue conversations for write (filtering/extraction is your private concern).""" def get_context(self, user_id, *, agent_name=None, thread_id=None, query=None): """Mandatory: return injection-ready text (how you retrieve and format is up to you).""" def search(self, query, top_k=5, *, user_id=None, agent_name=None, category=None): """Optional: return facts ranked by relevance (required for tool mode)."""

Only from_config + add + get_context are mandatory. Everything else is opt-in with sensible defaults:

Capability you wantMethod to overrideNotes
Tool mode (memory_search etc.)search() with supports_search = True, plus get_memory() and the fact CRUD methods belowFlag/implementation consistency is enforced at instantiation. Tool mode does not install MemoryMiddleware, so the per-turn add() writes stop unless you set requires_passive_writes_in_tool_mode = True (the add_nowait() flush before summarization still runs, and by default it calls add()); memory_add reads get_memory() before calling create_fact()
Inspect / clear / import memoryget_memory() / clear_memory() / import_memory()Unimplemented ops raise “not supported”. Return the DeerMem document shape, which the Gateway casts to MemoryResponse (version, lastUpdated, user, history, and facts whose entries carry id and content)
Fact-level CRUDcreate_fact() / update_fact() / delete_fact()Backs memory_add / memory_update / memory_delete in tool mode. create_fact returns (memory_data, fact_id); raise KeyError for an unknown fact id (HTTP 404)
Buffered writesadd_nowait() / shutdown_flush() / cancel_by_agent()The defaults only call add(), report the flush as done, and cancel nothing; override them if your add() buffers
Startup warm-up (indexing, encoders, …)warm(), returning True / False / NoneNone means nothing to warm; the log says “skipped” honestly
Releasing connections and other resourcesclose()Called on graceful shutdown
Cache reloadreload_memory()For caching backends after out-of-band edits

backend/packages/harness/deerflow/agents/memory/backends/README.md explains the DeerMem return shape and common pitfalls, and backends/noop/ is a working template to copy.

Step 3: Declare failure and conflict semantics

  • Read-failure policy is declared via read_failures_are_fatal_for_config(): permissive by default (failures return empty text); setting backend_config.failure_policy.read: fail_closed makes callers abort instead of degrading.
  • Raise MemoryConflictError for lost write races (the Gateway maps it to HTTP 409) and MemoryCorruptionError for unreadable storage (HTTP 500). Never rely on exception-text matching.

Step 4: Enable it

memory: manager_class: mybackend # i.e. backends/<folder name>

You can also skip the backends/ directory and use a dotted path directly: manager_class: mypackage.mymodule:MyMemoryManager. Either way, a resolution failure raises ValueError at startup — no silent fallback.

Available host hooks

from_config receives a set of host-default capabilities in **host_hooks — consume what you need:

HookWhat it does
callbacksA MemoryCallbacks instance: fired around your LLM calls; the default surfaces memory extraction as a dedicated Langfuse span
host_llm_factoryHost default model factory: use it for zero-config extraction when your backend has no dedicated model
should_keep_hidden_messageHidden-message filter: by default only messages carrying a human clarification response enter memory
trace_context_managerTrace context manager
extraction_callbackExtraction metrics callback (token usage, confidence filtering, rejection rates)

Consuming none of them is perfectly legal (the noop backend uses none).

Testing conventions

Backend-specific tests follow the naming convention backend/tests/test_<backend>_memory_backend.py; the existing deermem / mem0 / honcho / openviking tests are good templates.

Configuration

memory: enabled: true injection_enabled: true # Backend selector: deermem (default) | noop | openviking | a dotted path to a custom # MemoryManager subclass. Swap backend = drop a backends/<name>/ folder and # set this (see backend/.../agents/memory/backends/). manager_class: deermem # Operation mode: # middleware - default; passive background extraction after each turn # tool - experimental; model calls memory_* tools directly mode: middleware # DeerMem-private knobs. These live under backend_config (NOT at the top # level) because they are DeerMem-specific -- a different backend # self-interprets its own backend_config. Unknown keys are ignored. backend_config: # Data root. Empty = deer-flow base_dir; per-user memory at # {root}/users/{user_id}/memory.json. Absolute path = that root. storage_path: "" # Storage class (empty = FileMemoryStorage, the portable default; no importlib). storage_class: "" # Seconds to wait before processing queued memory updates (debounce) debounce_seconds: 30 # LLM for memory extraction. Omit all fields = the host factory injects the # app default model (mirrors the old model_name: null). Set explicitly to # use a different/cheaper model. model: # provider: openai # model: gpt-4o-mini # api_key: $OPENAI_API_KEY # base_url: # optional, for OpenAI-compatible gateways # temperature: # optional # Maximum number of facts to store max_facts: 100 # Minimum confidence score required to store a fact (0.0–1.0) fact_confidence_threshold: 0.7 # Maximum tokens to use for memory injection into system prompt max_injection_tokens: 2000

OpenViking backend

Set manager_class: openviking to send completed turns to an independent OpenViking server and recall its memory through the official langchain-openviking adapter. This first version supports one DeerFlow user with one OpenViking USER API key in middleware mode; DeerMem remains the default.

memory: enabled: true injection_enabled: true manager_class: openviking mode: middleware backend_config: base_url: http://openviking:1933 owner_user_id: default api_key_env: OPENVIKING_API_KEY failure_policy: read: fail_open write: log_and_drop retrieval: top_k: 8 score_threshold: 0.25 max_injection_chars: 12000

Use owner_user_id: default when DeerFlow authentication is disabled. Put the USER key in the server environment, not directly in config.yaml. Trusted-mode account headers, root-key memory access, and multi-user key provisioning are not part of this version. See docs/OPENVIKING.md for the full boundary and startup guide.

Global vs per-agent memory

DeerFlow supports two levels of memory:

  • Global memory: stored at {base_dir}/memory.json. Used when no specific agent is active or when the agent has no per-agent memory file.
  • Per-agent memory: stored at {base_dir}/agents/{agent_name}/memory.json. Used when a custom agent is active, keeping that agent’s learned knowledge separate.

The MemoryMiddleware automatically selects the correct memory file based on the active agent_name in the request configuration.

Agent names used for memory storage are validated against AGENT_NAME_PATTERN to ensure filesystem safety.

Storage location

By default, memory files are stored under the backend base directory:

  • Base directory: backend/.deer-flow/
  • Global memory: backend/.deer-flow/memory.json
  • Per-agent memory: backend/.deer-flow/agents/{agent_name}/memory.json

You can change the storage path with the storage_path field. Relative paths are resolved against the base directory. Use an absolute path to store memory in a custom location.

Custom storage backend

This section is about replacing DeerMem’s internal storage layer (lighter-weight — only changes where data lives). To hook up a fully independent memory system, see Implementing your own memory backend above.

The storage_class field allows you to replace the default file-based storage with a custom implementation. Any class that extends MemoryStorage (deerflow.agents.memory.backends.deermem.deermem.core.storage) and implements load(), reload(), and save() methods can be used:

memory: backend_config: storage_class: mypackage.storage.RedisMemoryStorage

The value is a dotted module.Class path and the class is constructed as cls(config) — see “Custom memory storage” under Customization for the full interface.

If the configured class cannot be imported, is not a MemoryStorage subclass, or cannot be constructed, memory does not fall back to FileMemoryStorage. Building the memory manager raises instead:

ValueError: backend_config.storage_class='mypackage.storage.NotThere' failed to load: No module named 'mypackage'. Refusing to silently fall back because memory is persistent state.

surfaced as a pydantic ValidationError from DeerMem’s constructor, because storage is wired in model_post_init().

Disabling memory

To disable memory entirely:

memory: enabled: false

To keep memory storage but prevent injection into the system prompt:

memory: enabled: true injection_enabled: false