15 KiB
mcp-memory
A high-performance, persistent Knowledge Graph and Context daemon for Antigravity, implementing the Model Context Protocol (MCP).
Overview
mcp-memory acts as the persistent "brain" for agy CLI agents. It tracks entities, relations, background tasks, engineering debt, architectural decisions, and error fixes across sessions.
To eliminate heavy Cross-OS I/O penalties when using WSL and Windows simultaneously, mcp-memory operates using a Dual-Transport Leader/Stub Architecture:
- The Server (
mcp-memory-server): Runs natively on the Windows host. It binds to0.0.0.0:3000, serving standard stdio to the primary Windowsagyinstance while simultaneously hosting an Axum HTTP and WebSocket server for secondary clients. - The Stub (
mcp-memory-stub): An ultra-lightweight proxy binary. WSLagyinstances run this native Linux stub, which transparently pipes stdio JSON-RPC traffic over the network to the Windows HTTP server (http://127.0.0.1:3000), completely bypassing WSL NTFS mounts. It features full MPSC queue buffering and a WebSocket reconnect handshake (notifications/tools/list_changed) so that tools automatically refresh seamlessly without disconnecting the CLI if the background server restarts.
Note for Users & LLMs: Please read the Strategic Guidelines and Effective Discourse Guide to learn how to perfectly trigger this server's advanced MCP tools.
Casing & Naming Standards
To prevent graph fragmentation and ensure optimal LLM tokenization and retrieval:
- Entity Types (
entity_type): Standardized asPascalCase(e.g.DatabaseTable,McpTool,ArchitectureComponent,File). - Relation Types (
relation_type): Standardized assnake_case(e.g.depends_on,calls,implements,uses). - Field Keys & Attributes: Standardized as
snake_case(e.g.file_path,git_commit,created_at).
Note: The server automatically normalizes and migrates incoming types to these canonical conventions on every read and write operation.
🛠️ Consolidated Smart MCP Tools
The server consolidates granular single-purpose tools into 11 concise, action-oriented smart domain handlers with zero prefix clutter:
tasks: Complete task lifecycle management (add,update,delete,list,set_criteria,verify).milestones: Milestone tracking (add,update,list).handoff_memos: Cross-session handoff notes (leave,read,clear).snippets: Reusable code snippet vault with BM25+Vector search (store,search,delete,tag).decisions: Architectural Decision Records (ADRs) (log,query,delete).tech_debt: Engineering technical debt backlog (log,resolve,list).environment: Infrastructure & tool fingerprints tracking (update_fingerprint,read_fingerprint,log_requirement,register,get_details).clipboard: Cross-OS clipboard management (read,write).hypotheses: Diagnostic hypothesis memory (log,query).agent_signals: Inter-agent signal bus (broadcast,query).process_logs: Process and daemon log management (watch,get,clear).
Key Features & Capabilities
📜 VCS-Agnostic Code Change Ledger & Recent Deltas (/api/ledger & memory://session/delta)
Maintains an audit ledger of all file modifications, commit hashes / SVN revisions (vcs_revision), repository branches, upstream URLs, and AI change summaries with deterministic length bounds. Fully agnostic across Git, Subversion (SVN), and Mercurial (Hg). Exposed via the Brain Monitor Web UI (/api/ledger) and accessible as a passive context resource (memory://session/delta).
💻 Terminal & Process Telemetry (/terminal/history)
Tracks active shell instances (PowerShell, Bash, Nushell, Zsh), command history, working directories, and exit codes in real time. Enables LLMs and the Brain Monitor UI to maintain total visibility over terminal execution contexts.
📋 Enriched Task Board, ADRs & Technical Debt Backlog
Supports structured priorities (low, medium, high, urgent), assigned subagent roles, automated verification commands, architectural decision alternatives and consequences, and granular technical debt tracking (line ranges, workarounds, effort estimates).
🕸️ Multi-Hop Subgraph Expansion (get_subgraph)
Performs a Breadth-First Search (BFS) around a target root entity node up to a specified depth (N hops), returning all connected sub-entities and relationships in a single call.
⚡ Automated Error Fix Auto-Matcher (suggest_error_fix)
Compares build and test stack traces against historical error resolutions using dense vector embeddings and signature matching, returning past solutions, modified files, and git commits.
💾 Memory State Checkpoints & Rollbacks (checkpoint_state / restore_state)
Saves point-in-time snapshots of graph entities, active tasks, and tech debt backlogs before risky operations, enabling seamless state restoration.
📊 Token Budgeting & RRF Search
- Token Budgeting: Supports
summary_level(compact|detailed|full) andmax_tokensparameters ontasks(list) andtech_debt(list). - Hybrid RRF Search:
omni_searchcombines Tantivy BM25 keyword matching with Dense Vector embeddings using Reciprocal Rank Fusion. - Session Delta Resource (
memory://session/delta): Delivers recent session changes in a compact context resource.
🏷️ Domain Tagging for Code Snippets (snippets)
Supports categorization tags (tags: Vec<String>) on code snippets for category-filtered searches and domain organization.
🧹 Self-Healing Graph Sweeper (sweep_graph_health)
Audits entity nodes for orphans and calculates name similarity to surface near-duplicate merge recommendations or auto-prune stale nodes.
🔗 Causal Lineage & Provenance Tracker (query_lineage)
Traces the full causal chain linking tasks, ADRs, audit ledger entries, git commits, and error fixes for any query.
🎯 Topological Unblocked Task Resolver (get_next_actionable_tasks)
Evaluates task dependency graphs and returns unblocked, ready-to-run tasks for subagent execution.
🧠 Chain-of-Thought & Diagnostic Hypothesis Memory (hypotheses)
Records structured diagnostic hypotheses, test evidence, and verification statuses (actions: log, query) to preserve reasoning across sessions.
📡 Inter-Agent Signal Bus (agent_signals)
Facilitates real-time peer-to-peer signal exchange between autonomous subagents (actions: broadcast, query) with TTL expiration and activity feeds.
🔒 Resilient Storage & Serde Parameter Tolerances
📑 Process & Daemon Log Management (process_logs)
Registers and tails live process and daemon log files with UTF-8 safe seeking (actions: watch, get, clear) to diagnose runtime behavior without reading multi-megabyte files into chat context.
- Explicit Fail-Fast Persistence Safety: Replaced unsafe silent fallback to temporary databases (
/tmp/mcp_store_fallback_*) with an explicit open retry and fail-fast panic unlessMCP_ALLOW_TMP_FALLBACK=1is explicitly set, preventing silent data loss. - Store Write Lock Minimization: Releases write lock immediately following in-memory mutation, serializing JSON payloads under read locks to allow non-blocking concurrent readers.
- Two-Phase Graph Condensation: Employs a non-destructive 2-phase commit in
condense_graph_worker(reading without clearing, inserting into the knowledge graph, and only pruning summarized records by timestamp/content upon verified success). - Redb Transient Lock Backoff: Added exponential backoff retry loop (3 attempts, 150ms delay) on Redb table lock acquisition to gracefully handle concurrent access contention.
- Offloaded Background Index Rebuilds: Heavy graph cloning and Tantivy re-indexing in
rebuild_indexare offloaded totokio::task::spawn_blockingto prevent starving Tokio async worker pools. - Watch-Based Non-Destructive Shutdown: Server cancellation signals utilize
tokio::sync::watchrather thanmpscto allow multi-consumer broadcast notifications. - Async Mutex Deadlock Elimination: Converted shared state and Neovim connection locks (
shutdown_tx,NVIM_CONN,ACTIVE_SOCKET,HEADLESS_PROC) totokio::sync::Mutexto prevent worker thread pool starvation across.awaitpoints. - Telemetry Session Deduplication & Channel Pruning: Added
LAST_SESSIONin-memory state deduplication for UDP telemetry writes (eliminating disk I/O thrashing) and distinguished WebSocketTrySendError::Fullbackpressure vsTrySendError::Closedclient pruning. - Graph Adjacency Indexing: Leverages
KnowledgeGraph::build_adjacency_mapto buildO(1)lookup adjacency lists for fast BFS shortest path graph queries. - Serde Parameter & Enum Ergonomics: Consolidated tools support flexible aliases (
source/from,target/to,relationType/relation_type,parent_id/parentId,camelCase/PascalCase/snake_case) so LLM tool invocations never fail due to parameter discrepancies. - Embedding Input Safeguards:
generate_embedding_asyncreturns explicit errors for empty string inputs instead of 0-length fallback vectors, guaranteeing vector dimension compatibility incosine_similarity. - Path Traversal Security Guards: Enforces path canonicalization (
validate_safe_path) to reject parent relative directory traversal (..) across process and file log endpoints. - Proactive Watcher Memory Eviction: Caps file watcher
last_processedmap size at 1,000 items and purges items older than 10 minutes to prevent long-running memory leaks. - Pre-cached Embedding Search: Reuses pre-computed snippet embeddings (
snippet.embedding), bypassing ONNX inference latency during in-memory semantic searches.
📡 Real-time WebSocket Memory Sync (ws://127.0.0.1:3000/ws)
Broadcasting event pipeline streams real-time graph, task, and activity mutations directly to the Brain Monitor UI.
Quick Start & Usage
1. Windows Installation (The Server & Stub)
To enforce strict process safety and eliminate file locks on Windows, the build, deploy, and execution lifecycle are entirely decoupled in the justfile.
The Golden Rule: You must gracefully stop the server before deploying a new binary. Deploy recipes only copy files; they do not kill processes.
The easiest way to manage this end-to-end (Stop -> Build -> Deploy -> Start) is using the chaining commands:
# For the main server:
just all-server-win
# For the lightweight stubs/nvim servers:
just all-stub-win
just all-nvim-win
If you want to perform these steps manually, follow this exact order:
just stop # 1. Gracefully shut down the background server (TCP 3000)
just build-win # 2. Compile the binaries
just deploy-win # 3. Move the executables into ~/.local/bin/
just start # 4. Spawns the daemon completely detached in the background
just verify # 5. Hits the /ping endpoint to ensure liveness
Auto-Start Configuration: To ensure the background server is always available, add this to your PowerShell profile:
if ($host.Name -eq 'ConsoleHost' -and -not (Get-Process mcp-memory-server -ErrorAction SilentlyContinue)) {
Start-Process -FilePath "C:\Users\reazul.ashraf\.local\bin\mcp-memory-server.exe" -WindowStyle Hidden -ErrorAction SilentlyContinue
}
Shutting Down & Managing:
just start
just stop
just restart
2. Config Setup (mcp_config.json)
Update your ~/.gemini/config/mcp_config.json:
{
"mcpServers": {
"memory": {
"command": "C:\\Users\\reazul.ashraf\\.local\\bin\\mcp-memory-stub.exe",
"args": []
}
}
}
Brain Monitor Dashboard
The server hosts a live, real-time SPA dashboard called the Brain Monitor.
To view the dashboard, open your browser at:
http://127.0.0.1:3000/
Dashboard Features:
- Interactive Knowledge Graph: Physics-simulated network graph with node-type coloring, drag-and-drop, and Inspector Panel.
- Kanban Board: Track active Tasks and trigger status transitions directly from the browser.
- Clipboard Inspector: Review OS-level clipboard image captures via
/api/clipboard/capture. - Live WebSocket Telemetry: Real-time UI updates triggered by server state changes.
- Code Change Ledger: Chronological audit trail of all code edits, commits, and summaries.
High-Performance Concurrency & Resilience Guarantees
- Atomic Store Write Lock Minimization:
Store::modifyandStore::modify_asyncrelease write lock guards immediately after applying state mutations, performing JSON serialization under read guards to prevent blocking concurrent readers during state serialization. - Async Commit Index Reader Auto-Reload:
MemoryIndex::commit()automatically reloads index searchers upon background commit completion, eliminating search latency and stale reader windows. - Async Channel Backpressure (
push_async):Store::modify_asyncusesDbWriteQueue::push_asyncwithtx.send(task).awaitbackpressure to guarantee database write persistence under heavy async write loads without dropping write transactions. - Atomic Search Index Swaps:
MemoryState::rebuild_indexconstructs and populates a newMemoryIndexinstance in isolation before performing an atomic pointer swap (*self.search_index.write().await = new_idx), eliminating transient empty search result windows. - Dynamic Character Micro-Batching:
generate_embeddings_asyncdynamically batches text payloads up to a 16,000 character budget insidespawn_blocking, eliminating heap spikes during high-volume vector indexing while keeping SIMD pipelines saturated. - Bounded Telemetry Detail Records: Activity and terminal telemetry buffers enforce a 4,000 character truncation ceiling on log details (
ActivityRecord,TerminalHistory) to prevent unbounded RAM growth under heavy RPC traffic. - Zero-Allocation Stream Formatting: Graph condensation loops (
condense_graph_worker) usestd::fmt::Writestring stream buffers to format subgraphs without allocating temporary string intermediates. - SIMD-Friendly Single-Pass Cosine Similarity:
cosine_similaritycalculates dot product and Euclidean norm squares in a single iterator fold pass over float vectors, enabling SIMD compiler auto-vectorization. - Safe Stream Decoding on Log Tails: Process log tailing (
process_logs, action:get) reads raw bytes and decodes using lossy UTF-8 conversion (String::from_utf8_lossy) to ensure resilience when seeking across multi-byte UTF-8 boundaries. - Serde Parameter & Enum Tolerance: All action enums (
HandoffMemoAction,HypothesisAction,AgentSignalAction,ProcessLogAction,SnippetSearchMode,Relation) support case-insensitive variants and field aliases (source/from,target/to,relationType/relation_type) to ensure seamless execution when LLMs pass varied string formatting.