197 lines
15 KiB
Markdown
197 lines
15 KiB
Markdown
# mcp-memory
|
|
|
|
A high-performance, persistent Knowledge Graph and Context daemon for Antigravity, implementing the Model Context Protocol (MCP).
|
|
|
|
## Overview
|
|
|
|
`mcp-memory` acts as the persistent "brain" for `agy` CLI agents. It tracks entities, relations, background tasks, engineering debt, architectural decisions, and error fixes across sessions.
|
|
|
|
To eliminate heavy Cross-OS I/O penalties when using WSL and Windows simultaneously, `mcp-memory` operates using a **Dual-Transport Leader/Stub Architecture**:
|
|
* **The Server (`mcp-memory-server`)**: Runs natively on the Windows host. It binds to `0.0.0.0:3000`, serving standard stdio to the primary Windows `agy` instance while simultaneously hosting an Axum HTTP and WebSocket server for secondary clients.
|
|
* **The Stub (`mcp-memory-stub`)**: An ultra-lightweight proxy binary. WSL `agy` instances run this native Linux stub, which transparently pipes stdio JSON-RPC traffic over the network to the Windows HTTP server (`http://127.0.0.1:3000`), completely bypassing WSL NTFS mounts. It features full MPSC queue buffering and a WebSocket reconnect handshake (`notifications/tools/list_changed`) so that tools automatically refresh seamlessly without disconnecting the CLI if the background server restarts.
|
|
|
|
> **Note for Users & LLMs**: Please read the [Strategic Guidelines](./instructions.md) and [Effective Discourse Guide](./EFFECTIVE_DISCOURSE.md) to learn how to perfectly trigger this server's advanced MCP tools.
|
|
|
|
---
|
|
|
|
## Casing & Naming Standards
|
|
|
|
To prevent graph fragmentation and ensure optimal LLM tokenization and retrieval:
|
|
* **Entity Types (`entity_type`)**: Standardized as **`PascalCase`** (e.g. `DatabaseTable`, `McpTool`, `ArchitectureComponent`, `File`).
|
|
* **Relation Types (`relation_type`)**: Standardized as **`snake_case`** (e.g. `depends_on`, `calls`, `implements`, `uses`).
|
|
* **Field Keys & Attributes**: Standardized as **`snake_case`** (e.g. `file_path`, `git_commit`, `created_at`).
|
|
|
|
*Note: The server automatically normalizes and migrates incoming types to these canonical conventions on every read and write operation.*
|
|
|
|
---
|
|
|
|
---
|
|
|
|
## 🛠️ Consolidated Smart MCP Tools
|
|
|
|
The server consolidates granular single-purpose tools into 12 concise, action-oriented smart domain handlers with zero prefix clutter:
|
|
|
|
* **`tasks`**: Complete task lifecycle management (`add`, `update`, `delete`, `list`, `set_criteria`, `verify`).
|
|
* **`milestones`**: Milestone tracking (`add`, `update`, `list`).
|
|
* **`sticky_notes`**: Ephemeral scratchpad notes with TTL (`add`, `read`, `delete`, `clear`).
|
|
* **`handoff_memos`**: Cross-session handoff notes (`leave`, `read`, `clear`).
|
|
* **`pinned_files`**: Working set file focus management (`pin`, `unpin`, `list`).
|
|
* **`context_workspaces`**: Workspace context state snapshots (`save`, `load`, `list`, `delete`, `diff`).
|
|
* **`pr_checklist`**: Pre-commit and PR checklist management (`add`, `get`, `clear`).
|
|
* **`snippets`**: Reusable code snippet vault with BM25+Vector search (`store`, `search`, `delete`, `tag`).
|
|
* **`decisions`**: Architectural Decision Records (ADRs) (`log`, `query`, `delete`).
|
|
* **`tech_debt`**: Engineering technical debt backlog (`log`, `resolve`, `list`).
|
|
* **`environment`**: Infrastructure & tool fingerprints tracking (`update_fingerprint`, `read_fingerprint`, `log_requirement`, `register`, `get_details`).
|
|
* **`clipboard`**: Cross-OS clipboard management (`read`, `write`, `toggle_watch`).
|
|
|
|
---
|
|
|
|
## Key Features & Capabilities
|
|
|
|
### 📜 VCS-Agnostic Code Change Ledger & Recent Deltas (`/api/ledger` & `memory://session/delta`)
|
|
Maintains an audit ledger of all file modifications, commit hashes / SVN revisions (`vcs_revision`), repository branches, upstream URLs, and AI change summaries with deterministic length bounds. Fully agnostic across Git, Subversion (SVN), and Mercurial (Hg). Exposed via the Brain Monitor Web UI (`/api/ledger`) and accessible as a passive context resource (`memory://session/delta`).
|
|
|
|
### 💻 Terminal & Process Telemetry (`/terminal/history`)
|
|
Tracks active shell instances (PowerShell, Bash, Nushell, Zsh), command history, working directories, and exit codes in real time. Enables LLMs and the Brain Monitor UI to maintain total visibility over terminal execution contexts.
|
|
|
|
### 📋 Enriched Task Board, ADRs & Technical Debt Backlog
|
|
Supports structured priorities (`low`, `medium`, `high`, `urgent`), assigned subagent roles, automated verification commands, architectural decision alternatives and consequences, and granular technical debt tracking (line ranges, workarounds, effort estimates).
|
|
|
|
### 🕸️ Multi-Hop Subgraph Expansion (`get_subgraph`)
|
|
Performs a Breadth-First Search (BFS) around a target root entity node up to a specified depth ($N$ hops), returning all connected sub-entities and relationships in a single call.
|
|
|
|
### ⚡ Automated Error Fix Auto-Matcher (`suggest_error_fix`)
|
|
Compares build and test stack traces against historical error resolutions using dense vector embeddings and signature matching, returning past solutions, modified files, and git commits.
|
|
|
|
### 💾 Memory State Checkpoints & Rollbacks (`checkpoint_state` / `restore_state`)
|
|
Saves point-in-time snapshots of graph entities, active tasks, and tech debt backlogs before risky operations, enabling seamless state restoration.
|
|
|
|
### 📊 Token Budgeting & RRF Search
|
|
* **Token Budgeting**: Supports `summary_level` (`compact` | `detailed` | `full`) and `max_tokens` parameters on `tasks` (list) and `tech_debt` (list).
|
|
* **Hybrid RRF Search**: `omni_search` combines Tantivy BM25 keyword matching with Dense Vector embeddings using Reciprocal Rank Fusion.
|
|
* **Session Delta Resource (`memory://session/delta`)**: Delivers recent session changes in a compact context resource.
|
|
|
|
### 🏷️ Domain Tagging for Code Snippets (`snippets`)
|
|
Supports categorization tags (`tags: Vec<String>`) on code snippets for category-filtered searches and domain organization.
|
|
|
|
### 🧹 Self-Healing Graph Sweeper (`sweep_graph_health`)
|
|
Audits entity nodes for orphans and calculates name similarity to surface near-duplicate merge recommendations or auto-prune stale nodes.
|
|
|
|
### 🔗 Causal Lineage & Provenance Tracker (`query_lineage`)
|
|
Traces the full causal chain linking tasks, ADRs, audit ledger entries, git commits, and error fixes for any query.
|
|
|
|
### 🎯 Topological Unblocked Task Resolver (`get_next_actionable_tasks`)
|
|
Evaluates task dependency graphs and returns unblocked, ready-to-run tasks for subagent execution.
|
|
|
|
### 🧠 Chain-of-Thought & Diagnostic Hypothesis Memory (`log_hypothesis` / `query_hypotheses`)
|
|
Records structured diagnostic hypotheses, test evidence, and verification statuses to preserve reasoning across sessions.
|
|
|
|
### 🔀 Context Workspace Diffing (`context_workspaces`)
|
|
Computes structured diffs of pinned files and active task IDs between two saved context workspaces.
|
|
|
|
### 🔒 Resilient Storage & Serde Parameter Tolerances
|
|
* **Explicit Fail-Fast Persistence Safety**: Replaced unsafe silent fallback to temporary databases (`/tmp/mcp_store_fallback_*`) with an explicit open retry and fail-fast panic unless `MCP_ALLOW_TMP_FALLBACK=1` is explicitly set, preventing silent data loss.
|
|
* **Store Write Lock Minimization**: Releases write lock immediately following in-memory mutation, serializing JSON payloads under read locks to allow non-blocking concurrent readers.
|
|
* **Two-Phase Graph Condensation**: Employs a non-destructive 2-phase commit in `condense_graph_worker` (reading without clearing, inserting into the knowledge graph, and only pruning summarized records by timestamp/content upon verified success).
|
|
* **Redb Transient Lock Backoff**: Added exponential backoff retry loop (3 attempts, 150ms delay) on Redb table lock acquisition to gracefully handle concurrent access contention.
|
|
* **Offloaded Background Index Rebuilds**: Heavy graph cloning and Tantivy re-indexing in `rebuild_index` are offloaded to `tokio::task::spawn_blocking` to prevent starving Tokio async worker pools.
|
|
* **Watch-Based Non-Destructive Shutdown**: Server cancellation signals utilize `tokio::sync::watch` rather than `mpsc` to allow multi-consumer broadcast notifications.
|
|
* **Async Mutex Deadlock Elimination**: Converted shared state and Neovim connection locks (`shutdown_tx`, `NVIM_CONN`, `ACTIVE_SOCKET`, `HEADLESS_PROC`) to `tokio::sync::Mutex` to prevent worker thread pool starvation across `.await` points.
|
|
* **Telemetry Session Deduplication & Channel Pruning**: Added `LAST_SESSION` in-memory state deduplication for UDP telemetry writes (eliminating disk I/O thrashing) and distinguished WebSocket `TrySendError::Full` backpressure vs `TrySendError::Closed` client pruning.
|
|
* **Graph Adjacency Indexing**: Leverages `KnowledgeGraph::build_adjacency_map` to build $O(1)$ lookup adjacency lists for fast BFS shortest path graph queries.
|
|
* **Serde Parameter & Enum Ergonomics**: Consolidated tools support flexible aliases (`source`/`from`, `target`/`to`, `relationType`/`relation_type`, `parent_id`/`parentId`, `camelCase`/`PascalCase`/`snake_case`) so LLM tool invocations never fail due to parameter discrepancies.
|
|
* **Embedding Input Safeguards**: `generate_embedding_async` returns explicit errors for empty string inputs instead of 0-length fallback vectors, guaranteeing vector dimension compatibility in `cosine_similarity`.
|
|
* **Path Traversal Security Guards**: Enforces path canonicalization (`validate_safe_path`) to reject parent relative directory traversal (`..`) across process and file log endpoints.
|
|
* **Proactive Watcher Memory Eviction**: Caps file watcher `last_processed` map size at 1,000 items and purges items older than 10 minutes to prevent long-running memory leaks.
|
|
* **Pre-cached Embedding Search**: Reuses pre-computed snippet embeddings (`snippet.embedding`), bypassing ONNX inference latency during in-memory semantic searches.
|
|
### 📡 Real-time WebSocket Memory Sync (`ws://127.0.0.1:3000/ws`)
|
|
Broadcasting event pipeline streams real-time graph, task, and activity mutations directly to the Brain Monitor UI.
|
|
|
|
---
|
|
|
|
## Quick Start & Usage
|
|
|
|
### 1. Windows Installation (The Server & Stub)
|
|
|
|
To enforce strict process safety and eliminate file locks on Windows, the build, deploy, and execution lifecycle are entirely decoupled in the `justfile`.
|
|
|
|
**The Golden Rule:** You must gracefully stop the server before deploying a new binary. Deploy recipes only copy files; they do not kill processes.
|
|
|
|
The easiest way to manage this end-to-end (Stop -> Build -> Deploy -> Start) is using the chaining commands:
|
|
```powershell
|
|
# For the main server:
|
|
just all-server-win
|
|
|
|
# For the lightweight stubs/nvim servers:
|
|
just all-stub-win
|
|
just all-nvim-win
|
|
```
|
|
|
|
If you want to perform these steps manually, follow this exact order:
|
|
```powershell
|
|
just stop # 1. Gracefully shut down the background server (TCP 3000)
|
|
just build-win # 2. Compile the binaries
|
|
just deploy-win # 3. Move the executables into ~/.local/bin/
|
|
just start # 4. Spawns the daemon completely detached in the background
|
|
just verify # 5. Hits the /ping endpoint to ensure liveness
|
|
```
|
|
|
|
**Auto-Start Configuration:** To ensure the background server is always available, add this to your PowerShell profile:
|
|
```powershell
|
|
if ($host.Name -eq 'ConsoleHost' -and -not (Get-Process mcp-memory-server -ErrorAction SilentlyContinue)) {
|
|
Start-Process -FilePath "C:\Users\reazul.ashraf\.local\bin\mcp-memory-server.exe" -WindowStyle Hidden -ErrorAction SilentlyContinue
|
|
}
|
|
```
|
|
|
|
**Shutting Down & Managing:**
|
|
```powershell
|
|
just start
|
|
just stop
|
|
just restart
|
|
```
|
|
|
|
### 2. Config Setup (`mcp_config.json`)
|
|
|
|
Update your `~/.gemini/config/mcp_config.json`:
|
|
```json
|
|
{
|
|
"mcpServers": {
|
|
"memory": {
|
|
"command": "C:\\Users\\reazul.ashraf\\.local\\bin\\mcp-memory-stub.exe",
|
|
"args": []
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## Brain Monitor Dashboard
|
|
|
|
The server hosts a live, real-time SPA dashboard called the **Brain Monitor**.
|
|
|
|
To view the dashboard, open your browser at:
|
|
`http://127.0.0.1:3000/`
|
|
|
|
### Dashboard Features:
|
|
* **Interactive Knowledge Graph:** Physics-simulated network graph with node-type coloring, drag-and-drop, and Inspector Panel.
|
|
* **Kanban Board:** Track active Tasks and trigger status transitions directly from the browser.
|
|
* **Clipboard Inspector:** Review OS-level clipboard image captures via `/api/clipboard/capture`.
|
|
* **Live WebSocket Telemetry:** Real-time UI updates triggered by server state changes.
|
|
* **Code Change Ledger:** Chronological audit trail of all code edits, commits, and summaries.
|
|
|
|
---
|
|
|
|
## High-Performance Concurrency & Resilience Guarantees
|
|
|
|
* **Atomic Store Write Lock Minimization**: `Store::modify` and `Store::modify_async` release write lock guards immediately after applying state mutations, performing JSON serialization under read guards to prevent blocking concurrent readers during state serialization.
|
|
* **Async Commit Index Reader Auto-Reload**: `MemoryIndex::commit()` automatically reloads index searchers upon background commit completion, eliminating search latency and stale reader windows.
|
|
* **Async Channel Backpressure (`push_async`)**: `Store::modify_async` uses `DbWriteQueue::push_async` with `tx.send(task).await` backpressure to guarantee database write persistence under heavy async write loads without dropping write transactions.
|
|
* **Atomic Search Index Swaps**: `MemoryState::rebuild_index` constructs and populates a new `MemoryIndex` instance in isolation before performing an atomic pointer swap (`*self.search_index.write().await = new_idx`), eliminating transient empty search result windows.
|
|
* **Dynamic Character Micro-Batching**: `generate_embeddings_async` dynamically batches text payloads up to a 16,000 character budget inside `spawn_blocking`, eliminating heap spikes during high-volume vector indexing while keeping SIMD pipelines saturated.
|
|
* **Bounded Telemetry Detail Records**: Activity and terminal telemetry buffers enforce a 4,000 character truncation ceiling on log details (`ActivityRecord`, `TerminalHistory`) to prevent unbounded RAM growth under heavy RPC traffic.
|
|
* **Zero-Allocation Stream Formatting**: Graph condensation loops (`condense_graph_worker`) use `std::fmt::Write` string stream buffers to format subgraphs without allocating temporary string intermediates.
|
|
* **SIMD-Friendly Single-Pass Cosine Similarity**: `cosine_similarity` calculates dot product and Euclidean norm squares in a single iterator fold pass over float vectors, enabling SIMD compiler auto-vectorization.
|
|
* **Safe Stream Decoding on Log Tails**: Log tail operations (`get_recent_logs`) read raw bytes and decode using lossy UTF-8 conversion (`String::from_utf8_lossy`) to ensure resilience when seeking across multi-byte UTF-8 boundaries.
|
|
* **Serde Parameter & Enum Tolerance**: All action enums (`StickyNoteAction`, `SnippetSearchMode`, `Relation`) support case-insensitive variants and field aliases (`source`/`from`, `target`/`to`, `relationType`/`relation_type`) to ensure seamless execution when LLMs pass varied string formatting.
|