refactor: consolidate nvim crates, extract server library, and update workspace dependencies

This commit is contained in:
Riz Ashraf committed 2026-10-04 01:42:59 +01:00
1 parent a083719cf1
commit 533adfd41b
53 files changed
+5967 -1230

No files matched your search

+89 -139
View File
@@ -1,60 +1,114 @@
# mcp-memory
A high-performance, persistent Knowledge Graph and Context daemon for Antigravity, implementing the Model Context Protocol (MCP).
`
## Overview
mcp-memory acts as the persistent "brain" for the agy CLI agents. It tracks entities, relations, background tasks, engineering debt, and architectural decisions across sessions.
`
To eliminate heavy Cross-OS I/O penalties when using WSL and Windows simultaneously, mcp-memory operates using a **Dual-Transport Leader/Stub Architecture**:
* **The Server (mcp-memory-server)**: Runs natively on the Windows host. It binds to 0.0.0.0:3000, serving standard stdio to the primary Windows agy instance while simultaneously hosting an Axum HTTP server for secondary clients.
* **The Stub (mcp-memory-stub)**: An ultra-lightweight proxy binary. WSL agy instances run this native Linux stub, which transparently pipes stdio JSON-RPC traffic over the network to the Windows HTTP server (http://127.0.0.1:3000), completely bypassing WSL NTFS mounts. It features full MPSC queue buffering and a WebSocket reconnect handshake (notifications/tools/list_changed) so that tools automatically refresh seamlessly without disconnecting the CLI if the background server restarts.
> **Note for Users & LLMs**: Please read the [Effective Discourse Guide](./EFFECTIVE_DISCOURSE.md) to learn which natural language phrases to use to perfectly trigger this server's advanced MCP tools.
`
A high-performance, persistent Knowledge Graph and Context daemon for Antigravity, implementing the Model Context Protocol (MCP).
## Overview
`mcp-memory` acts as the persistent "brain" for `agy` CLI agents. It tracks entities, relations, background tasks, engineering debt, architectural decisions, and error fixes across sessions.
To eliminate heavy Cross-OS I/O penalties when using WSL and Windows simultaneously, `mcp-memory` operates using a **Dual-Transport Leader/Stub Architecture**:
* **The Server (`mcp-memory-server`)**: Runs natively on the Windows host. It binds to `0.0.0.0:3000`, serving standard stdio to the primary Windows `agy` instance while simultaneously hosting an Axum HTTP and WebSocket server for secondary clients.
* **The Stub (`mcp-memory-stub`)**: An ultra-lightweight proxy binary. WSL `agy` instances run this native Linux stub, which transparently pipes stdio JSON-RPC traffic over the network to the Windows HTTP server (`http://127.0.0.1:3000`), completely bypassing WSL NTFS mounts. It features full MPSC queue buffering and a WebSocket reconnect handshake (`notifications/tools/list_changed`) so that tools automatically refresh seamlessly without disconnecting the CLI if the background server restarts.
> **Note for Users & LLMs**: Please read the [Strategic Guidelines](./instructions.md) and [Effective Discourse Guide](./EFFECTIVE_DISCOURSE.md) to learn how to perfectly trigger this server's advanced MCP tools.
---
## Casing & Naming Standards
To prevent graph fragmentation and ensure optimal LLM tokenization and retrieval:
* **Entity Types (`entity_type`)**: Standardized as **`PascalCase`** (e.g. `DatabaseTable`, `McpTool`, `ArchitectureComponent`, `File`).
* **Relation Types (`relation_type`)**: Standardized as **`snake_case`** (e.g. `depends_on`, `calls`, `implements`, `uses`).
* **Field Keys & Attributes**: Standardized as **`snake_case`** (e.g. `file_path`, `git_commit`, `created_at`).
*Note: The server automatically normalizes and migrates incoming types to these canonical conventions on every read and write operation.*
---
## Key Features & Capabilities
### 🕸️ Multi-Hop Subgraph Expansion (`get_subgraph`)
Performs a Breadth-First Search (BFS) around a target root entity node up to a specified depth ($N$ hops), returning all connected sub-entities and relationships in a single call.
### ⚡ Automated Error Fix Auto-Matcher (`suggest_error_fix`)
Compares build and test stack traces against historical error resolutions using dense vector embeddings and signature matching, returning past solutions, modified files, and git commits.
### 💾 Memory State Checkpoints & Rollbacks (`checkpoint_state` / `restore_state`)
Saves point-in-time snapshots of graph entities, active tasks, and tech debt backlogs before risky operations, enabling seamless state restoration.
### 📊 Token Budgeting & RRF Search
* **Token Budgeting**: Supports `summary_level` (`compact` | `detailed` | `full`) and `max_tokens` parameters on `list_active_tasks` and `list_tech_debt`.
* **Hybrid RRF Search**: `omni_search` combines Tantivy BM25 keyword matching with Dense Vector embeddings using Reciprocal Rank Fusion.
* **Session Delta Resource (`memory://session/delta`)**: Delivers recent session changes in a compact context resource.
### 🏷️ Domain Tagging for Code Snippets (`tag_snippet`)
Supports categorization tags (`tags: Vec<String>`) on code snippets for category-filtered searches and domain organization.
### 🧹 Self-Healing Graph Sweeper (`sweep_graph_health`)
Audits entity nodes for orphans and calculates name similarity to surface near-duplicate merge recommendations or auto-prune stale nodes.
### 🔗 Causal Lineage & Provenance Tracker (`query_lineage`)
Traces the full causal chain linking tasks, ADRs, audit ledger entries, git commits, and error fixes for any query.
### 🎯 Topological Unblocked Task Resolver (`get_next_actionable_tasks`)
Evaluates task dependency graphs and returns unblocked, ready-to-run tasks for subagent execution.
### 🧠 Chain-of-Thought & Diagnostic Hypothesis Memory (`log_hypothesis` / `query_hypotheses`)
Records structured diagnostic hypotheses, test evidence, and verification statuses to preserve reasoning across sessions.
### 🔀 Context Workspace Diffing (`diff_context_workspaces`)
Computes structured diffs of pinned files and active task IDs between two saved context workspaces.
### 📡 Real-time WebSocket Memory Sync (`ws://127.0.0.1:3000/ws`)
Broadcasting event pipeline streams real-time graph, task, and activity mutations directly to the Brain Monitor UI.
---
## Quick Start & Usage
### 1. Windows Installation (The Server & Stub)
To enforce strict process safety and eliminate file locks on Windows, the build, deploy, and execution lifecycle are entirely decoupled in the justfile.
To enforce strict process safety and eliminate file locks on Windows, the build, deploy, and execution lifecycle are entirely decoupled in the `justfile`.
**The Golden Rule:** You must gracefully stop the server before deploying a new binary. Deploy recipes only copy files; they do not kill processes.
The easiest way to manage this end-to-end (Stop -> Build -> Deploy -> Start) is using the chaining commands:
``powershell
```powershell
# For the main server:
just all-server-win
# For the lightweight stubs/nvim servers:
just all-stub-win
just all-nvim-win
``
```
If you want to perform these steps manually, you must follow this exact order to avoid NTFS locks:
``powershell
If you want to perform these steps manually, follow this exact order:
```powershell
just stop # 1. Gracefully shut down the background server (TCP 3000)
just build-win # 2. Compile the binaries
just deploy-win # 3. Move the executables into ~/.local/bin/
just start # 4. Spawns the daemon completely detached in the background
just verify # 5. Hits the /ping endpoint to ensure liveness
``
```
**Step 1:** To bypass Antigravity's lazy-loading and ensure the server is instantly available for WSL, configure your PowerShell profile to auto-start the background server when you open a terminal:
``powershell
# Add this to your PowerShell profile (ensuring it only fires on initial load, not background threads):
**Auto-Start Configuration:** To ensure the background server is always available, add this to your PowerShell profile:
```powershell
if ($host.Name -eq 'ConsoleHost' -and -not (Get-Process mcp-memory-server -ErrorAction SilentlyContinue)) {
Start-Process -FilePath "C:\Users\reazul.ashraf\.local\bin\mcp-memory-server.exe" -WindowStyle Hidden -ErrorAction SilentlyContinue
}
``
```
**Shutting Down & Managing:** If you need to stop, start, or restart the background daemon, NEVER use brute-force OS kill commands (e.g., Stop-Process, pkill). ALWAYS use the justfile wrappers, which trigger a graceful /shutdown over HTTP:
``powershell
**Shutting Down & Managing:**
```powershell
just start
just stop
just restart
``
Alternatively, you can gracefully shut down the server by invoking the executable with the --exit flag (mcp-memory-server.exe --exit) or hitting the HTTP endpoint (POST http://127.0.0.1:3000/shutdown).
`
**Step 2:** Update your Windows ~/.gemini/config/mcp_config.json to point the CLI to the ultra-lightweight stub (since the server is already running in the background):
`json
```
### 2. Config Setup (`mcp_config.json`)
Update your `~/.gemini/config/mcp_config.json`:
```json
{
"mcpServers": {
"memory": {
@@ -63,123 +117,19 @@ Alternatively, you can gracefully shut down the server by invoking the executabl
}
}
}
`
`
### 2. WSL / Linux Installation (The Stub)
Compile the ultra-lightweight stub as a native Linux binary directly from WSL (we do not use zigbuild):
```powershell
just deploy-stub-wsl
```
`
Update your WSL ~/.gemini/config/mcp_config.json:
`json
{
"mcpServers": {
"memory": {
"command": "/home/riz/.local/bin/mcp-memory-stub",
"args": [
"--target", "http://127.0.0.1:3000",
"--wake-cmd", "powershell.exe -NoProfile -WindowStyle Hidden -Command \"Start-Process -FilePath 'C:\\Users\\reazul.ashraf\\.local\\bin\\mcp-memory-server.exe' -WindowStyle Hidden\""
]
}
}
}
`
*Note: The --wake-cmd ensures that if you start WSL while Windows is completely asleep, the Linux stub will use WSL interop to silently spin up the Windows daemon in the background before connecting.*
## Native Local Ollama LLM Handshake
`mcp-memory-server` directly interfaces with local Ollama instances (e.g. `http://192.168.1.30:11434`) via a native compiled Rust client (`server/src/ollama.rs`):
* **Connection Handshake:** On startup and before executing LLM-enhanced tools, `mcp-memory-server` issues a **1.5-second health probe** (`GET /api/tags`).
* **Environment Configuration (`mcp_config.json`):**
* `OLLAMA_URL`: Target Ollama host URL (default: `http://192.168.1.30:11434`).
* `OLLAMA_CODER_MODEL`: Local coding model (e.g., `qwen2.5-coder:1.5b`).
* `OLLAMA_REASONING_MODEL`: Chain-of-thought reasoning model (e.g., `deepseek-r1:1.5b`).
* `OLLAMA_VISION_MODEL`: Multimodal vision model (e.g., `qwen3-vl:2b`).
* `OLLAMA_EMBED_MODEL`: Dense embedding model (e.g., `nomic-embed-text:latest`).
* **Graceful Degradation Guarantee:** If the local Ollama host is offline or unreachable, **no tool ever fails.** Every tool automatically falls back to pure Rust execution and local `redb`/`Tantivy` storage.
---
`
## Push Safety Gates
The daemon also operates as a global safety gate for Git. Before pushing code, run:
`ash
mcp-memory gate verify
`
This queries the daemon (via HTTP) to confirm if pre-push validation (like running tests via PrePushAuditor) has been cleared by the agent.
`
## Brain Monitor Dashboard
The server hosts a live, real-time SPA dashboard called the **Brain Monitor**.
To view the dashboard, simply navigate to the root endpoint in your browser while the server is running:
http://127.0.0.1:3000/
To view the dashboard, open your browser at:
`http://127.0.0.1:3000/`
### Dashboard Features:
* **Interactive Knowledge Graph:** A full physics-simulated network graph with node-type coloring, a drag-to-pan canvas, and an interactive Inspector Panel. Click any node to instantly view its stored observations.
* **Actionable Kanban Board:** Visually track all active agent Tasks. You can click 'Complete' directly from the UI to trigger a POST /api/tasks/{id}/complete REST call back to the daemon without needing the CLI.
* **📋 Inspect Clipboard (Image Capture):** Visually review OS-level clipboard images directly in the browser via the new /api/clipboard/capture endpoint, allowing the agent to dynamically 'see' screenshot errors or UI wireframes.
* **Sticky Notes & Search:** Browse your ephemeral notes and utilize the integrated fuzzy-search bar to locate graph entities instantly.
* **Native Dark Mode:** Fully styled for modern development environments.
## Rich Clipboard Tools
The memory server provides powerful, cross-OS native clipboard capabilities allowing the agent to inject and extract rich data. Because of the Leader/Stub architecture, all operations map directly to the Windows Host OS clipboard, regardless of whether the agent is running in Windows or WSL Ubuntu:
* **Text & HTML:** Read and write plain text or rich HTML formatting (using clipboard-win).
* **File Drops (CF_HDROP):** Read absolute file paths that were copied in Windows Explorer, or place file paths into the clipboard so the user can easily Ctrl+V them into IDEs or Explorer. (Note: Linux paths must be translated to \\wsl.localhost\... UNC paths first).
* **Images:** Read screenshots, and write generated images directly into the clipboard (using rboard).
You can also programmatically query the live backing APIs:
`curl http://127.0.0.1:3000/ping`
`curl http://127.0.0.1:3000/api/stats`
`curl http://127.0.0.1:3000/api/graph`
`curl http://127.0.0.1:3000/api/tasks`
## Further Reading
For a deep dive into the architecture, the **Redb LSM-tree** embedded database, `tantivy` indexing, and the HTTP SSE event loop, consult the design.md file in this repository.
## Neovim Integration
The linux-nvim and win-nvim MCP servers provide direct Msgpack-RPC communication with Neovim.
For this to work flawlessly across multiple Neovim instances (even split across Windows and WSL), you must load the provided gemini-integration.lua file in your Neovim init.lua:
`lua
dofile("C:/Users/reazul.ashraf/workspace/rust/mcp-memory/gemini-integration.lua")
`
### The "Last Focused Wins" Architecture
When you use the gemini-integration.lua script, Neovim acts as an active telemetry broadcaster.
Whenever you alt-tab into a Neovim window (FocusGained) or switch files (BufEnter):
1. **Fallback Sync:** Neovim instantly writes its unique Session ID (Named Pipe / Unix Socket) to ~/.gemini/active_nvim.txt.
2. **WebSocket Telemetry:** Neovim pushes a JSON payload containing the active filename, cursor row, and column via a connectionless UDP datagram to the Rust server's port 3002 listener.
3. **UI Broadcast:** The Rust server updates the global state and broadcasts this over WebSockets (/ws) so that the Brain Monitor Dashboard can animate your active file live in the UI!
### Interactive UDP UI (New)
The UI script (`gemini-ui.lua`) provides a deeply integrated, non-blocking pair-programming experience using zero-latency UDP:
* **Interactive Prompts:** The agent can trigger native `vim.ui.select` or `vim.ui.input` dialogs in your editor. Your responses are instantly routed back to the agent via UDP.
* **Smart Context (`<leader>ai`):** Highlighting code and pressing `<leader>ai` will package your prompt, file, and exact cursor/selection coordinates into a UDP packet and send it directly to the agent without spawning any subprocesses.
* **Ghost Text Diffs (Non-Destructive Review):** Instead of modifying your buffers directly, the agent uses `extmarks` to overlay proposed code changes as grayed-out "Ghost Text".
* Press `<leader>aa` (**A**gent **A**ccept) to apply the change and notify the agent.
* Press `<leader>ar` (**A**gent **R**eject) to dismiss the change and notify the agent.
### Neovim MCP Tools
The LLM agent interacts with your active Neovim session using a dedicated set of MCP tools. *(Note: /nvim/telemetry is strictly a one-way webhook for Neovim; the LLM uses the tools below to interact).*
* **vim_goto_line**: Open files and jump cursors directly from the LLM.
* **vim_set_diagnostics**: Push inline code review warnings as virtual text.
* **vim_get_active_buffer**: Read live, unsaved buffer contents.
* **vim_get_cursor**: Fetch precise line/column coordinates.
* **vim_get_visual_selection**: Read highlighted code blocks.
## Enhanced Developer Tools
- **AST Skeleton Extractor (
ead_file_skeleton)**: Uses ree-sitter to parse large code files (Rust, Python, TS/JS, Java, C, C++, Go) and return an AST structural outline containing only Imports, Structs, Enums, Traits, and Functions, massively saving LLM tokens.
- **Git Context (get_active_worktree_context)**: Uses native git2 C-bindings to retrieve the active branch, modified files, and a truncated local patch diff without shell parsing overhead.
- **Rolling Log Watcher**: Background background polling endpoints (watch_process_logs / get_recent_logs) to instantly debug daemon crashes.
- **Clipboard Watch Mode ( oggle_clipboard_watch_mode)**: Background daemon thread that auto-ingests your Ctrl+C clipboard activity directly into Knowledge Graph StickyNotes while you debug.
- **Ghost Text Previews (
vim_set_preview)**: Pushes proposed LLM code diffs directly into Neovim buffers as ephemeral virtual text.
- [Prompting Guide & Effective Discourse](./EFFECTIVE_DISCOURSE.md): Learn how to phrase prompts to get the most out of the agent and memory server.
## T2R & Token Efficiency Enhancements (V2)
- **AST Node Replacer (
eplace_ast_node)**: Uses ree-sitter to deterministically edit functions and structs without relying on exact line numbers or regex matching, ensuring zero syntax breaking edits.
- **Semantic Code Search (semantic_code_search)**: Integrates local vector embeddings to execute conceptual code searches instead of blind grep regexes, preventing hallucinated token consumption.
- **Interactive Terminal Integrations (
vim_send_to_terminal)**: Proxies shell execution into visible Neovim splits so human operators can watch agents compile code, debug output, and intervene interactively.
- **Bird's Eye Architecture View (
ead_directory_architecture)**: Generates a high-level summary of workspace directories using heuristic analysis to prevent LLMs from wasting tokens on reading dozens of files while exploring new repos.
* **Interactive Knowledge Graph:** Physics-simulated network graph with node-type coloring, drag-and-drop, and Inspector Panel.
* **Kanban Board:** Track active Tasks and trigger status transitions directly from the browser.
* **Clipboard Inspector:** Review OS-level clipboard image captures via `/api/clipboard/capture`.
* **Live WebSocket Telemetry:** Real-time UI updates triggered by server state changes.