# Integrating MCP Memory: A Strategy Guide for LLMs and Agents This guide documents the approach and rationale for integrating the `mcp-memory` server with agentic LLMs (like Antigravity). Because `mcp-memory` is a central hub for context, tasks, and environment state, it is critical that LLMs interact with it efficiently without exhausting their primary context window or causing workflow ambiguity. ## 1. The Core Philosophy: "The Central Brain" The `mcp-memory` server is the persistence layer for the AI development lifecycle. It holds: * **The Knowledge Graph:** Code changes, bug fixes, architecture decisions, and tech debt. * **Project State:** Milestones, tasks, acceptance criteria, and PR checklists. * **Environment State:** Handoff memos, standup reports, and environment fingerprints. * **Live UI Integrations:** Neovim buffer manipulation and user-action webhooks. **Rationale:** The LLM's context window is ephemeral and expensive. By pushing state to the `mcp-memory` server (via a local database and Tantivy index), the LLM can selectively retrieve only what it needs, when it needs it. ## 2. Global Rules vs. Subagents To maximize efficiency, we split interactions into two categories: **Synchronous Rules** (executed by the primary conversational agent) and **Asynchronous Subagents** (delegated background tasks). ### A. Synchronous Rules (The Primary Loop) The main LLM interacting with the user should be constrained by global system rules to ensure basic context synchronization. These actions must happen synchronously so the main agent never loses the plot. * **Context Initialization (`tasks` list, `get_preflight_context`):** Executed when a session starts. This gives the LLM immediate awareness of the current workflow. * **Context Switching (`manage_checkpoint`):** Executed when moving between branches or large features. This prevents context bleed between disparate tasks. * **End-of-Day Handoff (`add_session_summary`, `generate_standup_report`):** Triggered when the user logs off, seamlessly serializing the mental state of the LLM for tomorrow. ### B. Subagent Orchestration (The Background Team) Heavy or verbose interactions with the MCP server are delegated to specialized background subagents. This keeps the primary chat fast and focused on the code, while the "team" handles project management. #### 1. `MemoryLibrarian` (The Graph Curator) * **Role:** Analyzes git diffs and chat history to structure the Knowledge Graph. * **Tools:** `log_code_change`, `log_error_fix`, `create_entities`, `log_tech_debt`. * **Rationale:** Parsing diffs and determining entity relationships is token-heavy. Delegating this prevents the main agent from wasting reasoning cycles on database normalization. #### 2. `ScrumMaster` (The Project Manager) * **Role:** Manages the task lifecycle and acceptance criteria. * **Tools:** `add_task`, `update_task_status`, `add_milestone`, `verify_acceptance_criteria`. * **Rationale:** The main agent shouldn't have to repeatedly query "are we done yet?" The `ScrumMaster` runs alongside the session, validating criteria in the background and updating the board autonomously. #### 3. `DevOpsSRE` (The Environment Manager) * **Role:** Monitors dependencies and manages session transitions. * **Tools:** `update_env_fingerprint`, `leave_handoff_memo`. * **Rationale:** Prevents "it works on my machine" failures by passively updating fingerprints when build files (e.g., `Cargo.toml`) change. ## 3. Graceful Degradation & Server Resilience The `mcp-memory` server is a distinct background process (typically port 3000). The LLM ecosystem must handle server downtime gracefully: 1. **Event Webhooks:** If the server goes down, waiting webhook tasks (e.g., waiting for a user to save a file in Neovim) will drop. These *do not* self-heal. The LLM must recognize the dropped connection and prompt the user to retry the action. 2. **Persistent Storage:** Data (tasks, graph, ledger) is persisted to `mcp_store.redb`. When the server comes back online, no data is lost. The LLM can immediately resume querying. 3. **Subagent Fast-Failing:** If the `MemoryLibrarian` attempts to log a change while the server is offline, it will instantly fail. It is designed to abandon the background task and notify the primary agent. To recover, the primary agent can manually re-invoke the Librarian once the connection is restored, instructing it to analyze recent commits to backfill the graph. ## Conclusion By treating `mcp-memory` as the durable brain, and enforcing a strict division of labor between the primary agent loop and background subagents, we achieve a highly autonomous, highly resilient AI pair-programming environment that scales across long-running projects and multiple terminal sessions. ## 4. MCP Feature Differentiation (Cognitive Boundaries) The MCP protocol exposes three primary primitives. To prevent LLM confusion and API hallucination, the LLM must strictly adhere to the following interaction boundaries: ### A. Tools (For Stateful Mutation) * **When to use:** Use tools *exclusively* for mutating state (e.g., dd_task, log_code_change) or for highly targeted semantic searches (e.g., search_nodes, query_graph_path). * **LLM Awareness:** The LLM must not use tools to repeatedly poll for state changes. Tools represent active, expensive computing steps. ### B. Resources (For Passive Awareness) * **When to use:** Use URIs (e.g., memory://tasks/active, memory://session/delta) to read holistic project state. * **LLM Awareness:** The client integration should map these URIs to the LLM's context window. Instead of the LLM invoking a list_active_tasks tool (which costs a round-trip), the LLM should simply read the memory://tasks/active resource content if it needs to know what to do next. Resources are for passive, zero-cost reading. ### C. Prompts (For Macro-Workflows) * **When to use:** Use server-defined prompts to execute complex, multi-step routines that require bundled context. * **LLM Awareness:** Instead of the user or main agent trying to manually figure out the correct sequence of tools to end a session, the LLM should trigger the handoff_routine prompt. The server will respond with a strictly formatted message array that perfectly primes the LLM on exactly what to do next. Prompts act as "macro-instructions" to prevent the LLM from wandering off-script during complex transitions.