feat(embedding,vision): candle embeddings with offline fallback, on-demand clipboard vision capture, and concurrency audit
This commit is contained in:
1 parent
5bd8b1587a
commit
e4a0fe72df
47 files changed
+6292
-3503
No files matched your search
@@ -10,11 +10,20 @@ description: Strict guidelines for interacting with the mcp-memory server, ensur
|
||||
- If a port conflict occurs (e.g., a Rust panic `AddrInUse` during a `git push` gatekeeper check), **STOP** and immediately notify the user. Do not attempt to auto-resolve the conflict by killing the existing memory server process.
|
||||
|
||||
## 2. Proactive "Central Brain" Usage
|
||||
|
||||
> [!NOTE] Two-Tier Memory Architecture
|
||||
> 1. **Tier 1 (Static Markdown)**: Repository rules, coding style, tech constraints, and architectural boundaries belong in static git-tracked markdown (`rules/*.md`, `instructions.md`) and system prompts for 0ms latency and deterministic turn-0 enforcement.
|
||||
> 2. **Tier 2 (Structured DB & Telemetry)**: The MCP Memory server specializes in high-volume, dynamic data: file modification ledgers (`audit_ledger`), terminal command history, error resolutions (`log_error_fix`), active tasks, and preflight context aggregation.
|
||||
|
||||
The MCP Memory server is the central brain. You must be PROACTIVE, not reactive, in using it:
|
||||
- **Session Starts & Context Drops**: Always begin by calling `tasks` (action: "list"), `pinned_files` (action: "list"), and `sticky_notes` (action: "read").
|
||||
- **Sticky Notes**: Use `sticky_notes` (action: "add") for transient, session-scoped operational constraints (e.g., "Do not touch file X until Y is done").
|
||||
- **Error Fixes**: The moment a tricky, undocumented, or environment-specific bug is resolved (e.g., Bitbucket markdown rendering quirks, nuanced framework bugs), IMMEDIATELY call `log_error_fix`. Do not wait for the user to ask.
|
||||
- **Tech Debt**: If you notice an anti-pattern (e.g., nested `if` statements, arrow anti-pattern) but deliberately skip fixing it to focus on a feature, IMMEDIATELY call `log_tech_debt`.
|
||||
- **Sticky Notes**: Use `sticky_notes` (action: "add") for transient, session-scoped operational constraints (e.g., "Do not touch file X until Y is done"). Deletion supports both 1-based index (standard) and 0-based index 0.
|
||||
- **Error Fixes**: The moment a tricky, undocumented, or environment-specific bug is resolved (e.g., Bitbucket markdown rendering quirks, nuanced framework bugs), IMMEDIATELY call `log_error_fix`. Supply `repo_name`, `error_category`, and `stack_trace` so future searches can perform embedding-based match.
|
||||
- **Tech Debt**: If you notice an anti-pattern (e.g., nested `if` statements, arrow anti-pattern) but deliberately skip fixing it to focus on a feature, IMMEDIATELY call `tech_debt` (action: "log") with `description`, `file_path`, `line_range`, `workaround`, `effort_estimate`, and `severity`.
|
||||
- **Architectural Decisions (ADR)**: When selecting design patterns, crate choices, or system structure, call `decisions` (action: "log") with `author`, `affected_components`, `alternatives_considered`, `decision`, and `consequence`.
|
||||
- **Task Management**: When creating tasks, supply `priority` ('low'|'medium'|'high'|'urgent'), `assigned_agent` (e.g. subagent role), `verification_command` (automated test command), and `acceptance_criteria`.
|
||||
- **VCS & SVN Agnosticism**: Supply `vcs_type` ('git'|'svn'|'hg'), `vcs_revision` (git hash or svn revision like 'r12345'), and `upstream_url` to `log_code_change` and workspace tools.
|
||||
- **Terminal & Shell Context**: Terminal sessions and commands are automatically tracked in the server. Query `/terminal/history` or recent logs when analyzing shell execution context.
|
||||
|
||||
## 3. Delegation
|
||||
Continue to use the `MemoryLibrarian` subagent to log routine code changes (`log_code_change`) in the background to prevent cluttering the main conversation context.
|
||||
@@ -23,8 +32,10 @@ Continue to use the `MemoryLibrarian` subagent to log routine code changes (`log
|
||||
- **Batch Mutating Operations**: When creating or updating multiple graph entities, code snippets, or observations, always batch items into a single tool call array (e.g. `create_entities` with multiple items) to leverage the server's single-pass transaction flush.
|
||||
- **High-Signal Tool Confirmations**: Tool call execution responses return structured, informative summaries (entity names, types, created counts, and edge paths). Agents DO NOT need to invoke follow-up `open_nodes` calls purely to confirm successful creation.
|
||||
- **Tantivy Search Reader Refresh**: Search queries (`omni_search`, `search_nodes`) automatically reload pending commits prior to executing searches, ensuring immediate visibility of newly created items.
|
||||
- **Bounded Telemetry Buffers**: High-volume telemetry logs (`error_fixes` max 300, `ledger` max 500, `handoff_memos` max 200, `session_summaries` max 200, `agent_signals` max 500) enforce deterministic length caps to guarantee low memory footprints over long sessions.
|
||||
|
||||
## 5. Pure Native Rust Invariants & Subprocess Prohibitions
|
||||
## 5. Pure Native Rust Invariants & Security
|
||||
- **Zero External Subprocesses**: Native system handlers (`clipboard`, `ast`, `search`, `db`) MUST use pure native Rust crates (`arboard`, `tree-sitter`, `tantivy`, `psycopg`). Subprocess calls to `powershell.exe`, `wl-paste`, `xclip`, or `cmd.exe` are strictly banned in native handlers.
|
||||
- **Transient Lock Handling**: Transient OS handle collisions (e.g. Win32 OLE `OpenClipboard` locks) must be handled natively with retry loops and backoffs in Rust.
|
||||
|
||||
- **Atomic Serialization Scope**: All store updates (`Store::modify` / `modify_async`) perform state mutation and JSON serialization inside an atomic write lock scope to guarantee thread-safe `DbWriteQueue` synchronization.
|
||||
- **Path Traversal Guards**: AST and file handler operations enforce path canonicalization (`validate_safe_path`) to prevent directory traversal vulnerabilities (`..`).
|
||||
@@ -11,17 +11,17 @@ This isolates editor-control logic natively to whichever OS environment executio
|
||||
|
||||
## How the MCP Server Gets Called
|
||||
|
||||
The Antigravity CLI (`agy`) acts as the MCP Client and automatically manages the lifecycle of these servers.
|
||||
The Antigravity CLI (`agy`) acts as the MCP Client and automatically manages the lifecycle of these servers.
|
||||
|
||||
1. **Registration:** The servers are registered in the global configuration file:
|
||||
- WSL: `/home/riz/.gemini/config/mcp_config.json`
|
||||
- Windows: `C:\Users\reazul.ashraf\.gemini\config\mcp_config.json`
|
||||
|
||||
2. **Execution:**
|
||||
2. **Execution:**
|
||||
When `agy` starts up, it reads `mcp_config.json`. If it finds `"win-nvim": { "command": "C:\\Users\\reazul.ashraf\\.local\\bin\\mcp-memory-nvim.exe" }`, it will spawn that binary as a background subprocess using standard `stdio`.
|
||||
|
||||
3. **Communication:**
|
||||
- The LLM requests to use a tool (e.g., `nvim_goto_line`).
|
||||
- The LLM requests to use a consolidated tool (e.g., `nvim_view` with action `goto_line`, or `nvim_execute_lua`).
|
||||
- The `agy` CLI sends a JSON-RPC request to the `mcp-memory-nvim` subprocess via its `stdin`.
|
||||
- The Rust MCP Server receives the request, connects to the Neovim active socket/pipe (`~/.gemini/active_nvim.txt` or `\\.\pipe\nvim.*`), sends the Msgpack-RPC command, and writes the JSON-RPC response back to `stdout`.
|
||||
- The `agy` CLI reads the response from `stdout` and returns it to the LLM context.
|
||||
@@ -29,9 +29,12 @@ The Antigravity CLI (`agy`) acts as the MCP Client and automatically manages the
|
||||
## Capabilities & Requirements
|
||||
To use this architecture, Neovim must run the `gemini-integration.lua` script to broadcast its active socket to `~/.gemini/active_nvim.txt`.
|
||||
|
||||
The MCP server provides core tools including:
|
||||
1. **`nvim_goto_line`**
|
||||
2. **`nvim_set_diagnostics`**
|
||||
3. **`nvim_get_active_buffer`**
|
||||
4. **`nvim_get_cursor`**
|
||||
5. **`nvim_get_visual_selection`**
|
||||
The MCP server provides 7 cohesive domain tools:
|
||||
1. **`nvim_buffer`** (actions: `get_active`, `read`, `open`, `create_scratch`, `save`, `reload`, `close`, `list`, `search`)
|
||||
2. **`nvim_window`** (actions: `list`, `get_active`, `focus`, `split`, `close`)
|
||||
3. **`nvim_view`** (actions: `goto_line`, `get_cursor`, `get_viewport`, `get_selection`)
|
||||
4. **`nvim_diagnostics`** (actions: `get`, `set`, `set_quickfix`)
|
||||
5. **`nvim_visual`** (actions: `preview`, `extmark`, `highlight`, `clear_highlight`)
|
||||
6. **`nvim_execute_lua`** (direct Lua execution escape hatch)
|
||||
7. **`nvim_system`** (actions: `get_info`, `get_messages`, `send_to_terminal`)
|
||||
|
||||
@@ -1,7 +1,8 @@
|
||||
# Neovim MCP Enforcement Rule
|
||||
|
||||
When interacting with the user's Neovim editor (e.g., opening a file, moving the cursor, reading the active buffer, setting diagnostics), you MUST ALWAYS use the MCP tools provided by the `win-nvim` (Neovim) MCP server.
|
||||
When interacting with the user's Neovim editor (e.g., opening a file, moving the cursor, reading the active buffer, setting diagnostics), you MUST ALWAYS use the MCP tools provided by the `win-nvim` (Neovim) MCP server.
|
||||
|
||||
- You are strictly forbidden from using bash scripts, `nvim --server`, or other raw terminal/shell hacks to remote-control Neovim.
|
||||
- You must rely entirely on the MCP tool registry (e.g., `nvim_goto_line`, `nvim_get_active_buffer`, `nvim_get_cursor`, `nvim_get_visual_selection`, `nvim_set_diagnostics`).
|
||||
- You must rely entirely on the consolidated MCP tool registry (`nvim_buffer`, `nvim_window`, `nvim_view`, `nvim_diagnostics`, `nvim_visual`, `nvim_execute_lua`, `nvim_system`).
|
||||
- If the tool is eagerly loaded, use it natively as an agent tool. If lazy-loaded, invoke it via the `call_mcp_tool` mechanism.
|
||||
|
||||
@@ -1,10 +1,26 @@
|
||||
# Rust Guidelines & Quirks
|
||||
|
||||
## Concurrency & Locking
|
||||
- **Lock Poisoning Protection:** NEVER use `.unwrap()` when acquiring a `Mutex` or `RwLock` (e.g., `lock.write().unwrap()`). ALWAYS use `.unwrap_or_else(|e| e.into_inner())` to gracefully recover the underlying data from poisoned locks and prevent cascading panics across threads or async tasks.
|
||||
- **Panic-Free Architecture:** Avoid `.unwrap()` anywhere in production code. Use `.expect()` for startup initialization errors, and `.unwrap_or_else()`, `.unwrap_or_default()`, or proper `Result` propagation for runtime operations.
|
||||
## Concurrency & Async Locking
|
||||
- **Lock Poisoning & Async Safety:** Use `tokio::sync::RwLock` for state shared across Tokio async tasks (such as active WebSocket clients) to avoid blocking worker threads during broadcast fanouts. For synchronous locks, prefer non-poisoning structures or recover cleanly using `.unwrap_or_else(|e| e.into_inner())`.
|
||||
- **Atomic Store Write Lock Minimization:** Minimize write lock duration by executing JSON serialization under read lock guards, keeping write guards strictly to in-memory state mutations.
|
||||
- **Two-Phase Graph Condensation:** When performing summarization or condensation across stores (`condense_graph_worker`), implement a two-phase commit: read non-destructively and synthesize observations first, insert into the knowledge graph, and only prune summarized source records by timestamp/content after successful insertion.
|
||||
- **Watch-Based Non-Destructive Shutdown Channels:** Use `tokio::sync::watch` rather than `tokio::sync::mpsc` for cancellation signaling to allow multiple workers to observe shutdown state without consuming or starving sibling workers.
|
||||
- **Redb Transient Lock Resilience:** Implement retry loops with exponential backoff on table or database lock contention before aborting or panicking.
|
||||
- **Safe RPC Request Tracking:** Clean up pending request maps (`PENDING_REQUESTS.remove(&msgid)`) upon timeouts or channel disconnects to prevent orphan memory leaks.
|
||||
- **Panic-Free Architecture:** Avoid raw `.unwrap()` in production runtime paths. Use `.unwrap_or_default()`, or proper `Result` propagation for runtime operations.
|
||||
- **Background Worker Task Supervision:** Always track `tokio::task::JoinHandle` handles for background workers (`ttl_sweeper_worker`, `index_committer_worker`, `condense_graph_worker`) and log thread exit or panic events cleanly.
|
||||
- **Offload Heavy Index Rebuilds:** In `MemoryState::rebuild_index`, offload full graph cloning and Tantivy document re-indexing into `tokio::task::spawn_blocking` to avoid stalling async worker threads.
|
||||
|
||||
## Embedding & Memory Optimizations
|
||||
- **Dynamic Character Batching:** In embedding generation (`generate_embeddings_async`), dynamically chunk batches based on total character size (e.g. 16,384 chars) rather than static item counts to prevent OOM spikes on large files while maximizing SIMD throughput.
|
||||
- **SIMD Cosine Similarity:** Compute dot-product and norm accumulators in single-pass iterator folds to facilitate vector auto-vectorization across CPU instruction sets (`AVX2`/`NEON`).
|
||||
- **Memory Truncation Bounds:** Enforce a 4KB ceiling on telemetry detail strings (`ActivityRecord`, `TerminalHistory`) before queuing items into ring buffers to bound heap usage.
|
||||
|
||||
## IDE & Rust-Analyzer Quirks
|
||||
- **Boolean NOT Operator (E0600):** Avoid using the unary `!` operator on complex boolean expressions inside closures (e.g., `!(a == b && c == d)`). `rust-analyzer` may lose track of the type boundary and falsely report an E0600 error (`cannot apply unary operator ! to type bool`). Rewrite these expressions using De Morgan's laws (e.g., `a != b || c != d`).
|
||||
- **Option::None Shadowing:** If `rust-analyzer` throws a `non_snake_case` warning for `None` during pattern matching (often caused by wildcard imports like `use crate::models::*;` shadowing standard prelude variants), explicitly namespace the variant as `std::option::Option::None` to satisfy the LSP.
|
||||
- **Deep Cloning across Thread Boundaries:** When moving large structs (like entities with large text vectors) into a `tokio::task::spawn_blocking` closure for indexing or processing, construct the required primitive payloads or target structs on the main thread *before* the closure to avoid `.clone()`ing the entire massive struct across the `'static` boundary.
|
||||
- **Boolean NOT Operator (E0600):** Avoid using the unary `!` operator on complex boolean expressions inside closures (e.g., `!(a == b && c == d)`). Rewrite these expressions using De Morgan's laws (e.g., `a != b || c != d`).
|
||||
- **Option::None Shadowing:** If `rust-analyzer` throws a `non_snake_case` warning for `None` during pattern matching (often caused by wildcard imports like `use crate::models::*;`), explicitly namespace as `std::option::Option::None`.
|
||||
- **Deep Cloning across Thread Boundaries:** Construct target primitive payloads or target structs on the main thread *before* moving into `tokio::task::spawn_blocking` closures to avoid cloning massive structs across thread boundaries.
|
||||
|
||||
## Windows MSVC & Test Concurrency
|
||||
- **ONNX Runtime / Fastembed Concurrency Resilience:** ONNX Runtime (`ort.dll` via `fastembed`) model initialization is strictly managed via a thread-safe singleton (`OnceLock<Mutex<TextEmbedding>>`) behind a process-wide `INIT_MUTEX`. The historical `0xc0000374 STATUS_HEAP_CORRUPTION` crash under uncoordinated C-ABI initializations is fully resolved. Full parallel test execution (`cargo test --workspace` or `cargo nextest run --workspace`) across all CPU cores without `--test-threads=1` is safe, recommended, and standard across all platforms.
|
||||
|
||||
Reference in new issue
Block a user