refactor: apply 5-pass audit optimizations across mcp-memory codebase

This commit is contained in:
Riz Ashraf committed 2026-10-06 06:05:38 +01:00
1 parent 924b6d09fa
commit 5bd8b1587a
43 files changed
+1866 -1658

No files matched your search

+10
View File
@@ -53,6 +53,8 @@ The daemon has completely eliminated raw JSON file sprawl and fragmented delta-f
### Key Principles:
* **Embedded Database Engine:** All structured components (Tasks, Snippets, Tech Debt, Checklists, etc.) are stored as binary-encoded values inside a unified `redb` database file (`store.redb`).
* **ACID Compliance & File Locks:** The Windows daemon holds an exclusive read-write lock on the database file, guaranteeing zero data corruption, race conditions, or lock contention during concurrent access.
* **Atomic Write-Guard Scope:** Store modification methods (`Store::modify` and `Store::modify_async`) retain the write lock through both the in-memory mutation and JSON serialization phases, eliminating lock-release TOCTOU race conditions.
* **Store Quarantine Mode:** If deserialization fails during `Store::load_from_db`, the store flags `is_corrupted = true` and refuses to overwrite database keys with default values on subsequent writes.
* **Asynchronous Checkpointing:** The core Knowledge Graph (Entities, Relations, Observations) still utilizes a Write-Ahead Logging (WAL) pattern (`wal.jsonl`) and a master snapshot (`master.json`) to allow safe, lock-free memory mutations which are reconciled in the background.
## 6. Domain Models & Component Stores
@@ -67,6 +69,7 @@ Currently implemented persistent stores include:
## 7. Full-Text & Semantic Search Engine (Tantivy + FastEmbed)
To support blazing-fast, intelligent semantic retrieval across the sprawling knowledge graph, the daemon embeds **Tantivy** (a full-text search engine inspired by Apache Lucene) alongside **FastEmbed** (a local ONNX runtime for vector embeddings).
* **The `MemoryIndex`:** Whenever the graph or auxiliary stores mutate, a background thread dynamically rebuilds the Tantivy index (`tantivy_index/` dir) and computes semantic vectors.
* **Pre-cached Vector Embeddings:** `SearchService::semantic_search` reuses pre-cached snippet embedding vectors (`snippet.embedding`), bypassing redundant ONNX neural network inference calls during query execution.
* **Global Omni-Search:** This architecture powers the `omni_search` tool, allowing subagents to instantly fuzzy-search and semantically rank documents across Entities, Tasks, Snippets, Error Fixes, and ADRs simultaneously in milliseconds, without loading massive JSON arrays into RAM.
## 8. Webhook Telemetry & Passive Ingestion
@@ -187,3 +190,10 @@ It supports reading and writing rich formats natively to the Windows Host OS usi
* **File Drops (CF_HDROP):** The server can parse file lists copied from Windows Explorer, and can inversely synthesize file drops into the clipboard from absolute paths.
* **Images (CF_BITMAP):** The server natively rasterizes clipboard bitmaps to JPEG on read, and can write raw RgbaImage buffers back to the clipboard on write.
* **Developer Tooling:** read_file_skeleton (AST), get_active_worktree_context (Git), get_recent_logs, toggle_clipboard_watch_mode.
## 19. High-Performance Concurrency & Resilience Guarantees
* **Async Channel Backpressure (`push_async`)**: `Store::modify_async` uses `DbWriteQueue::push_async` with `tx.send(task).await` backpressure to guarantee database write persistence under heavy async write loads without dropping write transactions.
* **Atomic Search Index Swaps**: `MemoryState::rebuild_index` constructs and populates a new `MemoryIndex` instance in isolation before performing an atomic pointer swap (`*self.search_index.write().await = new_idx`), eliminating transient empty search result windows.
* **SIMD-Friendly Single-Pass Cosine Similarity**: `cosine_similarity` calculates dot product and Euclidean norm squares in a single linear pass over float vectors, enabling SIMD compiler auto-vectorization.
* **Safe Stream Decoding on Log Tails**: Log tail operations (`get_recent_logs`) read raw bytes and decode using lossy UTF-8 conversion (`String::from_utf8_lossy`) to ensure resilience when seeking across multi-byte UTF-8 boundaries.
* **Serde Parameter & Enum Tolerance**: All action enums (`StickyNoteAction`, `SnippetSearchMode`, `Relation`) support case-insensitive variants and field aliases (`source`/`from`, `target`/`to`, `relationType`/`relation_type`) to ensure seamless execution when LLMs pass varied string formatting.