refactor: consolidate hypotheses, agent_signals, and process_logs into action-based smart tools

This commit is contained in:
Riz Ashraf committed 2026-10-07 12:00:20 +01:00
1 parent 79209da711
commit 3b146f91c2
17 files changed
+438 -359

No files matched your search

+26 -34
View File
@@ -5,27 +5,21 @@
Our current MCP ecosystem is highly advanced, utilizing a **Dual-Transport Leader/Stub Architecture** (Windows Host + WSL Proxy) to completely eliminate cross-OS I/O latency.
### 1. Context & Token Optimization
* ␍ead_file_skeleton: Highly effective. Uses ree-sitter to extract ASTs (Rust, Python, TS). **Score: A+ (Massive token savings)**
* get_active_worktree_context: Native git2 integration. Bypasses shell parsing for clean JSON diffs. **Score: A**
* watch_process_logs / get_recent_logs: Direct file seeking. Prevents LLMs from reading multi-megabyte log files. **Score: A**
* `read_file_skeleton`: Highly effective. Uses tree-sitter to extract ASTs (Rust, Python, TS). **Score: A+ (Massive token savings)**
* `get_active_worktree_context`: Native git2 integration. Bypasses shell parsing for clean JSON diffs. **Score: A**
* `process_logs`: Direct file seeking and daemon log management (`watch`, `get`, `clear`). Prevents LLMs from reading multi-megabyte log files. **Score: A**
### 2. Neovim IDE Integration (
vim-core)
*
vim_set_preview,
vim_goto_line,
vim_set_diagnostics,
vim_execute_lua,
vim_get_active_buffer.
### 2. Neovim IDE Integration (nvim-core)
* `nvim_buffer`, `nvim_window`, `nvim_view`, `nvim_diagnostics`, `nvim_visual`, `nvim_execute_lua`, `nvim_system`.
* **Review:** Exceptional human QoL. The agent interacts with the code where the human's eyes actually are. Ghost text and diagnostic extmarks provide an IDE-like experience usually reserved for closed-source tools like Cursor. **Score: S-Tier**
### 3. Clipboard & Workflow
* write_clipboard, ␍ead_clipboard, oggle_clipboard_watch_mode.
* **Review:** Native cross-OS clipboard-win and rboard implementation. Auto-ingesting into StickyNotes bridges the gap between manual human research and the agent's context. **Score: A**
* `clipboard` (`read`, `write`).
* **Review:** Native cross-OS clipboard-win and arboard implementation with on-demand image grab. Bridges the gap between manual human research and the agent's context. **Score: A**
### 4. Graph & Memory Management
* create_entities, dd_sticky_note, save_context_workspace, handoff_routine.
* **Review:** Solid foundation for state persistence across branches and days. **Score: B+** (Could use more automated TTL/decay for outdated context).
* `create_entities`, `hypotheses`, `agent_signals`, `handoff_routine`.
* **Review:** Solid foundation for state persistence across branches and days. **Score: A**
---
@@ -33,29 +27,27 @@ vim_get_active_buffer.
To push the system to the absolute bleeding edge of autonomous coding, I propose the following 5 new tools/enhancements.
### 1. ␍eplace_ast_node (Robust Structural Editing)
* **The Problem:** The current ␍eplace_file_content uses exact string matching and line numbers. Line numbers change when humans edit simultaneously, and string matching fails on whitespace/indentation.
* **The Solution:** An MCP tool that takes (file_path, node_type, node_name, new_content). It uses ree-sitter to find the exact boundary of n execute(...) and replaces just that AST node.
### 1. `replace_ast_node` (Robust Structural Editing)
* **The Problem:** Standard text replacement uses exact string matching and line numbers. Line numbers change when humans edit simultaneously, and string matching fails on whitespace/indentation.
* **The Solution:** An MCP tool that takes `(file_path, node_type, node_name, new_content)`. It uses tree-sitter to find the exact boundary of `fn execute(...)` and replaces just that AST node.
* **Impact:** Zero LLM syntax/indentation errors. 100% robust edits. Drastically lowers Time-to-Resolve (T2R) by eliminating failed edit loops.
### 2. semantic_code_search (Local Vector Embeddings)
* **The Problem:** grep_search relies on exact regex. If the LLM guesses the wrong variable name, it wastes tokens searching and reading the wrong files.
* **The Solution:** We already have antivy and astembed in our Cargo.toml. We can index the AST blocks of the codebase in the background. The LLM can query *"Where is the auth token validated?"* and get the exact 3 relevant functions instantly.
### 2. `semantic_code_search` (Local Vector Embeddings)
* **The Problem:** Text search relies on exact regex. If the LLM guesses the wrong variable name, it wastes tokens searching and reading the wrong files.
* **The Solution:** Using Tantivy and BERT embeddings in our backend. We index the AST blocks of the codebase in the background. The LLM can query *"Where is the auth token validated?"* and get the exact 3 relevant functions instantly.
* **Impact:** Massive token cost reduction (no blind file reading). Instant T2R for codebase exploration.
### 3.
vim_send_to_terminal (Interactive Execution QoL)
* **The Problem:** When the agent runs a background terminal command (cargo build,
pm run dev), the output is hidden from the human, and interactive prompts cause the background task to hang indefinitely.
* **The Solution:** A tool that opens a Neovim :term split (or uses a mux pane) and sends the command there.
* **Impact:** Massive Human QoL. The human can watch the tests run natively, interact with prompts, see ANSI colors, and press <C-c> to kill it if it loops.
### 3. `nvim_system` terminal execution (Interactive Execution QoL)
* **The Problem:** When the agent runs a background terminal command (`cargo build`, `npm run dev`), the output is hidden from the human, and interactive prompts cause the background task to hang indefinitely.
* **The Solution:** Dispatch to Neovim terminal splits where the human can watch the tests run natively, interact with prompts, see ANSI colors, and interact seamlessly.
* **Impact:** Massive Human QoL.
### 4. ␍ead_directory_architecture (Bird's-Eye View)
* **The Problem:** ␍ead_file_skeleton works for one file. When entering a new repository, the LLM usually runs ls -R and then has to guess what files do based on their names.
* **The Solution:** A tool that scans a directory structure and uses basic heuristic parsing (or a tiny local embedding lookup) to return a JSON tree of files alongside a 1-sentence summary of what each file is responsible for.
### 4. `read_directory_architecture` (Bird's-Eye View)
* **The Problem:** Single file inspection works for one file. When entering a new repository, the LLM usually runs `ls -R` and then has to guess what files do based on their names.
* **The Solution:** A tool that scans a directory structure and returns a clean hierarchical tree alongside summaries of what each directory and key file is responsible for.
* **Impact:** Immediate holistic context. Eliminates the "exploration phase" token tax.
### 5. query_database_schema (Introspection)
* **The Problem:** Working with databases usually involves the LLM writing clunky bash scripts to run psql or sqlite3 to view table definitions, which often fail due to missing env vars or wrong dialects.
* **The Solution:** A direct MCP tool that parses the local .env, connects to the database (Postgres/SQLite), and returns a clean Markdown representation of the schema (Tables, Columns, Types, Foreign Keys).
* **Impact:** Prevents hallucinations about database structure. Fixes DB-related bugs significantly faster (T2R).
### 5. `query_database_schema` (Introspection)
* **The Problem:** Working with databases usually involves the LLM writing clunky scripts to view table definitions, which often fail due to missing env vars or wrong dialects.
* **The Solution:** A direct MCP tool that parses the local `.env`, connects to the database (PostgreSQL), and returns a clean Markdown representation of the schema (Tables, Columns, Types, Foreign Keys).
* **Impact:** Prevents hallucinations about database structure. Fixes DB-related bugs significantly faster (T2R).