@bturkis/code-search-cli
v0.2.0
Published
Semantic code-search MCP installer and indexer for AI coding agents: hybrid (vector + lexical RRF) ranking, gitignore-aware indexing, and optional cross-encoder rerank.
Readme
code-search-cli
Semantic code-search MCP installer, indexer, and token-savings dashboard for AI coding agents.
csearch installs a local MCP server that indexes your repositories and gives agents a semantic search tool before they read large files or run broad shell searches. It also installs IDE-specific guidance so supported agents prefer code-search during repository exploration.
Features
- Installs the bundled
code-searchMCP runtime under~/.csearch/code-search. - Bootstraps each repository with a local
.claude/mcp-code-searchindex directory. - Installs supported IDE adapters and agent guidance.
- Keeps indexes fresh with git hooks, Cursor edit hooks, and lazy freshness checks.
.gitignore-aware indexing. In a git repo, only tracked + untracked-but-not-ignored files are indexed (git ls-files), so build output, caches, and tool dumps never pollute the index. Falls back to a filtered filesystem walk outside git.- Hybrid ranking (RRF). Vector similarity and lexical scoring are combined with reciprocal rank fusion, so an exact identifier/symbol hit can outrank a semantically-adjacent-but-wrong chunk.
- No-hang search. A query embeds with a short fail-fast budget; if the embedding endpoint is down,
search_codedegrades to lexical-only results (with a diagnostics warning) instead of hanging the tool call. - Wrong-model guard. Some hosts (LM Studio) silently serve whatever embedding model is loaded when the requested one is missing. A dimension mismatch is detected and rejected so the index is never poisoned with vectors from the wrong model.
- Optional cross-encoder rerank. A second-stage
/v1/rerankpass (llama.cppllama-server --rerank, Jina/Cohere schema) can reorder the RRF top-N. Opt-in via arerankconfig block; any failure falls back to the RRF order. - Tracks global token savings and RTK-aware MCP response compression in
~/.csearch/metrics.jsonl. - Provides
csearch gain/csearch savingswith input tokens, output tokens, index savings, RTK-aware response savings, efficiency, client/IDE grouping, tool grouping, and query grouping. - Includes repair, reset, project discovery, and full reindex commands.
Installation
Requires Node.js 18 or newer.
npm install -g @bturkis/code-search-cliOr run directly:
npx @bturkis/code-search-cli helpQuick Start
csearch install
csearch uninstall --yes
csearch status
csearch gaincsearch install detects installed IDEs, asks which supported adapters to configure, checks the OS package manager preflight, detects embedding runners, asks which runner/model to use, installs or pulls what is missing, writes the embedding config, installs agent guidance, bootstraps the current repository, and runs an incremental index.
For non-interactive installs:
csearch install --all
csearch install --ide cursor
csearch install --ide claude-code
csearch uninstall --all --yesSupported IDE Adapters
Check detected IDEs:
csearch idesCurrently supported install targets:
- Cursor
- Claude Code
- OpenCode
Detected but not yet supported targets are listed as unsupported and are not modified.
Agent guidance is installed as managed blocks/files:
- Cursor:
~/.cursor/rules/csearch-code-search.mdc - Claude Code:
~/.claude/CLAUDE.md - OpenCode:
~/.config/opencode/AGENTS.md
Existing files are not overwritten. csearch only creates or updates the managed block between:
<!-- csearch managed guidance start -->
<!-- csearch managed guidance end -->The managed guidance is RTK-first: agents default to RTK for shell, file,
search, git, test, lint, and build work, and use the code-search MCP only when
they genuinely need semantic/conceptual discovery that grep cannot do. It
includes an RTK command map (read/search/list, git/PR, package managers/build,
tests, lint/format/types, infra/cloud/db, and output helpers) plus strict-usage
rules for rtk grep, rtk find, rtk read, and rtk git diff so agents do not
misuse them. It also tells agents to prefer rtk grep for exact known
strings/symbols, switch back to RTK once code-search has located the relevant
code, avoid mixing results across repositories, rerun search_code after
workspace/branch changes, and state explicitly when code-search is unavailable,
stale, or irrelevant. When using a global MCP registration, agents are
instructed to pass the current workspace path as project_root, and to stop if
returned diagnostics show a different repoRoot.
Every code-search MCP response passes through an RTK-aware output layer before
it reaches the model. Exact line-range results are passed through unchanged.
Large get_file responses are converted into an explicit structural map
(imports, exports, symbols, control/data-flow markers, omitted ranges, and exact
next-read hints), not a silent partial file; agents must request a narrow exact
range before making behavior, bug, or patch decisions. Repo diagnostics stay in
MCP metadata. Full raw files are still available with exact: true when truly
required.
RTK is the default shell output optimizer. The managed guidance tells agents to
run shell, file, git, test, lint, TypeScript, npm, and pnpm work through RTK so
output is compressed before it reaches the model, and to reach for code-search
only for semantic discovery. Agents are also instructed
not to answer from assumptions: they must trace important functions from
definition to callers and downstream callees with search_code, list_exports,
and focused get_file ranges, and call out any missing link instead of
guessing.
The guidance also tells agents to stay tied to the user's concrete goal, verify findings continuously against source/runtime/tests/diagnostics, and flag any unverified assumption or residual risk. For broad multi-area work, agents should use multiple focused subagents when supported, then the parent agent must review their outputs against code evidence, resolve contradictions, and rerun or verify any incomplete subagent result before producing the final answer.
Package Manager Preflight
Runner installation requires an OS package manager. Check it with:
csearch doctor package-managerIf no supported package manager is available, install stops and prints the official setup link:
- macOS: Homebrew, https://brew.sh/
- Windows: Windows Package Manager, https://learn.microsoft.com/windows/package-manager/winget/
- Linux:
apt,dnf,pacman, orzypper
Embedding Runner And Model Selection
install includes a runner/model wizard.
Supported runner targets:
- Ollama
- LM Studio
- Custom OpenAI-compatible remote endpoint
The installer first detects what is already installed and running. Installed runners are offered as-is. Missing supported runners are offered as install options when a supported OS package manager is available.
When a custom remote endpoint is selected, the installer fetches <base-url>/v1/models immediately after the URL is entered and asks you to choose from the returned model list.
Recommended model options include:
- Ollama:
qwen3-embedding:0.6b(recommended, 1024 dims) — strong multilingual + code retrieval - Ollama:
nomic-embed-text(768 dims) - Ollama:
mxbai-embed-large(1024 dims) - LM Studio:
text-embedding-qwen3-embedding-0.6b(recommended, 1024 dims) - LM Studio:
text-embedding-nomic-embed-text-v1.5(768 dims)
For Ollama, missing models are pulled automatically with ollama pull <model>.
Small-VRAM note (Ollama)
qwen3-embedding:0.6b advertises a 32k context. On a small GPU (for example a 6 GB laptop card) the stock context can spill to CPU or fail on long inputs. Derive a smaller-context, GPU-fit variant once:
printf 'FROM qwen3-embedding:0.6b\nPARAMETER num_ctx 4096\nPARAMETER num_batch 4096\n' > Modelfile
ollama create qwen3-embedding-4k -f ModelfileThen point the config at qwen3-embedding-4k. csearch's conservative tokenSafetyRatio (default 0.5) and batchSize (default 8) keep chunk sizes and batch load within a modest card's budget.
The selected embedding config is written to:
~/.csearch/code-search/config.jsonYou can change it later:
csearch config show
csearch config set-local
csearch config set-remote --url http://127.0.0.1:11434 --model qwen3-embedding:0.6b --dims 1024 --max-input-tokens 2048No machine-specific LAN/Tailscale IP addresses are shipped in the default config. Remote endpoints are only written when selected during install or set explicitly with config set-remote.
During indexing, csearch reads the loaded model's context metadata from the runner when available and splits chunks before sending them to LM Studio/Ollama. Because LM Studio/GGUF tokenizers can count code much higher than csearch's estimator, csearch applies a conservative tokenSafetyRatio (default 0.5) to the actual chunk budget, and sends embeddings in small batches (batchSize default 8) so a modest local GPU is not overloaded. Nomic v1/v1.5 is additionally capped at a 2048 context ceiling to avoid truncation warnings. Use --max-input-tokens only when the runner does not expose context metadata.
If LM Studio reloads a model during a long index run, csearch retries transient embed failures for up to 120 seconds before aborting. You can tune this with CODE_SEARCH_EMBED_RETRY_MS and CODE_SEARCH_EMBED_RETRY_DELAY_MS. Health probes are treated as ordering hints, not gates: if /v1/models is slow under load but /v1/embeddings still works, indexing continues.
The long retry budget applies to indexing only. A search_code query embeds with a short fail-fast budget (~4s, no retry), so a down or reloading endpoint never hangs an interactive search — it falls back to lexical-only results instead.
Search Ranking And Reliability
Hybrid ranking (RRF). Each result is scored by both vector similarity and a lexical match, then merged with reciprocal rank fusion (
k = 60). Embedding scores tend to cluster in a narrow band; fusing in the lexical rank lets an exact identifier/symbol hit rise above a semantically-close-but-wrong chunk.Lexical fallback. If the embedding endpoint is unreachable at query time,
search_codestill returns lexical-ranked results and adds avector search unavailable, lexical-only resultswarning to the response diagnostics.Wrong-model guard. The embeddings response is checked against the configured
dims. If a host silently serves a different model (a common LM Studio behavior when the requested model is not loaded), the dimension mismatch is rejected with an explicit error instead of writing corrupt vectors into the index.Optional cross-encoder rerank. Add a
rerankblock to the embedding config to enable a second-stage reranker over the RRF top-N:{ "embed": { /* ... */ }, "rerank": { "url": "http://127.0.0.1:8081", "model": "qwen3-reranker-0.6b", "path": "/v1/rerank", "topN": 40, "timeoutMs": 6000, "maxDocChars": 1200 } }The endpoint must speak the Jina/Cohere
/v1/rerankschema (for example llama.cppllama-server --rerank). Any rerank failure (endpoint down, timeout, bad response) falls back to the RRF order and records arerank skippeddiagnostics warning, so enabling it never makes search less reliable than stage one.
Commands
Install or Repair
csearch install
csearch install --all
csearch install --ide cursor
csearch fix --allinstall and fix sync the MCP runtime, configure the selected embedding runner/model, repair IDE integrations, install agent guidance, repair repo-local bootstrap files, repair git hooks, and run an incremental index.
Install prints progress for each step, including package-manager checks, runner/model setup, runtime sync, IDE configuration, repository bootstrap, and index refresh. It also reports whether RTK is installed. RTK is recommended for lower-token MCP response accounting and fallback shell output but is not required for code-search to work; when RTK is unavailable, code-search uses the built-in RTK-style compaction policy.
Automatic freshness is handled by two trigger types:
- Git hooks in each bootstrapped repository run incremental indexing after
post-commit,post-merge,post-checkout, andpost-rewrite. - Cursor
afterFileEdithooks run a debounced incremental index for edited files. The hook uses the neutral~/.csearch/code-searchruntime.
Uninstall
csearch uninstall --yes
csearch uninstall --all --yesuninstall removes the global runtime, global metrics, IDE MCP integrations, agent guidance, managed git hook blocks, repo-local .claude/mcp-code-search index artifacts, and generated repo .mcp.json code-search entries.
By default it cleans the current repository plus global runtime/IDE state. Use --all to also clean discovered repositories.
Status
csearch status
csearch doctor
csearch projectsstatus and doctor report the current repository's code-search MCP registration, MCP mode, edit hook, and index metadata. Cursor can run code-search as a global direct MCP server; in that mode tool calls accept project_root so the server can select the intended repo-local index even when ${workspaceFolder} is stale or ambiguous. If root selection is ambiguous or mismatched and project_root is missing, the MCP call fails closed instead of returning results from a potentially wrong project.
OpenCode support writes a local MCP entry to ~/.config/opencode/opencode.jsonc:
{
"mcp": {
"code-search": {
"type": "local",
"command": ["node", ".claude/mcp-code-search/server.mjs"],
"enabled": true
}
}
}The relative command intentionally goes through the repo-local shim, so each
OpenCode session uses the current project's own .claude/mcp-code-search
index.
projects lists discovered repo-local .claude/mcp-code-search indexes. Pass a scan root to narrow discovery, or set CODE_SEARCH_PROJECT_ROOTS to a path-delimited root list.
Token Savings
csearch gain
csearch savings
csearch gain --repoBy default, gain and savings read global stats from ~/.csearch/metrics.jsonl.
Metrics are recorded inside the MCP server before responses are returned to the IDE agent. Events are written for search_code, get_file, and list_exports, and include the MCP client name when the IDE provides it. A second RTK-aware compression metric records raw response tokens, output response tokens, and response tokens saved without increasing the tool-call count. Root diagnostics (rootSource, selected repoRoot, and warnings) are included so index-mixing reports can be audited.
Example output:
Code-Search Token Savings (Global: 1/1 projects)
==============================================================================
Total tool calls: 1 (search_code + get_file + list_exports)
Tokenizer: js-tiktoken/cl100k_base
AI input tokens: 8.3K
AI output tokens: 1.2K
Index saved: 6.5K
RTK-aware saved: 600 (1 compacted / 1 checked)
Tokens saved: 7.1K (85.5%)
Efficiency meter: █████████████████████████████░░░░░ 85.5%
By Tool
--------------------------------------------------------------------------------------------------------------------------------------
# Tool Count Idx In Raw Out RTK In RTK Out AI Out Saved Avg% Impact
--------------------------------------------------------------------------------------------------------------------------------------
1. get_file 1 8.3K 1.8K 1.8K 1.2K 1.2K 7.1K 85.5% ████████████
--------------------------------------------------------------------------------------------------------------------------------------In the tables, Saved is computed from the displayed data as Idx In - AI Out.
RTK-aware saved is only the response compaction delta (RTK In - RTK Out).
Use --repo to filter the report to the current repository. Legacy repo-local metrics are still read for repo-scoped reports.
Reset Metrics
csearch reset
csearch reset --global --yes
csearch metrics resetreset clears efficiency/gain statistics only: search counts, input tokens, output tokens, and saved token totals. Index artifacts are preserved.
Reindex
csearch reindex
csearch reindex --clean
csearch reindex --dry-runreindex rebuilds the current repository's index and preserves metrics. By default it is resumable: it keeps the existing index-meta.json / index-vectors.bin, reuses vectors for unchanged files, and only re-embeds what is new or changed. If a long run is interrupted (for example by an embedding endpoint reload), rerunning csearch reindex picks up where it left off instead of re-embedding everything. The rebuild writes atomically and refuses to reuse vectors when the embedder identity (provider/model/dims/maxInputTokens) changed, so switching embedding model automatically triggers a correct full rebuild without deleting first.
The file set follows .gitignore: inside a git repo, csearch indexes tracked + untracked-but-not-ignored files (git ls-files --cached --others --exclude-standard), so ignored build output, caches, and tool dumps are never embedded. Outside git it falls back to a filtered filesystem walk.
Use csearch reindex --clean to wipe the artifacts up front and rebuild from scratch (for example after index corruption).
Under the hood, full rebuilds checkpoint progress to disk every 25 freshly embedded files. Transient endpoint failures are retried for embed.retryMs (default 120s); a 4xx rejection (wrong model name, auth, or malformed request) fails fast with a configuration-focused message instead of the reload hint.
The health probe (/v1/models, healthTimeoutMs default 800ms) is treated as a hint for ordering candidates, never as a gate. Under heavy load a remote endpoint can answer embeddings while the health probe momentarily times out; in that case indexing keeps using the endpoint instead of aborting with a false "unreachable" error. Reachability is decided by the embeddings call itself: indexing only stops when the actual /v1/embeddings request keeps failing past embed.retryMs.
Runtime Layout
Global runtime and stats:
~/.csearch/
code-search/
metrics.jsonlPer-repository index:
<repo>/.claude/mcp-code-search/
index-meta.json
index-vectors.bin
server.mjs
config.jsonTroubleshooting
- Run
csearch fix --allif an IDE is not seeing the MCP server. - Start a new IDE agent session after changing MCP config; long-lived MCP processes can keep old runtime state.
- Run
csearch reindexif file counts or chunk counts look stale. - Run
csearch doctor package-managerbefore runner/model installation work. - On Windows, if
wingetreports an MS Store source error while installing Ollama, runwinget source update --name winget, then reruncsearch install. The installer pins runner installs to--source winget. - On Windows, after a fresh Ollama install, the current terminal may not have the updated PATH yet.
csearchchecks common Ollama install paths, but if model pull still cannot find Ollama, open a new terminal and reruncsearch install.
Development
npm install
npm test
npm run check