npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@muvon/octocode

v0.26.2

Published

AI-powered code indexer with semantic search and knowledge graphs

Readme

Structural Code Intelligence for AI Agents — MCP Server + Knowledge Graph + Semantic Search

CI Coverage Crates.io GitHub stars License Rust Release

Give your AI assistant a brain for your codebase. Octocode transforms your project into a navigable knowledge graph that Claude, Cursor, and other AI agents can search, understand, and navigate.

🚀 Quick Start🤖 MCP Integration📖 Documentation🌐 Website


🤖 Built for AI Agents

The Problem: AI assistants are blind to your codebase. They can't search your files, understand dependencies, or remember context across sessions.

The Solution: Octocode's MCP server gives AI agents:

  • 🔍 Semantic search — Find code by meaning, not keywords
  • 🕸️ Knowledge graph — Navigate imports, calls, and dependencies
  • 📝 Code signatures — View structure without reading entire files
  • 🧭 LSP precision — Go-to-definition, find-references, and hover docs via your language server

Works with: Claude Desktop • Cursor • Windsurf • Any MCP-compatible AI

Now your AI assistant can:

You: "Where is authentication handled?"
AI: *searches your codebase* "Authentication is in src/middleware/auth.rs,
    which imports jwt.rs for token validation and calls user_store.rs for lookup."

You: "What files depend on the payment module?"
AI: *queries knowledge graph* "src/api/handlers/payment.rs imports payment/mod.rs,
    which is also used by src/workers/refund.rs and src/cron/billing.rs"

You: "Find every call site of this function"
AI: *uses LSP find-references* "process_payment() is called from 4 places:
    checkout.rs:87, refund.rs:134, billing.rs:56, and tests/payment_test.rs:23"

🤔 Why Octocode?

Standard RAG treats your code as flat text chunks. It finds similar-sounding snippets but has no idea that auth_middleware.rs imports jwt.rs, calls user_store.rs, and is wired into router.rs. Octocode understands structure.

# Semantic search finds the right code
octocode search "authentication middleware"
→ src/middleware/auth.rs — Similarity: 0.9234

# The GraphRAG CLI queries the optional persisted graph
octocode config --graphrag-enabled true
octocode index
octocode graphrag get-relationships --node-id src/middleware/auth.rs
Outgoing:
  imports → jwt (src/auth/jwt.rs): token validation logic
  calls   → user_store (src/db/user_store.rs): user lookup by token
Incoming:
  imports ← router (src/router.rs): wires auth into the request pipeline

Octocode uses tree-sitter AST parsing to build a live graph of files, symbols, imports, calls, inheritance, and implementations. The MCP graphrag tool builds this graph lazily from the current source tree, without an index, embeddings, or an LLM. Optional indexed GraphRAG adds semantic file discovery, descriptions, and broader architectural relationships.

🔬 How It Works

Current Source → Tree-sitter AST → Live Symbol Graph ──────────────→ MCP `graphrag`
                                           ↑                              ↑
Indexed Code → Embeddings + Optional LLM → Persisted File Enrichment ─────┘
  1. Live AST Graph — tree-sitter extracts file and symbol nodes plus deterministic contains, imports, calls, extends, and implements relationships directly from current source
  2. Always-on Graph Navigation — MCP graph lookup, relationship traversal, path finding, and overview work with [graphrag].enabled = false
  3. Optional Enrichment — enabling indexed GraphRAG overlays semantic file matches, LLM descriptions, and broader file-level architectural relationships; symbols are never embedded or LLM-generated
  4. Hybrid Search — semantic similarity + BM25 full-text search + reranking handles meaning-based code retrieval separately
  5. MCP Server — exposes semantic_search, view_signatures, graphrag, and structural_search to any MCP-compatible client

✨ What Makes It Different

| | Standard RAG | Doc Lookup Tools | Octocode | |---|---|---|---| | Indexes | Text chunks | External library docs | Your codebase structure (AST) | | Understands | Similar text | API specs & usage | Functions, imports, dependencies | | Cross-file | No | No | Yes — navigates the dependency graph | | Relationships | No | No | imports, calls, implements, extends... | | AI integration | Varies | MCP | Native MCP server + LSP |

Doc tools give AI the manual for libraries you use. Octocode gives AI the blueprint of how you put them together.

Built with Rust for performance. Local-first for privacy. Open source (Apache 2.0) for transparency.

📊 Retrieval Quality

Octocode ships a reproducible retrieval benchmark (benchmark/): 127 curated code-search queries with line-range ground truth, run against octocode's own source (pinned at b1771ba so annotations never drift). The numbers below use a fully local, no-API-key stack — jina-embeddings-v2-base-code via fastembed, no reranker — so they are a floor, not a ceiling:

| Config | Hit@5 | Hit@10 | MRR | NDCG@10 | Recall@10 | |---|---|---|---|---|---| | Dense vector only | 0.598 | 0.717 | 0.485 | 0.528 | 0.671 | | Hybrid, default RRF weights (0.7/0.3) | 0.598 | 0.717 | 0.485 | 0.528 | 0.671 | | Hybrid, keyword-tuned (0.3/0.7) | 0.732 | 0.835 | 0.572 | 0.620 | 0.807 |

Tilting RRF fusion toward the BM25/keyword signal — which carries disproportionate weight for code's exact identifiers — lifts Hit@5 by +22% and Recall@10 by +20% at zero added cost.

The benchmark also flags what doesn't help here (full 6-variant matrix in benchmark/RESULTS.md): a generic local cross-encoder reranker (bge-reranker-base) actually regressed results (Hit@5 0.732 → 0.598) — code retrieval needs a code-aware reranker (e.g. voyage:rerank-2.5), not an off-the-shelf one.

git clone https://github.com/Muvon/octocode && cd octocode
git worktree add /tmp/corpus b1771ba   # pin the corpus to the ground-truth commit
CORPUS=/tmp/corpus python3 benchmark/run_matrix.py   # set OCTO_BIN to use a custom binary

See benchmark/README.md for methodology and metric definitions.

🚀 Quick Start

1. Install

# Universal installer (Linux, macOS, Windows)
curl -fsSL https://raw.githubusercontent.com/Muvon/octocode/master/install.sh | sh

# macOS with Homebrew
brew install muvon/tap/octocode
# From crates.io
cargo install octocode

# Or from source (latest)
cargo install --git https://github.com/Muvon/octocode

# Download binary from releases
# https://github.com/Muvon/octocode/releases

See Installation Guide for platform-specific instructions.

2. Set Up API Keys

# Optional: embedding provider (defaults to local FastEmbed — no key needed)
export VOYAGE_API_KEY="your-voyage-api-key"

# Optional: LLM for commit messages, code review
export OPENROUTER_API_KEY="your-openrouter-api-key"

Get your Voyage API key: voyageai.com (free tier available)

Octocode supports multiple embedding providers:

# OpenAI
export OPENAI_API_KEY="your-key"
octocode config --code-embedding-model "openai:text-embedding-3-small"

# Jina AI
export JINA_API_KEY="your-key"
octocode config --code-embedding-model "jina:jina-embeddings-v3"

# Google
export GOOGLE_API_KEY="your-key"
octocode config --code-embedding-model "google:text-embedding-005"

See API Keys guide for all supported providers.

3. Index Your Codebase

cd /your/project
octocode index
# → ✓ Indexing complete! 342 of 342 files processed (342 new, 0 unchanged)

4. Search Your Code

# Natural language search
octocode search "authentication middleware"

# Multi-query for broader results
octocode search "auth" "middleware" "session"

# Filter by language
octocode search "database connection pool" --language rust

# Search commit history
octocode search "authentication refactor" --mode commits

5. Connect Your AI Assistant

Add to your MCP client config (Claude Desktop, Cursor, Windsurf):

{
  "mcpServers": {
    "octocode": {
      "command": "octocode",
      "args": ["mcp", "--path", "/your/project"]
    }
  }
}

Done! Your AI assistant now understands your codebase structure.

🔌 MCP Server Integration

Octocode includes a built-in MCP server that exposes your codebase as tools to AI assistants. This is the primary way to use Octocode — give your AI assistant direct access to search and navigate your code.

Available Tools

| Tool | What It Does | |------|--------------| | semantic_search | Find code by meaning — "authentication flow", "error handling", "database queries" | | view_signatures | View file structure — function signatures, class definitions, imports | | graphrag | Always-on file/symbol graph — search nodes, inspect relationships, and find paths without indexing | | structural_search | AST pattern matching — find .unwrap() calls, new instantiations, specific patterns | | lsp_goto_definition | Jump to a symbol's definition (requires --with-lsp) | | lsp_find_references | Find all usages of a symbol across the workspace (requires --with-lsp) | | lsp_hover | Type info and documentation for a symbol (requires --with-lsp) | | lsp_document_symbols / lsp_workspace_symbols / lsp_completion | File symbols, workspace-wide symbol search, completions (requires --with-lsp) |

Enable the LSP tools by starting the server with your language server:

octocode mcp --path /your/project --with-lsp="rust-analyzer"

Conversational AI Examples

Once connected, your AI assistant can answer questions about your codebase:

You: "Where is user authentication implemented?"
AI: *uses semantic_search* "Found in src/auth/login.rs. The authenticate() function
    validates credentials against the database, generates a JWT token, and stores
    the session in Redis."

You: "What files depend on the payment module?"
AI: *uses graphrag* "src/api/handlers/payment.rs imports payment/mod.rs, which is also
    used by src/workers/refund.rs and src/cron/billing.rs. The payment module exports
    process_payment() and validate_transaction() functions."

You: "Show me all error handling in the API layer"
AI: *uses structural_search* "Found 23 error handling patterns in src/api/:
    - 15 use Result<T, ApiError> with explicit error types
    - 8 use .unwrap() (potential panics in handlers/user.rs:42, handlers/auth.rs:87)
    - 3 use .expect() with custom messages"

Quick Setup

Octomind (Recommended) — Zero setup, Octocode pre-configured:

curl -fsSL https://raw.githubusercontent.com/muvon/octomind/master/install.sh | bash
octomind run developer:rust

Claude Code (CLI) — Command-line setup:

claude mcp add octocode -- octocode mcp --path /path/to/your/project

Claude Desktop / Cursor / Windsurf — Add to config:

{
  "mcpServers": {
    "octocode": {
      "command": "octocode",
      "args": ["mcp", "--path", "/path/to/your/project"]
    }
  }
}

Config locations:

  • Claude Desktop: ~/Library/Application Support/Claude/claude_desktop_config.json (macOS)
  • Cursor: ~/.cursor/mcp.json or Settings → MCP Servers
  • Windsurf: Settings → MCP

📖 Complete MCP Client Setup Guide — Detailed instructions for 15+ clients including VS Code (Cline/Continue), Zed, Replit, and more.

🌐 Supported Languages

17 languages with full tree-sitter AST parsing:

| Language | Extensions | Features | |----------|------------|----------| | Rust | .rs | Full AST parsing, pub/use detection, module structure | | Python | .py | Import/class/function extraction, docstring parsing | | TypeScript/JavaScript | .ts, .tsx, .js, .jsx | ES6 imports/exports, type definitions | | Go | .go | Package/import analysis, struct/interface parsing | | PHP | .php | Class/function extraction, namespace support | | C++ | .cpp, .cc, .cxx, .c++, .c, .h, .hpp, .hxx, .cppm, .ixx, .mxx, .ccm, .cxxm | Include analysis, class/function extraction, C++20 module support | | Ruby | .rb | Class/module extraction, method definitions | | Elixir | .ex, .exs | Module/protocol extraction, function and macro definitions | | Java | .java | Import analysis, class/method extraction | | Swift | .swift | Class/struct/protocol extraction, import analysis | | Svelte | .svelte | Component structure, script/style block extraction | | Lua | .lua | Function and table extraction | | CSS | .css, .scss, .sass | Rule and selector extraction | | JSON | .json | Structure analysis, key extraction | | Bash | .sh, .bash | Function and variable extraction | | Markdown | .md, .markdown | Document section indexing, header extraction |

📚 Documentation

🔒 Privacy & Security

  • 🏠 Local-first — fully local embedding via fastembed (no API key required); cloud providers optional
  • 🔐 Secure — API keys stored locally, env vars supported
  • 🚫 Respects .gitignore — Never indexes sensitive files
  • 🛡️ MCP security — Local-only server, no external network for search
  • 📤 Cloud-safe — cloud providers receive only the code chunks being embedded; use local models for fully offline indexing

We measure semantic search quality using a hand-annotated ground truth dataset of 254 queries (127 code + 127 docs) with precise line-range annotations. Each query has 1–3 expected results scored by relevance.

These numbers use the full cloud stack — contextual retrieval, Voyage reranker, RaBitQ quantization — on commit b1771ba with benchmark config. For the fully local baseline and the complete variant matrix, see Retrieval Quality above and benchmark/RESULTS.md.

| Metric | Score | |--------|-------| | Hit@5 | 0.929 (118/127) | | Hit@10 | 0.953 (121/127) | | MRR | 0.776 | | NDCG@10 | 0.801 | | Recall@5 | 0.902 | | Recall@10 | 0.921 |

Missed queries (6 of 127):

| # | Query | Expected | Got (top 1) | |---|-------|----------|-------------| | 43 | how to set up MCP proxy for managing multiple repositories | doc/MCP_INTEGRATION.md:286-311 | doc/MCP_INTEGRATION.md:286-4 | | 51 | what are the prerequisites before using octocode | doc/GETTING_STARTED.md:6-12 | doc/CONTRIBUTING.md:7-33 | | 59 | what to do when hitting API rate limits | doc/GETTING_STARTED.md:209-216 | doc/PERFORMANCE.md:304-356 | | 75 | typical performance metrics for small medium and large projects | doc/PERFORMANCE.md:4-13 | doc/PERFORMANCE.md:414-14 | | 112 | how to install octocode on different operating systems | INSTALL.md:4-14 | INSTALL.md:49-70 | | 115 | how to fix macOS Gatekeeper blocking the binary | INSTALL.md:199-206 | INSTALL.md:198-119 |

| Metric | Score | |--------|-------| | Hit@5 | 0.992 (126/127) | | Hit@10 | 0.992 (126/127) | | MRR | 0.895 | | NDCG@10 | 0.906 | | Recall@5 | 0.962 | | Recall@10 | 0.974 |

Missed queries (1 of 127):

| # | Query | Expected | Got (top 1) | |---|-------|----------|-------------| | 105 | how does the system ensure two developers get the same database path | src/storage.rs:60-83 | src/mcp/proxy.rs:631-644 |

Metrics: Hit@k (did the answer appear?), MRR (how high?), NDCG@10 (are best results ranked first?), Recall@k (how many found?). See benchmark/ for methodology, scoring script, and the full dataset.

🤝 Community & Support

⚖️ License

Apache License 2.0 — See LICENSE for details.


Built with 🦀 Rust by Muvon in Hong Kong

⭐ Star🍴 Fork📣 Share