npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

magector

v2.23.0

Published

Semantic code search for Magento 2 — index, search, MCP server

Downloads

4,315

Readme

Magector

Technology-aware MCP server for Magento 2 and Adobe Commerce with intelligent indexing and search.

Magector is a Model Context Protocol (MCP) server that deeply understands Magento 2 and Adobe Commerce. It builds a semantic vector index of your entire codebase — 18,000+ files across hundreds of modules — and exposes 48 tools that let AI assistants search, navigate, and understand the code with domain-specific intelligence. Instead of grepping for keywords, your AI asks "how are checkout totals calculated?" and gets ranked, relevant results in under 50ms, enriched with Magento pattern detection (plugins, observers, controllers, DI preferences, layout XML, and 20+ more).

Rust Node.js Magento Adobe Commerce Accuracy License: MIT


Why Magector

Magento 2 and Adobe Commerce have 18,000+ PHP, XML, JS, PHTML, and GraphQL files spread across hundreds of modules. The codebase relies heavily on indirection — plugins intercept methods defined in other modules, observers react to events dispatched elsewhere, di.xml rewires interfaces to concrete classes, and layout XML stitches blocks and templates together. No single file tells the full story.

Generic search tools — grep, IDE search, or the keyword matching built into AI assistants — can't bridge this gap. They find literal strings but can't connect "how does checkout calculate totals?" to TotalsCollector.php when the word "totals" appears in hundreds of unrelated files.

Magector solves this with three layers of intelligence:

  1. Semantic vector index — every file is embedded into a 384-dimensional space (ONNX, all-MiniLM-L6-v2) where meaning matters more than keywords. A search for "payment capture" returns CaptureOperation.php because the embeddings are close, not because the file contains the word "capture".

  2. Magento technology awareness — 20+ pattern detectors identify plugins, observers, controllers, blocks, cron jobs, GraphQL resolvers, DI preferences, layout XML, and more. Every search result is enriched with what kind of Magento component it is, so the AI client understands the code's role in the system.

  3. Adaptive learning (SONA) — Magector tracks which results you actually use and adjusts future rankings with MicroLoRA feedback, getting smarter over time without any API calls.

The result: your AI assistant calls one MCP tool and gets ranked, pattern-enriched results in 10-45ms — instead of burning tokens grepping through dozens of wrong files. High relevance accuracy means the AI reads fewer, more targeted files, which optimizes context window usage, reduces API costs, and accelerates development cycles.

| Approach | Semantic matches | Magento-aware | Speed (18K files) | |----------|:---------------------:|:---------------------------:|:-----------------:| | grep / ripgrep | No | No | 100-500ms | | IDE search | No | No | 200-1000ms | | GitHub search | Partial | No | 500-2000ms | | Magector | Yes | Yes | 10-45ms |


Features

  • Semantic search -- find code by meaning, not exact keywords
  • 99.2% accuracy -- validated with 101 E2E test queries across 16 tool categories, plus 557 Rust-level test cases
  • Hybrid search -- combines semantic vector similarity with keyword re-ranking for best-of-both-worlds results
  • Structured JSON output -- results include file path, class name, methods list, role badges, and content snippets for minimal round-trips
  • Persistent serve mode -- keeps ONNX model and HNSW index resident in memory, eliminating cold-start latency
  • Incremental re-indexing -- background file watcher detects changes and updates the index without restart (tombstone + compact strategy)
  • ONNX embeddings -- native 384-dim transformer embeddings via ONNX Runtime
  • 36K+ vectors -- indexes the complete Magento 2 / Adobe Commerce codebase including framework internals
  • Magento-aware -- understands controllers, plugins, observers, blocks, resolvers, repositories, and 20+ Magento patterns
  • Adobe Commerce compatible -- works with both Magento Open Source and Adobe Commerce (B2B, Staging, and all Commerce-specific modules)
  • AST-powered -- tree-sitter parsing for PHP and JavaScript extracts classes, methods, namespaces, and inheritance
  • Cross-tool discovery -- tool descriptions include keywords and "See also" references so AI clients find the right tool on the first try
  • SONA feedback learning -- self-adjusting search that learns from MCP tool call patterns (e.g., search → find_plugin refines future rankings for similar queries)
  • SONA v2 with MicroLoRA + EWC++ -- rank-2 low-rank adapter (1536 params, ~6KB) adjusts query embeddings based on learned patterns; Elastic Weight Consolidation prevents catastrophic forgetting during online learning
  • Diff analysis -- risk scoring and change classification for git commits and staged changes
  • Complexity analysis -- cyclomatic complexity, function count, and hotspot detection across modules
  • Fast -- 10-45ms queries via persistent serve process, batched ONNX embedding with adaptive thread scaling
  • LLM description enrichment -- generate natural-language descriptions of di.xml files using Claude, stored in SQLite, and prepend them to embedding text so descriptions influence vector search ranking (not just post-retrieval display)
  • MCP server -- 48 tools integrating with Claude Code, Cursor, and any MCP-compatible AI tool
  • Clean architecture -- Rust core handles all indexing/search, Node.js MCP server delegates to it

Architecture

flowchart LR
  subgraph node ["Node.js Layer"]
    direction TB
    G["CLI<br/>init · index · search · describe"]
    E["MCP Server<br/>48 tools · LRU cache"]
    F["Persistent Serve Process"]
    G --> F
    E --> F
  end

  F -->|"stdin/stdout JSON"| rust

  subgraph rust ["Rust Core"]
    direction TB
    A["AST Parser<br/>PHP · JS · XML"]
    B["Pattern Detection<br/>20+ Magento patterns"]
    B2["Description Enrichment<br/>LLM-powered di.xml summaries"]
    C["ONNX Embedder<br/>all-MiniLM-L6-v2 · 384d"]
    D["HNSW Vector Search<br/>hybrid reranking · SONA"]
    A --> B --> B2 --> C --> D
  end

  style rust fill:#f4a460,color:#000
  style node fill:#68b684,color:#000

Indexing Pipeline

flowchart LR
  A["Source File"] --> B["AST Parser"]
  B --> C["Pattern Detection"]
  C --> D["Text Enrichment"]
  D --> D2{"Descriptions DB?"}
  D2 -->|Yes| D3["Prepend LLM Description"]
  D2 -->|No| E["ONNX Embedding"]
  D3 --> E
  E --> F[("HNSW Index")]
  A --> G["Metadata"] --> F

Search Pipeline

flowchart LR
  Q["Query"] --> E1["Synonym Enrichment"]
  E1 --> E2["ONNX Embedding"]
  E2 --> H["HNSW Search"]
  H --> R["Hybrid Reranking"]
  R --> SA["SONA Adjustment"]
  SA --> J["Structured JSON"]

Components

| Component | Technology | Purpose | |-----------|-----------|---------| | Embeddings | ort (ONNX Runtime) | all-MiniLM-L6-v2, 384 dimensions | | Vector search | exact cosine scan up to 300k vectors, hnsw_rs above + hybrid reranking | Nearest neighbors + keyword boosting | | PHP parsing | tree-sitter-php | Class, method, namespace extraction | | JS parsing | tree-sitter-javascript | AMD/ES6 module detection | | Pattern detection | Custom Rust | 20+ Magento-specific patterns | | CLI | clap | Command-line interface (index, search, serve, validate) | | Unified metadata | rusqlite (bundled SQLite) | LLM descriptions, method-chain enrichment, process state, cache — all in .magector/data.db | | SONA | Custom Rust | Feedback learning with MicroLoRA + EWC++ | | MCP server | @modelcontextprotocol/sdk | AI tool integration with structured JSON output | | Config data | JSON exports in .magector/config-data/ | One-time core_config_data exports per environment for config tracing | | PHP scan snapshot | JSON in .magector/php-scan.json (0600) | Class hierarchy and event dispatch sites read once per root and shared by its MCP instances, checked by file stamps |


Security

Magector operates on source code indexed from potentially-untrusted vendor/ dependencies and is driven by an LLM that may be manipulated via prompt injection in indexed comments, docblocks, or markdown. The following hardening applies as of v2.15.1:

Indexed markdown

The only markdown in the index is the project's own module READMEs, app/code/<Vendor>/<Module>/README.md: written next to that code and trusted like its comments. Markdown from vendor/ packages and anywhere else is never indexed, so magento_search does not hand third-party prose to the LLM.

Path traversal protection

All tools that accept a path argument (magento_read, magento_grep, magento_ast_search, magento_find_dataobject_issues) route the input through safePath() / safeRelPath() helpers in src/mcp-server.js. These:

  1. Resolve the argument against MAGENTO_ROOT with path.resolve() (normalizes .., symlinks are not followed during validation).
  2. Reject any resolved path that does not lie inside MAGENTO_ROOT.

This prevents a hostile vendor/ comment from instructing the LLM to e.g. magento_read ../../home/user/.ssh/id_rsa. Both the standalone case handlers and their magento_batch counterparts share the same chokepoint.

Shell injection hardening in auto-update

src/update.js fetches the latest field from the npm registry and re-execs itself with the new version string. Previously this was interpolated into a shell command; a tampered registry response could inject shell metacharacters. As of v2.15.1:

  • The re-exec passes argv as an array to a no-shell spawner (no intermediate shell).
  • A semver-strict isSafeVersion() validator rejects any version string containing metacharacters or that does not match X.Y.Z / X.Y.Z-prerelease form.
  • Fails closed: the auto-update is silently skipped rather than run a malformed version.

Unix socket permissions

The serve-proxy Unix socket at .magector/serve.sock is created with chmod 0600 immediately after listen(). On multi-user systems, another local account can no longer connect and query the vector index (which would leak indexed source snippets). The chmod is best-effort on platforms that don't support it (logged to .magector/magector.log).

Reporting vulnerabilities

If you find a security issue, please open an issue on the GitHub repo and mark it as security-related. Do not post reproducers that leak actual source contents from private codebases.


Quick Start

Prerequisites

1. Initialize in Your Project

cd /path/to/your/magento2  # or Adobe Commerce project
npx magector init

This single command handles the entire setup:

flowchart LR
  A["npx magector init"] --> B["Verify<br/>Project"]
  B --> C["Download<br/>ONNX Model"]
  C --> D["Index<br/>Codebase"]
  D --> E["Detect IDE<br/>Cursor · Claude Code"]
  E --> E2["API Key<br/>(optional)"]
  E2 --> F["Write MCP<br/>Config"]
  F --> G["Update<br/>.gitignore"]

2. Search

npx magector search "product price calculation"
npx magector search "checkout totals collector" -l 20

3. Re-index After Changes

npx magector index

4. IDE Setup Only (Skip Indexing)

npx magector setup

CLI Reference

Rust Core CLI

magector-core <COMMAND>

Commands:
  index       Index a Magento codebase
  search      Search the index semantically
  serve       Start persistent server mode (stdin/stdout JSON protocol)
  describe    Generate LLM descriptions for di.xml files (requires ANTHROPIC_API_KEY)
  validate    Run validation suite (downloads Magento if needed)
  download    Download Magento 2 Open Source
  stats       Show index statistics
  embed       Generate embedding for text

index

magector-core index [OPTIONS]

Options:
  -m, --magento-root <PATH>          Path to Magento root directory
  -d, --database <PATH>              Index database path [default: ./.magector/index.db]
  -c, --model-cache <PATH>           Model cache directory [default: ./models]
      --descriptions-db <PATH>       Path to descriptions SQLite DB (descriptions are prepended to embeddings)
  -v, --verbose                      Enable verbose output

When --descriptions-db is provided (or auto-detected as data.db next to the index), descriptions are prepended to the embedding text as "Description: {text}\n\n" before the raw file content. This places semantic terms within the 256-token ONNX window, significantly improving retrieval of di.xml files for natural-language queries.

search

magector-core search <QUERY> [OPTIONS]

Options:
  -d, --database <PATH>   Index database path [default: ./.magector/index.db]
  -l, --limit <N>         Number of results [default: 10]
  -f, --format <FORMAT>   Output format: text, json [default: text]

describe

magector-core describe [OPTIONS]

Options:
  -m, --magento-root <PATH>   Path to Magento root directory
  -o, --output <PATH>         Output SQLite database [default: ./.magector/data.db]
      --force                 Re-describe all files (ignore cache)

Generates natural-language descriptions of di.xml files using the Anthropic API (Claude Sonnet). Requires ANTHROPIC_API_KEY environment variable. Descriptions are stored in a SQLite database and used during indexing to enrich embeddings. Only files with changed content hashes are re-described (incremental by default).

serve

magector-core serve [OPTIONS]

Options:
  -d, --database <PATH>            Index database path [default: ./.magector/index.db]
  -c, --model-cache <PATH>         Model cache directory [default: ./models]
  -m, --magento-root <PATH>        Magento root (enables file watcher)
      --descriptions-db <PATH>     Path to descriptions SQLite DB
      --watch-interval <SECS>      File watcher poll interval [default: 60]

Starts a persistent process that reads JSON queries from stdin and writes JSON responses to stdout. Keeps the ONNX model and HNSW index resident in memory for fast repeated queries.

When --magento-root is provided, a background file watcher polls for changed files every --watch-interval seconds and incrementally re-indexes them without restart. Modified and deleted files are soft-deleted (tombstoned) in the HNSW index; new vectors are appended. When tombstoned entries exceed 20% of total vectors, the index is automatically compacted by rebuilding the HNSW graph.

Protocol (one JSON object per line):

// Request:
{"command":"search","query":"product price","limit":10}

// Response:
{"ok":true,"data":[{"id":123,"score":0.85,"metadata":{...}}]}

// Stats request:
{"command":"stats"}

// Watcher status:
{"command":"watcher_status"}
// Response:
{"ok":true,"data":{"running":true,"tracked_files":18234,"last_scan_changes":3,"interval_secs":60}}

// Descriptions (all LLM descriptions from SQLite DB):
{"command":"descriptions"}
// Response:
{"ok":true,"data":{"app/code/Magento/Catalog/etc/di.xml":{"hash":"...","description":"...","model":"claude-sonnet-4-5-20250929","timestamp":1769875137},...}}

// Describe (generate descriptions + auto-reindex affected files):
{"command":"describe"}
// Response:
{"ok":true,"data":{"files_found":371,"described":5,"skipped":366,"errors":0,"described_paths":["app/code/..."]}}

// SONA feedback:
{"command":"feedback","signals":[{"type":"refinement_to_plugin","query":"checkout totals","timestamp":1700000000000}]}
// Response:
{"ok":true,"data":{"learned":1}}

// SONA status:
{"command":"sona_status"}
// Response:
{"ok":true,"data":{"learned_patterns":5,"total_observations":12}}

Node.js CLI

npx magector init [path]        # Full setup: index + IDE config
npx magector index [path]       # Index (or re-index) Magento codebase
npx magector search <query>     # Search indexed code
npx magector describe [path]    # Generate LLM descriptions for di.xml files
npx magector stats              # Show indexer statistics
npx magector setup [path]       # IDE setup only (no indexing)
npx magector snapshot save <file> [path]   # Index, then archive the index + PHP scan
npx magector snapshot load <file> [path]   # Restore an archive into a root
npx magector snapshot info <file>          # Version, root and entries of an archive
npx magector mcp                # Start MCP server
npx magector help               # Show help

The describe command and magento_describe MCP tool require an Anthropic API key. During npx magector init, you are prompted to paste your key (optional). If provided, it is stored in the MCP config file as the ANTHROPIC_API_KEY environment variable so the MCP server can use it automatically. You can also set it manually later by adding "ANTHROPIC_API_KEY": "sk-..." to the env section in .mcp.json or ~/.cursor/mcp.json.

Environment Variables

| Variable | Description | Default | |----------|-------------|---------| | MAGENTO_ROOT | Path to Magento installation | Current directory | | MAGECTOR_DB | Path to index database | $MAGENTO_ROOT/.magector/index.db (index <path>: <path>/.magector/index.db) | | MAGECTOR_BIN | Path to magector-core binary | Auto-detected | | MAGECTOR_MODELS | Path to the ONNX model directory; a missing model is downloaded here | ~/.magector/models/ | | MAGECTOR_INDEX_TIMEOUT | Indexing wall-clock timeout in milliseconds. Override for very large codebases or CPU-constrained environments. | 14400000 (4 h) | | MAGECTOR_THREADS | Max ONNX intra-op + rayon parsing threads. Equivalent to the --threads CLI flag. | Half of CPU cores | | OMP_NUM_THREADS | Fallback thread limit if MAGECTOR_THREADS is not set (de facto standard for ONNX/OpenMP). | — | | MAGECTOR_BATCH_SIZE | Embedding batch size (higher = faster, more RAM). Equivalent to --batch-size. | 256 | | MAGECTOR_MAX_OUTPUT_CHARS | Cap on one MCP tool answer, in characters; a longer answer is cut at a line boundary with a note to narrow the query. | 40000 (~10k tokens) | | MAGECTOR_FILE_LIST_TTL_MS | How long the list of the modules' etc/ files is reused before it is listed again (files added mid-session show up after this); also how often composer's PSR-4 map and classmap are checked for a composer dump-autoload. Modules themselves are rediscovered as soon as app/etc/config.php or the composer registrations change. 0: always fresh. | 2000 | | MAGECTOR_SERVE_WAIT_MS | How long a search waits for the serve process to become ready before it runs a single-shot search instead. | 20000 | | MAGECTOR_PHP_LIST_TTL_MS | How long the list of all PHP files (event dispatchers) is reused before the tree is walked again. | 30000 | | MAGECTOR_AUTO_INDEX | 0: the MCP server never starts an index (none, or an incompatible one) — for CI and agent jobs that bring their own index. The structural tools work without one; semantic search reports it is missing. | 1 (index in the background) | | MAGECTOR_PREWARM_PHP | 1: read the PHP class hierarchy and the event dispatch sites in the background after the MCP server starts, in every instance, so find_event_dispatchers and find_implementors answer in milliseconds; 0: never — their first call reads the tree (seconds on a large project). Unset: in the primary instance only (the one that owns the serve process), except with MAGECTOR_AUTO_INDEX=0, which keeps background CPU off. | primary instance (off with MAGECTOR_AUTO_INDEX=0) | | MAGECTOR_PHP_SNAPSHOT | 0: every MCP instance reads the PHP tree itself. Default: the instance that reads it writes .magector/php-scan.json; the other instances of the same root load it, checked against the tree by the mtime and size of every file it read, instead of reading ~100k files again. | on | | MAGECTOR_PHP_SCAN_WAIT_MS | How long an instance waits for another instance's PHP scan (.magector/php-scan.lock) before it reads the tree itself. | 180000 | | MAGECTOR_PREWARM_CONFIG | 1: prepare the configuration notice of the DI / event answers (every di.xml / events.xml read, each area merged; 0.2–0.6 s on 300–600 modules) in the background, 1.5 s after the MCP server starts with no tool call in flight, so the first DI answer does not wait for it; 0: never — the first DI or event answer prepares it. 1 runs it in every instance. Unset: in the primary instance only, except with MAGECTOR_AUTO_INDEX=0, which keeps background CPU off. Only on a Magento install (app/etc/config.php under MAGENTO_ROOT). | primary instance (off with MAGECTOR_AUTO_INDEX=0) | | MAGECTOR_PHP | Command that runs PHP 8.1+ able to load the Magento root, for the native magento_validate_config (the PHP program is piped to its stdin), e.g. docker exec -i -u www-data <container> php, warden env exec -T php-fpm php, ddev exec php. Only an explicit command is used — never php from PATH: the native check loads the project's autoloader and bootstraps Magento (which writes its DI config to the configured cache; with var/.regenerate present the areas are not read, as bootstrapping would delete generated/ and var/cache). Unset: built-in check. | — | | MAGECTOR_PHP_TIMEOUT_MS | Time limit of one native check (it runs asynchronously; the server keeps answering). | 120000 | | MAGECTOR_PHP_ROOT | The Magento root as MAGECTOR_PHP sees it (inside the container) | MAGENTO_ROOT | | ANTHROPIC_API_KEY | API key for description generation (describe command) | — |

These defaults apply to the Node.js CLI and the MCP server. The Rust core's own -d flag (see above) defaults to ./.magector/index.db in its working directory.

Running in containers

The Linux binaries (x64 and arm64) link libstdc++ statically and need glibc 2.34 or newer: they run in Warden and DDEV PHP containers (CentOS Stream 9 / RHEL 9), Debian 12 and Ubuntu 22.04+. Alpine (musl) is not supported — use a glibc Node image such as node:22-bookworm-slim.

Keep the index and the model beside the code, so they survive container restarts and work offline. Set both variables for every magector command — index, search and the MCP server:

export MAGENTO_ROOT=/var/www/html MAGECTOR_MODELS=/var/www/html/.magector/models
npx -y [email protected] index

Re-running index embeds only files whose content changed: a checkout or docker cp that only rewrote timestamps costs a content-hash comparison, not a re-embed, and a run that changes nothing leaves index.db untouched. Set MAGECTOR_NO_UPDATE=1 in images and CI so the CLI does not re-run itself as the latest npm version.

Build the index once, restore it where it is used

On slow disks (CI runners, network volumes) indexing and the PHP scan that find_implementors / find_event_dispatchers need take longest. Build both where the code is built — an image build — and restore them in the run:

# image build: index, read the PHP tree, write one archive
npx magector snapshot save /opt/magector.tar.gz /var/www/html
# run: restore it (the old index is kept as index.db.bak)
npx magector snapshot load /opt/magector.tar.gz /var/www/html

The archive (tar, gzip for .gz / .tgz) holds index.db, its manifest, the SONA state, the descriptions DB and .magector/php-scan.json. A restore writes everything beside the live files and replaces them only when every entry is complete and matches its sha256 — computed while the archive is read, so nothing is read twice; a snapshot of another Magector version is refused. It can be restored into another path: the vector index stores paths relative to the root, and the PHP scan is rewritten for the new root.

Whether anything is read again depends on the files' mtimes. In an image they are the build's, so neither load --update nor the MCP server reads a file again; load says so (it compares mtime and size, without reading). After a fresh checkout with new mtimes, --update hashes those files to find that their content is the same, and the MCP server reads the PHP tree again (a changed file with classes makes the scan unusable). save --no-index archives the existing index as it is. The models are not in the archive: keep MAGECTOR_MODELS in the image.

Constraining CPU usage during indexing

Indexing a large enterprise codebase (~80K files) can saturate CPU during PHASE 2 (ONNX embedding generation). To keep a developer machine responsive while indexing, lower the thread count:

npx magector index --threads 2                  # use only 2 cores for both parsing and embedding
MAGECTOR_THREADS=2 npx magector index           # equivalent via env var
OMP_NUM_THREADS=2 npx magector index            # also honored as a fallback

The --threads flag and MAGECTOR_THREADS / OMP_NUM_THREADS env vars constrain both the rayon thread pool used by PHASE 1 (parallel AST parsing) and the ONNX intra-op thread pool used by PHASE 2 (embedding inference). The active thread source is logged at startup so you can verify it took effect:

INFO Rayon global pool: 2 threads (available: 16)
INFO ONNX intra_threads: 2 (available: 16, source: --threads flag)

For very large or CPU-constrained runs, you may also need to extend the wall-clock timeout (default 4 hours):

MAGECTOR_INDEX_TIMEOUT=28800000 npx magector index --threads 2   # 8 h timeout, 2 threads

Resume after timeout or interrupt

Indexing writes a crash-safe checkpoint to disk every 50 batches (~12,800 files). If the process is killed or times out mid-run, just re-run npx magector index — it auto-resumes from the last checkpoint:

npx magector index
# ♻️  Resuming from previous run: 38400 vectors across 12200 files already indexed
# ✓ Found 79771 total files; 12200 already indexed, 67571 remaining to process

The indexer collects already-embedded file paths from the existing DB, filters them out of file discovery, preserves the existing HNSW state, and only parses/embeds the files that aren't in the DB yet. Partial resume also picks up new files added to the tree since the previous run.

To force a full rebuild (e.g. after a schema change or if you want to discard stale vectors), pass --force:

npx magector index --force

MCP Server Tools

The MCP server exposes 48 tools for AI-assisted Magento 2 and Adobe Commerce development. All search tools return structured JSON with file paths, class names, methods, role badges, and content snippets -- enabling AI clients to parse results programmatically and minimize file-read round-trips.

Output Format

All search tools return structured JSON:

{
  "results": [
    {
      "rank": 1,
      "score": 0.892,
      "path": "vendor/magento/module-catalog/Model/ProductRepository.php",
      "module": "Magento_Catalog",
      "className": "ProductRepository",
      "namespace": "Magento\\Catalog\\Model",
      "methods": ["save", "getById", "getList", "delete", "deleteById"],
      "magentoType": "repository",
      "fileType": "php",
      "badges": ["repository"],
      "snippet": "class ProductRepository implements ProductRepositoryInterface..."
    }
  ],
  "count": 1
}

Key fields:

  • methods -- list of method names in the class (avoids needing to read the file)
  • badges -- role indicators: plugin, controller, observer, repository, graphql-resolver, model, block
  • snippet -- first 300 characters of indexed content for quick assessment

How exact are the results?

Every answer is one of four kinds. The tables below say which kind each tool returns, and where it can return less than the codebase contains — the case to watch when you search for everything a change affects.

| Kind | Meaning | Safe to conclude "nothing else"? | |------|---------|----------------------------------| | Exact | Structural, from the configuration / source files, resolved the way Magento resolves them at runtime | Yes, within the limits listed below | | Superset (marked) | Every declaration, including ones that do not take effect — each marked: disabled, superseded, module disabled, does not run (with the reason) | Yes — drop what is marked | | Fuzzy | Substring match on a short (non-namespaced) name; may include unrelated types | Yes, but filter the noise — pass a FQCN for an exact answer | | Semantic | Ranked candidates from the vector index | No — use them to find a starting point |

DI and events (pass a fully qualified name)

| Tool | Returns | Can return less when … | |------|---------|------------------------| | magento_find_plugin | DI Plugin Registrations: superset (marked) — declarations on the class, its parents and interfaces, the real class of a virtual type; Effective state: exact merge by plugin name in module load order; per plugin method: does not run when the target method is final / static / non-public / never intercepted / missing, the class is final or implements NoninterceptableInterface. First block: semantic | a parent or interface cannot be read (no PHP file); plugins are added at runtime without di.xml | | magento_find_preference | Exact: effective preference per area, superseded and ignored declarations, the class finally instantiated. Then semantic related files | app/etc/config.php is missing (load order is then approximated from <sequence>) | | magento_find_observer | Superset (marked): every declaration incl. ones that only disable or modify an observer, with area; Effective state: exact merge by observer name | — | | magento_find_event_flow, magento_find_event_dispatchers | Observers: as find_observer. Dispatchers: literal dispatch('event') calls, and names built from a property or class constant of the class that runs the code ($this->_eventPrefix . '_save_after'), resolved per concrete subclass through every trait and parent, from a di.xml constructor argument (area, virtual type); a name partly known only at runtime is listed as possible (controller_action_predispatch_*) | a part computed by a method call or a request value stays *; a dispatch added to a file that had none shows in the next session | | magento_trace_dependency, magento_find_di_wiring | Exact: preferences, plugins (incl. inherited), virtual types resolving to the class (transitively), DI arguments that inject it (through virtual types, preferences, Factory, Proxy); plus the effective states above | the class is only type-hinted in a constructor without a di.xml argument (see impact_analysis) | | magento_impact_analysis | DI references: exact (as above); API exposure: exact (webapi.xml services incl. interface → preference, schema.graphqls resolvers); PHP files: exact FQCN occurrences + semantic candidates; runtime callers: constructor-typed properties | the class is reached through an untyped variable, ObjectManager, or a factory result stored in a local variable | | magento_find_implementors | Exact (FQCN): everything that is instanceof the type — direct implementors, extending interfaces, their implementors and all subclasses, transitively, with the path; preferences per area. Short name: fuzzy (implements naming the short name) | a class in the chain has no readable PHP file (e.g. generated code); a group holds more than 50 classes (the first 50 are listed, the rest counted) |

With a short name (no namespace) the DI tools fall back to fuzzy matching — useful for exploring, not for a complete impact list.

Every tool answer is capped at 40,000 characters (MAGECTOR_MAX_OUTPUT_CHARS); a longer answer ends with Output truncated — narrow the query (full class name, targetMethod, a namespace) rather than read on.

Ambiguity is reported as a problem of the project, not of the tool: when two modules declare the same preference, plugin or observer, neither depends on the other (no <sequence>, no composer require) and swapping them changes the result, the outcome depends on incidental module order and can change with an update.

Other structural answers

| Tool | Returns | Can return less when … | |------|---------|------------------------| | magento_find_table_usage | Superset: every PHP file and db_schema.xml with the table name as a string literal (ResourceModel _init, getTableName, setup patches) + semantic | the table name is built dynamically | | magento_find_controller | Exact: routes.xml → module → controller class (admin routes under Controller/Adminhtml); then semantic | the route is registered by a custom router | | magento_find_api | Exact: webapi.xml of the enabled modules merged as Magento merges it (route by url + method, service class / method attribute by attribute, ACL by ref) — every route whose URL contains the query, whose service method or class equals it, or whose service resolves to the given class; the class that runs (preference); then semantic | routes added at runtime (Magento_WebapiAsync's /async/… variants are noted, not listed) | | magento_find_graphql | Exact over etc/schema.graphqls of the enabled modules, cut into types and merged as GraphQlReader does (types of one name merged in module order, extend type, interface fields copied into object types, a type written in a # comment is read — as Magento reads it); resolver, @cache identity, typeResolver; the schema readers registered besides it are named; then semantic | fields added by the other schema readers (CatalogGraphQl's EAV attribute readers: attributes flagged for GraphQL in the database) — named, not listed | | magento_find_cron | Exact: crontab.xml merged by group + job, then the crontab defaults of config.xml (schedule, config path, run model) over it | schedules or jobs saved in the admin (core_config_data) — noted | | magento_find_db_schema | Exact: db_schema.xml of the enabled modules + app/etc/db_schema.xml merged by table / column / referenceId; columns, keys and indexes with the names Magento creates, disabled elements marked, foreign keys to the table; legacy setup scripts: semantic | tables created by legacy setup scripts or at runtime; a table prefix (names shown without it) | | magento_module_structure | Exact: every file of the module directory (from the module index, not guessed from the name), state and load position, and what the module declares in the merged configuration | — | | magento_find_class | A virtual type name resolves exactly to its di.xml declaration and real class; PHP classes: semantic + filesystem fallback | — | | magento_trace_config | system.xml definition and PHP readers (constant or literal path) | the path is concatenated at runtime | | magento_grep, magento_ast_search | Exact text / AST matches | — |

Configuration Magento rejects

The DI and event tools read what the files say. When Magento cannot load a file, or reads a value differently than written, they say so first: the answers of find_plugin, find_observer, find_preference, find_di_wiring, find_event_flow, trace_dependency, impact_analysis, find_class and batch start with a notice listing the rejected files of enabled modules, and values in the answer's files that do not take effect as written. magento_validate_config gives the details, with Magento's own messages:

| Engine | Checks | Verified | |--------|--------|----------| | native (MAGECTOR_PHP only) | Magento's classes on every file (Config\Dom — not well-formed XML fails in every mode; the DI / events converters and argument interpreters under Magento's ErrorHandler), then Magento's readers (ObjectManager\Config\Reader\Dom, Event\Config\Reader) on every area in production and developer mode — the merged configuration, as Magento validates it; other files against the schema they declare | is Magento | | built-in (no PHP) | per file: the first libxml error (message and line), the converter and argument-interpreter rules ported from Magento, values read differently than written (observer disabled="1", non-integer sortOrder, text next to <item>s); per area: its files merged as Config\Dom merges them, then the same rules — a file another file completes loads, two files that load alone can fail together | against libxml 2.9.14 / Mage-OS 2.4.9 (scripts/verify-magento, mode config): same first fatal error on 2,000 files with syntax edits (1,384 not well-formed); same converter verdict on 3,000 files with value edits (2,518 converter exceptions; const names native only); nothing reported on the 3,075 unmodified config files of the project; merge (mode merge): the same merged document and verdict as Magento's readers on every DI / events area of a Mage-OS 2.4.9 and a Magento 2.4.5 project (30 areas) and on 15,008 merges of mutated files with their originals |

The built-in check does not validate schemas (developer mode) and cannot check const / init_parameter arguments (PHP's defined()). The native check runs on PHP 8.1–8.4 (on Magento 2.4.5 with PHP 8.4 it gives the same result as with PHP 8.1).

The results are static analysis of the files. For a running installation, the object manager configuration read at runtime remains the reference — note that bin/magento dev:di:info lists plugins disabled with disabled="true" as active.

Search Tools

| Tool | Description | |------|-------------| | magento_search | Semantic search -- find any PHP class, method, XML config, template, GraphQL schema, or module README section by natural language | | magento_find_class | Find PHP class, interface, abstract class, or trait by name | | magento_find_method | Find method implementations across the codebase |

Magento-Specific Finders

| Tool | Description | |------|-------------| | magento_find_config | Find XML configuration (di.xml, events.xml, routes.xml, system.xml, webapi.xml, module.xml, layout) | | magento_find_template | Find PHTML template files for frontend or admin rendering | | magento_find_plugin | Find interceptor plugins (before/after/around methods) and di.xml declarations. Resolves plugin PHP files and extracts interceptor method signatures (v2.5) | | magento_find_fieldset | Find fieldset.xml definitions controlling data copy between entities (order→quote, quote→order). Shows fields per aspect (to_order, to_edit) (v2.5) | | magento_find_observer | Find event observers and events.xml declarations | | magento_find_preference | Find DI preference overrides -- which class implements an interface | | magento_find_controller | Find MVC controllers by frontend or admin route path | | magento_find_block | Find Block classes for view rendering | | magento_find_graphql | Find GraphQL schema definitions, resolvers, types, queries, and mutations | | magento_find_api | Find REST/SOAP API endpoints in webapi.xml | | magento_find_cron | Find cron job definitions in crontab.xml | | magento_find_db_schema | Find database table definitions in db_schema.xml (declarative schema) |

Flow & Dependency Tracing

| Tool | Description | |------|-------------| | magento_trace_flow | Trace execution flow from an entry point (route, API, GraphQL, event, cron) -- maps controller → plugins → observers → templates with code snippets (v2.5) | | magento_trace_shipping_chain | Trace the complete shipping rate chain: carriers → collectRates plugins → rate modifiers → totals collectors → fieldset mappings (v2.5) | | magento_trace_dependency | Trace DI graph for a class/interface -- preferences, plugins, virtualTypes, argument overrides (parses all di.xml, no index needed) | | magento_find_event_flow | Trace complete event chain: dispatchers → observers → handler PHP classes (parses events.xml + vector search) | | magento_find_event_dispatchers | Where an event is dispatched: the site, and for a name built from the class that runs the code the class and where its value comes from; * in the name, match: exact / strict / wildcard (v2.3) | | magento_find_layout | Find layout XML files by handle or content -- lists blocks, containers, and referenceBlock declarations | | magento_trace_data_flow | Trace how a data attribute flows: find all setters (magic setter, setData, addData) and getters (magic getter, getData) across PHP and XML (v2.3) | | magento_trace_call_chain | Trace internal method call chain: follows $this->method(), $this->dep->method(), and dispatch() calls to build an execution tree (v2.2) |

Auto-detects entry type from pattern (/V1/... → API, snake_case → event, camelCase → GraphQL, path/segments → route), or override with entryType. Use depth: "shallow" (entry + config + plugins) or depth: "deep" (adds observers, layout, templates, DI preferences).

Impact & Testing

| Tool | Description | |------|-------------| | magento_impact_analysis | Analyze impact of changing a class -- finds use statements, DI references, direct instantiations, and type hints across the codebase | | magento_find_test | Find PHPUnit tests for a given class/method -- searches Test/ directories for coverage, mocks, and assertions | | magento_find_implementors | Find all classes implementing a given PHP interface -- scans implements keywords and di.xml <preference> declarations (v2.2) | | magento_find_callers | Find all call sites of a method across PHP and XML files -- ->method() and ::method() calls (v2.2) | | magento_find_di_wiring | Complete DI picture for a class: preferences, plugins, constructor args, virtual types, and argument overrides from di.xml (v2.2) |

Diagnostics

| Tool | Description | |------|-------------| | magento_error_parser | Parse Magento error messages and map to root cause, affected files, and fix suggestions (10 known patterns) | | magento_performance_profile | Profile a Magento subsystem (checkout_totals, order_place, product_save, etc.) for performance bottlenecks -- plugins, observers, and complexity hotspots | | magento_validate_config | Configuration XML the way Magento loads it: files it rejects in every mode, developer-mode (schema) failures per area, values read differently than written — with Magento's messages. Native through PHP (MAGECTOR_PHP), otherwise built-in. See Configuration Magento rejects |

Analysis Tools

| Tool | Description | |------|-------------| | magento_analyze_diff | Analyze git diffs for risk scoring and change classification | | magento_complexity | Analyze cyclomatic complexity, function count, and line count |

Utility Tools

| Tool | Description | |------|-------------| | magento_module_structure | Show complete module structure -- controllers, models, blocks, plugins, observers, configs | | magento_index | Trigger re-indexing of the codebase (also kicks off background enrichment) | | magento_describe | Generate LLM descriptions for di.xml files (requires ANTHROPIC_API_KEY), stored in .magector/data.db, auto-reindexes affected files | | magento_stats | View index statistics | | magento_batch | Execute multiple tool queries in parallel in one MCP roundtrip. Supports all search, find, grep, read, and null-risk tools. Use to avoid N×3-5s roundtrip overhead. | | magento_grep | Exact text/regex search across PHP/XML/PHTML files (grep -rn -E internally). Supports filesOnly mode (like grep -l), context lines, ignoreCase, include patterns. (v2.9) | | magento_read | Read a specific file with optional methodName extraction (~10× fewer tokens than reading the whole file) and startLine/endLine range. (v2.10) | | magento_trace_api | Trace REST/GraphQL API endpoint from URL to implementation: webapi.xml → service interface → DI preference → method body. One call replaces 4-5 grep+read steps. (v2.11) | | magento_trace_config | Trace a config path end-to-end: system.xml admin definition → PHP classes that consume the value → actual DB values from config-data exports. Accepts exact path or keyword search. (v2.17) | | magento_find_trigger | Find database triggers across the codebase | | magento_find_table_usage | Find all PHP code referencing a specific database table |

Null-Safety Analysis (v2.12–v2.15)

| Tool | Description | |------|-------------| | magento_ast_search | Structural PHP code search using tree-sitter. Named patterns: dataobject-set-null (detect setX(null) anti-pattern), unchecked-method-chain (detect $this->dep->method() chains). Pattern arg is an enum, not free-text. Executed in Rust serve process — no external dependency. (v2.16) | | magento_enrich | Build the method-chain enrichment index. Scans all vendor/ PHP files for ->firstMethod()->secondMethod() chains and detects null guards in surrounding code. Stores results in .magector/data.db (SQLite, via Rust serve). Runs automatically after magento_index. (v2.13, moved to Rust v2.16) | | magento_find_null_risks | Query the enrichment index for method chains without null guards. O(1) SQLite query instead of file scanning. Pass firstMethod to filter (e.g., "getPayment" → all ->getPayment()->anything() without null guard). Requires magento_enrich. (v2.13) | | magento_find_dataobject_issues | Detect setX(null) anti-pattern on Magento DataObject subclasses. setX(null) stores ['x' => null] in _data — hasX() (via array_key_exists) returns true even when the value is null, creating false-positive guard conditions. Use during field-lifecycle audits or when debugging "value persists but shouldn't" bugs. Uses tree-sitter. (v2.15, tree-sitter v2.16) |

Search Enhancements (v2.1)

  • Hybrid BM25+vector search -- combines text frequency scoring with semantic vector similarity for better exact class name matches
  • Query expansion -- automatically expands queries with Magento domain synonyms (plugin → interceptor, checkout → cart/quote/totals, etc.)
  • Module filtering -- moduleFilter parameter on magento_search to limit results by vendor/module pattern. Accepts a single string or array of strings. Supports wildcards, e.g., "Vendor_*" or ["Acme_PaymentGateway", "Acme_FreeShipping"]
  • Non-blocking reindex -- old index stays usable during background rebuild; new index is built to a temp path and swapped in atomically on completion

Deep Code Analysis (v2.2)

  • magento_find_implementors -- find all classes implementing a PHP interface (PHP implements + di.xml <preference>)
  • magento_find_callers -- find all call sites of a method across PHP and XML files
  • magento_find_di_wiring -- complete DI picture: preferences, plugins, constructor args, virtual types, argument overrides
  • magento_trace_call_chain -- trace internal method execution chain: $this->method(), $this->dep->method(), and dispatch() calls with event→observer resolution

Data Flow & Event Tracing (v2.3)

  • magento_trace_data_flow -- trace all setters and getters for a data attribute (magic methods, setData/getData, addData, constants, XML references). Answers "who writes/reads custom_discounted_price_incl_tax on Quote\Address?"
  • magento_find_event_dispatchers -- where an event is dispatched: literal names and names built from the class that runs the code (catalog_product_save_after → AbstractModel::afterSave() for Catalog\Model\Product, $_eventPrefix in Product.php), one line per site and class. Complements magento_find_event_flow.
  • magento_find_plugin area context -- enriched output shows DI area (frontend/adminhtml/global/graphql) and explicit di.xml plugin registrations when targetClass is provided

Tool Cross-References

Each tool description includes "See also" hints to help AI clients chain tools effectively:

graph LR
  cls["find_class"] --> plg["find_plugin"]
  cls --> prf["find_preference"]
  cls --> mtd["find_method"]
  cfg["find_config"] --> obs["find_observer"]
  cfg --> prf
  cfg --> api["find_api"]
  plg --> cls
  plg --> mtd
  tpl["find_template"] --> blk["find_block"]
  blk --> tpl
  blk --> cfg
  dbs["find_db_schema"] --> cls
  gql["find_graphql"] --> cls
  gql --> mtd
  ctl["find_controller"] --> cfg
  trc["trace_flow"] -.-> ctl
  trc -.-> plg
  trc -.-> obs
  trc -.-> tpl
  trc -.-> api
  trc -.-> gql
  dep["trace_dependency"] --> prf
  dep --> plg
  evf["find_event_flow"] --> obs
  imp["impact_analysis"] --> dep
  imp --> cls
  tst["find_test"] --> cls
  err["error_parser"] --> dep
  lay["find_layout"] --> blk

  style cls fill:#4a90d9,color:#fff
  style mtd fill:#4a90d9,color:#fff
  style cfg fill:#e8a838,color:#000
  style plg fill:#d94a4a,color:#fff
  style obs fill:#d94a4a,color:#fff
  style prf fill:#e8a838,color:#000
  style api fill:#e8a838,color:#000
  style tpl fill:#68b684,color:#000
  style blk fill:#68b684,color:#000
  style dbs fill:#9b59b6,color:#fff
  style gql fill:#9b59b6,color:#fff
  style ctl fill:#4a90d9,color:#fff
  style trc fill:#2ecc71,color:#000

Query Examples

magento_search("how are checkout totals calculated")
magento_search("product price with tier pricing and catalog rules")
magento_find_class("ProductRepositoryInterface")
magento_find_method("getById")
magento_find_config("di.xml plugin for ProductRepository")
magento_find_plugin({ targetClass: "Topmenu" })
magento_find_observer("sales_order_place_after")
magento_find_preference("StoreManagerInterface")
magento_find_api("/V1/orders")
magento_find_controller("catalog/product/view")
magento_find_graphql("placeOrder")
magento_find_db_schema("sales_order")
magento_find_cron("indexer")
magento_find_block("cart totals")
magento_find_template("minicart")
magento_analyze_diff({ commitHash: "abc123" })
magento_complexity({ module: "Magento_Catalog", threshold: 10 })
magento_describe()
magento_trace_flow({ entryPoint: "checkout/cart/add", depth: "deep" })
magento_trace_flow({ entryPoint: "/V1/products" })
magento_trace_flow({ entryPoint: "placeOrder", entryType: "graphql" })
magento_trace_flow({ entryPoint: "sales_order_place_after" })
magento_trace_data_flow({ attributeKey: "custom_discounted_price_incl_tax", modelClass: "Quote\\Address" })
magento_find_event_dispatchers({ eventName: "custom_discount_rule_validation_before" })
magento_find_implementors({ interfaceName: "ProductRepositoryInterface" })
magento_find_callers({ methodName: "collectTotals", className: "TotalsCollector" })
magento_find_di_wiring({ className: "CartManagementInterface" })
magento_trace_call_chain({ className: "Magento\\Quote\\Model\\QuoteManagement", methodName: "submit" })

Supported Platforms

Pre-built binaries are provided for the following platforms:

| Platform | Architecture | Package | |----------|-------------|---------| | macOS | ARM64 (Apple Silicon) | @magector/cli-darwin-arm64 | | Linux | x86_64 | @magector/cli-linux-x64 | | Linux | ARM64 | @magector/cli-linux-arm64 | | Windows | x86_64 | @magector/cli-win32-x64 |

Note: macOS Intel (x86_64) is not supported as a pre-built binary. Intel Mac users can build from source.


Validation

Magector is validated at two levels:

  1. E2E MCP accuracy tests -- 101 queries across 16 tool categories via stdio JSON-RPC
  2. Rust-level validation -- 557 test cases across 50+ categories against Magento 2.4.7

E2E Accuracy (MCP Tools)

---
config:
  themeVariables:
    pie1: "#4caf50"
    pie2: "#f44336"
---
pie title Test Pass Rate (101 queries)
  "Passed (101)" : 101
  "Failed (0)" : 0

| Metric | Value | |--------|-------| | Grade | A+ (99.2/100) | | Pass rate | 101/101 (100%) | | Precision | 98.7% | | MRR | 99.3% | | NDCG@10 | 98.7% | | Index size | 35,795 vectors | | Query time | 10-45ms |

Integration Tests

66 integration tests covering MCP protocol compliance, tool schemas, tool calls (including magento_describe), analysis tools, and stdout JSON integrity.

Running Tests

# E2E accuracy tests (101 queries, requires indexed codebase)
npm run test:accuracy
npm run test:accuracy:verbose

# Integration tests (66 tests)
npm test

# SONA/MicroLoRA benefit evaluation (180 queries, baseline vs post-training)
npm run test:sona-eval
npm run test:sona-eval:verbose

# Rust validation (557 test cases)
cd rust-core && cargo run --release -- validate -m ./magento2 --skip-index

Project Structure

magector/
├── src/                          # Node.js source
│   ├── cli.js                    # CLI entry point (npx magector <command>)
│   ├── mcp-server.js             # MCP server (48 tools, structured JSON output)
│   ├── binary.js                 # Platform binary resolver
│   ├── model.js                  # ONNX model resolver/downloader
│   ├── init.js                   # Full init command (index + IDE config)
│   ├── magento-patterns.js       # Magento pattern detection (JS)
│   ├── templates/                # IDE rules templates
│   │   ├── cursorrules.js        # .cursorrules content
│   │   └── claude-md.js          # CLAUDE.md content
│   └── validation/               # JS validation suite
│       ├── validator.js
│       ├── benchmark.js
│       ├── test-queries.js
│       ├── test-data-generator.js
│       └── accuracy-calculator.js
├── tests/                        # Automated tests
│   ├── mcp-server.test.js        # Integration tests (64 tests)
│   ├── mcp-accuracy.test.js      # E2E accuracy tests (101 queries)
│   ├── mcp-sona.test.js          # SONA feedback integration tests (8 tests)
│   ├── mcp-sona-eval.test.js     # SONA/MicroLoRA benefit evaluation (180 queries)
│   ├── describe-benefit-eval.test.js  # Description enrichment benefit evaluation
│   └── results/                  # Test result artifacts
│       ├── accuracy-report.json
│       └── sona-eval-report.json
├── platforms/                    # Platform-specific binary packages
│   ├── darwin-arm64/             # macOS ARM (Apple Silicon)
│   ├── linux-x64/                # Linux x64
│   ├── linux-arm64/              # Linux ARM64
│   └── win32-x64/                # Windows x64
├── rust-core/                    # Rust high-performance core
│   ├── Cargo.toml
│   ├── src/
│   │   ├── main.rs               # Rust CLI (index, search, serve, validate)
│   │   ├── lib.rs                # Library exports
│   │   ├── indexer.rs             # Core indexing with progress output
│   │   ├── embedder.rs            # ONNX embedding (MiniLM-L6-v2)
│   │   ├── vectordb.rs            # HNSW vector database + hybrid search + tombstones
│   │   ├── watcher.rs             # File watcher for incremental re-indexing
│   │   ├── ast.rs                 # Tree-sitter AST (PHP + JS)
│   │   ├── magento.rs             # Magento pattern detection (Rust)
│   │   ├── describe.rs            # LLM description generation + SQLite storage
│   │   ├── sona.rs                # SONA feedback learning + MicroLoRA + EWC++
│   │   └── validation.rs          # 557 test cases, validation framework
│   └── models/                   # ONNX model files (auto-downloaded)
│       ├── all-MiniLM-L6-v2.onnx
│       └── tokenizer.json
├── .github/
│   └── workflows/
│       └── release.yml           # Cross-compile + publish CI
├── scripts/
│   └── setup.sh                  # Claude Code MCP setup script
├── config/
│   └── mcp-config.json           # MCP server configuration template
├── package.json
├── .gitignore
├── LICENSE
└── README.md

How It Works

1. Indexing

Magector scans every .php, .js, .xml, .phtml, and .graphqls file in a Magento 2 or Adobe Commerce codebase, plus each module README at app/code/<Vendor>/<Module>/README.md:

  1. AST parsing -- Tree-sitter extracts class names, namespaces, methods, inheritance, and interface implementations from PHP and JavaScript files
  2. Pattern detection -- Identifies Magento-specific patterns: controllers, models, repositories, plugins, observers, blocks, GraphQL resolvers, admin grids, cron jobs, and more
  3. Search text enrichment -- Combines AST metadata with Magento pattern keywords to create semantically rich text representations
  4. Description enrichment -- If a descriptions SQLite DB is present, LLM-generated natural-language descriptions are prepended to the embedding text as "Description: {text}\n\n", placing semantic DI concepts (preferences, plugins, virtual types, subsystem names) within the 256-token ONNX window
  5. Embedding -- ONNX Runtime generates 384-dimensional vectors using all-MiniLM-L6-v2
  6. Indexing -- Vectors are stored in index.db; up to 300k of them a search compares the query with every vector (exact, a few ms, nothing to build at start-up), above that an HNSW graph finds approximate nearest neighbors

2. Searching

  1. Query text is enriched with pattern synonyms (e.g., "controller" adds "action execute http request dispatch")
  2. The enriched query is embedded into the same 384-dimensional vector space
  3. The nearest vectors are found by cosine similarity (every vector compared up to 300k, the HNSW graph above)
  4. Hybrid reranking boosts results with keyword matches in path and search text
  5. SONA adjustment -- MicroLoRA adapts the query embedding based on learned patterns; EWC++ prevents forgetting earlier learning
  6. Results are returned as structured JSON with file path, class name, methods, role badges, and content snippet

3. Persistent Serve Mode

The MCP server spawns a persistent Rust process (magector-core serve) that keeps the ONNX model and HNSW index loaded in memory. Queries are sent as JSON over stdin and responses returned via stdout -- eliminating the ~2.6s cold-start overhead of loading the model per query. Up to 300k vectors it is ready about a second after it starts. If it is not ready, a search waits for it up to MAGECTOR_SERVE_WAIT_MS, then runs a single-shot magector-core search (asynchronous, at most 20 s), so a tool answers inside the MCP client's 60 s. One serve process per project: the first MCP instance starts it and the others join it over .magector/serve.sock — also instances in other containers sharing .magector (an MCP gateway starting one per session), since the lock records the PID namespace of its holder.

flowchart LR
  subgraph startup ["Startup (once)"]
    S1["Load Model"] --> S2["Load Index"] --> S3["Ready Signal"]
  end
  startup --> query
  subgraph query ["Per Query (10-45ms)"]
    Q1["stdin JSON"] --> Q2["Embed"] --> Q3["HNSW Search"] --> Q4["Rerank"] --> Q5["stdout JSON"]
  end
  subgraph fallback ["Fallback"]
    F1["execFileSync ~2.6s"]
  end

  style startup fill:#e8f4e8,color:#000
  style query fill:#e8e8f4,color:#000
  style fallback fill:#f4e8e8,color:#000

4. File Watcher (Incremental Re-indexing)

When the serve process is started with --magento-root, a background thread polls the filesystem for changes every 60 seconds (configurable via --watch-interval). Changed files are incrementally re-indexed without restarting the server.

Since hnsw_rs does not support point deletion, Magector uses a tombstone strategy: old vectors for modified/deleted files are marked as tombstoned and filtered out of search results. New vectors are appended. When tombstoned entries exceed 20% of total vectors, the HNSW graph is automatically rebuilt (compacted) to reclaim memory and restore search performance.

flowchart LR
  W1["Sleep 60s"] --> W2["Scan Filesystem"] --> W3{"Changes?"}
  W3 -->|No| W1
  W3 -->|Yes| W4["Tombstone Old Vectors"] --> W5["Parse + Embed New Files"] --> W6["Append to HNSW"] --> W7{"Tombstone > 20%?"}
  W7 -->|Yes| W8["Compact / Rebuild HNSW"] --> W9["Save to Disk"]
  W7 -->|No| W9
  W9 --> W1

  style W4 fill:#f4e8e8,color:#000
  style W5 fill:#e8f4e8,color:#000
  style W8 fill:#e8e8f4,color:#000

5. MCP Integration

The MCP server delegates all search/index operations to the Rust core binary. Analysis tools (diff, complexity) use ruvector JS modules directly.

sequenceDiagram
  participant Dev
  participant AI
  participant MCP
  participant Rust
  participant HNSW

  Dev->>AI: "checkout totals?"
  AI->>MCP: magento_search(...)
  MCP->>Rust: JSON query
  Rust->>HNSW: embed + search
  HNSW-->>Rust: candidates
  Rust-->>MCP: JSON results
  MCP-->>AI: paths, methods, badges
  AI-->>Dev: TotalsCollector.php

6. SONA Feedback Learning

The MCP server tracks sequences of tool calls and sends feedback signals to the Rust process. Over time, this adjusts search result rankings based on observed usage patterns.

How it works: The Node.js SessionTracker watches for follow-up tool calls after magento_search. If a user searches and then immediately calls magento_find_plugin, SONA learns that similar queries should boost plugin results. The learned weights are persisted to a .sona file alongside the index.

| MCP Call Sequence | Signal | Effect on Future Searches | |---|---|---| | magento_search → magento_find_plugin (within 30s) | refinement_to_plugin | Boosts plugin results | | magento_search → magento_find_class (within 30s) | refinement_to_class | Boosts class matches | | magento_search → magento_find_config (within 30s) | refinement_to_config | Boosts config/XML results | | magento_search → magento_find_observer (within 30s) | refinement_to_observer | Boosts observer results | | magento_search → magento_find_controller (within 30s) | refinement_to_controller | Boosts controller results | | magento_search → magento_find_block (within 30s) | refinement_to_block | Boosts block results | | magento_search → magento_trace_flow (within 30s) | trace_after_search | Boosts controller results | | magento_search(Q1) → magento_search(Q2) (within 60s) | query_refinement | Tracked for analysis |

Characteristics:

  • Score adjustments are capped at ±0.15 to avoid overwhelming semantic similarity
  • Learning rate decays with repeated observations (diminishing returns)
  • Learned weights are keyed by normalized, order-independent query term hashes
  • Always active -- no feature flags or build-time opt-in required
  • Persisted via bincode to <db_path>.sona (e.g., .magector/index.db.sona)

SONA v2: MicroLoRA + EWC++

SONA v2 adds embedding-level adaptation via a MicroLoRA adapter and Elastic Weight Consolidation:

| Component | Parameters | Purpose | |-----------|-----------|---------| | MicroLoRA | 1536 (rank-2, 2×384×2) | Adjusts query embeddings before HNSW search | | EWC++ | Fisher matrix (384 values) | Prevents catastrophic forgetting during online learning |

  • adjust_query_embedding() applies the LoRA transform + L2 normalization before vector search; cosine similarity guard (≥0.90) skips destructive adjustments
  • learn_with_embeddings() updates LoRA weights from feedback signals with EWC regularization (λ=2000) and decaying learning rate
  • 3-tier scoring with negative learning: positive signals boost the followed feature type, mild negative learning (0.1×) demotes unrelated types
  • V1→V2 persistence format is backward-compatible (auto-upgrades on load)
cd rust-core && cargo build --release

7. LLM Description Enrichment

Magector can generate natural-language descriptions of di.xml files using the Anthropic API and embed them directly into the vector index. This significantly improves search ranking for semantic queries about dependency injection.

Workflow:

# 1. Generate descriptions (one-time, incremental — only re-describes changed files)
ANTHROPIC_API_KEY=sk-... npx magector describe /path/to/magento

# 2. Re-index with descriptions embedded into vectors
npx magector index /path/to/magento

Or via the MCP tool: magento_describe() generates descriptions and auto-reindexes affected files in one step.

How it works: Each di.xml file is sent to Claude Sonnet with a prompt optimized for semantic search retrieval. The resulting description (~70 words) is stored in a SQLite database (.magector/data.db). During indexing, descriptions are prepended to the embedding text as "Description: {text}\n\n" before the raw file content, placing semantic terms (preferences, plugins, virtual types, subsystem names) within the ONNX model's 256-token window.

Measured impact (A/B experiment, 25 queries, Magento 2.4.7, 17,891 vectors, 371 described files):

| Metric | Without Descriptions | With Descriptions | Delta | |--------|---------------------|-------------------|-------| | Precision@K | 1.6% | 20.3% | +18.7% | | MRR | 0.031 | 0.330 | +0.30 | | NDCG@10 | 0.037 | 0.369 | +0.33 | | di.xml results/query | 0.2 | 3.0 | +2.8 | | Query win rate | — | — | 76% |


Magento Patterns Detected

mindmap
  root((Patterns))
    PHP
      Controller