@shikhil9447/ai-context-engine
v0.1.3
Published
Agent-independent AI Context Engine: project discovery, static indexing, branch-aware knowledge, and context retrieval for existing AI coding agents (Claude Code, Cursor, Antigravity, VS Code). No AI API key required for the core workflow.
Maintainers
Readme
AI Context Engine
An agent-independent context engine that helps existing AI coding agents (Claude Code, Cursor, Antigravity, VS Code extensions) understand a project faster — without changing how you talk to them. You keep saying things like "Implement voucher type and series number for Proforma Invoice" to your normal agent; this tool builds and maintains the project understanding underneath that.
Don't make the AI remember more. Make the AI need to remember less. Actual source code is always the source of truth. AI knowledge is an accelerator, not the source of truth.
See AI_IMPLEMENTATION.md for the full implementation record — what's built, what's a deliberate stub, and why — and validation/v1-final-audit.md for a requirement-by-requirement audit against the V1 spec.
Install & usage
Published on npm — no local checkout needed:
npx @shikhil9447/ai-context-engine init # discover + index the current project
npx @shikhil9447/ai-context-engine refresh # re-index (alias: index)
npx @shikhil9447/ai-context-engine doctor # diagnose project/index/knowledge/lock health(The CLI command itself is ai-context; the npm package name is scoped as @shikhil9447/ai-context-engine because the unscoped name was already taken.) This is a dev-time tool, not a runtime dependency — don't add it to your application's package.json. To hack on this repo directly instead, npm run build then run node dist/cli/index.js <command>.
No AI API key is required for any of this. init/refresh/doctor are fully functional offline, using only static analysis of your own repositories. An AI API key only matters if you explicitly opt into the optional semantic analysis provider (see below) — the normal workflow never needs one.
Claude Code skill
The package ships a ready-made Claude Code skill — it teaches Claude when to run init/refresh/doctor and how to read the external state before working on a task, without you having to explain any of this yourself each time.
Fastest way to get it: copy the block below straight into a new file at .claude/skills/ai-context-engine/SKILL.md in your project (or ~/.claude/skills/ai-context-engine/SKILL.md for every project) — no install needed just to grab the skill itself.
---
name: ai-context-engine
description: Sets up, refreshes, or diagnoses the AI Context Engine (@shikhil9447/ai-context-engine) for a coding project, and consults its external project intelligence before implementing substantial tasks. Use this skill whenever the user asks to "set up ai-context-engine", "init the context engine", "run ai-context init/refresh/doctor", asks whether a project "has context engine set up" or "has AGENTS.md configured for AI context", or reports stale/missing project context. Also consult this skill proactively — without being asked — at the start of any non-trivial implementation task (new feature, bug fix touching unfamiliar code, cross-repo change) in a project where an AGENTS.md, CLAUDE.md, or similar file contains an "ai-context-engine:start" marker, or where `~/.ai-engineer/projects/` contains a matching project directory: read the external context there before exploring the codebase from scratch, since it already has an evidence-backed index of files, symbols, routes, database tables, and relationships that would otherwise take many tool calls to rediscover.
---
# AI Context Engine
`@shikhil9447/ai-context-engine` gives coding agents fast, evidence-backed understanding of a project — indexed files/symbols/imports, HTTP routes, database tables, cross-file relationships, and derived knowledge claims — stored **outside** the repository. It does not replace you; it accelerates you. Source code is always the final authority: if the external state disagrees with what you read in a file, trust the file and treat the external state as stale.
This skill covers two situations: **setting up or maintaining** the tool, and **using its output** when working on a project that already has it.
## Setting up / maintaining a project
Three commands, run from the project root. No AI API key is required for any of them.
```bash
npx @shikhil9447/ai-context-engine init # one-time setup
npx @shikhil9447/ai-context-engine refresh # re-index after changes (incremental; alias: "index")
npx @shikhil9447/ai-context-engine doctor # health check
```
- **`init`**: discovers the project's repositories (single or multi-repo), tech stack, and Git branch(es); indexes files/symbols/imports/routes/DB tables; derives evidence-backed knowledge (entrypoints, module structure, route/table surface); and writes or updates a small instruction block in whichever agent config file already exists (`CLAUDE.md`, `.cursorrules`, etc.) or `AGENTS.md` as a fallback. Run this once per project.
- **`refresh`**: re-indexes incrementally (only changed files are reprocessed). Suggest this after a substantial round of edits, or if `doctor` reports a stale index.
- **`doctor`**: reports whether the project is initialized, whether the index/knowledge are stale, missing repositories, branch-policy violations, lock issues, and whether an optional semantic analysis provider is configured. Run this when something seems off, or when the user asks "is this set up correctly?"
For a multi-repo project, an optional `ai-context.yaml` in the project root can declare repositories explicitly (useful when auto-discovery's one-level-deep heuristic isn't enough):
```yaml
project:
name: MyProject
repositories:
- id: frontend
path: ../my-frontend
type: frontend
- id: backend
path: ../my-backend
type: backend
- id: common
path: ../my-common
type: shared
branch_policy: dev-only
```
Deeper semantic knowledge (architecture judgment, business concepts, workflow narratives — beyond the mechanical facts `init` always derives) is available by opting into `analysis.provider: anthropic` in that same config file plus an `ANTHROPIC_API_KEY` environment variable. This is entirely optional — mention it if the user wants richer knowledge, but don't suggest it's required.
## Using an already-set-up project
Before diving into unfamiliar code for a real task, check whether the project has this set up: look for an `ai-context-engine:start` marker in `AGENTS.md`/`CLAUDE.md`/similar, or a matching directory under `~/.ai-engineer/projects/`. If it's there, that instruction block tells you the exact external state path — read from it selectively, the same way its own generated instructions describe:
- `config.yaml` — which repositories make up this project
- `metadata/snapshot.json` — per-repository tech stack and current Git branch/dirty state
- `indexes/files.json`, `indexes/symbols.json`, `indexes/imports.json` — what exists in the codebase, with file+line evidence
- `indexes/routes.json`, `indexes/db-tables.json` — detected HTTP routes and database tables/models
- `relationships/relationships.json` — evidence-backed relationships between files/repos (e.g. cross-repo `depends_on`)
- `knowledge/<category>/*.json` — derived claims, each with `evidence`, `confidence`, and `status`. **Check `status` before trusting a claim** — `STALE` means the source it was based on has since changed; treat it as a hint to verify, not as fact.
- `decisions/*.json` — recorded architectural/design decisions with context and rejected alternatives, if any exist
Pull in only what's relevant to the task at hand — don't read every file in the external state for a small change. This mirrors the tool's own design principle: point at what's relevant, don't dump everything.
**Branch awareness**: in a multi-repo project, never assume one repository's branch applies to another — `metadata/snapshot.json` records each repository's branch independently. A `knowledge` entry scoped to `BRANCH` or `CLIENT` is relevant only on matching branches; treat it as reference material otherwise, not something to copy blindly onto a different branch/client.
**Keep it current**: if you make substantial changes to indexed files during the task, suggest (or run) `npx @shikhil9447/ai-context-engine refresh` afterward so the index reflects the new state for the next task.If you're already installing the package anyway, the same file also lives at skills/ai-context-engine/SKILL.md inside it:
npm install --no-save @shikhil9447/ai-context-engine # or: npx to fetch it temporarily
mkdir -p .claude/skills/ai-context-engine
cp node_modules/@shikhil9447/ai-context-engine/skills/ai-context-engine/SKILL.md .claude/skills/ai-context-engine/What init / refresh do
- Discover the project root and any nested Git repositories (or read an explicit
ai-context.yaml). - Detect technology stacks (Node/TypeScript/JavaScript/React/Python) per repository via marker files.
- Detect Git state per repository: current branch, dirty state, recent commits, base-branch guess. Never assumes one repo's branch equals another's; warns on
branch_policy: dev-onlyviolations. - Apply security exclusions (
.env*, keys, credentials,node_modules,dist, DB dumps, etc.) before anything is indexed or read. - Index: files, a lightweight JS/TS symbol index (exported functions/classes), import edges (resolved against the file index), HTTP routes (Express-style), and database tables/models (SQL migrations, Sequelize, Mongoose, Prisma) — all evidence-backed (file + line).
- Derive relationships: intra-repo
importsedges and cross-repodepends_onedges (frompackage.json), each with confidence and evidence. - Record knowledge: mechanical facts (entrypoints, tech stack, top-level module directories, route/table surface) always; plus, if you've configured a semantic
AnalysisProvider(see below), deeper architecture/workflow/business-concept claims. Every entry is aKnowledgeEntrywith evidence, confidence, and status. - Update agent instructions: writes a small, idempotent, marker-delimited block into whichever agent config file already exists in the project (
CLAUDE.md,.cursorrules, etc., orAGENTS.mdas a fallback) pointing at the external state — never dumping the knowledge base itself. - Write everything to external state at
~/.ai-engineer/projects/<project-id>/— never into the application repo.
refresh does all of the above incrementally: files are hashed, and only new/changed files have their symbols and import edges recomputed — unchanged files' prior entries are carried over.
External state layout
~/.ai-engineer/projects/<project-id>/
├── config.yaml # project + repositories
├── .lock # advisory lock while init/refresh is running (auto-reclaimed if stale)
├── metadata/snapshot.json # per-repo tech + git snapshot
├── indexes/
│ ├── files.json # path, extension, size, language, content hash
│ ├── symbols.json # exported functions/classes, file+line evidence
│ ├── imports.json # import/require edges, resolved where possible
│ ├── routes.json # detected HTTP routes
│ ├── db-tables.json # detected DB tables/models
│ └── manifest.json # counts, timestamp, incremental-refresh stats
├── relationships/relationships.json # evidence-backed relationships between files/repos
├── knowledge/{architecture,modules,workflows,business,constraints,shared,branches}/
│ # one JSON file per KnowledgeEntry, each with evidence + confidence + status
├── decisions/ # one JSON file per DecisionRecord (context, alternatives, consequences)
├── tasks/{active,completed}/ # one JSON file per task: requirements, checkpoints, decisions
├── history/YYYY-MM-DD.jsonl # append-only log of init/refresh/knowledge-conflict events
├── validation/ # reserved for future validation reports
└── agent/ # reserved for future agent-specific stateMulti-repository / explicit config
Drop an ai-context.yaml in the project root to declare repositories explicitly (useful when auto-discovery's one-level-deep heuristic isn't enough, or to set a branch_policy):
project:
name: ERPForce
repositories:
- id: frontend
path: ../erpforce-fe
type: frontend
- id: backend
path: ../erpforce-be
type: backend
- id: common
path: ../erpforce-common-hub-fe
type: shared
branch_policy: dev-only # warns in init/doctor if worked on outside "dev"
exclude:
- some-extra-dir-to-skipKnowledge, retrieval, and task state (programmatic API)
Beyond the CLI, src/core/index.ts exports a programmatic API meant to be driven by whichever agent is doing the work (Section 3 of the spec deliberately keeps the CLI to init/refresh/doctor — no task-entry command):
retrieveContext(query, sources)— keyword/relationship/branch-aware relevance ranking over files, symbols, relationships, and knowledge. Returns ranked items with human-readablereasons.createTask/updateRequirement/completeTask— persistent, resumable task state with requirement traceability.completeTaskthrows if any requirement is unimplemented or unvalidated.createKnowledgeEntry/checkKnowledgeFreshness— evidence-enforced knowledge with staleness detection against current source.runReflection(task, repoSnapshots)— mechanically-checkable subset of the Section 36 reflection checklist (requirement completion, branch correctness, working-tree cleanliness); semantic questions (duplicate logic, missed consumers, etc.) are returned withanswer: nullrather than silently skipped, since they need a reasoning agent.runSemanticReflection(mechanicalReport, provider, input)— additive, async: fills in thenullsemantic questions using aReflectionProvider, without ever overwriting a mechanical answer. With no provider (nullReflectionProvider), the report is unchanged.writeKnowledgeWithConflictDetection(store, candidates, headCommitByRepo)— writes reflection/analysis-discovered knowledge; if a candidate shares evidence with an existing entry but asserts a different claim, the old entry is markedSTALE(never silently overwritten) rather than ignored or duplicated.createDecisionRecord/supersedeDecision— dedicated decision records (context, alternatives considered, rejected alternatives, consequences), separate fromKnowledgeEntry.createConstraint— first-class negative knowledge ("Common package only uses dev.") — an evidence-enforcedKnowledgeEntrywithcategory: "constraints".findSimilarImplementations(name, symbols)— surfaces existing symbols whose name shares word-tokens with a proposed one, so a task like "add voucher series to Proforma Invoice" can findgenerateInvoiceSeriesNumberas prior art before writing something new.retrieveContext'smaxApproxTokens— context budgeting: trims the ranked result set to an approximate token budget, highest-scored items first.
Analysis providers (semantic knowledge/reflection)
Deep architecture/workflow/business-concept analysis and semantic reflection need a model in the loop. V1 does not require one — AnalysisProvider/ReflectionProvider are plain interfaces, and the default is always the dependency-free, no-network nullProvider/nullReflectionProvider:
interface AnalysisProvider {
name: string;
analyze(input: AnalysisInput): Promise<CreateKnowledgeInputLike[]>;
}A real, opt-in implementation ships in the box: AnthropicAnalysisProvider / AnthropicReflectionProvider, backed by the Anthropic Messages API. To enable it:
# ai-context.yaml
analysis:
provider: anthropic # default: "heuristic" (no semantic provider)
model: claude-sonnet-5 # optional; also settable via AI_CONTEXT_MODELexport ANTHROPIC_API_KEY=sk-...With no key set, AnthropicAnalysisProvider.analyze() simply returns [] — mechanical/heuristic knowledge always runs regardless, so the tool is fully useful offline. Every candidate the provider returns is schema-validated and cross-checked against the actual indexed files — a claim whose evidence cites a file that isn't in the index is dropped, never trusted at face value. doctor reports whether the configured provider is ready, without ever making a live API call or printing the credential itself.
To use a different vendor, implement AnalysisProvider/ReflectionProvider yourself and pass it via runInit({ ..., analysisProvider }) / runRefresh({ ..., analysisProvider }) — the core has no vendor-specific code outside analysis/resolve-provider.ts.
Security
.env*, *.pem/*.key/*.pfx/*.p12, credential/secret/password-named files, SSH private keys, .npmrc, DB dumps, and vendor/build directories (node_modules, dist, build, coverage, .git, .venv, etc.) are excluded before any file is opened — not just omitted from output. Configurable per-project via exclude: in ai-context.yaml.
Troubleshooting
Run ai-context doctor. It reports:
- whether the project is initialized and
config.yamlis valid - missing/non-Git repositories and
branch_policyviolations - whether the index is missing or stale (>7 days), and whether every configured repository was actually indexed
- whether any knowledge entries have gone stale relative to current source
- whether the configured analysis provider is ready (config-only check, never a live API call)
- whether a lock from a previous (possibly crashed) run is still present, and whether it's stale enough to be auto-reclaimed
Development
npm run build # tsc -> dist/
npm run dev # tsx src/cli/index.ts (no build step)
npm test # vitest run
npm run test:watchV2 extension points
Not implemented by design — see AI_IMPLEMENTATION.md Section 28 and validation/v1-final-audit.md for the full list and reasoning: full API/DB call-chain resolution beyond existence+evidence, wiring retrieveContext/task-state into a live agent turn (no target agent exposes a documented local API for this), embeddings/vector retrieval, file watchers, ERPForce validation against a real checkout, and cross-project/organization-wide knowledge.
Testing
npm test86 tests across discovery, Git/branch detection, security exclusions, static indexing, API/DB graph detection, knowledge lifecycle & staleness, knowledge-conflict detection, decision records, context retrieval & budgeting, similar-implementation search, task lifecycle, agent adapters, the concurrency lock, mechanical + semantic reflection, the analysis-provider pipeline (including fail-safe and no-hallucinated-evidence behavior, via a stubbed fetch — no real network calls), and full init → doctor → refresh end-to-end flows.
