@dtranllc/kb-genie
v1.1.4
Published
Cursor plugin that turns raw technical documents into an agent-optimized knowledge base. This package is a companion CLI that initializes a knowledge-base folder template.
Downloads
44
Maintainers
Readme
KB Genie — Agentic Workflow Plugin
Transform raw technical documents into an agent-optimized hierarchical knowledge base with semantic chunks, structured summaries, a living concept wiki, and quality-gated indexing.
What It Does
- Converts raw documents to clean Markdown — normalizes heading hierarchy, removes headers/footers/page numbers/OCR noise, preserves code blocks and tables
- Produces structured document summaries — extracts key claims, methods, results, limitations, relevance assessment, and search tags into YAML frontmatter
- Splits documents into semantic chunks — divides content on heading boundaries (not arbitrary token limits), enriches each chunk with summary, keywords, entities, semantic key, and potential questions
- Maintains a living concept wiki — extracts concepts from all documents, creates new wiki entries and updates existing ones as new documents are ingested
- Keeps an authoritative master index — single YAML catalog listing every processed document with metadata, file paths, chunk counts, and concept links
- Builds a knowledge graph (optional) — extracts entities and relations from chunks to produce a structured graph in JSON format
- Runs quality spot-checks — samples 20% of documents (minimum 3), validates all metadata fields, checks for near-duplicate chunks, verifies index completeness
Agents
| Agent | Model | Role |
|-------|-------|------|
| @kb-orchestrator | sonnet | Plans ingestion runs, spawns workers, monitors progress, maintains index.yaml, produces status reports |
| @kb-ingestion | sonnet | Detects new/changed files in raw/, converts to clean Markdown, extracts bibliographic metadata |
| @kb-summarizer | sonnet | Produces structured document-level summaries with YAML frontmatter |
| @kb-chunker | sonnet | Splits Markdown into semantic chunks with rich metadata enrichment |
| @kb-concept-distiller | sonnet | Maintains living wiki under concepts/ — creates and updates concept pages |
| @kb-indexer | fast | Keeps index.yaml authoritative — scans outputs, updates catalog |
| @kb-graph-builder | sonnet | Extracts entities and relations from chunks to produce knowledge graph JSON |
| @kb-critic | sonnet | Quality spot-checks of summaries, chunks, concept pages, and index |
Skills
Canonical skill files live only under plugins/kb-genie/skills/.
| Skill | When to use |
|-------|-------------|
| kb-genie | Chat starts with Genie,, Genie!, or Genie: — answer from the knowledge base via kb-rag only |
| kb-ingest | Run the full ingestion pipeline on raw/ |
| kb-check | Quality spot-check after an ingestion run |
| kb-index-rebuild | Rebuild stale or incomplete index.yaml |
Commands
| Command | What it does |
|---------|--------------|
| /kb-ingest | Run the full ingestion pipeline |
| /kb-check | Quality spot-check |
| /kb-index-rebuild | Rebuild index.yaml from output directories |
| /kb-retrieve | Retrieve ranked, cited context via kb-rag |
Quick Start
Install the Cursor plugin (Import from Repo)
This repository is a Cursor Team Marketplace. Cursor reads .cursor-plugin/marketplace.json and loads the plugin from plugins/kb-genie/.
- Push this repository to GitHub.
- In Cursor, open Dashboard → Plugins → Team Marketplaces.
- Choose Import from Repo and paste the GitHub URL.
- Review the parsed
kb-genieplugin, then save the marketplace.
Test locally (optional)
Copy the plugin directory into Cursor’s local plugins folder. Use a real directory — Cursor rejects external symlinks.
mkdir -p ~/.cursor/plugins/local
rm -rf ~/.cursor/plugins/local/kb-genie
cp -R plugins/kb-genie ~/.cursor/plugins/local/kb-genieThen restart Cursor or run Developer: Reload Window.
For Genie chat and /kb-retrieve, install the bundled kb-rag CLI:
pip install -e plugins/kb-genie/skills/kb-genie/tools/kb-ragKnowledge-base folder template
npx @dtranllc/kb-genie only copies a knowledge-base/ folder template. It does not install the Cursor plugin.
npx @dtranllc/kb-genie initIngest Documents
- Create a knowledge base directory and place your documents in
raw/:
knowledge-base/
└── raw/
├── whitepaper-2026.pdf
├── spec-api-v2.docx
└── notes-meeting-2026.md- Open Cursor and invoke the orchestrator:
@kb-orchestrator
Knowledge base root: /path/to/knowledge-base/
Ingest all new files in raw/The orchestrator will run all 7 specialist agents automatically, validate outputs at each phase, and produce a final status report.
Individual Agent Invocation
You do not have to run the full pipeline every time:
@kb-chunker
Knowledge base root: /path/to/knowledge-base/
Documents: whitepaper-2026, spec-api-v2@kb-critic
Knowledge base root: /path/to/knowledge-base/@kb-indexer
Knowledge base root: /path/to/knowledge-base/Knowledge Base Folder Structure
knowledge-base/
├── raw/ # IMMUTABLE originals (never modified by agents)
├── markdown/ # Clean full-document Markdown
├── summaries/ # Document-level summaries (YAML frontmatter)
├── chunks/ # Semantic chunks (one .md per chunk)
├── concepts/ # Living wiki (one .md per concept)
├── graphs/ # Knowledge graph JSON (optional)
├── index.yaml # Master catalog
├── logs/
│ └── ingestion-runs/
│ ├── run-YYYYMMDD-NNN.log
│ └── run-YYYYMMDD-NNN-quality.md
└── tasks/
├── pending/
├── in-progress/
└── completed/CLI Reference
npx @dtranllc/kb-genie # Show usage
npx @dtranllc/kb-genie init # Copy knowledge-base template to current directory
npx @dtranllc/kb-genie info # Show agent inventoryPipeline
Input: knowledge base root directory + (optional) file list
↓
[Stage 1] kb-orchestrator → run log + task queue
↓
[Stage 2] kb-ingestion → markdown/<doc_id>.md + summaries/<doc_id>.meta.yaml
↓
[Stage 3] kb-summarizer → summaries/<doc_id>.md
↓
[Stage 4] kb-chunker → chunks/<chunk_id>.md
↓
[Stage 5] kb-concept-distiller → concepts/<concept-slug>.md
↓
[Stage 6] kb-indexer → index.yaml
↓
[Stage 7] kb-graph-builder → graphs/knowledge-graph.json (optional, parallel)
[Stage 7] kb-critic → run log quality report (parallel)
↓
Final Status Report → UserDependencies
# Required CLI tools for document conversion:
# Pandoc — universal document converter (macOS: brew install pandoc, Linux: apt-get install pandoc)
# pdftotext — PDF text extraction (macOS: brew install poppler, Linux: apt-get install poppler-utils)
# Optional: Python for embedding-based near-duplicate detection
pip install sentence-transformers numpy
# Optional: kb-rag CLI for Genie chat and /kb-retrieve
pip install -e plugins/kb-genie/skills/kb-genie/tools/kb-ragQuality Gates
| Gate | Requirement | |------|-------------| | Every chunk | non-empty summary | | Every chunk | non-empty semantic_key | | Every chunk | at least one potential_question | | Every concept page | cites at least one source chunk | | Every concept page | non-empty Definition section | | Every document summary | non-empty relevance_to_software | | Near-duplicate chunks | cosine similarity > 0.95 → merge or flag | | index.yaml | lists every processed document |
License
MIT
