@dikolab/kbdb
v0.16.2
Published
A searchable second brain for AI agents: a file-based knowledge base with ranked keyword and semantic (hybrid) search, as a CLI and MCP server.
Maintainers
Readme
@dikolab/kbdb
A searchable second brain for AI agents. A file-based knowledge base with ranked keyword and semantic (hybrid) search -- learn your documents, then recall the relevant knowledge. No external server. Runs as a CLI and MCP server.
Docs | GitLab | NPM | JSR | License: AGPL-3.0
Status: Beta -- actively developed. Core features (search, recall, MCP) are stable and tested.
What is kbdb?
kbdb gives AI agents a persistent, searchable second brain. Point it at your Markdown docs and it indexes them into a file-based knowledge base -- then agents (and you) recall the most relevant knowledge by ranked keyword and semantic search, not exact-key lookup. It is a living store: agents learn new facts, update them, and recall them across sessions.
No external server to install, no cloud account -- just files on disk. It runs anywhere Node.js or Deno runs, and works as an MCP server, so agents like Claude can plug it in as a memory tool.
How search works: kbdb uses keyword search by default -- synonyms are expanded, terms are ranked by relevance, and headings carry 2× weight in scoring. When an exact query finds nothing, kbdb automatically loosens the match so you still get the best available results.
Want smarter results? Use --algo hybrid to blend
keyword matching with similarity search -- finding
results even when different words describe the same
concept. The default TF-IDF embedding provider
works offline with zero setup. Swap it for a
third-party provider (local ONNX model or remote
API) in worker.toml when you need richer
embeddings.
Knowledge stays fresh: Re-learn a file and kbdb
replaces the old version automatically.
Near-duplicate detection warns you when you are
learning something you already have -- by embedding
similarity, so it catches the same fact reworded, not
just the same bytes. kbdb contradictions reports
sections that cover the same ground so you can read
them together. Integrity checks verify checksums,
orphans and references. Confidence scores help agents
tell strong matches from weak ones.
Getting Started
What You Need
One of these (pick whichever you already have):
- Node.js version 20 or newer -- Download
- Deno version 2.6 or newer --
Download
(2.6 is the floor: the storage engine loads its
WebAssembly through source-phase imports, which
is what lets it run offline after one
deno install. Older Deno fails with a misleadingModule not foundnaming a.wasmfile that is present.)
That's it. No database server. No extra tools.
Install
Using Node.js:
CLI build hosted on NPM.
npm install -g @dikolab/kbdbUsing Deno:
CLI build hosted on JSR.
deno install -Agf jsr:@dikolab/kbdb/cliSee the CLI Installation Guide for prerequisites and verification steps.
Try It Out
1. Create a knowledge base
kbdb db init --db ./my-kbThis creates a .kbdb folder that holds all your
data.
2. Feed it your docs
kbdb learn ./docsPoint it at a folder of Markdown files. kbdb reads
them, breaks them into sections, and builds a
search index. Add --tags design,v2 to tag
sections for scoping, --replace to update
existing sections from the same source, or
--level 2 to set the hierarchical depth
(1 = broadest, 6 = narrowest). When learning a
directory, level is auto-detected from folder
depth.
3. Search
kbdb search "how does auth work"Results are ranked by relevance with snippets
showing where your terms matched. Output defaults
to --format rec (recfile: one field: value per
line) for easy grepping. Other formats: json
(machine-readable), text (numbered list), and
mcp (JSON-RPC 2.0 envelope). Use --offset to
page through large result sets.
To try hybrid search (keyword + AI similarity):
kbdb search "how does auth work" --algo hybridTip:
--dbis optional for the CLI. kbdb walks up from your working directory to the nearest.kbdbfolder, so commands just work anywhere inside a project. Point at a specific base with--db <dir>(the parent of.kbdb), or setKBDB_DB_DIR. Only themcpserver requires an explicit--db-- it never searches the working directory.
Search across bases: enrich results with
read-only knowledge from other databases using
--other-db <dir> (repeatable), or add --cascade
to also pull from .kbdb folders in parent
directories:
kbdb search "how does auth work" \
--other-db ~/shared-kb --cascadeEvery result carries a source_db field -- the
database root it came from -- which you can paste
straight back into --db or --other-db.
Scripting: Add
--format jsonto get structured JSON output for parsing. Use--non-interactiveor setKBDB_NON_INTERACTIVE=1to suppress prompts in CI pipelines.
4. Recall context
kbdb recall <kbid> --depth 1Start with a search result's kbid and expand context progressively: depth 0 gives the section content, depth 1 adds parent documents and back-references, depth 2 adds siblings and forward references, depth 3 includes full text of referenced sections.
Knowledge Base
Build, search, and maintain your knowledge store.
- Import Markdown and plain text files with tags and source tracking
- Smart updates -- re-learning a file supersedes the old version instead of duplicating it
- History -- a superseded section is retired, not
deleted:
kbdb historywalks the chain from either end, and an old kb-id still resolves - Search with three algorithms: keyword (default), AI similarity, or hybrid (both)
- Auto-fallback -- if your exact query finds nothing, kbdb loosens the match automatically
- Recall sections with progressive context --
from a quick summary to full related content,
or as deep as a
--max-tokensbudget allows - Measure whether retrieval is actually any
good --
kbdb evalscores Recall@k, MRR and nDCG@k against your own dataset, and exits non-zero when a change makes ranking worse - Neighbourhood --
kbdb neighbourhoodsays what relates to a section and how: eight typed edges, seven of them recorded facts and one inferred - Consolidate --
kbdb consolidateproposes groups of sections that could become one. It proposes only; you write the merge and apply it yourself - Export -- snapshot your knowledge base for backup
- Verify database integrity and clean up stale data
- Rebuild indexes if anything goes wrong
See the Knowledge Base Guide for the full walkthrough, including export and backup.
Agent Tooling
Integrate kbdb with AI agents and custom tools.
MCP quick-start (Claude CLI):
claude mcp add kbdb -- \
npx @dikolab/kbdb mcp --db /path/to/projectSee the MCP Installation Guide for Claude Code, VS Code, and Claude Desktop config files, plus troubleshooting.
- MCP server with 30 tools -- search, recall, learn, revise, gaps, contradictions, export, skill/agent search, and more
- Skills -- store reusable prompt templates with fill-in-the-blank arguments
- Agents -- create AI agent profiles that combine a persona with skills
- Auto-capture -- the MCP server can proactively suggest knowledge to store from your conversations
- Daemon resilience -- configurable request timeout and automatic retry with daemon respawn
- Worker daemon lifecycle management -- stop and restart the background process
- Granular Deno permissions -- the daemon
runs with scoped permissions instead of
--allow-all - Path confinement -- the daemon rejects
path traversal (
..) in export/import
See the Agent Tooling Guide for MCP setup, skills, agents, and the library API.
For Developers
Library API
Use kbdb programmatically in your Node.js or Deno project:
import { createWorkerClient } from '@dikolab/kbdb';
// Spawns a background worker if not already running
const client = await createWorkerClient({
contextPath: '/path/to/.kbdb',
requestTimeoutMs: 30_000,
});
const results = await client.search({
query: 'authentication',
limit: 10,
offset: 0,
});
console.log(results.items);
client.disconnect();Pass contextPath (the .kbdb directory itself)
or dbPath (the parent directory -- kbdb discovers
.kbdb inside it).
See the Library API Reference for the full API.
Development Setup
git clone https://gitlab.com/diko316/knowledge-base-db.git
cd knowledge-base-db
npm install
npm testDocker
A Docker setup is included with all build tools (Node.js and Deno):
HOST_UMASK=$(umask) docker compose run --rm tool shRun make benchmark to measure search and rebuild
latency at scale -- results are written to
docs/benchmark/benchmark.md
automatically.
See the Makefile for all available build targets.
Contributing
- Fork the repository
- Create a feature branch
- Make your changes and add tests
- Run
npm testandnpm run lint - Open a merge request
Documentation
- CLI Installation Guide -- prerequisites, npm/JSR install, verification
- Knowledge Base Guide -- importing, searching, recall, export
- Agent Tooling Guide -- MCP, skills, agents, library API
- CLI Reference -- full command list with examples
- MCP Installation Guide -- Claude CLI, Claude Code, VS Code, Claude Desktop
- MCP Server Guide -- setup, tools, environment config
- Search and Ranking -- how search works under the hood
- Storage Architecture -- file formats and directory layout
- Benchmark Results -- search and rebuild latency at scale
- Release Notes
- Architecture Overview
The search engine
Storage, indexing and ranking come from
@dikolab/vdb,
kbdb's sibling project by the same author. Its
documentation covers the retrieval side in depth:
- vdb Overview -- storage model, partitions, BM25F, vector and hybrid search
- vdb Examples -- worked queries and ranking behaviour
Support
kbdb is free, AGPL-licensed software. If it earns a place in your workflow, you can support ongoing development via PayPal.
License
This project is dual-licensed:
- Open source under the
GNU Affero General Public License v3.0
(
AGPL-3.0-only) - Commercial license available for closed-source or SaaS use
Versions <= 0.5.0 remain under the ISC license.
See LICENSING.md for details and contact information.
