rag-lab
v0.3.7
Published
RAG Engineering Academy and local retrieval experimentation lab
Downloads
154
Readme
rag.lab — RAG Engineering Academy
rag.lab is a complete learning platform for retrieval-augmented generation and production AI agents. The Academy starts at http://127.0.0.1:5173/ and teaches from first principles through principal-level architecture defense with an Understand → Try it yourself → Prove it loop.
The interface is dark-only and intentionally uses a practice loop:
- Predict what will happen.
- Run the experiment with actual inputs.
- Inspect the intermediate result, source spans, prompt, timings, or vectors.
- Change one thing.
- Write a finding before continuing.
There are no quiz-only completion badges. A lesson checkpoint requires a correct concept check and a written explanation or experiment observation. Academy notes, checkpoints, interview answers and traces are kept in browser storage and can be exported as JSON; database evaluation runs are stored in PostgreSQL.
What is included
- 48 guided lessons in eight chapters, from “Why an AI needs your documents” through ingestion, PostgreSQL retrieval, evaluation, production, advanced retrieval, reliable agents, and principal-level ownership.
- A senior/principal interview room with architecture, SQL, security, evaluation, operations, product judgment, behavioral, and coding prompts. Kaizan preparation is explicitly based on public product information and does not claim access to private architecture or interview questions.
- A production desk with tenant isolation, deletion, freshness, structured-output validation, agent state, human approval, capacity, migration, recovery, and release exercises.
- A real SciFact/BEIR corpus of 5,183 scientific abstracts, 809 training queries and 300 held-out test queries with published relevance labels. See
data/ATTRIBUTION.mdfor origin, license and checksum. - 2,400 separately labeled synthetic client-service records across 100 fictional accounts for client-intelligence workflows, stale revisions, account filters, blockers, and adversarial source text. These are not Kaizan customer data.
- An editable chunking and lexical-search sandbox with exact source offsets.
- A live RAG pipeline where every stage is runnable against local PostgreSQL and LM Studio: create an immutable source snapshot, compare fixed/sentence/paragraph and semantic chunking, embed in resumable batches, build HNSW or IVFFlat, promote or roll back, and inspect stored vectors and context.
- Retrieval experiments for lexical, exact vector, HNSW/IVFFlat ANN and hybrid RRF, plus rewrite, multi-query and HyDE transformations, parent/neighbour evidence expansion, MMR, model reranking and the Qwen3 binary reranker. Every run stores candidates, exclusions, prompt, citations, model calls, timings, SQL and
EXPLAIN ANALYZE. - Repeatable retrieval evaluations, live answer evaluations, exportable traces, and failure drills for tenant isolation, model outages, atomic resume, incomplete promotion and snapshot deletion.
- Real local model calls through an OpenAI-compatible endpoint for embeddings, grounded generation, and bounded agent search refinement.
- PostgreSQL + pgvector with immutable embedding recipes, model/dimension checks, current-version filtering, and account filtering.
- Lexical, vector, and hybrid RRF retrieval, source-backed prompts, citation label checks, and inspectable traces.
- PostgreSQL-backed document browsing, qrels, corpus hashes, saved evaluation runs, per-query recall/precision/MRR/nDCG, and exact/ANN plan inspection.
- A bounded read-only agent that may make one refined search, validates its decision schema, exposes a trace, and never sends external messages.
- A local experiment journal and exportable project evidence.
The Academy explicitly marks teaching fixtures and simplified demonstrations. It does not claim that a tutorial completion is production certification. Real production work still needs authenticated identity, document-level authorization, source ownership, deletion propagation, domain-reviewed labels, incident response, and deployment tests.
Run it
Prerequisites:
- Bun 1.4+
- PostgreSQL with pgvector 0.8+
- LM Studio (or another local OpenAI-compatible server)
The workspace runs directly on your native PostgreSQL installation and local model server. It has no Docker or Compose dependency.
Start PostgreSQL and the local model server, then:
bun run migrate
bun run data:import
bun run devOpen http://127.0.0.1:5173/.
bun run data:download downloads and verifies the public SciFact archive before extraction. bun run data:import validates its 5,183 unique IDs, imports qrels, and generates the separate synthetic client corpus. The current local database has both corpora imported. bun run scripts/embed-corpus.ts creates a resumable local embedding recipe, embeds both corpora in committed batches using the local OpenAI-compatible endpoint, and builds HNSW; override EMBEDDING_MODEL and QUERY_PREFIX when using another model.
The Academy can be browsed without models, but live generation and vector experiments need a local server at http://127.0.0.1:1234/v1. Open Connections, click Discover, and choose an embedding model plus a generation model. The app does not download model files. If your server needs a key, set MODEL_API_KEY for the local server process; keys are never placed in browser storage or traces.
Open Live RAG pipeline to work end to end. Copy up to 1,000 SciFact or synthetic client records (or import your own text/Markdown), edit the sources, and create a frozen snapshot. Then embed it with the selected local embedding model, build an ANN index, promote it, and run retrieval-only or grounded answers with the selected local generation model. The snapshot and run history make each comparison reproducible; changing a setting creates a new experiment rather than mutating the evidence underneath you.
Verification
bun run typecheck
bun run test
bun run test:integration
bun run test:academy
bun run test:pipeline
bun run buildUnit tests cover source-preserving chunk offsets, vector validation, retrieval and RRF, metrics, prompt evidence mapping, curriculum reachability, and the Academy’s benchmark calculations. The integration check verifies the local model adapter, input validation and origin protection. test:academy verifies corpus filters, plans, resumable embedding storage, HNSW, citation labels, bounded agent behavior and API validation. docs/tenant-isolation.sql is a rollback-only restricted-role RLS exercise.
Research and boundaries
docs/RESEARCH.md is the research map behind the course. It links the original RAG paper, pgvector and PostgreSQL documentation, BEIR/SciFact, Sentence Transformers, contextual retrieval, HyDE, ColBERT, GraphRAG, LangGraph, LlamaIndex, Haystack, Docling, Ragas, OpenTelemetry and MCP, and records which ideas are implemented versus exercises or researched options.
bun run test:pipeline runs the deterministic end-to-end contract against a temporary local database. bun run scripts/verify-live-pipeline.ts exercises the same routes with the configured LM Studio models, writes real snapshots and traces, and requires PostgreSQL plus the local model server to be running. Set LIVE_BUILD_ID to reuse an existing snapshot or REMAINING_ONLY=1 to resume only unfinished embedding work.
The Academy distinguishes retrieval quality from answer faithfulness, authorization and operations. It does not certify production safety, reproduce Kaizan’s private stack, or present synthetic client data as customer data.
Sources used in the curriculum
The Academy links directly to primary references while keeping the experiments local: pgvector, PostgreSQL full-text search, Anthropic contextual retrieval, Microsoft GraphRAG query modes, and LangGraph agentic RAG.
