@namahapdf/docindex
v0.3.0
Published
Make a folder of PDFs greppable. Extracts real, layout-aware text to plain files with page addresses, so ripgrep (and your agent) can find things without loading whole documents.
Readme
@namahapdf/docindex
Make a folder of PDFs greppable.
npx @namahapdf/docindex .Indexing 47 documents in /work/contracts
[1/47] vendor-msa.pdf … 118 pages (2.4s)
...
→ indexed 47 documents (3,114 pages) in 12s → .pdfindex/ (18.2 MB, 3% of source)
try: rg "your search" .pdfindex/
tip: add .pdfindex/ to .gitignore — it rebuilds from the PDFs in seconds$ rg -B5 "termination" .pdfindex/
vendor-msa.txt
=== vendor-msa.pdf | page 47 ===
12. Termination
Either party may terminate this Agreement...Why
grep is astonishingly token-efficient for code — not because it's smart, but
because it returns addresses. Search costs ~200 tokens, you get back
file.ts:412, you read 50 lines. You never load the codebase.
PDFs have no addresses. So the only move is "load the whole document": ~300k tokens for a 500-page spec, and flatly impossible for a folder of fifty. This gives them addresses and writes them to disk as plain text, so the tools you already have can find things.
No install to search. No index format to learn. No MCP server required. No network calls — nothing leaves your machine.
The extraction is the point
Almost everything in this space is pdftotext underneath, which interleaves
two-column layouts, shreds tables and drops rotated pages. An index over bad
extraction is worse than useless: it points you at garbage confidently, and you
can't tell.
This uses the NamahaPDF engine — real glyph geometry, column-aware reading order,
faithful tables, page /Rotate applied so sideways pages read upright.
Usage
docindex <folder> [options]
--out <dir> where to write the index (default <folder>/.pdfindex)
--max-pages <n> stop each document after n pages
--force re-extract everything, ignoring content hashes
--retry re-attempt documents that previously crashed the indexer
--compress write .txt.gz/.md.gz — searchable only with `rg -z`
--format <fmt> output format: txt (default) or md
--quiet only print the summary
-h, --help this message--format md writes .md instead of .txt — the content is identical either
way (the address line was always followed by real #/## headings and
|-pipe tables; --format md only changes the file extension so
markdown-aware tooling recognizes it). Switching --format between runs
forces a full re-index rather than mixing both extensions in one index
directory.
Re-runs are incremental: unchanged files (by SHA-256) are skipped, and two copies of the same document share one text file.
Exit codes: 0 clean, 1 a document failed to extract, 2 bad usage.
What lands on disk
.pdfindex/
vendor-msa.txt
specs__spec-v3.txt
index.jsonEach .txt is the document's text with an address line before every page:
=== specs/spec-v3.pdf | page 47 ===index.json holds what grep can't give: per-document outline (headings + page),
page count, source path, content hash, extractor version, and warnings — e.g.
"no text layer on 12 of 312 pages — likely scanned".
Does it cost a lot of disk?
Not much — text is the small part of a PDF. Measured:
| corpus | index size | |---|---| | 36 mixed business/academic PDFs (56 MB) | 1.8% of source | | 13 ISO/PDF-association specs (25 MB) | 12% of source |
The spread is the point: image-heavy and scanned documents are large and yield
little text, while a dense typographic spec is nearly all words. Even the
worst case is an eighth of the source, and the CLI prints the real number for
your folder so you never have to guess. Add .pdfindex/ to .gitignore; it's
a cache and rebuilds in seconds.
--compress exists but is off by default on purpose: plain rg can't see inside
a .gz, so a compressed index turns a missing hit into a silent one — a far
worse failure than 3% of disk.
It never touches your PDFs
Your documents are opened read-only and copied, never modified, moved or
deleted. Every write this tool makes goes into the index directory, and the
only files it ever deletes are text files it wrote itself and no longer needs
(when you delete or rename a source PDF, its stale .txt goes with it).
If you point --out at a folder that already holds your own work, two rules keep
it safe:
- It will not overwrite a file it did not write. A
report.pdfnext to a hand-writtenreport.txtgets indexed asreport-<hash>.txt; if that name is taken too, the document is reported as a conflict and skipped rather than guessed at. A document missing from the index is recoverable; your overwritten notes are not. - It only deletes plain filenames directly inside the index directory. The
names come from
index.json, which is an ordinary file you can edit, so a mangled or hand-copied entry can't be turned into a delete somewhere else.
Nothing is uploaded and no network request is made. Deleting the whole index directory is always safe — it rebuilds from your PDFs.
Scanned PDFs
A scan with no text layer produces no text, and the CLI says so per document rather than silently emitting nothing. OCR is a paid NamahaPDF feature, not part of this tool.
Limitations
- PDF only in v1.
.xlsx/.pptxare reported as skipped, not indexed. - No OCR, no ranking, no search server — that's
rg's job. - Documents are processed sequentially.
Programmatic use
import { indexFolder } from '@namahapdf/docindex';
const summary = await indexFolder({ root: './contracts' });
console.log(summary.pages, summary.warnings);License
Free to use, unlimited, for personal and commercial work — no key, no watermark, no telemetry. See LICENSE. The rest of the NamahaPDF SDK is commercially licensed.
