npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@namahapdf/docindex

v0.3.0

Published

Make a folder of PDFs greppable. Extracts real, layout-aware text to plain files with page addresses, so ripgrep (and your agent) can find things without loading whole documents.

Readme

@namahapdf/docindex

Make a folder of PDFs greppable.

npx @namahapdf/docindex .
Indexing 47 documents in /work/contracts
[1/47] vendor-msa.pdf … 118 pages (2.4s)
...
→ indexed 47 documents (3,114 pages) in 12s → .pdfindex/ (18.2 MB, 3% of source)
   try: rg "your search" .pdfindex/
   tip: add .pdfindex/ to .gitignore — it rebuilds from the PDFs in seconds
$ rg -B5 "termination" .pdfindex/
vendor-msa.txt
=== vendor-msa.pdf | page 47 ===
12. Termination
Either party may terminate this Agreement...

Why

grep is astonishingly token-efficient for code — not because it's smart, but because it returns addresses. Search costs ~200 tokens, you get back file.ts:412, you read 50 lines. You never load the codebase.

PDFs have no addresses. So the only move is "load the whole document": ~300k tokens for a 500-page spec, and flatly impossible for a folder of fifty. This gives them addresses and writes them to disk as plain text, so the tools you already have can find things.

No install to search. No index format to learn. No MCP server required. No network calls — nothing leaves your machine.

The extraction is the point

Almost everything in this space is pdftotext underneath, which interleaves two-column layouts, shreds tables and drops rotated pages. An index over bad extraction is worse than useless: it points you at garbage confidently, and you can't tell.

This uses the NamahaPDF engine — real glyph geometry, column-aware reading order, faithful tables, page /Rotate applied so sideways pages read upright.

Usage

docindex <folder> [options]

  --out <dir>       where to write the index      (default <folder>/.pdfindex)
  --max-pages <n>   stop each document after n pages
  --force           re-extract everything, ignoring content hashes
  --retry           re-attempt documents that previously crashed the indexer
  --compress        write .txt.gz/.md.gz — searchable only with `rg -z`
  --format <fmt>    output format: txt (default) or md
  --quiet           only print the summary
  -h, --help        this message

--format md writes .md instead of .txt — the content is identical either way (the address line was always followed by real #/## headings and |-pipe tables; --format md only changes the file extension so markdown-aware tooling recognizes it). Switching --format between runs forces a full re-index rather than mixing both extensions in one index directory.

Re-runs are incremental: unchanged files (by SHA-256) are skipped, and two copies of the same document share one text file.

Exit codes: 0 clean, 1 a document failed to extract, 2 bad usage.

What lands on disk

.pdfindex/
  vendor-msa.txt
  specs__spec-v3.txt
  index.json

Each .txt is the document's text with an address line before every page:

=== specs/spec-v3.pdf | page 47 ===

index.json holds what grep can't give: per-document outline (headings + page), page count, source path, content hash, extractor version, and warnings — e.g. "no text layer on 12 of 312 pages — likely scanned".

Does it cost a lot of disk?

Not much — text is the small part of a PDF. Measured:

| corpus | index size | |---|---| | 36 mixed business/academic PDFs (56 MB) | 1.8% of source | | 13 ISO/PDF-association specs (25 MB) | 12% of source |

The spread is the point: image-heavy and scanned documents are large and yield little text, while a dense typographic spec is nearly all words. Even the worst case is an eighth of the source, and the CLI prints the real number for your folder so you never have to guess. Add .pdfindex/ to .gitignore; it's a cache and rebuilds in seconds.

--compress exists but is off by default on purpose: plain rg can't see inside a .gz, so a compressed index turns a missing hit into a silent one — a far worse failure than 3% of disk.

It never touches your PDFs

Your documents are opened read-only and copied, never modified, moved or deleted. Every write this tool makes goes into the index directory, and the only files it ever deletes are text files it wrote itself and no longer needs (when you delete or rename a source PDF, its stale .txt goes with it).

If you point --out at a folder that already holds your own work, two rules keep it safe:

  • It will not overwrite a file it did not write. A report.pdf next to a hand-written report.txt gets indexed as report-<hash>.txt; if that name is taken too, the document is reported as a conflict and skipped rather than guessed at. A document missing from the index is recoverable; your overwritten notes are not.
  • It only deletes plain filenames directly inside the index directory. The names come from index.json, which is an ordinary file you can edit, so a mangled or hand-copied entry can't be turned into a delete somewhere else.

Nothing is uploaded and no network request is made. Deleting the whole index directory is always safe — it rebuilds from your PDFs.

Scanned PDFs

A scan with no text layer produces no text, and the CLI says so per document rather than silently emitting nothing. OCR is a paid NamahaPDF feature, not part of this tool.

Limitations

  • PDF only in v1. .xlsx / .pptx are reported as skipped, not indexed.
  • No OCR, no ranking, no search server — that's rg's job.
  • Documents are processed sequentially.

Programmatic use

import { indexFolder } from '@namahapdf/docindex';

const summary = await indexFolder({ root: './contracts' });
console.log(summary.pages, summary.warnings);

License

Free to use, unlimited, for personal and commercial work — no key, no watermark, no telemetry. See LICENSE. The rest of the NamahaPDF SDK is commercially licensed.