npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@nalinor/mupdf4llm

v0.3.1

Published

TypeScript port of pymupdf4llm — convert PDFs to LLM-ready Markdown using the mupdf WASM package, with optional OCR of table cells.

Readme

mupdf4llm

npm CI Docs License: AGPL v3

TypeScript/Bun port of pymupdf4llm on top of the official mupdf WASM package. Converts PDFs into LLM-ready Markdown with reading-order text, headers, bullets, inline styling (bold, italic, monospaced), tables, images, and per-word coordinates. It tracks pymupdf4llm.to_markdown(doc) closely — a set of synthetic fixtures match exactly, but real-world PDFs can diverge; the parity & limits guide documents what differs and why.

📖 Full documentation: iamnalinor.github.io/mupdf4llm

Install

npm i @nalinor/mupdf4llm
# or
bun add @nalinor/mupdf4llm

Requires Node 20+ or Bun ≥ 1.0. The only runtime dependency is mupdf (WASM, no native build).

The package is ESM only (import). From CommonJS code, load it with const { toMarkdown } = await import("@nalinor/mupdf4llm"): its mupdf dependency uses top-level await, which require() cannot load.

Quick start

import { toMarkdown } from "@nalinor/mupdf4llm";
import { readFileSync } from "node:fs";

const md = await toMarkdown(readFileSync("paper.pdf"));
console.log(md);

Tables in scans, with OCR of every cell (npm i ppu-paddle-ocr onnxruntime-node once):

const md = await toMarkdown(readFileSync("scan.pdf"), { tableStrategy: "pixels" });

Per-page chunks for RAG:

import { toMarkdownPages } from "@nalinor/mupdf4llm";

const chunks = await toMarkdownPages(readFileSync("paper.pdf"), {
  extractWords: true,
});

See the quick start guide for more, including options, table strategies, image extraction, and the LlamaIndex adapter.

What's in scope

  • Reading-order text extraction with header inference (IdentifyHeaders, TocHeaders)
  • Multi-column layout detection
  • Five table strategies: lines_strict, lines, text, explicit, and pixels for scans
  • OCR of table cells (textSource), RapidOCR by default via optional peers, or any engine you plug in
  • Image extraction (writeImages) and inline base64 embedding (embedImages)
  • Per-word coordinates (extractWords)
  • Form-field extraction (getKeyValues)
  • Page rotation handling
  • LlamaIndex adapter at the @nalinor/mupdf4llm/llama subpath

What's NOT available (and why)

  • pymupdf.layout features (to_text, to_json, layout-mode to_markdown) — require Artifex's closed-source pymupdf-layout ONNX wheel (Polyform Noncommercial license, no JS distribution).
  • Whole-page OCR — the official mupdf WASM bundle ships without Tesseract/Leptonica. Only table cells are OCR'd; see the OCR guide.

Detailed write-up: parity and limits.

Development

bun install
pip install pymupdf4llm   # required for the parity test suite
bun test                  # parity + unit tests
bun run lint              # prettier --write + eslint + tsc
bun run docs:dev          # local doc preview at http://localhost:5173
bun run build             # emits dist/{index,llama}.{js,d.ts}

See CONTRIBUTING.md for release workflow and project layout.

License

AGPL-3.0-or-later — inherited from PyMuPDF / pymupdf4llm and MuPDF. See LICENSE. Commercial licensing of MuPDF is available from Artifex Software.

Credits

  • Artifex Software — author of MuPDF and the mupdf npm package
  • pymupdf4llm authors — the Python algorithms this port reimplements