npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

leanpdf

v0.5.0

Published

Small, low-memory, streaming PDF toolkit for browsers, Node and Bun. First tool: shrink PDFs by recompressing their images, copying everything else byte for byte.

Readme

leanpdf

A small, streaming PDF toolkit for browsers, Node and Bun. It compresses images, reads text and metadata, renders pages to a canvas, edits, merges and decrypts. Anything it doesn't change is copied byte for byte.

  • Low memory. Files are read in slices and written in one pass. A 600 MB PDF compresses in about 300 MB of RSS.
  • Small. No dependencies, and every feature is a separate module. Compressing in the browser is 45 KB minified (17 KB gzipped); it uses the browser's own JPEG and zlib.
  • Conservative. Unusual objects are left alone, images are only replaced when they get at least 10% smaller, and damaged files are repaired as they're read.

Compressing, compared with the usual JavaScript options at default settings (details, charts):

| File | leanpdf (Node) | pdf-lib + sharp | Ghostscript WASM | |---|--:|--:|--:| | brochure.pdf (41 MB) | 3.2 MB (−92%), 211 MB RSS | 3.1 MB (−92%), 228 MB RSS | 1.3 MB (−97%), 198 MB RSS | | scan.pdf (28 MB) | 3.4 MB (−88%), 173 MB RSS | 28 MB (±0%), 144 MB RSS | 4.7 MB (−84%), 207 MB RSS | | report.pdf (2.0 MB) | 0.5 MB (−75%), 130 MB RSS | 1.3 MB (−35%), 125 MB RSS | 0.2 MB (−91%), 158 MB RSS | | large.pdf (601 MB) | 13 MB (−98%), 298 MB RSS | 275 MB (−54%), 1085 MB RSS | 7.6 MB (−99%), 846 MB RSS |

The other features on large.pdf (601 MB), wall time and peak RSS:

| Feature | leanpdf | pdf-lib | PDF.js | MuPDF.js | |---|--:|--:|--:|--:| | Info | 0.15 s, 68 MB | 1.08 s, 669 MB | 3.78 s, 1319 MB | 4.13 s, 1875 MB | | Text extraction | 0.18 s, 72 MB | – | 4.04 s, 1319 MB | 4.70 s, 1875 MB | | Select pages | 0.45 s, 86 MB | 2.73 s, 892 MB | – | 5.98 s, 1875 MB | | Rotate all pages | 1.55 s, 82 MB | 6.18 s, 1270 MB | – | out of memory | | Decrypt | 5.59 s, 123 MB | – | – | out of memory |

(– means the library has no such feature.)

Install

npm install leanpdf          # or: bun add leanpdf
npm install sharp            # only to compress in Node and Bun, or with the CLI
npm install @napi-rs/canvas  # only to render in Node and Bun, or with the CLI

ESM only. Runs in modern browsers, Node 18.17+ and Bun (sharp and @napi-rs/canvas need what they need; both ship prebuilt binaries).

Compressing

In the browser

import { compressPdfBlob } from 'leanpdf';

const { blob, report } = await compressPdfBlob(file, { maxWidth: 1600, maxHeight: 1600, jpegQuality: 0.75 });

The result's unchanged parts are slices of the input, so they never enter JS memory. Run it in a Web Worker for large files. To stream straight to disk, give compressPdf a WritableStream:

import { compressPdf, BrowserImageCodec } from 'leanpdf';

const handle = await showSaveFilePicker({ suggestedName: 'compressed.pdf' });
const report = await compressPdf(file, await handle.createWritable(), {
  codec: new BrowserImageCodec(),
  onProgress: ({ processedObjects, totalObjects }) => console.log(processedObjects / totalObjects),
});

In Node and Bun

import { compressPdfFile } from 'leanpdf/node';
import { SharpImageCodec } from 'leanpdf/sharp';

const report = await compressPdfFile('in.pdf', 'out.pdf', { codec: new SharpImageCodec(), maxWidth: 2000, maxHeight: 2000 });

The output is written to a temporary file and renamed into place, so the input may be the output.

From the command line

leanpdf compress in.pdf out.pdf --max 1600 --quality 0.75

| Option | Default | | |---|---|---| | --max <px> | 1600 | Max image width and height | | --max-width, --max-height | | Set them separately | | -q, --quality <0-1> | 0.75 | JPEG quality | | --min-bytes <n> | 20000 | Skip image streams smaller than this | | --min-savings <r> | 0.9 | Keep a result only if new ≤ old × r | | -j, --concurrency <n> | 1 | Images processed in parallel | | --no-gray | | Encode grayscale images as RGB | | --progressive | | Write progressive JPEGs | | --streams | | Also Flate-compress uncompressed streams | | --strip | | Also remove metadata and drop unused objects | | --gc | | Also drop unused objects | | --json | | Print the report (plus time and peak RSS) as JSON |

The CLI has a command for every feature; see Command line.

Reading documents

openPdf reads the cross-reference index; everything else is read on demand.

import { openPdf, getInfo, getPages, getOutline, getFormFields, listAttachments, attachmentStream } from 'leanpdf';

const doc = await openPdf(file); // a Blob/File, its bytes (Uint8Array, ArrayBuffer), or any RandomAccessSource
const info = await getInfo(doc); // { title, author, creationDate, pageCount, tagged, signed, ... }
const pages = await getPages(doc); // [{ width, height, rotate, label }]
const outline = await getOutline(doc); // [{ title, pageIndex, url, children }]
for (const a of await listAttachments(doc)) {
  const stream = attachmentStream(doc, a); // ReadableStream: big files never sit in memory whole
}

| Function | Returns | |---|---| | getInfo(doc) | version, page count, title, author, dates, XMP, flags (encrypted, tagged, forms, JavaScript, signed) | | getPages(doc) | width, height, rotation and label of each page | | getOutline(doc) | bookmark tree with target pages and URLs | | getLinks(doc) | links per page, with a URL or target page | | getFormFields(doc) | form fields: name, type, value, options | | listAttachments(doc) | embedded files: name, size, type, dates | | readAttachment(doc, a) / attachmentStream(doc, a) | a file's bytes, whole or streamed | | listImages(doc) | images with their pages, size, color space and encoding | | extractImage(doc, num) | JPEG or JPEG 2000 bytes as stored, or 8-bit pixels | | extractText(doc, opts?) | text per page, as an async iterator | | extractAllText(doc, opts?) | all text, pages separated by form feeds |

On encrypted documents everything except getInfo and getPages throws PdfEncryptedError.

extractText is meant for search indexing: words and lines in drawing order, glyphs without a Unicode mapping left out.

Rendering

import { openPdf, renderPage } from 'leanpdf';

const doc = await openPdf(await file.arrayBuffer()); // or openPdf(file) for files too large to hold
const canvas = new OffscreenCanvas(1, 1); // or a <canvas> element; it is resized to the page
const { width, height, warnings } = await renderPage(doc, 0, canvas, { scale: 2 }); // 144 dpi

The browser does the drawing through Canvas 2D, on the main thread or in a worker. leanpdf interprets the page and turns embedded fonts (TrueType, OpenType, CFF, Type 1, Type 3) into outlines; fonts that aren't embedded use a similar system font. Images stream from the file straight to about the size they're drawn, a page's JPEGs decode in parallel, and decoded images are cached per document. JPEGs go to the browser's decoder, except CMYK ones, which browsers invert: leanpdf decodes those itself.

A File is read in pieces through a 4 MB cache of 64 KB blocks, so a page takes a few reads, and each read from a File is a round trip to the browser (slowest on phones). When the file fits in memory, opening its bytes saves those too.

| Option | Default | | |---|---|---| | scale | 1 | Pixels per point (1 = 72 dpi) | | width, height | | Fit the page in this many pixels instead | | background | '#fff' | CSS color painted first; null for transparent | | annotations | true | Draw annotation appearances (form fields, stamps, highlights) | | signal | | AbortSignal |

Not supported yet: JBIG2 images (listed in warnings), non-embedded CJK fonts with predefined CMaps, knockout groups, overprint and ICC profiles.

In Node and Bun

Node has no canvas of its own; leanpdf/canvas renders with @napi-rs/canvas (Skia, the engine Chrome draws with), an optional dependency like sharp:

import { writeFile } from 'node:fs/promises';
import { openPdf } from 'leanpdf';
import { NodeFileSource } from 'leanpdf/node';
import { renderPageImage } from 'leanpdf/canvas';

const doc = await openPdf(await NodeFileSource.open('in.pdf'));
const { data } = await renderPageImage(doc, 0, { scale: 2, format: 'png' }); // or 'jpeg', 'webp' with quality
await writeFile('page-1.png', data);

renderPage(doc, index, canvas?, options) from leanpdf/canvas draws onto a @napi-rs/canvas Canvas instead (a new one when omitted) and returns it with the result. Pages render as in the browser; fonts that aren't embedded come from the system's (on Linux, Liberation or the URW fonts stand in for Helvetica, Times and Courier).

Editing

Every write goes through rewritePdf, which runs a list of plugins in one pass. It reads any input (a File, bytes, ...) and writes to a WritableStream or a sink; for a Blob, pass a BlobPartsSink and read its .blob (Inputs and outputs):

import { rewritePdf, BlobPartsSink, compressImages, stripMetadata, removeJavaScript, selectPages, removeUnused, BrowserImageCodec } from 'leanpdf';

const output = new BlobPartsSink();
const report = await rewritePdf(file, output, [
  compressImages({ codec: new BrowserImageCodec() }),
  stripMetadata(),
  removeJavaScript(),
  selectPages([2, 0, 1]), // keep and reorder (0-based)
  removeUnused(),
]);
const blob = output.blob;

| Plugin | What it does | |---|---| | compressImages(options) | recompress images (what compressPdf runs) | | stripMetadata() | remove document info, XMP and page thumbnails | | removeJavaScript() | remove scripts and script actions | | removeAttachments() | remove embedded files | | rotatePages(deg \| fn) | rotate all pages, or per page | | selectPages(indices) | keep and reorder pages; bookmarks and links to removed pages go too | | recompressStreams() | Flate-compress uncompressed streams | | repairStreams() | fix broken stream lengths | | removeUnused() | drop objects nothing refers to |

To write your own, see the Plugin type and src/features/rotate.ts.

Merging

import { mergePdfs } from 'leanpdf';

const report = await mergePdfs([fileA, fileB], output, { pages: [undefined, [0, 2]] }); // all of A, then B's pages 1 and 3

Only the objects the selected pages need are copied. Bookmarks, form fields and layers are merged; document JavaScript, attachments, structure tags and XMP are dropped.

Decrypting

import { decryptPdf, PdfPasswordError } from 'leanpdf';

const report = await decryptPdf(file, output, { password: 'secret' }); // { method: 'AES-256 (R6)', password: 'user' }

RC4 and AES-128/256 (revisions 2–6), with the user or owner password; the empty password is tried by default. A wrong one throws PdfPasswordError. Public-key encryption isn't supported.

Command line

compress and images --extract need sharp, render needs @napi-rs/canvas. Page numbers are 1-based (3,1-2,5-).

| Command | | |---|---| | compress <in> <out> | recompress images (options above) | | info <in> | metadata, flags and page sizes | | text <in> [-o out.txt] [--pages 1-3] | extract text | | images <in> [--extract <dir>] | list or save images | | attachments <in> [--extract <dir>] | list or save embedded files | | outline <in>, links <in>, fields <in> | bookmarks, links, form fields | | clean <in> <out> | remove metadata, JavaScript, attachments and unused objects | | pages <in> <out> <pages> | keep and reorder pages | | rotate <in> <out> <degrees> [--pages 1,3] | rotate pages clockwise | | merge <out> <in>[:pages]... | concatenate PDFs | | decrypt <in> <out> [--password <pw>] | remove encryption | | render <in> <out.png> [--pages 1-3] [--dpi 144] | render pages to PNG, JPEG or WebP (by extension); several pages become out-1.png, out-2.png, ...; --width/--height fit pages instead, -q sets JPEG and WebP quality | | repair <in> <out> | rebuild the index and fix broken streams |

Every command takes --json and --quiet.

API

compressPdf(input, output, options): Promise<CompressReport>

Writes a new PDF to output and closes it, or aborts it on failure. input and output are as in Inputs and outputs.

| Option | Default | | |---|---|---| | codec | required | An ImageCodec | | maxWidth, maxHeight | 1600 | Images are downscaled to fit, never enlarged | | jpegQuality | 0.75 | 0..1 | | preserveGray | true | Ask the codec to keep grayscale images grayscale | | minImageBytes | 20 000 | Skip image streams smaller than this | | minSavingsRatio | 0.9 | Keep a replacement only if newSize <= oldSize * ratio | | concurrency | 1 | Images in flight at once. Memory grows with it. | | signal | | AbortSignal | | onProgress | | ({ processedObjects, totalObjects, bytesSaved }) => void |

interface CompressReport {
  imagesSeen: number;
  imagesRecompressed: number;
  imagesSkipped: Record<string, number>; // reason -> count, see below
  inputBytes: number;
  outputBytes: number;
  signaturesInvalidated: boolean;        // the input was digitally signed
  xrefRepaired: boolean;                 // damaged cross-reference data was rebuilt
  warnings: string[];
}

Errors (PdfError subclasses): PdfEncryptedError, PdfPasswordError, PdfFormatError, SourceReadError. Also RangeError for bad options, and the signal's reason when aborted.

Entry points and bundle sizes

| Import | Contents | |---|---| | leanpdf | everything that runs in a browser | | leanpdf/node | NodeFileSource, NodeFileSink, compressPdfFile | | leanpdf/sharp | SharpImageCodec (the only module that imports sharp) | | leanpdf/canvas | renderPage, renderPageImage for Node and Bun (the only module that imports @napi-rs/canvas) |

Bundlers keep only what you import. Minified sizes, including the core each needs (bun run size). Decoders for images browsers can't decode (JPEG 2000, CMYK JPEG, fax; 30 KB) load with import() the first time a page needs one, so bundlers that split code keep them out of these:

| Import | Minified | Gzipped | |---|--:|--:| | compressPdfBlob (core, Blob I/O, browser codec) | 44.8 KB | 17.2 KB | | openPdf | 22.6 KB | 8.9 KB | | openPdf + getInfo | 30.4 KB | 12.0 KB | | openPdf + getOutline, getLinks, getFormFields | 31.9 KB | 12.3 KB | | openPdf + extractText | 48.0 KB | 20.5 KB | | openPdf + renderPage | 126.5 KB | 53.0 KB | | rewritePdf + all editing plugins | 50.6 KB | 19.1 KB | | mergePdfs | 41.5 KB | 16.0 KB | | decryptPdf | 43.8 KB | 17.4 KB | | everything | 221.5 KB | 88.5 KB |

Inputs and outputs

Everything that reads a PDF takes a PdfInput: a Blob or File (read in pieces, never whole), the file's bytes (Uint8Array or ArrayBuffer), or any RandomAccessSource, such as await NodeFileSource.open(path) from leanpdf/node.

Everything that writes takes a PdfOutput: a WritableStream<Uint8Array> (from showSaveFilePicker, Writable.toWeb(fs.createWriteStream(path)), a TransformStream, ...) or an OutputSink:

  • new BlobPartsSink() builds a Blob (read .blob afterwards). Unchanged ranges of a Blob input stay slices of it, so they never enter JS memory.
  • await NodeFileSink.create(path) (from leanpdf/node) writes a file.

The output is closed when the function succeeds and aborted when it fails. To read or write anything else, implement these:

interface RandomAccessSource {
  readonly size: number;
  read(offset: number, length: number): Promise<Uint8Array>;
}
interface OutputSink {
  write(chunk: Uint8Array): Promise<void>;
  copyRange(source: RandomAccessSource, offset: number, length: number): Promise<void>;
  close(): Promise<void>;
  abort?(reason?: unknown): Promise<void>;
}

Codecs

interface ImageCodec {
  recompress(input: ImageInput, opts: RecompressOptions): Promise<ImageOutput | null>;
}

The input is the original JPEG bytes or 8-bit gray/RGB pixels. Return a JPEG, or null to keep the original.

  • BrowserImageCodec uses createImageBitmap and OffscreenCanvas. Browsers only encode RGB JPEGs, so gray images become RGB.
  • SharpImageCodec uses sharp with mozjpeg and keeps gray images gray.

What gets recompressed

An image XObject is a candidate when all of these hold:

| | | |---|---| | Filters | none, /DCTDecode, /FlateDecode, or [/FlateDecode /DCTDecode] | | Color space | /DeviceGray, /DeviceRGB, or /ICCBased with /N 1 or 3 | | Bits per component | 8 | | /Decode | absent or the identity | | Flate predictors | none, TIFF 2, or PNG 10–15 (with /Colors and /Columns matching the image) | | JPEG data | 8-bit, 1 or 3 components, matching the dictionary | | Size | at least minImageBytes of encoded data |

Other dictionary entries of a rewritten image are kept. Soft masks (image transparency) are downscaled to the same box, losslessly.

Skipped images are counted in report.imagesSkipped:

| Reason | Meaning | |---|---| | small | below minImageBytes | | cmyk, indexed, separation, deviceN, lab, colorSpace, noColorSpace | unsupported color space (CMYK includes 4-component JPEGs) | | jbig2, ccitt, filter | unsupported filter or filter chain | | jpx | JPEG 2000 with its own alpha channel (/SMaskInData) | | bitsPerComponent, decode, predictor | not 8-bit, non-identity /Decode, unsupported predictor | | imageMask, colorKeyMask | stencil masks and images with a /Mask color-key array | | softMask | a soft mask that already fits, or isn't gray Flate data | | matte | a pre-blended image or its /Matte soft mask (their dimensions must stay equal) | | jpegTransform, jpegUnsupported, jpegMismatch, jpegInvalid | JPEG variants the codecs can't take, or broken headers | | external, malformed, tooLarge | /F external streams, broken dictionaries, over 2²⁹ samples | | decodeError | the stream data couldn't be decoded completely | | codecDeclined, codecError, codecOutputInvalid | the codec declined, failed or returned a bad JPEG | | noGain | not enough smaller |

Benchmarks

Four generated documents (bench/corpus.ts): a photo brochure, 300 dpi scans, a text-heavy report, and a 600 MB file of distinct photos. Each tool runs in its own process on a 4-core Linux container, and peak RSS includes the runtime itself (about 45 MB for Node). There are charts on the website.

Compression

pdf-lib uses sharp for images at the same settings as leanpdf; Ghostscript uses its /ebook preset; MuPDF.js and qpdf are lossless. PSNR compares rendered pages with the input.

brochure.pdf (41 MB)

| Tool | Output | Saved | Wall time | CPU time | Peak RSS | Valid (qpdf) | PSNR | |---|--:|--:|--:|--:|--:|:-:|--:| | leanpdf (Node, sharp) | 3.2 MB | 92% | 5.1 s | 5.7 s | 211 MB | yes | 48.6 dB | | leanpdf (Bun, sharp) | 3.2 MB | 92% | 5.2 s | 6.2 s | 185 MB | yes | 48.6 dB | | pdf-lib 1.17 + sharp (Node) | 3.1 MB | 92% | 7.3 s | 8.0 s | 228 MB | yes | 48.6 dB | | Ghostscript WASM, /ebook (Node) | 1.3 MB | 97% | 10.2 s | 10.6 s | 198 MB | yes | 42.8 dB | | MuPDF.js 1.28, lossless (Node) | 41 MB | 0% | 0.6 s | 0.5 s | 210 MB | yes | ∞ | | Ghostscript 10 native, /ebook | 1.3 MB | 97% | 4.5 s | 4.4 s | 30 MB | yes | 42.8 dB | | qpdf 11 native, lossless | 41 MB | 0% | 0.2 s | 0.1 s | 30 MB | yes | ∞ |

scan.pdf (28 MB)

| Tool | Output | Saved | Wall time | CPU time | Peak RSS | Valid (qpdf) | PSNR | |---|--:|--:|--:|--:|--:|:-:|--:| | leanpdf (Node, sharp) | 3.4 MB | 88% | 5.3 s | 6.1 s | 173 MB | yes | 30.5 dB | | leanpdf (Bun, sharp) | 3.4 MB | 88% | 5.4 s | 6.6 s | 149 MB | yes | 30.5 dB | | pdf-lib 1.17 + sharp (Node) | 28 MB | -0% | 0.6 s | 0.7 s | 144 MB | yes | ∞ | | Ghostscript WASM, /ebook (Node) | 4.7 MB | 84% | 6.6 s | 6.8 s | 207 MB | yes | 34.3 dB | | MuPDF.js 1.28, lossless (Node) | 28 MB | 0% | 0.6 s | 0.6 s | 175 MB | yes | ∞ | | Ghostscript 10 native, /ebook | 4.7 MB | 84% | 2.7 s | 2.6 s | 29 MB | yes | 34.3 dB | | qpdf 11 native, lossless | 16 MB | 45% | 2.6 s | 2.5 s | 30 MB | yes | ∞ |

report.pdf (2.0 MB)

| Tool | Output | Saved | Wall time | CPU time | Peak RSS | Valid (qpdf) | PSNR | |---|--:|--:|--:|--:|--:|:-:|--:| | leanpdf (Node, sharp) | 0.5 MB | 75% | 0.9 s | 1.0 s | 130 MB | yes | 67.1 dB | | leanpdf (Bun, sharp) | 0.5 MB | 75% | 0.8 s | 1.1 s | 133 MB | yes | 67.1 dB | | pdf-lib 1.17 + sharp (Node) | 1.3 MB | 35% | 1.0 s | 1.2 s | 125 MB | yes | 71.8 dB | | Ghostscript WASM, /ebook (Node) | 0.2 MB | 91% | 3.1 s | 3.6 s | 158 MB | yes | 50.4 dB | | MuPDF.js 1.28, lossless (Node) | 2.0 MB | 0% | 0.2 s | 0.2 s | 85 MB | yes | ∞ | | Ghostscript 10 native, /ebook | 0.2 MB | 91% | 0.7 s | 0.7 s | 35 MB | yes | 50.4 dB | | qpdf 11 native, lossless | 2.6 MB | -31% | 0.2 s | 0.1 s | 29 MB | yes | ∞ |

large.pdf (601 MB)

| Tool | Output | Saved | Wall time | CPU time | Peak RSS | Valid (qpdf) | PSNR | |---|--:|--:|--:|--:|--:|:-:|--:| | leanpdf (Node, sharp) | 13 MB | 98% | 32.0 s | 36.6 s | 298 MB | yes | – | | leanpdf (Bun, sharp) | 13 MB | 98% | 32.7 s | 39.7 s | 301 MB | yes | – | | pdf-lib 1.17 + sharp (Node) | 275 MB | 54% | 32.3 s | 33.3 s | 1085 MB | yes | – | | Ghostscript WASM, /ebook (Node) | 7.6 MB | 99% | 62.0 s | 61.7 s | 846 MB | yes | – | | MuPDF.js 1.28, lossless (Node) | out of memory | | 12.8 s | 12.8 s | 1929 MB | | | | Ghostscript 10 native, /ebook | 7.6 MB | 99% | 32.7 s | 32.3 s | 36 MB | yes | – | | qpdf 11 native, lossless | 889 MB | -48% | 42.1 s | 40.6 s | 84 MB | yes | – |

  • leanpdf needs about 300 MB for the 600 MB file. pdf-lib needs over 1 GB, Ghostscript WASM 850 MB, and MuPDF.js runs out of memory.
  • Ghostscript makes the smallest files by re-rendering everything at 150 dpi, with lower quality.
  • pdf-lib's script can't decode PNG-predicted images, so the scans don't shrink.

Other features

The same files, through the libraries that offer each feature: pdf-lib, PDF.js (its Node build) and MuPDF.js, with native qpdf for reference. Each cell is wall time and peak RSS; jobs on the smaller files ran three times and the median counts. Written files pass qpdf --check with the right page count. Rendering runs in headless Chromium with software rasterization: every page at 144 dpi from the file in memory, with the first three pages compared with MuPDF's rendering (the differences are mostly fonts, since the corpus doesn't embed them).

Info: metadata and page sizes

| Tool | brochure.pdf | scan.pdf | report.pdf | large.pdf | |---|--:|--:|--:|--:| | leanpdf (Node) | 0.12 s, 59 MB | 0.11 s, 58 MB | 0.12 s, 60 MB | 0.15 s, 68 MB | | pdf-lib 1.17 (Node) | 0.23 s, 106 MB | 0.20 s, 95 MB | 0.14 s, 69 MB | 1.08 s, 669 MB | | PDF.js 6.3 (Node) | 0.40 s, 199 MB | 0.39 s, 153 MB | 0.29 s, 118 MB | 3.78 s, 1319 MB | | MuPDF.js 1.28 (Node) | 0.19 s, 195 MB | 0.17 s, 157 MB | 0.12 s, 79 MB | 4.13 s, 1875 MB |

Text extraction, all pages

| Tool | brochure.pdf | scan.pdf | report.pdf | large.pdf | |---|--:|--:|--:|--:| | leanpdf (Node) | 0.17 s, 67 MB | 0.11 s, 59 MB | 0.24 s, 72 MB | 0.18 s, 72 MB | | PDF.js 6.3 (Node) | 0.45 s, 200 MB | 0.38 s, 153 MB | 0.52 s, 148 MB | 4.04 s, 1319 MB | | MuPDF.js 1.28 (Node) | 0.27 s, 228 MB | 0.22 s, 164 MB | 0.19 s, 96 MB | 4.70 s, 1875 MB |

Select pages: keep every other page

| Tool | brochure.pdf | scan.pdf | report.pdf | large.pdf | |---|--:|--:|--:|--:| | leanpdf (Node) | 0.22 s, 62 MB | 0.13 s, 61 MB | 0.14 s, 62 MB | 0.45 s, 86 MB | | pdf-lib 1.17 (Node) | 0.27 s, 126 MB | 0.24 s, 109 MB | 0.16 s, 67 MB | 2.73 s, 892 MB | | MuPDF.js 1.28 (Node) | 0.23 s, 210 MB | 0.24 s, 173 MB | 0.14 s, 85 MB | 5.98 s, 1875 MB | | qpdf 11 native | 0.07 s, 29 MB | 0.06 s, 29 MB | 0.03 s, 29 MB | 0.80 s, 29 MB |

Rotate all pages

| Tool | brochure.pdf | scan.pdf | report.pdf | large.pdf | |---|--:|--:|--:|--:| | leanpdf (Node) | 0.16 s, 62 MB | 0.14 s, 61 MB | 0.14 s, 61 MB | 1.55 s, 82 MB | | pdf-lib 1.17 (Node) | 0.29 s, 147 MB | 0.26 s, 120 MB | 0.17 s, 67 MB | 6.18 s, 1270 MB | | MuPDF.js 1.28 (Node) | 0.28 s, 210 MB | 0.31 s, 173 MB | 0.14 s, 86 MB | out of memory | | qpdf 11 native | 0.09 s, 29 MB | 0.06 s, 29 MB | 0.03 s, 29 MB | 0.54 s, 51 MB |

Merge

| Tool | brochure + scan + report | all 4 files | |---|--:|--:| | leanpdf (Node) | 0.21 s, 65 MB | 1.91 s, 96 MB | | pdf-lib 1.17 (Node) | 0.44 s, 216 MB | 7.71 s, 1415 MB | | MuPDF.js 1.28 (Node) | 0.53 s, 439 MB | out of memory | | qpdf 11 native | 0.13 s, 29 MB | 0.69 s, 54 MB |

Decrypt (AES-256, user password)

| Tool | brochure.pdf | scan.pdf | report.pdf | large.pdf | |---|--:|--:|--:|--:| | leanpdf (Node) | 0.28 s, 96 MB | 0.26 s, 93 MB | 0.19 s, 69 MB | 5.59 s, 123 MB | | MuPDF.js 1.28 (Node) | 1.00 s, 197 MB | 0.71 s, 175 MB | 0.19 s, 85 MB | out of memory | | qpdf 11 native | 0.30 s, 29 MB | 0.15 s, 29 MB | 0.05 s, 29 MB | 2.71 s, 54 MB |

Render all pages (Chromium, 144 dpi)

| File | Renderer | All pages | Per page (median) | Pixels unlike MuPDF | |---|---|--:|--:|--:| | brochure.pdf (12 pages) | leanpdf (Chromium) | 0.84 s | 69 ms | 2.1% | | brochure.pdf (12 pages) | PDF.js 6.3 (Chromium) | 3.35 s | 298 ms | 1.6% | | scan.pdf (12 pages) | leanpdf (Chromium) | 1.55 s | 128 ms | <0.1% | | scan.pdf (12 pages) | PDF.js 6.3 (Chromium) | 2.43 s | 196 ms | <0.1% | | report.pdf (40 pages) | leanpdf (Chromium) | 0.49 s | 9 ms | 3.0% | | report.pdf (40 pages) | PDF.js 6.3 (Chromium) | 0.86 s | 17 ms | 2.2% | | large.pdf (79 pages) | leanpdf (Chromium) | 9.82 s | 78 ms | <0.1% | | large.pdf (79 pages) | PDF.js 6.3 (Chromium) | 23.8 s | 200 ms | <0.1% |

  • On the 600 MB file, leanpdf needs at most 125 MB for any of these. pdf-lib, PDF.js and MuPDF.js load the whole file first (up to 1.9 GB), and MuPDF.js runs out of memory writing it.
  • On the small files, most of the time is Node starting (about 0.1 s).
  • Native qpdf writes big files 2-3 times faster than leanpdf, in similar memory.
  • leanpdf renders 1.5-4 times faster than PDF.js. Both match MuPDF on images; on text pages 2-3% of pixels differ because each uses its own substitute fonts.

To run them (needs qpdf, gs and Chromium; the corpus is generated on first run):

bun install --cwd bench && bun run build
bun bench/run.ts && bun bench/features.ts
cp bench/.out/results.json bench/.out/features.json bench/ && bun bench/readme.ts

How it works

  1. Index. Read the cross-reference chain (tables, streams, incremental updates) into typed arrays, 13 bytes per object. Broken offsets are found again by scanning; an unusable index is rebuilt.
  2. Scan. Read each object's dictionary only, to find masks, signatures and old xref data.
  3. Write. Copy unchanged objects as byte ranges, write what plugins changed, then one new cross-reference section.

Known limitations

  • Encrypted PDFs must be decrypted first.
  • Rewriting a signed PDF invalidates its signatures (the report says so).
  • Linearization is lost.
  • Not recompressed: CMYK, indexed, spot-color and calibrated images; JBIG2, CCITT; bit depths other than 8 (except JPEG 2000); inline images.
  • No deduplication or font subsetting.

Development

bun install
bun run typecheck
bun test                     # unit, corpus/e2e (qpdf + MuPDF rendering), fuzz, codec contract, browser
bun run size                 # bundle size per import; fails if compression exceeds 50 KB
bun run build                # tsc -> dist/ (ESM + .d.ts)
bun run test:large           # > 500 MB memory-bound check
bun run site:dev             # the website, with the in-browser app, at http://localhost:5173
bun run site:build           # static site in site/dist (see site/README.md for Cloudflare Pages)
bun bench/run.ts             # benchmarks (bun install --cwd bench first)
bun test/browser/corpus.ts <dir> --png out/   # render a folder of PDFs, compare with MuPDF

For a large corpus, pdf.js's test PDFs work well: clone mozilla/pdf.js and point the corpus runner at test/pdfs.

Browser tests need Chromium (bunx playwright-core install chromium); validation tests need qpdf.

Releasing

Set the version in package.json, add a CHANGELOG.md section and push to main. Then push a vX.Y.Z tag, or run the Publish workflow on main with dry-run off. It tests, publishes to npm with provenance and creates the GitHub release.