npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

wacz-pdf

v1.1.0

Published

Render the HTML pages in a WACZ web archive to PDF, one per page, via wabac.js replay in headless Chromium

Readme

wacz-pdf

Turn the pages inside a WACZ web archive into PDF files, one PDF per page.

Pages come from the crawler's own page list (pages/pages.jsonl and pages/extraPages.jsonl), falling back to the CDX index for archives written without one. Each page is then handled according to what it is:

  • HTML is replayed through wabac.js, Webrecorder's own replay engine, running as a service worker inside headless Chromium, so subresources (CSS, JS, images, fonts) are served from the archive and the page renders faithfully. The replayed page is printed with page.pdf().
  • PDFs are copied straight out of the archive. An archived PDF is already a PDF; replaying one would only screenshot Chromium's PDF viewer. Copying the original bytes is lossless — text layer, fonts, vectors, bookmarks and page count all survive — and needs no browser at all.

That second path is not a niche case. In one government-site crawl, 758 of the 925 pages were PDFs: they extract in well under a second, against minutes of Chromium time for a far worse result.

Install

Rendering uses Playwright's Chromium, so download it once (a one-time ~150 MB fetch into Playwright's shared cache):

npx playwright install chromium

Then either run without installing:

npx wacz-pdf archive.wacz -o pdfs

or install globally for a wacz-pdf command on your PATH:

npm install -g wacz-pdf
wacz-pdf archive.wacz -o pdfs

Requires Node.js 18+.

Usage

# Turn every page in the archive into a PDF under ./pdfs/
wacz-pdf archive.wacz -o pdfs

# Just list the pages found, with what each one is (no output written)
wacz-pdf archive.wacz --list

# Only the archived PDFs, extracted without starting a browser
wacz-pdf archive.wacz -o pdfs --include '\.pdf$'

# A4, landscape, print stylesheet, first 10 pages only
wacz-pdf archive.wacz -o pdfs --format A4 --landscape --print-media --limit 10

(With npx, prefix each command with npx , e.g. npx wacz-pdf archive.wacz --list.)

Options

| Flag | Description | | --- | --- | | -o, --out <dir> | Output directory (default pdfs) | | --format <name> | Paper size: Letter, A4, Legal, … (default Letter) | | --landscape | Landscape orientation | | --single-page | One continuous page per article, sized to content (no pagination) | | -j, --concurrency <n> | Render n pages in parallel (default 1; auto = cores − 2) | | --print-media | Use print CSS instead of screen CSS (screen is the default) | | --include <re> | Keep only pages whose URL matches this regex (repeatable) | | --exclude <re> | Skip pages whose URL matches this regex (repeatable) | | --list | List the pages found (timestamp, kind, URL, title), write nothing | | --limit <n> | Process at most n pages | | --no-extract | Replay archived PDFs in the browser instead of copying them out (rarely what you want) | | --inject <js> | Run JS in each page before printing; @file reads from a file (repeatable) |

URL filters use JavaScript regex syntax, matched case-insensitively against the full URL. --include is applied before --exclude. Example: only article pages, minus tag listings: --include '/\\d{4}/' --exclude '/tag/'.

How it works

WACZ ─► read page list + CDX                             src/wacz.ts
     ├─ PDF  ─► copy bytes out of the WARC ─► file       src/extract.ts, src/warc.ts
     └─ HTML ─► serve sw.js + WACZ (HTTP Range)          src/server.ts
                headless Chromium + wabac service worker
                └─ replay in an <iframe> ─► page.pdf()   src/render.ts

The two passes share one output sequence, so filenames stay in page order no matter which pass wrote them. The replay server and browser only start if there is HTML to print.

Replay content is loaded inside an iframe rather than the top frame: wabac serves iframe requests as rewritten replay content, whereas a top-frame navigation returns its interactive replay UI instead.

Where the page list comes from

pages/*.jsonl is the crawler's own record of what it treated as a page. The CDX is a poor substitute for it: "every text/html 200 in the archive" also means every iframe, ad frame and XHR-fetched fragment, and it cannot tell a PDF the crawler navigated to from one it merely happened to fetch. So the CDX fallback is deliberately narrower — HTML 200s only, the conservative list.

The CDX is read either way, because it is the only place that records where each resource's bytes live (filename/offset/length). That is what makes direct extraction possible.

Extracting a resource

WACZ requires archive/*.warc.gz be Stored (uncompressed) inside the zip, so a byte range in the zip is a byte range in the file, and each WARC record is its own gzip member. One archived resource therefore costs a single ranged read plus one gunzip, whatever the archive's size. From there src/warc.ts strips the WARC and HTTP headers and returns the entity body, handling chunked transfer-encoding, Content-Encoding, and revisit records (which hold no payload and are resolved back to the original capture by digest).

Two details matter for correctness:

  • Cut the body at the HTTP Content-Length. Otherwise the WARC record's 4-byte separator comes along and the bytes no longer match their digest.
  • Verify before decoding. The recorded digest covers the body as stored, i.e. before Content-Encoding is reversed.

Every extraction is checked against the digest in the CDX, so a mismatch fails that page rather than writing a corrupt file.

Reading the WACZ

All of this rests on a small, self-contained ZIP reader (src/zipread.ts) that parses the archive's central directory and reads entries — or slices of them — by byte range, so multi-gigabyte archives never have to be loaded whole. This is deliberately hand-rolled rather than pulled from a library:

  • There is no wacz package on npm. @harvard-lil/js-wacz is for creating and validating WACZ files, not enumerating pages to render.
  • @webrecorder/wabac (already a dependency, for replay) exposes a ZipRangeReader, but its loaders are browser-oriented (fetch, Blob, FileSystemFileHandle) with no Node filesystem loader — using it here would mean writing an fs-backed loader anyway, i.e. re-implementing what zipread.ts already does, while coupling to an undocumented internal API.

So the reader stays local: a couple hundred lines of synchronous, dependency- free, Zip64-aware code scoped exactly to the need.

Injecting JavaScript

Archived pages sometimes capture a modal ("register or sign in") overlaying the content, with the page scroll-locked behind it. The content is still in the DOM — the overlay is just painted on top. --inject runs a snippet inside the replayed page, after it loads but before it's printed, so you can clean it up:

# inline
wacz-pdf archive.wacz --inject \
  "document.querySelectorAll('.modal,[role=dialog]').forEach(e=>e.remove());\
   document.documentElement.style.overflow='auto'"

# or from a file (repeatable), e.g. a reusable cleanup.js
wacz-pdf archive.wacz --inject @cleanup.js

See examples/dismiss-modal.js for a starting-point script that removes a "register / sign in to keep reading" modal and unlocks page scrolling so the underlying article prints.

Notes:

  • The script runs in the replayed page's own context (the iframe), via Playwright's evaluation channel, so it works even when the archived page sets a restrictive Content-Security-Policy.
  • It's best-effort: a script that throws logs nothing and does not fail the page's render, so write defensively (e.g. optional chaining).

Development

The source is TypeScript under src/, compiled to dist/ with tsc.

git clone <repo> && cd wacz-pdf
npm install
npx playwright install chromium

npm run build     # compile src/ -> dist/
npm run lint      # eslint
npm test          # build, then unit + end-to-end (renders a fixture to PDF)
npm run test:unit # build, then unit only (no browser needed)

Run the local build with node dist/cli.js <archive.wacz> … (or npm link once for a global wacz-pdf that points at your working copy).

The end-to-end test needs the Playwright Chromium browser (npx playwright install chromium); set WACZ_PDF_SKIP_E2E=1 to skip it. CI (GitHub Actions) runs lint, build, and the full test suite on push and PRs.

Performance

Extracting archived PDFs is essentially free — a ranged read and a gunzip per file, no browser — so an archive that is mostly PDFs finishes in seconds.

Rendering is the expensive half, and is CPU-bound: each page is a real Chromium render. By default pages render one at a time. Pass -j <n> (or -j auto) to render several in parallel, where each runs in its own tab/renderer process, so it scales roughly linearly up to your physical core count.

All tabs share a single browser and a single wabac service worker, so the archive's index is loaded only once no matter how many workers you use. The trade-off at high concurrency is memory: N tabs accumulate N× the renderer state over a long run.

Known limitations

  • JS-heavy / SPA pages may render partially if not all their requests were captured.
  • Pages captured multiple times are deduplicated to the most recent capture.
  • Pages that are neither HTML nor PDF (Word documents, plain text, …) are reported and skipped.
  • Pages the crawler recorded with a non-2xx status are dropped. Older pages.jsonl files record no status, so nothing is dropped for those.
  • ZipNum-clustered CDX indexes are read whole (fine for typical archives; very large indexes are loaded into memory).
  • A revisit whose original capture lives in a different WACZ of a multi-part crawl can't be resolved, and fails that page.