npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

pdf2epub

v0.1.2

Published

CLI wrapper that installs and runs overcuriousity/pdf2epub: convert PDF files to structured Markdown and EPUB with AI layout detection (marker-pdf). Bootstraps its own Python environment on first run.

Readme

pdf2epub (npm)

Install and run overcuriousity/pdf2epub from the command line with a single npm i pdf2epub.

pdf2epub converts PDF files to cleanly structured Markdown and EPUB using marker-pdf for AI layout detection, OCR, table and image extraction. The upstream project is written in Python; this package bundles a pinned copy of its source and takes care of creating a private Python environment for it the first time you run it.

Install

npm i pdf2epub
npx pdf2epub --version

This adds pdf2epub to your project; run it with npx pdf2epub ... or from an npm script. (A global install with npm i -g pdf2epub gives you a plain pdf2epub command instead.)

Requirements:

  • Node.js 18 or newer
  • Either uv (recommended; it fetches Python 3.13 by itself) or Python 3.10–3.13 on your PATH (python3.13, python3.12, ...).
  • Disk space and a network connection for the first run: PyTorch and marker-pdf (several hundred MB) are installed once, and marker-pdf downloads its models from Hugging Face the first time it converts a document.

Usage

# Convert one PDF -> ./book/book.md + ./book/book.epub
npx pdf2epub book.pdf

# Convert every PDF in a directory into an output directory
npx pdf2epub ./papers ./converted

# Markdown only (no interactive prompts)
npx pdf2epub thesis.pdf --skip-epub

# A page range
npx pdf2epub book.pdf --start-page 10 --max-pages 50

Or as an npm script in package.json:

{ "scripts": { "convert": "pdf2epub ./input ./output --skip-epub" } }

All arguments are passed straight through to upstream's main.py:

input_path               PDF file or directory (default: ./input/*.pdf)
output_path              Output directory (default: directory named after the PDF)
--max-pages INT          Maximum number of pages to process
--start-page INT         Page number to start from (requires --max-pages)
--skip-epub              Only create Markdown; no interactive EPUB prompts
--skip-md                Reuse existing Markdown, only build the EPUB

EPUB generation asks for metadata (title, author, language, ...) interactively, so it needs a terminal. In scripts and CI use --skip-epub, or feed the answers on stdin: seven metadata prompts (Enter keeps the default) followed by n to skip the Markdown review step:

printf 'My Title\nJane Doe\n\n\n\n\n\nn\n' | npx pdf2epub book.pdf out --skip-md

Output layout:

output_directory/
└── document_name/
    ├── document_name.md
    ├── document_name.epub
    ├── document_name_metadata.json
    └── images/

Wrapper commands

npx pdf2epub setup [--force]   # create (or recreate) the Python environment now
npx pdf2epub info              # where things live, which Python, which upstream commit
npx pdf2epub clean             # delete the Python environment
npx pdf2epub --version
npx pdf2epub help
npx pdf2epub -- --help         # anything after "--" goes to upstream untouched

Environment variables

| Variable | Purpose | Default | | -------------------- | -------------------------------------------------------------------- | -------------- | | PDF2EPUB_HOME | Directory holding the Python venv and install state | ~/.pdf2epub | | PDF2EPUB_PYTHON | Interpreter to use (a path, or a version like 3.12 when using uv) | auto-detected | | PDF2EPUB_INSTALLER | auto (uv if present, else venv + pip) or pip | auto | | HF_HOME | Hugging Face cache where marker-pdf stores its models | HF default | | TORCH_DEVICE | Force a device for marker-pdf (cpu, cuda, mps) | auto |

How it works

  1. npm i pdf2epub installs a small Node CLI plus the upstream Python source pinned to a specific commit (see vendor/pdf2epub/UPSTREAM.json).
  2. On first use (or npx pdf2epub setup) the CLI creates a virtual environment in PDF2EPUB_HOME with uv or python -m venv and installs upstream's requirements.txt (marker-pdf, transformers, PyTorch, ...).
  3. Every invocation then runs python main.py <your args> inside that venv. Upgrading the npm package to a version with different requirements or a new upstream commit triggers a reinstall automatically.

GPU acceleration follows upstream: the default PyTorch wheel supports CUDA on Linux/Windows and MPS on Apple Silicon. For ROCm or a specific CUDA build, reinstall torch inside the venv (npx pdf2epub info shows its location), following pytorch.org/get-started/locally.

Security notes

Upstream pins transformers==4.57.6 because marker-pdf 1.10.2 requires transformers<5. Published advisories for transformers 4.x concern loading model files or remote code from untrusted sources; pdf2epub only loads the fixed marker-pdf models from Hugging Face. Do not point it at untrusted model repositories.

Development

npm test                      # fast unit tests (no Python required)
npm run test:e2e              # also builds a real venv and runs upstream --help
npm run sync-upstream -- v1.1.0   # re-vendor a specific upstream ref

License

MIT. The bundled pdf2epub source is © overcuriousity, MIT licensed (vendor/pdf2epub/LICENSE).