@dudko.dev/pdf-to-md-cli
v0.3.5
Published
Command-line PDF → Markdown converter
Maintainers
Readme
@dudko.dev/pdf-to-md-cli
pdf2md — PDF to Markdown on the command line. The parsing is a WebAssembly
module running locally; nothing is uploaded anywhere.
npx @dudko.dev/pdf-to-md-cli report.pdf
npm install -g @dudko.dev/pdf-to-md-cli && pdf2md report.pdfThe hosted service
pdf2md.dudko.dev does the same in a browser, with
the file never leaving the tab — worth a look before installing anything. For a
pipeline that is not this command, the HTTP API is at
pdf2md.dudko.dev/api and an MCP endpoint for
assistants at https://pdf2md.dudko.dev/mcp; both need a free account at
auth.dudko.dev.
A paid account carries commercial permission for that service, not for this package: running the software yourself needs the written licence below.
Pictures rather than documents:
vectorize.dudko.dev, with
@dudko.dev/vectorize-cli
on the command line.
Use
pdf2md report.pdf # Markdown to stdout
pdf2md report.pdf -o report.md
pdf2md *.pdf -o out/ # each input named after itself
cat report.pdf | pdf2md - # from stdin
pdf2md report.pdf -o out/ --images # plus out/report.assets/img-1.png…, linked from the Markdown
pdf2md scan.pdf -o out/ --images --jpx keep # JPEG 2000 kept as .jp2 rather than decoded to PNG
pdf2md report.pdf -o out/ --images --no-vectors # pictures only; no SVG of the charts and diagrams
pdf2md report.pdf --json # the whole result, not just the Markdown
pdf2md report.pdf --detect # type, pages, which pages need OCR (JSON; add --json to pretty-print)
pdf2md report.pdf --text # plain text, no Markdown structure
pdf2md report.pdf --pages 1,3,5-7 --profile compact
pdf2md secret.pdf --password hunter2
pdf2md paper.pdf --page-markers --strip-headers-footers--help lists everything, including the --no-* flags that turn off individual
detectors (headings, lists, code, bold, italic, URL rewriting, hyphenation
repair) and --underline, which turns on the one that is off by default: <u>
around text with a line under it.
Exit codes
| Code | Meaning | | --- | --- | | 0 | everything converted | | 1 | bad usage, or nothing could be read | | 2 | nothing could be parsed | | 3 | some files converted, some did not |
Made for pipelines: 3 is the one that means "look at stderr".
Not OCR
A scanned PDF is reported as Scanned with no Markdown, and the command exits
2. Run it through an OCR tool first, then convert the result.
Licence
Free for noncommercial use under PolyForm Noncommercial 1.0.0, with a 32-day trial for evaluation at work. Commercial use needs a licence: [email protected].
Your rights come from LICENSE and, for commercial use, from a written
agreement. No web page grants them; a paid account at
pdf2md.dudko.dev is permission to use that service.
The notices for the third-party code compiled into the module are in
THIRD-PARTY-NOTICES.md in @dudko.dev/pdf-to-md-core, which this package installs.
