@lucidaquarian/pdf-merge
v1.6.1
Published
Merge PDFs from Base64 strings, Buffers, file paths, or URLs — and assemble PDFs from text (.txt) and image (.png/.jpg) inputs — in Node.js and the browser, with page selection, document metadata, text & image watermarks, page numbers, AbortSignal cancell
Downloads
552
Maintainers
Keywords
Readme
@lucidaquarian/pdf-merge
Merge PDFs in Node.js or TypeScript — from Base64, Buffers, file paths, or URLs — with per-input page selection, document metadata, text & image watermarks, page numbers,
AbortSignalcancellation, hardened URL fetching, typed errors, a CLI, and dual ESM/CommonJS support. Built onpdf-lib.
Why this library?
Most PDF-merge packages on npm cover the basics — take an array of
buffers, concatenate them. @lucidaquarian/pdf-merge adds the
production-grade extras that real applications usually have to bolt on
themselves:
- Four input sources in one API — Base64,
Buffer/Uint8Array, file paths, and URLs — so you don't have to pre-decode or pre-fetch. - Assemble from mixed formats — drop
.png,.jpg, and.txtinputs in alongside PDFs; images become one page each and text is rendered with full Unicode via a bundled, lazy-loaded font (see Mixed-format inputs). - Per-input page selection via array (
[1, 3, 5]) or range string ("1-3,5,8-10") — pick exactly the pages you want from each source. - Document metadata — stamp title, author, subject, keywords, creator, and creation/modification dates on the merged output.
- Watermarks and page numbers — overlay configurable text or an
image / logo (PNG / JPEG) on every page, plus continuous page numbering
with
{current}/{total}templates. AbortSignalcancellation — cancel mid-merge from React effects, request handlers, or batch jobs; in-flight URL fetches abort cleanly.- Hardened URL fetching — protocol allowlist, response size cap,
per-request timeout, redirect-protocol checks, bounded concurrency, and URL
sanitization in error messages so signed-URL tokens never reach your logs.
Opt into
blockPrivateHosts/allowedHostsfor SSRF protection when URLs are untrusted (see SSRF hardening). - Typed errors with input-index correlation —
instanceof PdfFetchErrortells you exactly which URL failed and gives you the sanitized.urland.index. - Bundled
pdf-mergeCLI (see CLI for the recommended invocation forms). - Dual ESM and CommonJS build — works in modern bundlers, Deno, Bun,
and legacy
require()consumers from a single install. - Order-preserving — output pages always follow the input array order.
- Zero runtime config — sensible defaults, fully configurable per call.
Install
npm install @lucidaquarian/pdf-mergeRequires Node.js 18+ (uses the global fetch and AbortController), and
works in modern browsers via bundlers — see Browser usage.
Quick start
import {
mergeBase64PDFs,
mergePdfBuffers,
mergePdfFiles,
mergePdfUrls,
} from '@lucidaquarian/pdf-merge';
// From Base64 strings → Base64 string
const b64Out = await mergeBase64PDFs([pdfA_b64, pdfB_b64]);
// From raw bytes → Uint8Array (no Base64 round-trip)
const bytesOut = await mergePdfBuffers([bufA, bufB]);
// From disk → Uint8Array
const fileOut = await mergePdfFiles(['./a.pdf', './b.pdf']);
// From URLs → Base64 string
const urlOut = await mergePdfUrls(['https://example.com/a.pdf', 'https://example.com/b.pdf']);
// With watermark, page numbers, metadata, and cancellation
const controller = new AbortController();
const annotated = await mergePdfFiles(['./cover.pdf', './body.pdf'], {
metadata: { title: 'Annual Report 2026', author: 'Operations' },
watermark: { text: 'CONFIDENTIAL', opacity: 0.15, rotate: 45 },
pageNumbers: { format: 'Page {current} of {total}' },
signal: controller.signal,
});Merged file info (page count & size)
Every merge function has a *WithInfo sibling that returns the merged PDF
together with its total page count and byte size — computed during the
merge, so there's no need to re-parse the output:
import { mergePdfBuffersWithInfo } from '@lucidaquarian/pdf-merge';
const { pdf, pageCount, byteLength } = await mergePdfBuffersWithInfo([bufA, bufB]);
// pdf -> Uint8Array (exactly what mergePdfBuffers returns)
// pageCount -> total pages in the merged document
// byteLength -> size of the merged PDF in bytesThe variants — mergeBase64PDFsWithInfo, mergePdfBuffersWithInfo,
mergePdfFilesWithInfo (Node only), and mergePdfUrlsWithInfo — take the
same arguments and options as the plain functions and return a
MergeResult<T>:
interface MergeResult<T> {
pdf: T; // string for the Base64/URL variants, Uint8Array otherwise
pageCount: number;
byteLength: number; // size of the PDF bytes (for Base64/URL variants,
// the decoded size — not the Base64 string length)
}The plain functions (mergePdfBuffers, …) are unchanged: each simply returns
the pdf field of its *WithInfo counterpart.
Mixed-format inputs (text & images)
mergePdfFiles and mergePdfBuffers accept .txt, .png, and
.jpg/.jpeg inputs alongside PDFs. Each non-PDF input is converted to a
PDF and merged in order — so you can staple a cover image and a text note onto
a PDF in one call:
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
// Format is auto-detected from the file extension.
const out = await mergePdfFiles(['cover.png', 'report.pdf', 'notes.txt']);With buffers, PNG/JPEG are detected from their magic bytes, but plain text
has no signature, so mark it with format: 'txt':
import { mergePdfBuffers } from '@lucidaquarian/pdf-merge';
const out = await mergePdfBuffers([
pdfBytes, // auto-detected PDF
{ data: pngBytes, format: 'png' }, // (or just pngBytes — magic bytes)
{ data: new TextEncoder().encode('Hello, 世界 needs a font override'), format: 'txt' },
]);Images become one page each. By default each image is fitted to a US
Letter page, the same size and margins as converted text pages: 0.75 in,
which also keeps page numbers clear of the image.
The page turns landscape for a landscape image, and the image is scaled down
to fit, centered, with its aspect ratio kept. Small images such as logos are
never enlarged. The full-resolution image is embedded either way, so zooming
in on the PDF still shows every pixel. Configure this via options.image:
await mergePdfFiles(['report.pdf', 'photo.jpg'], {
image: {
fit: 'page', // default; 'natural' sizes the page to the image itself
pageSize: 'a4', // 'letter' (default), 'a4', or [width, height] in points
margin: 0, // points; default 54 (0.75in); 0 fills the page edge to edge
},
});With fit: 'natural', the page is sized to the image at one pixel per point
(1/72 inch), the behaviour before 1.6.0. A modern photo then produces a very
large page; a 12-megapixel photo becomes about 56 × 42 inches, which looks
heavily zoomed in next to normal pages. Natural pages are capped at the PDF
format's 200-inch maximum.
Text is rendered as multi-page US-Letter pages with full Unicode
via a bundled DejaVu Sans font (Latin,
Cyrillic, Greek, and many symbols). The font is lazy-loaded — PDF- and
image-only merges never pull it in. Line breaks are kept (LF, CRLF, CR,
vertical tab and the Unicode line/paragraph separators), a form feed starts a
new page, tabs expand to four spaces, and long lines wrap. Configure layout, or
supply your own font for scripts DejaVu doesn't cover (e.g. CJK/emoji), via
options.text:
import { readFile } from 'node:fs/promises';
await mergePdfFiles(['notes.txt'], {
text: {
fontSize: 12, // default 11
margin: 54, // points; default 54 (0.75in)
lineHeight: 1.4, // × font size; default 1.35
font: await readFile('./NotoSansCJK.ttf'), // override the bundled font
},
});Notes and limits:
- Detection override: a per-input
formatfield ('pdf' | 'txt' | 'png' | 'jpg' | 'jpeg') overrides auto-detection. - Page selection (
pages) applies to PDF inputs only; it's ignored for converted text/image inputs. - Size caps: text inputs are capped at 200 KB (→
PdfMergeError), and conversion cost is linear in that size, so even pathological inputs finish within a few seconds. Images have no file-size cap, but an image is rejected if the dimensions declared in its header exceedmaxImagePixels(default 100 megapixels). For PNGs this blocks decompression bombs; for JPEGs (never decoded) it prevents an absurd header from producing an out-of-spec page. Raise the option for legitimately huge scans, or lower it for untrusted uploads. See PNG decompression bombs. Unknown formats and malformed images throw typed errors; a truncated or malformed PNG is rejected withInvalidPdfFormatErrorbefore decoding. Animated PNGs embed their single frame as before; ones with more than one frame are rejected, aspdf-libcannot embed them. - Characters missing from the font render as its fallback glyph rather than
throwing — pass
options.text.fontfor full coverage of other scripts. - Package size: bundling the default font adds ~1 MB to the install, but it's lazy-loaded, so it never affects PDF/image-only merges at runtime.
Browser usage
@lucidaquarian/pdf-merge ships a browser-safe build. When you import it in a
bundler (Vite, webpack, Next.js, Rollup, esbuild, …), the package's browser
export condition
automatically selects it — no config needed:
import { mergeBase64PDFs, mergePdfBuffers, mergePdfUrls } from '@lucidaquarian/pdf-merge';
// In the browser you usually already have the bytes (a File, a fetch body, …):
const bytes = new Uint8Array(await file.arrayBuffer());
const merged = await mergePdfBuffers([bytes, otherBytes]); // -> Uint8ArrayYou can also import the browser entry explicitly (handy for tooling that doesn't apply export conditions, and for browser-accurate types):
import { mergeBase64PDFs } from '@lucidaquarian/pdf-merge/browser';What works in the browser: mergeBase64PDFs, mergePdfBuffers, and
mergePdfUrls, plus all options (page selection, metadata, watermark, page
numbers, AbortSignal).
Differences from Node:
mergePdfFilesis not available — there's no filesystem in the browser. It's exported only from the Node build.mergePdfUrlsuses the browser'sfetch, so requests are subject to CORS — the PDF host must sendAccess-Control-Allow-Origin. Redirects are handled by the browser, not the library.- The SSRF host guard is limited.
blockPrivateHosts/allowedHostscan only check the initial URL's plainly-formatted literal IP /localhostin the browser — there's no DNS resolution and redirects can't be intercepted, so hostnames and obfuscated IP forms (e.g.http://127.1/, decimal, or hex addresses) are not caught client-side. In the browser, CORS is the effective SSRF boundary; the full DNS + per-redirect guard (and DNS resolution of obfuscated forms) runs in Node only.
Recipes
Copy-paste starting points for the things people most often need.
Merge every PDF in a folder
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
import { promises as fs } from 'fs';
import path from 'path';
const dir = './invoices';
const files = (await fs.readdir(dir))
.filter((f) => f.toLowerCase().endsWith('.pdf'))
.sort() // readdir order is not guaranteed — sort for predictable page order
.map((f) => path.join(dir, f));
const merged = await mergePdfFiles(files);
await fs.writeFile('./all-invoices.pdf', merged);Merge and stream back from an Express / Fastify handler
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
app.get('/report.pdf', async (req, res) => {
const merged = await mergePdfFiles(['./cover.pdf', './body.pdf']);
res.setHeader('Content-Type', 'application/pdf');
res.setHeader('Content-Disposition', 'inline; filename="report.pdf"');
res.end(Buffer.from(merged)); // merged is a Uint8Array
});Merge Base64 PDFs and return Base64 (e.g. serverless / JSON APIs)
import { mergeBase64PDFs } from '@lucidaquarian/pdf-merge';
// Inputs and output are Base64 strings — no filesystem needed.
const mergedB64 = await mergeBase64PDFs([pdfA_b64, pdfB_b64]);
return { statusCode: 200, body: JSON.stringify({ pdf: mergedB64 }) };Fetch PDFs from authenticated URLs (signed S3, bearer tokens)
import { mergePdfUrls } from '@lucidaquarian/pdf-merge';
const merged = await mergePdfUrls(
[
// Per-URL header — e.g. a bearer token for one specific source.
{ url: 'https://api.example.com/a.pdf', headers: { Authorization: 'Bearer TOKEN_A' } },
// A pre-signed S3 URL needs no header — the signature is in the query string.
'https://bucket.s3.amazonaws.com/b.pdf?X-Amz-Signature=...',
],
{
// A default header applied to every request (per-URL headers win on conflict).
headers: { 'User-Agent': 'my-app/1.0' },
// Only allow https so a malicious input can't downgrade to http.
allowedProtocols: ['https:'],
timeoutMs: 10_000,
concurrency: 4,
},
);Pick specific pages from each source
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
const merged = await mergePdfFiles([
'./cover.pdf', // all pages
{ path: './body.pdf', pages: '2-9' }, // pages 2 through 9
{ path: './appendix.pdf', pages: [1, 3, 5] }, // pages 1, 3 and 5
]);Add a "DRAFT" watermark and continuous page numbers
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
const merged = await mergePdfFiles(['./a.pdf', './b.pdf'], {
watermark: {
text: 'DRAFT',
opacity: 0.2,
rotate: 45,
color: { r: 0.8, g: 0.1, b: 0.1 },
},
pageNumbers: {
format: 'Page {current} of {total}',
position: 'bottom-center',
},
});Prepend a clickable table of contents
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
// Local files are labelled by their basename automatically.
const merged = await mergePdfFiles(['./cover.pdf', './report.pdf', './appendix.pdf'], {
tableOfContents: {},
});
// Or set your own labels per document with `title`:
const withTitles = await mergePdfFiles(
[
{ path: './cover.pdf', title: 'Cover' },
{ path: './report.pdf', title: 'Q3 Report' },
{ path: './appendix.pdf', title: 'Appendix A' },
],
{ tableOfContents: { heading: 'Contents' } },
);Each entry links to the page that document starts on, and the printed numbers are the physical page positions — so the TOC page is page 1 and the numbers line up with what a PDF viewer shows. See Table of contents for all options.
Cancel a long merge from a React effect
import { useEffect, useState } from 'react';
import { mergePdfUrls } from '@lucidaquarian/pdf-merge';
function useMergedPdf(urls: string[]) {
const [pdf, setPdf] = useState<string | null>(null);
useEffect(() => {
const controller = new AbortController();
mergePdfUrls(urls, { signal: controller.signal })
.then(setPdf)
.catch((err) => {
if (err?.name !== 'AbortError') throw err; // ignore user-cancels
});
// Aborts the in-flight fetch + merge if the component unmounts or urls change.
return () => controller.abort();
}, [urls]);
return pdf;
}Note: this library currently targets Node.js — it uses Node's
Bufferinternally. To run it in a browser bundle you'll need aBufferpolyfill (most bundlers can inject one); a dedicated browser build is planned. TheAbortSignalcancellation pattern above applies during SSR / server-side rendering as-is.mergePdfFilesis Node-only regardless, since it reads from disk.
Set document metadata (title, author, keywords)
import { mergePdfFiles } from '@lucidaquarian/pdf-merge';
const merged = await mergePdfFiles(['./a.pdf', './b.pdf'], {
metadata: {
title: 'Q3 Financial Report',
author: 'Finance Team',
subject: 'Quarterly results',
keywords: ['finance', 'q3', '2026'],
creator: 'my-app',
},
});Merge from the command line (no code)
# One-shot with npx (use the full scoped name)
npx @lucidaquarian/pdf-merge cover.pdf body.pdf appendix.pdf -o out.pdf
# Pick pages per input with the :range suffix
npx @lucidaquarian/pdf-merge cover.pdf 'body.pdf:2-9' -o out.pdf
# Mix local files and URLs, add metadata
npx @lucidaquarian/pdf-merge \
cover.pdf https://example.com/body.pdf \
-o report.pdf --title 'Report' --author 'Ops'API
mergeBase64PDFs(inputs, options?): Promise<string>
Merges Base64-encoded PDFs and returns the merged document as Base64.
- Accepts plain Base64 or
data:application/pdf;base64,…prefixed strings. inputsis(string | { data: string; pages?: PageSelector; title?: string })[].optionsacceptsmetadata,watermark,pageNumbers,tableOfContents,ignoreEncryption,maxImagePixels, andsignal. See Document metadata, Watermark, Page numbers, Table of contents, Encrypted source PDFs, PNG decompression bombs, and Cancellation.
const merged = await mergeBase64PDFs([
reportA_b64, // all pages of A
{ data: reportB_b64, pages: [1, 3, 5] }, // pages 1, 3, 5 of B
{ data: reportC_b64, pages: '1-3,7' }, // pages 1, 2, 3, 7 of C
]);mergePdfBuffers(inputs, options?): Promise<Uint8Array>
Merges raw PDF bytes (from fs.readFile, an S3 SDK, an HTTP body, multer,
etc.) and returns the merged document as a Uint8Array. Skips the Base64
round-trip, saving ~33 % memory vs mergeBase64PDFs.
inputsis(Buffer | Uint8Array | { data: Buffer | Uint8Array; pages?: PageSelector; title?: string; format?: InputFormat })[]. Non-PDF inputs (.txt,.png,.jpg) are converted and merged in place; see Mixed-format inputs.optionsacceptsmetadata,watermark,pageNumbers,tableOfContents,text,image,ignoreEncryption,maxImagePixels, andsignal. See Document metadata, Watermark, Page numbers, Table of contents, Mixed-format inputs (fortextandimage), Encrypted source PDFs, PNG decompression bombs, and Cancellation.
const merged = await mergePdfBuffers([bufA, { data: bufB, pages: [4, 2] }]);mergePdfFiles(inputs, options?): Promise<Uint8Array>
Reads PDFs from the given file paths in parallel and merges them. Returns a
Uint8Array — write it straight to disk.
inputsis(string | { path: string; pages?: PageSelector; title?: string })[]. In a generated table of contents, entries are labelled by the file's basename unless you pass atitle. Non-PDF files (.txt,.png,.jpg) are converted and merged in place; see Mixed-format inputs.optionsacceptsmetadata,watermark,pageNumbers,tableOfContents,text,image,ignoreEncryption,maxImagePixels, andsignal. See Document metadata, Watermark, Page numbers, Table of contents, Mixed-format inputs (fortextandimage), Encrypted source PDFs, PNG decompression bombs, and Cancellation.
import { promises as fs } from 'fs';
const merged = await mergePdfFiles([
'./cover.pdf',
{ path: './body.pdf', pages: '2-9' },
]);
await fs.writeFile('./out.pdf', merged);mergePdfUrls(urls, options?): Promise<string>
Fetches PDFs from the given URLs (concurrently, with a bounded pool) and returns the merged document as Base64.
urls is (string | { url: string; headers?: Record<string,string>; pages?: PageSelector; title?: string })[].
Options:
| Option | Default | Description |
| --- | --- | --- |
| timeoutMs | 5000 | Per-request timeout in milliseconds, covering the host check, the request and the body download. Infinity means no timeout. Must be a positive number. |
| maxBytesPerUrl | 100 * 1024 * 1024 | Maximum bytes accepted from any single response. Infinity removes the cap. Must be a positive number; NaN (for example from an unset environment variable) is rejected rather than silently disabling the cap. |
| allowedProtocols | ['http:', 'https:'] | URL protocols allowed. Pass ['https:'] to harden further. |
| blockPrivateHosts | false | Reject URLs (and redirect targets) that resolve to loopback/private/link-local IPs (127.0.0.1, 10.x, 169.254.169.254, …). Enable when URLs are untrusted. See SSRF hardening. |
| allowedHosts | — | Case-insensitive hostname allowlist. When set, any URL or redirect target whose host isn't listed is rejected. See SSRF hardening. |
| concurrency | 8 | Maximum number of URL fetches running in parallel. |
| headers | — | Default headers applied to every fetch. Per-URL headers (object form) override on key conflict. |
| metadata | — | Document metadata to stamp on the merged output. See Document metadata. |
| watermark | — | Text or image / logo watermark to draw on every page of the merged output. See Watermark. |
| pageNumbers | — | Stamp continuous page numbers. See Page numbers. |
| tableOfContents | — | Prepend a clickable table-of-contents page. See Table of contents. |
| signal | — | AbortSignal to cancel the merge. See Cancellation. |
| ignoreEncryption | false | Allow merging source PDFs that declare encryption metadata. See Encrypted source PDFs. |
| maxImagePixels | 100_000_000 | Pixel-count cap for a PNG image watermark, guarding against decompression bombs. See PNG decompression bombs. |
const merged = await mergePdfUrls(
[
'https://example.com/a.pdf',
{ url: 'https://example.com/b.pdf', headers: { Authorization: 'Bearer token-b' }, pages: '1-3' },
],
{
timeoutMs: 8000,
maxBytesPerUrl: 20 * 1024 * 1024,
allowedProtocols: ['https:'],
concurrency: 4,
headers: { 'User-Agent': 'my-app/1.0' },
},
);Document metadata
Every merge function accepts an optional second options argument with a
metadata field. Stamp the merged document with whatever properties your
downstream system surfaces (Finder, Explorer, document-management systems,
email clients, etc.):
await mergePdfFiles(
['./cover.pdf', './body.pdf'],
{
metadata: {
title: 'Annual Report 2026',
author: 'Operations',
subject: 'Year-end summary',
keywords: ['annual', 'operations', '2026'],
creator: 'my-app/1.0',
creationDate: new Date(),
modificationDate: new Date(),
},
},
);The MergePdfUrlsOptions interface extends the same shape, so all four
functions accept { metadata } the same way. Every field is optional — omit
what you don't want to set.
The
Producerfield is hard-coded by pdf-lib on save and cannot be customized at this layer.
Watermark
Every merge function accepts an optional watermark field on the same
options argument. The watermark is stamped on every page of the merged
output. It can be a text label or an image / logo — provide exactly
one of text or image. The feature is fully opt-in — when watermark is
omitted, no extra drawing is performed and the original byte-identical fast
path is preserved.
position (shared by both kinds) accepts a named placement — 'center'
(default), 'top-left', 'top-right', 'bottom-left', 'bottom-right' —
or an explicit { x, y } in PDF points measured from the bottom-left of the
page. Named placements are computed per page on the page as a viewer shows it,
so they work on any page size, on pages rotated with /Rotate (the mark reads
upright), and on pages whose visible area (CropBox) does not start at the
origin. An explicit { x, y } is taken as-is in the page's own, unrotated
coordinate space.
Text watermark
Drawn with the built-in Helvetica font (no additional fonts are embedded).
await mergePdfFiles(
['./report.pdf'],
{
watermark: {
text: 'CONFIDENTIAL',
opacity: 0.18, // 0–1, default 0.2
fontSize: 90, // points, default 48
color: { r: 0.7, g: 0.1, b: 0.1 }, // RGB 0–1, default mid-gray
rotate: 45, // degrees CCW, default 0
position: 'center', // see above
},
},
);Invalid input throws PdfMergeError: empty text, opacity outside 0–1,
non-positive font size, or color channels outside 0–1.
The text watermark is drawn with the built-in Helvetica font, which only supports Latin-1 (WinAnsi) text. Watermark text containing characters outside that set — CJK, emoji, most non-Latin scripts — throws a
PdfMergeError. The same applies topageNumbers.format.
Image / logo watermark
Stamp a PNG or JPEG on every page. Supply the image as raw bytes
(Uint8Array / ArrayBuffer) or a Base64 string (a bare payload or a
data:image/...;base64, URI). The format is auto-detected from the image's
magic bytes; set format to 'png' or 'jpg' to override. The image is
embedded once and reused across pages, so the size overhead stays flat
regardless of page count.
import { readFile } from 'node:fs/promises';
await mergePdfFiles(
['./report.pdf'],
{
watermark: {
image: await readFile('./logo.png'), // Uint8Array | ArrayBuffer | Base64
// format: 'png', // optional — auto-detected when omitted
width: 120, // points; height derived from aspect ratio
// height: 60, // set either/both; omit both for natural size
// scale: 0.5, // uniform factor on the natural pixel size
opacity: 0.15, // 0–1, default 0.2
rotate: 0, // degrees CCW, default 0
position: 'bottom-right',
},
},
);Sizing precedence: if width and/or height are given they win (the missing
dimension is derived from the image's aspect ratio); otherwise scale
multiplies the natural pixel size; otherwise the image is drawn at its natural
pixel size. For a named position the image is centered on the anchor; for an
explicit { x, y } the point is the image's bottom-left corner.
Invalid input throws PdfMergeError: a missing/mistyped image, an invalid
Base64 string, an unrecognized format, an undetectable image format, or a
non-positive width / height / scale. Setting both text and image
(or neither) also throws.
A watermark image whose header declares more than maxImagePixels pixels
(default 100 megapixels) is rejected before anything is decoded — for PNGs
this guards against decompression bombs; for JPEGs it keeps an absurd header
from being drawn at an out-of-spec size. Pass a higher maxImagePixels in the
merge options for a legitimately huge image. A malformed or truncated PNG, or
an animated one with more than one frame, throws InvalidPdfFormatError. See
PNG decompression bombs.
Page numbers
Stamp continuous page numbering on every page of the merged output.
Common use case: stitch several PDFs together and number 1..N across
the result. Available on the same options argument as everything else.
await mergePdfFiles(
['./cover.pdf', './body.pdf', './appendix.pdf'],
{
pageNumbers: {
format: 'Page {current} of {total}', // tokens: {current}, {total}
startAt: 1, // first-page value; default 1
position: 'bottom-center', // see below
fontSize: 10, // default 10
color: { r: 0.4, g: 0.4, b: 0.4 }, // default mid-gray
},
},
);position accepts a named placement — 'bottom-center' (default),
'top-left', 'top-center', 'top-right', 'bottom-left',
'bottom-right' — or an explicit { x, y } in PDF points from the
bottom-left of the page. As with watermarks, named placements follow the page
as displayed (rotation and visible area), and the number reads upright; an
explicit { x, y } is in the page's own, unrotated coordinates. A page
selected more than once (pages: [1, 1]) shows only its own number on each
copy.
Invalid input throws PdfMergeError (empty format, non-integer
startAt, non-positive fontSize, color channels outside 0–1,
unknown position).
Table of contents
Prepend a printed contents page that lists every source document and the page it starts on. Each row is a clickable internal link that jumps to that document, and rows spill onto additional pages automatically when there are too many to fit. Works in both Node and the browser.
await mergePdfFiles(
['./cover.pdf', './report.pdf', './appendix.pdf'],
{
tableOfContents: {
heading: 'Table of Contents', // default; pass '' to omit
fontSize: 12, // entry rows, default 12
headingFontSize: 20, // default 20
color: { r: 0, g: 0, b: 0 }, // RGB 0–1, default black
pageSize: [612, 792], // default: match the first content page
},
},
);Pass an empty object (tableOfContents: {}) to enable it with the defaults.
Entry labels come from each source's title. For mergePdfFiles, the
label defaults to the file's basename when no title is given; for the
other merge functions it falls back to "Document N". Set a title on the
object form of any input to control the label:
await mergePdfBuffers(
[
{ data: coverBytes, title: 'Cover' },
{ data: reportBytes, title: 'Q3 Report' },
],
{ tableOfContents: {} },
);Page numbering. The listed numbers are the physical page positions
in the final PDF, so the TOC page itself is page 1 and the numbers match
what a PDF viewer shows. When you also enable pageNumbers, the stamped
numbers cover the TOC page(s) too, so both stay in agreement. The returned
pageCount (and the *WithInfo variants) likewise include the TOC page(s).
Page size. Margins scale down on small pages so the layout stays usable.
If the size inherited from the first content page is too small to hold the
heading and one row, the TOC page falls back to US Letter (612×792); an
explicit pageSize that small is rejected instead.
Invalid input throws PdfMergeError (a non-object tableOfContents,
non-string heading, non-positive fontSize / headingFontSize, color
channels outside 0–1, a malformed or too-small pageSize). As with
watermarks, labels and the heading must be Latin-1 (WinAnsi) renderable.
Cancellation with AbortSignal
Every merge function accepts options.signal: AbortSignal. The merge
aborts cleanly at the next source-iteration boundary, and in-flight
HTTP fetches in mergePdfUrls are cancelled immediately, as are pending
host checks (DNS lookups) when blockPrivateHosts or allowedHosts is set.
An already-aborted signal fails the call before any work starts.
const controller = new AbortController();
// Cancel after 30 seconds, or whenever the user clicks "stop".
setTimeout(() => controller.abort(new Error('took too long')), 30_000);
try {
await mergePdfUrls(urls, { signal: controller.signal, timeoutMs: 60_000 });
} catch (err) {
if (err instanceof Error && err.message === 'took too long') {
// user-cancelled — clean up state and move on
} else {
throw err;
}
}If you call controller.abort(reason) with a reason, that exact value
is thrown. Without a reason, the merge throws an AbortError-shaped
DOMException. Either way, signal.aborted checks in caller code
behave correctly (this matches the Node fetch convention).
Encrypted source PDFs
Some PDFs declare encryption metadata but contain readable content
streams — a quirk of older generators. By default mergePdfBuffers and
the other merge functions reject these (matching pdf-lib's behavior).
Pass options.ignoreEncryption: true to opt in:
await mergePdfBuffers([legacyReport], { ignoreEncryption: true });This skips pdf-lib's encryption check at load time. It does not decrypt the document — truly password-protected files will still fail because their content streams cannot be read without the key.
Page selection
Every merge function accepts an object form per input that lets you pick which pages to keep from that document. Pages are 1-indexed to match what you see in a PDF viewer.
PageSelector is either:
- a
number[]— explicit page numbers, e.g.[1, 3, 5]. Duplicates are kept (the page appears multiple times in the output). - a
string— comma-separated ranges, e.g."1-3,5,8-10". Descending ranges ("5-1") reverse the page order.
Out-of-range, zero, negative, or unparseable selectors throw PdfMergeError.
Inputs without a pages field — including all plain-string / plain-Buffer
inputs — behave exactly as they always have and include every page.
Errors
All errors extend PdfMergeError, so a single catch (err: PdfMergeError)
covers everything.
| Error | When it's thrown | Notable fields |
| --- | --- | --- |
| PdfMergeError | Empty input, invalid input shape or option (including wrong metadata types), bad URL, disallowed protocol, invalid page selector. | — |
| InvalidPdfFormatError | Input is not valid Base64, not a parseable PDF, fails the %PDF- header check, or is an image that cannot be embedded (including a truncated or malformed PNG, or an animated PNG with more than one frame). | — |
| PdfFetchError | HTTP failure, timeout (including a DNS lookup that times out), oversize response, redirect to a disallowed protocol, a URL with an embedded username or password, or non-PDF response body. | .url (credentials/query stripped), .index (failing input position) |
import { PdfFetchError } from '@lucidaquarian/pdf-merge';
try {
await mergePdfUrls(urls);
} catch (err) {
if (err instanceof PdfFetchError) {
console.error(`URL #${err.index} failed: ${err.url}`);
} else {
throw err;
}
}Security model
mergePdfUrls is the higher-risk function — it dereferences caller-supplied
URLs. Defaults are chosen to be safe out of the box:
- Protocol allowlist. Only
http:andhttps:are accepted;file:,data:,ftp:, etc. are rejected before any socket is opened. - Response size cap.
maxBytesPerUrlis enforced both against the declaredContent-Lengthand during streaming, so a hostile endpoint that serves an unbounded body cannot exhaust the process heap. - URL sanitization in errors. Credentials (
user:pass@) and query strings are stripped from URLs before they appear in error messages or onPdfFetchError.url, so signed-URL tokens and HTTP Basic passwords don't leak into logs. A URL with an embedded username or password is rejected before any request (fetchrefuses them); send credentials in anAuthorizationheader instead. - Rejected responses are closed. An error status, an oversized response or a rejected redirect has its body discarded at once, so the connection doesn't keep streaming.
- Redirect-protocol check. If a redirect lands on a disallowed protocol, the request is rejected.
- Per-request timeout. Default
5000 msviaAbortController. - Bounded concurrency.
concurrency(default8) caps simultaneous in-flight fetches.
SSRF hardening
By default the protections above are protocol- and size-based: a URL that
points at an internal host over http:/https: (e.g. http://10.0.0.1/… or
the cloud metadata endpoint http://169.254.169.254/…) is still fetched.
If you pass untrusted / user-supplied URLs, opt into host-level blocking:
const merged = await mergePdfUrls(userUrls, {
blockPrivateHosts: true, // reject loopback/private/link-local/reserved IPs
allowedHosts: ['cdn.example.com'], // and/or restrict to an explicit allowlist
allowedProtocols: ['https:'], // often worth pairing with https-only
});blockPrivateHostsresolves each host's DNS and rejects it if the URL — or any redirect target — islocalhostor an address in a loopback, private (10/8,172.16/12,192.168/16), CGNAT (100.64/10), link-local (169.254/16, incl. cloud metadata), or reserved range, for both IPv4 and IPv6 (including IPv4-mapped addresses).allowedHostsrestricts fetches to an explicit hostname allowlist.- When either guard is enabled, redirects are followed manually and each
hop is validated before it is contacted, so a redirect to an internal host
is never reached. Credential headers (
Authorization,Cookie) are dropped on cross-origin redirects.
Residual caveat: the host is resolved and checked, then the request re-resolves DNS to connect — a narrow TOCTOU window remains against an attacker actively rebinding DNS between the two lookups. For the strongest guarantee, also restrict egress at the network layer.
Resource limits for untrusted PDFs
PDFs store their content in compressed streams, and parsing one decompresses
those streams into memory. A hostile PDF can be a decompression bomb — a
small file whose internal streams inflate to something far larger — so
processing fully untrusted input carries a memory/CPU exhaustion (DoS) risk.
This is inherent to PDF parsing (it happens inside pdf-lib), not specific to
this library.
What the library already bounds:
maxBytesPerUrl(default ~100 MiB) caps the downloaded size of each URL response, andtimeoutMscaps how long each fetch may take.
What it does not bound (and can't, at this layer):
- The decompressed / in-memory size after
pdf-libparses a document. A file undermaxBytesPerUrlcan still expand to much more in memory.
If you merge PDFs from untrusted users, add an outer limit yourself:
- Lower
maxBytesPerUrlfor untrusted sources, and pre-check the size of any Buffer/file inputs before passing them in. - Run the merge in a worker thread or child process with a bounded heap
(e.g.
node --max-old-space-size=512), so a bomb kills the worker instead of the host process, and enforce a wall-clock timeout around the whole merge.
Image inputs: PNG decompression bombs
Image inputs and image watermarks are a distinct case, and this one is
bounded for you. pdf-lib decodes a PNG into raw pixels, allocating
width × height × 4 bytes based on the dimensions declared in the header —
before it has read any pixel data. A 68-byte PNG that claims to be
100 000 × 100 000 px would therefore force a multi-gigabyte allocation and
crash the process.
The library rejects such images from the header alone, before pdf-lib ever
sees them, via maxImagePixels (default 100 megapixels, i.e.
100_000_000). It applies to converted image inputs and to image
watermarks. Legitimate large scans that exceed it can be allowed explicitly:
await mergePdfFiles(inputs, { maxImagePixels: 250_000_000 });JPEGs are embedded without decoding (pdf-lib copies the DCT stream as-is),
so they carry no such allocation. Their frame-header dimensions are still
checked against the same cap, because they size the output page: an absurd
header such as 65535 × 65535 would otherwise yield a page far beyond the PDF
specification's 14400-point maximum, which some viewers reject.
Malformed PNGs are rejected before decoding. pdf-lib's PNG decoder
never returns on some damaged files — a cut-off upload, an incomplete
compressed stream, or a text chunk missing its separators — which would hang
the process. It also accepts a header placed anywhere in the file and decodes
animation frames at whatever size they declare, which could slip a huge image
past maxImagePixels. Every PNG is therefore checked first: the header must
come first and appear once, every chunk the decoder reads must fit inside the
file, and the compressed image data must be complete and exactly the size the
header declares. Animated PNGs (APNG) are handled as pdf-lib handles them: a
single-frame one embeds its frame, which must lie within the image and pass
the same data checks; one with more than one frame is rejected up front
(pdf-lib refuses those, but only after decoding every frame). A PNG that fails throws InvalidPdfFormatError.
Harmless quirks that browsers tolerate (bad checksums, a missing end marker,
trailing bytes) are still accepted. The check costs one extra pass over the
compressed data, about 130 ms for a 12-megapixel image.
Budget for the cap you choose. The cap prevents crashes, not cost. A
valid PNG under the limit still has to be decoded: measured at 40 MP, a
152 KB file took ~8 s and ~160 MB resident (about 4 bytes per pixel, roughly
double that transiently during decoding). At the 100 MP default that is on the
order of 20 s and 400–800 MB per image. If you accept untrusted uploads,
especially in a multi-tenant service, set maxImagePixels well below the
default — 20–30 MP covers ordinary scans and photographs — and treat the
memory as roughly maxImagePixels × 8 bytes per concurrent image.
Dependency maintenance status
The PDF engine, pdf-lib (with its
@pdf-lib/fontkit and @pdf-lib/upng companions), has not had a release
since late 2021. It is stable and has no published advisories, but it is no
longer actively maintained, so upstream security fixes should not be expected —
which is why this library applies its own input guards (the SSRF checks, size
caps, and PNG pixel limit above) rather than relying on the engine to defend
itself. Treat this as a known risk in your own dependency review.
What this library does not do for you:
- Authentication. Pass any required tokens via the
headersoption (or per-URLheadersfor signed requests).
CLI
The package ships a small pdf-merge binary. There are two recommended
ways to invoke it:
One-shot via npx (no install)
Always pass the full scoped package name, otherwise npx will try to
resolve an unrelated pdf-merge package from the registry:
npx @lucidaquarian/pdf-merge cover.pdf body.pdf appendix.pdf -o out.pdfInstalled (global or as a dev dependency)
Once installed, you can call the unscoped pdf-merge directly — npm puts
the bin on your PATH:
# global install
npm install -g @lucidaquarian/pdf-merge
pdf-merge cover.pdf body.pdf appendix.pdf -o out.pdf
# or, within a project
npm install --save-dev @lucidaquarian/pdf-merge
npx pdf-merge cover.pdf body.pdf appendix.pdf -o out.pdfMix local files and URLs in one call, and attach a page selector to any input
with a trailing colon. If a URL's query string itself ends in a colon and
digits (…?t=10:3), the CLI reads that as a page selector, as it always has,
and prints a warning; write the colon as %3A when it belongs to the URL:
npx @lucidaquarian/pdf-merge \
cover.pdf \
'body.pdf:2-9' \
https://example.com/appendix.pdf \
-o annual-report.pdf \
--title 'Annual Report 2026' \
--author 'Operations' \
--keywords 'annual,operations,2026'Add --toc to prepend a clickable table-of-contents page; entries are
labelled by each input's filename:
npx @lucidaquarian/pdf-merge cover.pdf report.pdf appendix.pdf -o out.pdf --tocRun pdf-merge --help for the full option list. Metadata flags
(--title, --author, --subject, --creator, --keywords), --toc,
and URL options (--concurrency, --timeout, --https-only) are all
supported.
Module format
The package ships both ESM and CommonJS builds via a conditional exports
map, so all of the following work without any tooling tweaks:
import { mergePdfFiles } from '@lucidaquarian/pdf-merge'; // ESM / bundlersconst { mergePdfFiles } = require('@lucidaquarian/pdf-merge'); // CommonJSTypeScript definitions are shipped from the CJS build and resolve automatically for both consumers.
Backward compatibility
Page selection, per-URL headers, and the buffer/file/concurrency options were
added as additive unions — every previous call signature still
type-checks and produces byte-identical output. If you don't pass a pages
field, the merge runs through the original code path unchanged. A regression
test pins this.
Development
npm install
npm run build # tsc → dist/cjs + dist/esm
npm test # unit + integration tests
npm run test:cli # rebuilds and exercises the CLI binary + dual build
npm run test:e2e # end-to-end harness against real generated PDFsThe unit suite spins up a local HTTP server and exercises ordering, 404
handling, invalid Base64, empty input, timeouts, non-PDF responses, the
security defaults, page selection, header precedence, and concurrency
capping. The e2e harness in test/e2e.ts generates real multi-page A4 PDFs
and exercises every public method end-to-end, including round-tripping each
merged output back through pdf-lib.
