@transkripid/pdf-text-replace
v1.1.0
Published
Find and replace text in PDF files with preserved formatting
Maintainers
Readme
pdf-text-replace
A small TypeScript library for finding and replacing text inside PDF files without re-rendering them. Your original fonts, colors, layout, and page structure stay intact — only the matched text is swapped out.
Think of it as String.prototype.replace() for PDFs.
Why this exists
Most PDF libraries either rebuild the document from scratch (losing formatting) or require you to know exactly which object holds which text. This library does neither: it parses the PDF's content streams in place, finds text-drawing operators (Tj / TJ), and surgically rewrites them. The result is a file that looks identical to the original except for the words you changed.
It is intentionally narrow in scope so it can be fast, dependency-light, and predictable.
Installation
npm install @transkripid/pdf-text-replace
# or
pnpm add @transkripid/pdf-text-replace
# or
yarn add @transkripid/pdf-text-replaceRequires Node.js 18+.
Quick start
import { fromBuffer } from '@transkripid/pdf-text-replace';
import { readFileSync, writeFileSync } from 'node:fs';
const input = readFileSync('document.pdf');
const output = fromBuffer(input)
.replace('John Doe', 'Jane Smith')
.replace('[email protected]', '[email protected]')
.replace(/\d{4}-\d{4}-\d{4}/g, 'XXXX-XXXX-XXXX')
.toBuffer();
writeFileSync('modified.pdf', output);That's the whole library. Queue up as many replacements as you want, then call .toBuffer() to produce the modified PDF.
Streaming API
For larger files or pipeline-style processing, use the Node.js Transform stream:
import { createPDFTransformStream } from '@transkripid/pdf-text-replace';
import { createReadStream, createWriteStream } from 'node:fs';
import { pipeline } from 'node:stream/promises';
const transform = createPDFTransformStream()
.replace('CONFIDENTIAL', 'PUBLIC')
.replace(/Draft v\d+/g, 'Final');
await pipeline(
createReadStream('input.pdf'),
transform,
createWriteStream('output.pdf')
);Note: the stream buffers the full PDF before processing (PDF cross-reference tables require random access), so this is a convenience wrapper rather than true incremental streaming.
API
fromBuffer(input, options?)
Create a chainable PDF context from a buffer.
| Parameter | Type | Description |
|---|---|---|
| input | Buffer \| Uint8Array | The PDF file contents |
| options.optimize | boolean | Merge queued replacements into a single combined regex (default: true) |
Returns a PDFContext with .replace() and .toBuffer().
.replace(search, replacement)
Queue a replacement. Returns the context for chaining.
| Parameter | Type | Description |
|---|---|---|
| search | string \| RegExp | Pattern to find. Plain strings are matched literally |
| replacement | string | Text to substitute. Unicode is transliterated to ASCII (see below) |
.toBuffer()
Apply all queued replacements and return the result as a Buffer.
Returns the input unchanged when:
- No replacements were queued
- None of the patterns matched anything
- An error occurred mid-processing (parsing, decompression, etc.)
This library favors graceful degradation over exceptions — it will never throw on a malformed PDF, it just hands the original back.
createPDFTransformStream(options?)
Returns a Node.js Transform stream with the same .replace() chaining API. Same options as fromBuffer.
How it works
- Locate content streams — scan the PDF for
stream/endstreamblocks and detect whether they'reFlateDecode-compressed. - Decompress and scan — for each text-bearing stream, extract the bytes between
(...)Tjand[...]TJtext-showing operators. - Apply replacements — perform the substitutions on the literal strings, preserving the surrounding operators.
- Adjust spacing — when the replacement text has a different width, the horizontal scaling operator (
Tz) is updated so the surrounding layout doesn't break. - Rebuild — re-compress modified streams, fix up each stream's
/Length, and rewrite the cross-reference (xref) table so byte offsets stay valid.
The optimizer (on by default) combines all queued .replace() calls into a single regex pass over each stream, which is roughly an order of magnitude faster than running them serially when you have many replacements.
Unicode handling
PDF fonts shipped with WinAnsiEncoding can't render arbitrary Unicode glyphs. Rather than silently producing tofu boxes, this library transliterates non-ASCII replacement text to its closest ASCII equivalent using any-ascii:
| Input | Becomes |
|---|---|
| 银宵 (Chinese) | YinXiao |
| 스트레이 (Korean) | seuteulei |
| Привет (Cyrillic) | Privet |
| José García (accented) | Jose Garcia |
If preserving the original glyphs matters to you, this library is not the right tool — you'd need one that can embed new fonts.
Limitations
This library deliberately handles the common case well rather than every PDF in the wild:
- ✅ Standard text PDFs using
WinAnsiEncoding - ✅ FlateDecode-compressed content streams
- ❌ CID fonts and Identity-H encoded text (most Asian-language PDFs)
- ❌ Text split across multiple operators (matching is per-operator)
- ❌ Scanned / image-based PDFs (no OCR — there's no text to replace)
- ❌ Forms, annotations, and metadata fields (only page content streams)
If a replacement can't be applied for any of these reasons, you'll get the original PDF back unchanged.
Development
pnpm install # install dependencies
pnpm test # run tests in watch mode
pnpm test:run # run tests once
pnpm typecheck # type-check without emitting
pnpm build # build ESM + CJS bundles into dist/Tests use real PDF fixtures in tests/fixtures/. See AGENTS.md for more on the codebase layout and conventions.
License
MIT
