@secretpdf/pdf-anonymizer
v0.0.1
Published
Finds and replaces words directly in PDF content streams (binary text replacement), with an optional black-box redaction overlay. Ships as a library - no CLI, no filesystem/network access.
Maintainers
Readme
PdfAnonymizer
Finds and replaces words directly inside PDF files — a real binary text
replacement in the PDF's content streams, not just a black box / redaction
overlay drawn on top of the original text. After processing, the original
words are gone from the file's bytes and the replacement values are there
instead (verifiable by decompressing the content stream or opening the PDF
and copy-pasting the text). Optionally, on top of that real replacement, it
can also draw a black rectangle over each replaced span for a familiar
"redacted" look - see shouldRedact below.
This package is a library only - there's no CLI or file I/O here. It has
no dependency on the filesystem, environment variables, or the network (it's
built on Uint8Array, not Node's Buffer), so it works identically whether
you import it into a Node CLI, a server endpoint, or a fully-offline browser
page that never sends the PDF anywhere. It's written in TypeScript and ships
compiled output plus full .d.ts type declarations.
Install
npm install @secretpdf/pdf-anonymizerThat's the whole setup for a consuming project - main/types/exports in
package.json point at pre-built dist/ output (JS + full .d.ts
declarations), so there's nothing else to configure, and pdf-lib is
pulled in automatically as a dependency.
Developing this repo
npm install
npm run build # compiles lib/**/*.ts -> dist/ (also runs automatically
# via the "prepare" script on install/pack/publish)
npm test # runs the test/ suite (Node's built-in test runner, via tsx)Usage
import { anonymizePdf, loadPluginRules } from '@secretpdf/pdf-anonymizer';
const pluginRules = loadPluginRules({
'us-phonenumbers': { enabled: true, options: {} },
});
const result = await anonymizePdf(
input,
{
'John Doe': 'XXXXX XXXXX',
'[email protected]': '[email protected]',
'123-45-6789': 'XXX-XX-XXXX',
},
{ shouldRedact: false, pluginRules }
);
// result: { bytes, replacements, streamsScanned, streamsChanged, warnings }inputaccepts any of: aUint8Array/ NodeBuffer, anArrayBuffer, a browserBloborFile, or a base64 string (a plain base64 string, or adata:...;base64,...URL - thedata:prefix is stripped automatically).- The second argument is the literal, exact-match find/replace
pairsobject. options.shouldRedactis optional (defaults tofalse) and applies to every replacement, frompairsand from plugins alike. Whentrue, in addition to the real text replacement, a solid black rectangle is drawn over each replaced span, sized from the replacement text's own glyph widths in the document's font. This is a belt-and-suspenders visual cue on top of the real replacement (which already removes the original bytes) - not a substitute for it.options.pluginRulesis optional and lets you match thingspairscan't express as a literal string - patterns like phone numbers, where the exact text varies from document to document. Build it withloadPluginRules(pluginsConfig), where each key ofpluginsConfigis a built-in plugin name;enabledturns it on andoptionsconfigures it. See Plugins below.result.bytesis aUint8Array- write it directly withfs.writeFileSync(path, result.bytes), wrap it withnew Blob([result.bytes])for a browser download link, or pass it throughbytesToBase64().
Matching is exact and case-sensitive. If one key is a substring of
another (e.g. "John" and "John Doe"), the longer key always wins at that
position, so "John Doe" won't get partially replaced by a shorter rule.
When a plugin match and a pairs match start at the same spot, the longer
one wins (a plugin can override a shorter literal pair, or vice versa).
import {
anonymizePdf,
anonymizePdfBytes, // strict Uint8Array-in/Uint8Array-out core, skips input normalization
loadPluginRules,
REGISTRY, // the built-in plugin registry, keyed by name
toBytes, // normalizes any supported input shape to a Uint8Array
bytesToBase64,
base64ToBytes,
bytesToBlob, // (bytes, type = 'application/pdf') => Blob, browser-only
type AnonymizeOptions,
type AnonymizeResult,
type WordMap,
type PluginRule,
type PdfInput,
} from '@secretpdf/pdf-anonymizer';In a browser (bundled - pdf-lib and this library are plain npm packages,
no build-time file-system or network access is ever required at runtime):
import { anonymizePdf, loadPluginRules } from '@secretpdf/pdf-anonymizer';
fileInput.addEventListener('change', async () => {
const file = fileInput.files[0]; // a File, which is also a Blob
const pluginRules = loadPluginRules({ 'us-phonenumbers': { enabled: true } });
const result = await anonymizePdf(file, { 'John Doe': 'XXXXX XXXXX' }, { pluginRules });
const url = URL.createObjectURL(new Blob([result.bytes], { type: 'application/pdf' }));
// everything above ran entirely in the tab - the PDF was never uploaded anywhere
});Plugins
Plugins contribute pattern-based (regexp) find/replace rules alongside your
literal pairs, for values that aren't fixed strings. They're plain
TypeScript modules under lib/plugins/ and are opted into via the
pluginsConfig object passed to loadPluginRules().
us-phonenumbers
Matches common US phone-number formats: (123) 456-7890, 123-456-7890,
123.456.7890, 123 456 7890, and with a leading 1/+1. By default, only
punctuated numbers are matched (a bare 10-digit run is too easy to confuse
with an invoice/tracking number); each match is masked digit-by-digit
((XXX) XXX-XXXX), keeping the original punctuation so the replacement is
the same shape as the original.
"us-phonenumbers": {
"enabled": true,
"options": {
"matchBareDigits": false,
"maskChar": "X",
"replacement": null
}
}matchBareDigits(defaultfalse): also match unpunctuated 10-digit runs.maskChar(default"X"): character used to mask each digit.replacement(default unset): if set, used verbatim as the full replacement instead of digit-masking (e.g."REDACTED").
email-addresses
Matches email addresses ([email protected], etc.), guarded
on both ends so it can't grab a sub-run out of a longer valid local-part, and
requiring a letters-only TLD of 2+ characters to avoid false positives on
things like decimal version numbers. By default each match is masked
character-by-character (letters/digits only), keeping @, ., and the rest
of the local-part punctuation intact, so [email protected] becomes
[email protected] - same shape, PII gone.
"email-addresses": {
"enabled": true,
"options": {
"maskChar": "X",
"replacement": null
}
}maskChar(default"X"): character used to mask each letter/digit.replacement(default unset): if set, used verbatim as the full replacement instead of masking (e.g."REDACTED").
dates
Matches 2024-01-05 (ISO), 1/5/2024 / 01-05-24 / 1.5.2024 (numeric,
same separator required on both sides), January 5, 2024 / Jan 5th 2024
(month first), and 5 January 2024 / 5th of Jan, 2024 (day first).
Numeric dates require all three components and a digit boundary, so a
fraction like 3/4 or part of a longer number is never matched; textual
dates require a 4-digit year, so a bare January 5 (which could just be
prose) is left alone. Same default masking behavior as the other plugins:
January 5, 2024 becomes XXXXXXX X, XXXX.
"dates": {
"enabled": true,
"options": {
"maskChar": "X",
"replacement": null
}
}maskChar(default"X"): character used to mask each letter/digit.replacement(default unset): if set, used verbatim as the full replacement instead of masking (e.g."REDACTED").
urls
Matches https://..., http://..., and bare www.... addresses. The match
is greedy over the usual URL character set but always backtracks to end on
an alphanumeric/path-safe character, so trailing sentence punctuation right
after a URL ("...see https://example.com." or "(https://example.com)")
isn't swept into the match. Same default masking behavior as the other
plugins: https://example.com/a?b=1 becomes XXXXX://XXXXXXX.XXX/X?X=X.
"urls": {
"enabled": true,
"options": {
"maskChar": "X",
"replacement": null
}
}maskChar(default"X"): character used to mask each letter/digit.replacement(default unset): if set, used verbatim as the full replacement instead of masking (e.g."REDACTED").
gender-age
Matches combined gender/age markers: M(65), F (32), (65)M, Male(65),
Female (32), (65)Female, case-insensitively, with an optional space
around the parens/letter. Only M/F are matched by default (plus the full
words), and only the parenthesized shapes - specific enough to be low-risk.
Guarded so MR(65), (65)MRS, and MRS(40) don't match. Same default
masking behavior as the other plugins: Male(65) becomes XXXX(XX).
"gender-age": {
"enabled": true,
"options": {
"extraLetters": "",
"matchLooseSeparators": false,
"maskChar": "X",
"replacement": null
}
}extraLetters(default""): extra single-letter gender codes to also match, e.g."XOU".matchLooseSeparators(defaultfalse): also match unpunctuatedM/65/M-65/65 Mstyle pairs (no parens). Off by default since that shape collides more easily with unrelated content (model numbers, temperatures, etc.) - only turn it on if you've checked your documents don't have that kind of false positive.maskChar(default"X"): character used to mask each letter/digit.replacement(default unset): if set, used verbatim as the full replacement instead of masking (e.g."REDACTED").
us-names
Matches US given/family names against the ~36k-word list in
lib/plugins/names-dataset.json, rather than a fixed shape - this is a
dictionary lookup, not a pattern, so it's a fundamentally different (and
noisier) kind of match than the other plugins. Only capitalized, Title-case
words (Heather, not heather or HEATHER) are matched by default, and
only if they aren't in a built-in exclude list of common look-alikes (month
names, day names, and everyday words like "Will", "May", "Hope", "Grace",
"Name", "Page" that are also dataset entries).
This will produce more false positives than the other plugins - any
capitalized word that happens to be a real first/last name somewhere will
match, including in running prose. Review the output and extend
excludeWords for your document type before trusting this on anything
sensitive. Same default masking as the others: John Smith becomes
XXXX XXXXX (each name masked independently).
"us-names": {
"enabled": true,
"options": {
"excludeWords": [],
"useDefaultExcludeWords": true,
"matchAllCaps": false,
"maskChar": "X",
"replacement": null
}
}excludeWords(default[]): extra words (case-insensitive) to never treat as a name, on top of the built-in list.useDefaultExcludeWords(defaulttrue): setfalseto ignore the built-in exclude list and rely only onexcludeWords.matchAllCaps(defaultfalse): also match ALL-CAPS tokens (names printed in headers/labels/IDs) - higher false-positive risk than Title-case-only matching, since short acronyms collide with short names more easily.maskChar(default"X"): character used to mask each letter.replacement(default unset): if set, used verbatim as the full replacement instead of masking (e.g."REDACTED").
Adding a plugin
A plugin is a function (options) => [{ regex, replace }, ...] registered
in lib/plugins/index.ts's REGISTRY. regex is normally a RegExp with
the sticky (y) flag, since it's tested at each scan position rather than
searched forward - but it only needs to duck-type that interface
({ lastIndex, exec(text) }), so a dictionary-backed matcher like
us-names can return null from a hand-written exec() to reject a
shape-only match that isn't actually in its wordlist. replace is either a
fixed string or (matchedText, execResult) => string for a computed/
format-preserving replacement.
How it works
- Each PDF is parsed with
pdf-liband every content stream that can contain visible text is located: page content streams, Form XObjects (recursively), and annotation appearance streams. - Each content stream is decompressed and tokenized by hand (see
lib/tokenizer.ts) into real PDF syntax (strings, names, numbers, operators) so text-showing operators (Tj,TJ,',") can be found precisely. - For each text-showing operator, the string operand(s) are decoded to
actual text:
- Simple fonts (Type1/TrueType with WinAnsi/MacRoman/Standard encoding, one byte per character) decode directly.
- Composite fonts (Type0 / Identity-H, one glyph = one embedded
subset font — how PDF viewers/exporters like Chrome, LibreOffice and
Word usually embed text) are decoded via the font's
ToUnicodeCMap.
- Your
pairsand enabled plugin rules are applied to the decoded text (see Plugins below). - If anything changed, the new text is re-encoded back into raw PDF string
bytes (hex or literal, matching the original token) and spliced directly
into the stream — untouched parts of the file are left byte-for-byte
identical. The stream's compression is dropped (stored raw) since it just
changed;
pdf-librecomputes/Lengthautomatically on save. - If
shouldRedactis on, for each replaced span the tool also tracks the text matrix, font size, and character spacing in effect at that point in the stream (the same state PDF viewers maintain betweenBT/ET), plus the font's glyph widths (/Widthsor/W), to compute the on-page box the replacement text occupies. It then splices inET ... re f ... BToperators drawing a black rectangle over just that box, immediately before the replaced text — closing and reopening the text object because PDF doesn't allow path-painting operators inside one.
Limitations (by design, for safety)
- Embedded subset fonts can only draw glyphs they contain. Producers
like Word/LibreOffice/Chrome typically embed a subset font containing
only the characters actually used in the document. If your replacement
text uses a character that never appeared anywhere the font was used
(e.g. replacing digits-only text with letters), there's no glyph to draw
it with. The tool detects this and skips that specific replacement
(leaving the original text untouched there) rather than producing a
corrupted/garbled PDF, and prints a warning. Prefer replacement values
that reuse characters already present in the document (e.g.
X/0fillers only work ifX/0already appear somewhere). - Word splitting across kerning fragments. Some producers split a single
word into several kerned fragments inside one
TJarray (e.g.(J)(o)(h)(n)with spacing numbers between each). This is handled: the whole array is decoded as one logical string before matching. However, matching across separateTj/'/"operator calls (rather than fragments within the sameTJarray) is not attempted — that would require full text-layout reconstruction, out of scope here. - Scanned/rasterized text (an image of text, no real text layer) can't be edited this way — there is no text to replace. Consider OCR + reprint if you need that.
- Matching is case-sensitive exact substring matching; no regex/fuzzy matching.
- Redaction rectangles are an approximation. Box width is computed from
the font's declared glyph widths (falling back to a generous estimate when
a font omits them) and ignores any explicit kerning adjustments inside a
TJarray; box height comes from the font's ascent/descent (or a default if absent). A small margin is added on every side to bias towards over-covering rather than under-covering, but pathological/rotated/skewed layouts may still end up slightly mis-sized. The real text replacement (which actually removes the original bytes) is unaffected by this either way.
