npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@secretpdf/pdf-anonymizer

v0.0.1

Published

Finds and replaces words directly in PDF content streams (binary text replacement), with an optional black-box redaction overlay. Ships as a library - no CLI, no filesystem/network access.

Readme

PdfAnonymizer

Finds and replaces words directly inside PDF files — a real binary text replacement in the PDF's content streams, not just a black box / redaction overlay drawn on top of the original text. After processing, the original words are gone from the file's bytes and the replacement values are there instead (verifiable by decompressing the content stream or opening the PDF and copy-pasting the text). Optionally, on top of that real replacement, it can also draw a black rectangle over each replaced span for a familiar "redacted" look - see shouldRedact below.

This package is a library only - there's no CLI or file I/O here. It has no dependency on the filesystem, environment variables, or the network (it's built on Uint8Array, not Node's Buffer), so it works identically whether you import it into a Node CLI, a server endpoint, or a fully-offline browser page that never sends the PDF anywhere. It's written in TypeScript and ships compiled output plus full .d.ts type declarations.

Install

npm install @secretpdf/pdf-anonymizer

That's the whole setup for a consuming project - main/types/exports in package.json point at pre-built dist/ output (JS + full .d.ts declarations), so there's nothing else to configure, and pdf-lib is pulled in automatically as a dependency.

Developing this repo

npm install
npm run build   # compiles lib/**/*.ts -> dist/ (also runs automatically
                 # via the "prepare" script on install/pack/publish)
npm test        # runs the test/ suite (Node's built-in test runner, via tsx)

Usage

import { anonymizePdf, loadPluginRules } from '@secretpdf/pdf-anonymizer';

const pluginRules = loadPluginRules({
  'us-phonenumbers': { enabled: true, options: {} },
});

const result = await anonymizePdf(
  input,
  {
    'John Doe': 'XXXXX XXXXX',
    '[email protected]': '[email protected]',
    '123-45-6789': 'XXX-XX-XXXX',
  },
  { shouldRedact: false, pluginRules }
);
// result: { bytes, replacements, streamsScanned, streamsChanged, warnings }
  • input accepts any of: a Uint8Array / Node Buffer, an ArrayBuffer, a browser Blob or File, or a base64 string (a plain base64 string, or a data:...;base64,... URL - the data: prefix is stripped automatically).
  • The second argument is the literal, exact-match find/replace pairs object.
  • options.shouldRedact is optional (defaults to false) and applies to every replacement, from pairs and from plugins alike. When true, in addition to the real text replacement, a solid black rectangle is drawn over each replaced span, sized from the replacement text's own glyph widths in the document's font. This is a belt-and-suspenders visual cue on top of the real replacement (which already removes the original bytes) - not a substitute for it.
  • options.pluginRules is optional and lets you match things pairs can't express as a literal string - patterns like phone numbers, where the exact text varies from document to document. Build it with loadPluginRules(pluginsConfig), where each key of pluginsConfig is a built-in plugin name; enabled turns it on and options configures it. See Plugins below.
  • result.bytes is a Uint8Array - write it directly with fs.writeFileSync(path, result.bytes), wrap it with new Blob([result.bytes]) for a browser download link, or pass it through bytesToBase64().

Matching is exact and case-sensitive. If one key is a substring of another (e.g. "John" and "John Doe"), the longer key always wins at that position, so "John Doe" won't get partially replaced by a shorter rule. When a plugin match and a pairs match start at the same spot, the longer one wins (a plugin can override a shorter literal pair, or vice versa).

import {
  anonymizePdf,
  anonymizePdfBytes, // strict Uint8Array-in/Uint8Array-out core, skips input normalization
  loadPluginRules,
  REGISTRY, // the built-in plugin registry, keyed by name
  toBytes, // normalizes any supported input shape to a Uint8Array
  bytesToBase64,
  base64ToBytes,
  bytesToBlob, // (bytes, type = 'application/pdf') => Blob, browser-only
  type AnonymizeOptions,
  type AnonymizeResult,
  type WordMap,
  type PluginRule,
  type PdfInput,
} from '@secretpdf/pdf-anonymizer';

In a browser (bundled - pdf-lib and this library are plain npm packages, no build-time file-system or network access is ever required at runtime):

import { anonymizePdf, loadPluginRules } from '@secretpdf/pdf-anonymizer';

fileInput.addEventListener('change', async () => {
  const file = fileInput.files[0]; // a File, which is also a Blob
  const pluginRules = loadPluginRules({ 'us-phonenumbers': { enabled: true } });
  const result = await anonymizePdf(file, { 'John Doe': 'XXXXX XXXXX' }, { pluginRules });
  const url = URL.createObjectURL(new Blob([result.bytes], { type: 'application/pdf' }));
  // everything above ran entirely in the tab - the PDF was never uploaded anywhere
});

Plugins

Plugins contribute pattern-based (regexp) find/replace rules alongside your literal pairs, for values that aren't fixed strings. They're plain TypeScript modules under lib/plugins/ and are opted into via the pluginsConfig object passed to loadPluginRules().

us-phonenumbers

Matches common US phone-number formats: (123) 456-7890, 123-456-7890, 123.456.7890, 123 456 7890, and with a leading 1/+1. By default, only punctuated numbers are matched (a bare 10-digit run is too easy to confuse with an invoice/tracking number); each match is masked digit-by-digit ((XXX) XXX-XXXX), keeping the original punctuation so the replacement is the same shape as the original.

"us-phonenumbers": {
  "enabled": true,
  "options": {
    "matchBareDigits": false,
    "maskChar": "X",
    "replacement": null
  }
}
  • matchBareDigits (default false): also match unpunctuated 10-digit runs.
  • maskChar (default "X"): character used to mask each digit.
  • replacement (default unset): if set, used verbatim as the full replacement instead of digit-masking (e.g. "REDACTED").

email-addresses

Matches email addresses ([email protected], etc.), guarded on both ends so it can't grab a sub-run out of a longer valid local-part, and requiring a letters-only TLD of 2+ characters to avoid false positives on things like decimal version numbers. By default each match is masked character-by-character (letters/digits only), keeping @, ., and the rest of the local-part punctuation intact, so [email protected] becomes [email protected] - same shape, PII gone.

"email-addresses": {
  "enabled": true,
  "options": {
    "maskChar": "X",
    "replacement": null
  }
}
  • maskChar (default "X"): character used to mask each letter/digit.
  • replacement (default unset): if set, used verbatim as the full replacement instead of masking (e.g. "REDACTED").

dates

Matches 2024-01-05 (ISO), 1/5/2024 / 01-05-24 / 1.5.2024 (numeric, same separator required on both sides), January 5, 2024 / Jan 5th 2024 (month first), and 5 January 2024 / 5th of Jan, 2024 (day first). Numeric dates require all three components and a digit boundary, so a fraction like 3/4 or part of a longer number is never matched; textual dates require a 4-digit year, so a bare January 5 (which could just be prose) is left alone. Same default masking behavior as the other plugins: January 5, 2024 becomes XXXXXXX X, XXXX.

"dates": {
  "enabled": true,
  "options": {
    "maskChar": "X",
    "replacement": null
  }
}
  • maskChar (default "X"): character used to mask each letter/digit.
  • replacement (default unset): if set, used verbatim as the full replacement instead of masking (e.g. "REDACTED").

urls

Matches https://..., http://..., and bare www.... addresses. The match is greedy over the usual URL character set but always backtracks to end on an alphanumeric/path-safe character, so trailing sentence punctuation right after a URL ("...see https://example.com." or "(https://example.com)") isn't swept into the match. Same default masking behavior as the other plugins: https://example.com/a?b=1 becomes XXXXX://XXXXXXX.XXX/X?X=X.

"urls": {
  "enabled": true,
  "options": {
    "maskChar": "X",
    "replacement": null
  }
}
  • maskChar (default "X"): character used to mask each letter/digit.
  • replacement (default unset): if set, used verbatim as the full replacement instead of masking (e.g. "REDACTED").

gender-age

Matches combined gender/age markers: M(65), F (32), (65)M, Male(65), Female (32), (65)Female, case-insensitively, with an optional space around the parens/letter. Only M/F are matched by default (plus the full words), and only the parenthesized shapes - specific enough to be low-risk. Guarded so MR(65), (65)MRS, and MRS(40) don't match. Same default masking behavior as the other plugins: Male(65) becomes XXXX(XX).

"gender-age": {
  "enabled": true,
  "options": {
    "extraLetters": "",
    "matchLooseSeparators": false,
    "maskChar": "X",
    "replacement": null
  }
}
  • extraLetters (default ""): extra single-letter gender codes to also match, e.g. "XOU".
  • matchLooseSeparators (default false): also match unpunctuated M/65 / M-65 / 65 M style pairs (no parens). Off by default since that shape collides more easily with unrelated content (model numbers, temperatures, etc.) - only turn it on if you've checked your documents don't have that kind of false positive.
  • maskChar (default "X"): character used to mask each letter/digit.
  • replacement (default unset): if set, used verbatim as the full replacement instead of masking (e.g. "REDACTED").

us-names

Matches US given/family names against the ~36k-word list in lib/plugins/names-dataset.json, rather than a fixed shape - this is a dictionary lookup, not a pattern, so it's a fundamentally different (and noisier) kind of match than the other plugins. Only capitalized, Title-case words (Heather, not heather or HEATHER) are matched by default, and only if they aren't in a built-in exclude list of common look-alikes (month names, day names, and everyday words like "Will", "May", "Hope", "Grace", "Name", "Page" that are also dataset entries).

This will produce more false positives than the other plugins - any capitalized word that happens to be a real first/last name somewhere will match, including in running prose. Review the output and extend excludeWords for your document type before trusting this on anything sensitive. Same default masking as the others: John Smith becomes XXXX XXXXX (each name masked independently).

"us-names": {
  "enabled": true,
  "options": {
    "excludeWords": [],
    "useDefaultExcludeWords": true,
    "matchAllCaps": false,
    "maskChar": "X",
    "replacement": null
  }
}
  • excludeWords (default []): extra words (case-insensitive) to never treat as a name, on top of the built-in list.
  • useDefaultExcludeWords (default true): set false to ignore the built-in exclude list and rely only on excludeWords.
  • matchAllCaps (default false): also match ALL-CAPS tokens (names printed in headers/labels/IDs) - higher false-positive risk than Title-case-only matching, since short acronyms collide with short names more easily.
  • maskChar (default "X"): character used to mask each letter.
  • replacement (default unset): if set, used verbatim as the full replacement instead of masking (e.g. "REDACTED").

Adding a plugin

A plugin is a function (options) => [{ regex, replace }, ...] registered in lib/plugins/index.ts's REGISTRY. regex is normally a RegExp with the sticky (y) flag, since it's tested at each scan position rather than searched forward - but it only needs to duck-type that interface ({ lastIndex, exec(text) }), so a dictionary-backed matcher like us-names can return null from a hand-written exec() to reject a shape-only match that isn't actually in its wordlist. replace is either a fixed string or (matchedText, execResult) => string for a computed/ format-preserving replacement.

How it works

  1. Each PDF is parsed with pdf-lib and every content stream that can contain visible text is located: page content streams, Form XObjects (recursively), and annotation appearance streams.
  2. Each content stream is decompressed and tokenized by hand (see lib/tokenizer.ts) into real PDF syntax (strings, names, numbers, operators) so text-showing operators (Tj, TJ, ', ") can be found precisely.
  3. For each text-showing operator, the string operand(s) are decoded to actual text:
    • Simple fonts (Type1/TrueType with WinAnsi/MacRoman/Standard encoding, one byte per character) decode directly.
    • Composite fonts (Type0 / Identity-H, one glyph = one embedded subset font — how PDF viewers/exporters like Chrome, LibreOffice and Word usually embed text) are decoded via the font's ToUnicode CMap.
  4. Your pairs and enabled plugin rules are applied to the decoded text (see Plugins below).
  5. If anything changed, the new text is re-encoded back into raw PDF string bytes (hex or literal, matching the original token) and spliced directly into the stream — untouched parts of the file are left byte-for-byte identical. The stream's compression is dropped (stored raw) since it just changed; pdf-lib recomputes /Length automatically on save.
  6. If shouldRedact is on, for each replaced span the tool also tracks the text matrix, font size, and character spacing in effect at that point in the stream (the same state PDF viewers maintain between BT/ET), plus the font's glyph widths (/Widths or /W), to compute the on-page box the replacement text occupies. It then splices in ET ... re f ... BT operators drawing a black rectangle over just that box, immediately before the replaced text — closing and reopening the text object because PDF doesn't allow path-painting operators inside one.

Limitations (by design, for safety)

  • Embedded subset fonts can only draw glyphs they contain. Producers like Word/LibreOffice/Chrome typically embed a subset font containing only the characters actually used in the document. If your replacement text uses a character that never appeared anywhere the font was used (e.g. replacing digits-only text with letters), there's no glyph to draw it with. The tool detects this and skips that specific replacement (leaving the original text untouched there) rather than producing a corrupted/garbled PDF, and prints a warning. Prefer replacement values that reuse characters already present in the document (e.g. X/0 fillers only work if X/0 already appear somewhere).
  • Word splitting across kerning fragments. Some producers split a single word into several kerned fragments inside one TJ array (e.g. (J)(o)(h)(n) with spacing numbers between each). This is handled: the whole array is decoded as one logical string before matching. However, matching across separate Tj/'/" operator calls (rather than fragments within the same TJ array) is not attempted — that would require full text-layout reconstruction, out of scope here.
  • Scanned/rasterized text (an image of text, no real text layer) can't be edited this way — there is no text to replace. Consider OCR + reprint if you need that.
  • Matching is case-sensitive exact substring matching; no regex/fuzzy matching.
  • Redaction rectangles are an approximation. Box width is computed from the font's declared glyph widths (falling back to a generous estimate when a font omits them) and ignores any explicit kerning adjustments inside a TJ array; box height comes from the font's ascent/descent (or a default if absent). A small margin is added on every side to bias towards over-covering rather than under-covering, but pathological/rotated/skewed layouts may still end up slightly mis-sized. The real text replacement (which actually removes the original bytes) is unaffected by this either way.