@pawanosman/textdiff
v0.1.1
Published
Unicode-aware word, grapheme, and line diffs with exact UTF-16 positions and patch application.
Maintainers
Readme
TextDiff
Unicode-aware text diffs for Node.js and browsers, with no runtime dependencies. Compare whole words, grapheme clusters (such as emoji), or lines, then apply the changes to reproduce the new text exactly.
Install
npm install @pawanosman/textdiff
# or
pnpm add @pawanosman/textdiffNode.js 20 or newer is supported. Both ES modules and CommonJS include TypeScript declarations.
Usage
import { applyTextDiffs, getTextDiffs } from "@pawanosman/textdiff";
const oldText = "The quick brown fox";
const newText = "The fast dark wolf";
const diffs = getTextDiffs(oldText, newText);
// [
// {
// oldText: "quick brown fox",
// position: { startIndex: 4, endIndex: 19 },
// newText: "fast dark wolf",
// changeType: "replace"
// }
// ]
console.log(applyTextDiffs(oldText, diffs) === newText); // trueCommonJS:
const { getTextDiffs, applyTextDiffs } = require("@pawanosman/textdiff");The default export is an object containing the same two functions.
Choose the comparison unit
// Keep emoji sequences and combining marks together.
const emojiDiffs = getTextDiffs("Hello 👩💻!", "Hello 👨💻!", {
granularity: "grapheme",
});
// Compare lines, preserving their line endings.
const lineDiffs = getTextDiffs("first\r\nold\r\n", "first\r\nnew\r\n", {
granularity: "line",
});
// Use locale-aware word boundaries and keep separate changes separate.
const wordDiffs = getTextDiffs("The quick brown fox", "The fast dark fox", {
locale: "en",
mergeAdjacentChanges: false,
});API
export type ChangeType = "insert" | "delete" | "replace" | "spell-correction";
export interface TextDiff {
oldText: string;
position: { startIndex: number; endIndex: number };
newText: string;
changeType: ChangeType;
}
export interface TextDiffOptions {
locale?: string | string[];
granularity?: "word" | "grapheme" | "line";
mergeAdjacentChanges?: boolean;
}
export function getTextDiffs(
oldText: string,
newText: string,
options?: TextDiffOptions,
): TextDiff[];
export function applyTextDiffs(oldText: string, diffs: readonly TextDiff[]): string;getTextDiffs(oldText, newText, options?)
Returns ordered, non-overlapping changes in the coordinates of oldText. Equal inputs return an empty array.
| Option | Default | Behavior |
| --- | --- | --- |
| granularity | "word" | Compare words, grapheme clusters, or lines. Spaces and punctuation are preserved in every mode. |
| locale | Runtime default | A locale or preference list passed to Intl.Segmenter for word and grapheme boundaries. |
| mergeAdjacentChanges | true | Merge changes separated only by unchanged spaces or punctuation in word mode. In grapheme and line modes, only touching changes merge. Set to false to retain individual change ranges. |
Change types:
insert: an empty original range gains text.delete: an original range is removed.replace: an original range is replaced.spell-correction: a heuristic label for a small edit between single words in word mode. It does not use a dictionary or verify spelling; a similar word can receive this label. It is applied like any replacement. The heuristic compares NFC-normalized Unicode code points without changing the returned text. Words over 256 code points after normalization, or over 512 UTF-16 code units before normalization, skip this classification and usereplace.
Positions and exact text preservation
startIndex is inclusive and endIndex is exclusive. Both are UTF-16 code unit offsets, matching JavaScript's String.prototype.slice, rather than counts of visible characters. Insertion ranges have equal start and end indices. Every diff satisfies:
oldText.slice(diff.position.startIndex, diff.position.endIndex) === diff.oldTextAll positions refer to the original input, even when earlier changes insert or delete text. Use applyTextDiffs to apply a complete result.
Inputs are not normalized. For example, "é" and "e\u0301" compare as different strings even though they can look identical. Whitespace, punctuation, emoji, combining marks, and line endings remain exactly as supplied. If your application wants Unicode normalization, normalize both inputs before comparison; returned offsets will then refer to those normalized strings.
applyTextDiffs(oldText, diffs)
Returns the text produced by applying all changes to oldText. It does not mutate the input or diff array. It validates that ranges are ordered, non-overlapping, and within the original string, and that each diff's oldText matches its original range. Change types must match the supplied text: insertions have empty oldText, deletions have empty newText, and replacements have both. No-op entries are rejected. Invalid patches throw instead of silently applying to the wrong text.
const original = "Hello world";
const updated = "Hello, world!";
const changes = getTextDiffs(original, updated);
applyTextDiffs(original, changes); // "Hello, world!"Multiple insertions at the same offset are applied in array order. Line mode recognizes CRLF, LF, and lone CR endings and includes each ending in its line token.
Runtime and segmentation
Word mode uses Intl.Segmenter when available. Without it, word mode falls back to a Unicode regular expression that groups letters, numbers, and combining marks; boundaries for languages without spaces can be less precise. Grapheme mode requires Intl.Segmenter and throws if it is unavailable. Line mode does not require it.
Malformed locale identifiers throw in every mode, including fallback word segmentation. Locale and ICU data supplied by the runtime can affect word boundaries, so use a consistent runtime and locale when reproducible segmentation matters. Browser use requires a bundler or a browser-compatible ES module setup with Unicode property escape support.
Development
npm install
npm run check # Type-check, build both module formats, and run tests
npm run example # Build and run the examples
npm pack --dry-run # Inspect package contentsThe exact sequence matcher avoids allocating a full quadratic diff matrix. Large inputs with many differences can still require substantial computation; select line mode when comparing large documents by line.
