npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

toxicity

v0.1.1

Published

Fast bad-word detection for real-time game chat. Resistant to leetspeak, spacing and homoglyph evasion, with low false positives.

Readme

toxicity

Fast bad-word detection for real-time game chat.

  • Evasion-resistant. Leetspeak ($h1t), spaced-out letters (f u c k), joiners (f-u-c-k), elongation (fuuuck), self-censoring (f***), Unicode lookalikes (fuck, 𝐟𝐮𝐜𝐤, fück, Cyrillic сunt) and zero-width characters are all resolved before matching, not enumerated in the word list.
  • Low false positives. Boundaries are checked against the original text, so class, assassin, scunthorpe, the pen is blue, co-op and 455 stay clean.
  • One pass over the message. All terms are matched with a single Aho-Corasick automaton compiled to a dense table, not one regular expression per word.
  • Severity tiers. Slurs and strong profanity, mild insults, and policy words (strip, how to kill) are separate tiers, so a game picks its own line.
  • Zero dependencies, ESM and CommonJS, works in Node and the browser.

Install

npm install toxicity

Usage

import { Toxicity } from "toxicity";

const filter = new Toxicity();

filter.isProfane("what the f u c k");     // true
filter.isProfane("nice class choice");    // false

filter.censor("what the f*ck, idiot");    // "what the ****, *****"
filter.censor("what the fuck", { replace: "koffing" });  // "what the koffing"

filter.find("you F.U.C.K.ing noob");
// [{ term: "fucking", tier: 1, start: 4, end: 15, text: "F.U.C.K.ing" }]

find returns offsets into the original string, so a match spans the separators and repeated letters it swallowed. Matches never overlap; the longest term wins. When a wildcard could stand for several letters, term names one of the readings that matched, while text is always what the user actually typed.

API

new Toxicity(options?)

| Option | Default | | |---|---|---| | words | bundled vocabulary | term list, as a multi-line string or an array of lines | | allow | bundled allowlist | words that suppress a match inside them | | tiers | [1, 2] | which severity tiers are enabled | | wildcards | true | treat * inside a word as any letter or nothing | | leet | true | resolve leetspeak and symbol substitutions | | homoglyphs | true | map Cyrillic and Greek lookalikes to Latin letters | | strict | false | read punctuation the way the word list spells it (below) |

The automaton is built on the first call, in about 5 ms. Call build() to do it up front.

Methods

  • isProfane(text) returns a boolean, stopping at the first match.
  • find(text) returns { term, tier, start, end, text }[] in original coordinates.
  • censor(text, strategy?) replaces every match. Strategies: "asterisks" (default, same length), "grawlix", "keepFirst", { replace: "koffing" }, or a callback (match, text) => string.
  • add(words, tier?), remove(words), allow(words) edit the lists; the automaton is rebuilt on the next call.

The bundled lists are exported as words and allowlist, so they can be filtered or extended:

import { Toxicity, words } from "toxicity";

const custom = new Toxicity({
  words: words.split("\n").filter((line) => line !== "damn"),
});

Strict mode

The word list this package is built from spells words with punctuation that means something else in ordinary chat: ) for a d (asshea)), a trailing ! for an i (pak!), # and ^ for any letter (f###), and digits before a short suffix (45s). Reading those by default would flag smileys and numbers, so they are off unless you ask:

const filter = new Toxicity({ strict: true });

Strict mode raises recall on the source list from 99.2% to 99.5%, and on the raw list including its junk lines from 96.7% to 99.0%. It costs about a third of the throughput and introduces false positives that the default avoids: 45s reads as a word, and letters wrapped in carets or underscores can spell one. Use it where a false positive is cheap and a miss is not, such as usernames, or where the text is already suspected of being obfuscated.

Smileys, hashtags and mentions stay clean in both modes.

Neither mode reaches 100%, and the last half percent is deliberate. Most of what remains is a spelling whose leet reading is an ordinary word: @unt reads as "aunt", s h!t splits the way "was hit" does. Catching those means flagging the innocent word, which is the trade this package is built to avoid.

Term syntax

One term per line. # starts a comment, and # tier 2 sets the tier of the lines that follow.

| Line | Matches | |---|---| | fuck | anywhere, including inside a word (fuckface, motherfucker) | | \|ass\| | only as a whole word (not class, not assassin) | | \|shit | at a word start, with any suffix (shitty, not bullshit) | | job\| | at a word end | | blow job | as one word or split there (blowjob, blow-job, blow job) |

Terms are normalised the same way as input text, so writing them in plain letters is enough. A doubled letter is significant: ass does not match as, while asssss still matches ass.

Allowlist entries suppress any term fully inside them. scunthorpe protects the town, pen is protects the phrase without protecting penis.

Tiers

| Tier | Contents | On by default | |---|---|---| | 1 | slurs, strong profanity, explicit sexual terms | yes | | 2 | mild profanity and insults (damn, crap, idiot) | yes | | 3 | policy words that are ordinary English (strip, organ, how to kill) | no |

new Toxicity({ tiers: [1] });          // slurs only
new Toxicity({ tiers: [1, 2, 3] });    // strictest, for younger audiences

Accuracy

The bundled vocabulary is derived from a 38,000-entry list of profanity and its leetspeak variants. Four suites guard every change:

  • Recall. Every entry of the source list, fed in as a message, must still be detected. 99.2% are by default and 99.5% in strict mode; the rest are frozen in test/fixtures/recall-misses.txt, so a change that loses coverage fails the build.
  • Precision. No word among the 30,000 most common English words, and no line of a 300-line corpus of ordinary game chat, may be flagged with the default tiers.
  • Reference corpora. Every string the test suites of obscenity, @2toad/profanity and censor-sensor assert is clean, including all 67 whitelist entries of obscenity's English preset, must stay clean here; every string they call profane must be caught. Those whitelists are the most valuable false-positive data available, because each entry is a case someone reported as wrongly flagged.
  • Documented differences. Where this library deliberately disagrees with one of them, the disagreement is a test of its own in test/reference.test.ts, not a silent gap.

Benchmarks

Node 22 on an M-series Mac, npm run bench. Operations per second, higher is better.

| message | toxicity | obscenity | @2toad/profanity | censor-sensor | bad-words | |---|---:|---:|---:|---:|---:| | short, clean | 3,053,000 | 59,000 | 5,033,000 | 75,000 | 3,400 | | short, profane | 2,626,000 | 104,000 | 5,486,000 | 83,000 | 3,500 | | leetspeak | 766,000 | 62,000 | 4,381,000 | 74,000 | 3,500 | | 250 characters, clean | 137,000 | 18,000 | 998,000 | 37,000 | 2,900 | | self-censored (f***) | 40,000 | 56,000 | 4,291,000 | 72,000 | 3,400 | | censor a leet message | 271,000 | 52,000 | 2,664,000 | — | 170 |

Only the first four rows compare like with like. On the f*** row the other libraries are fast because they find nothing.

On a realistic mix of 1,000 chat lines averaging 11 characters, 5% of them profane (npm run bench:profile):

| | per message | throughput | |---|---:|---:| | isProfane | 0.59 µs | 1.7M messages/s | | find | 0.50 µs | 2.0M messages/s | | censor | 0.53 µs | 1.9M messages/s |

Building the automaton takes about 8 ms on first use and holds roughly 2.5 MB, shared by every call on that instance.

@2toad/profanity is faster because it runs one regular expression over the raw text: it detects none of the leetspeak, spacing or wildcard evasions above unless the exact spelling is already in its word list. Against obscenity, which does resolve them, toxicity is 8 to 50 times faster. Messages made mostly of asterisks are the one shape where the wildcard search costs more than a regular expression scan.

How it works

  1. Normalise. One pass turns the message into letter codes, recording for each letter its run length, its span in the original string, whether a word starts there, and which other letters it could be (1 is i or l, * is anything). Numbers are left alone, so 455 is not ass, and punctuation ending a word is a boundary, so ass! is still ass rather than assi.
  2. Match. All terms live in one Aho-Corasick automaton compiled to a dense transition table, so scanning costs one lookup per letter regardless of list size. Messages with ambiguous characters take a second pass that explores the alternatives.
  3. Accept. A candidate is checked against the boundaries of the original text, the run lengths a term requires, and the allowlist. A match spanning a space must start and end on word boundaries, and its fragments must be short, so f u c k matches but pen is and was hit do not.

Credits

The matching model owes a lot to obscenity: the normalise-then-match pipeline, the boundary syntax and the separate censoring pass are its ideas. The severity tiers come from censor-sensor.

License

MIT