npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

ingot-scan

v0.1.6

Published

Is the benchmark you are quoting inside the data you trained on? Contamination scanning with the matching text shown, on your machine, nothing uploaded.

Readme

Ingot

What's inside AI training data?

See which exact words — yours, or a benchmark's — are inside the data an AI learned from. The words themselves, side by side with their surroundings, never a score.

Run it in your browser · new here? how it works in five minutes · three ways contamination scanning silently fails · the registry · how every number was measured

Ingot answers that with evidence rather than a score: it shows the matching text, side by side with its surroundings, so you can judge each match yourself. It runs on your machine — in a browser tab or on the command line — and nothing is uploaded, because there is nowhere to upload it to.

npx ingot-scan contaminate --index gsm8k --corpus your-corpus.jsonl

No clone, no benchmark download, no account. Or open the web scanner, drop a file, and watch the network panel stay empty.

Apache-2.0. Zero runtime dependencies: Node 24 executes the TypeScript directly, with no build step and nothing installed at scan time.

Read this before the numbers

  • Exact matching only. Paraphrased contamination is not counted and is invisible to the headline. On lightly edited copies, recall is 81.5% at n=10 and falls as n grows — and that is against random word dropping, not against someone deliberately rewriting.
  • A match is not a verdict. Canonical text with one natural phrasing looks exactly like leakage in any count. Every one of the first six registry findings turned out to be canonical — prime sequences, the ten digits, "I Have a Dream". Only reading the words tells you which it is, which is why Ingot always shows them.
  • Canonical is not the same as harmless. A test item written from a public web page can sit in a training corpus as ordinary web text: nobody leaked it, and a model trained on that corpus still saw it. The canonicality label says where text came from, not that seeing it was free.
  • Questions are indexed; answers are not. Published indexes exclude answer text, so a corpus that reproduces every solution while paraphrasing its questions scans clean.
  • "Verbatim" means after normalization. Matching runs on lowercased tokens with punctuation stripped. For prose the difference is cosmetic; for code it is not — a HumanEval match is a run of identifiers and words, with operators and structure invisible.
  • It does not stop a determined vendor. Anyone can run Ingot on their own corpus, see what matched, and edit until nothing does. docs/threat-model.md says so plainly.
  • The provenance scanner's detection floor is 50% contamination, with a 13% false positive rate. That is not a good number yet, and the four reasons are below.

What it does

Ingot indexes the benchmark and streams the corpus once. lm-evaluation-harness does the reverse, indexing the corpus, and reports nine days on the Pile.

  7.0 MB/sec single-threaded, end to end  →  20 GB in about 48 minutes

Measured on the real thing: 21.4 GB of gzipped C4 shards on disk, GSM8K at n=10, 51 minutes wall clock, decompression and parsing and hashing included. Reproduce with node scripts/pretraining-scan.ts.

An earlier version of this README said 37.1 MB/sec and 20 GB in 9.0 minutes. That was wrong — not the measurement, the extrapolation. 37.1 MB/sec is the scan kernel: tokenize, roll the n-gram, look it up, over text already read off disk and already JSON-parsed, which is what scripts/bench-scan.ts measures and says it measures. Quoting it as a scanning rate silently dropped decompression, line splitting, JSON.parse, corpus hashing and match bookkeeping — together the majority of the work. scripts/bench-pipeline.ts measures the whole path and attributes the cost stage by stage.

The kernel figure is still the right one for judging optimisation work, and it is still where the remaining headroom is. It is the wrong one for answering "how long will my scan take", which is the only question a reader was asking.

Published indexes carry one-way hashes and item ids, never benchmark text, so an index can be distributed for a benchmark whose licence forbids redistributing the data — and a user checks their corpus without downloading the benchmark at all. MMLU is 5.35 MB and loads in under a second.

Every report ends with a receipt: scanner version, index format, benchmark identity hash, n, stride, gram count, corpus name, size, document count, corpus hash, and the exact command. A third party reproduces any published number from it without asking us for anything.

Both surfaces produce the same self-contained HTML report--out report.html on the command line, a download button in the browser. No scripts, no external requests: it opens from an email attachment on a machine with no network, which is the point of an artifact you can hand to a reviewer.

Quickstart

gsm8k and humaneval ship with the package, so this needs nothing else:

npx ingot-scan contaminate --index gsm8k --corpus mine.jsonl

Corpus format is JSONL, one record per line. The text field is auto-detected as text, response, output, completion, answer or content, or pass --text-field. --index also takes a path to any published .idx.bin.gz, which is how MMLU and anything you build yourself are used.

From a clone, with no install step at all:

node --test test/*.test.ts        # 128 tests; the defect log is docs/measurements.md

node scripts/fetch-benchmarks.ts  # public benchmarks, normalised
node scripts/build-web.ts         # browser bundle + publishable indexes
npx --yes serve web               # the scanner, at the root URL

node src/cli.ts contaminate --index web/indexes/mmlu.idx.bin.gz --corpus mine.jsonl

The n sweep: 13 is a poor default

Everyone matches on 13-grams because GPT-3 (Brown et al., 2020) used it. We could find no systematic re-derivation since. Running 8 through 13 on a 1,000-item benchmark with 300 items planted in 6,000 documents, verbatim and again with roughly one word in eleven dropped:

| n | verbatim recall | edited-copy recall | false positives | unscannable items | |---|---|---|---|---| | 8 | 100% | 90.9% | 5 / 1000 | 2 / 1000 | | 9 | 100% | 85.9% | 3 / 1000 | 2 / 1000 | | 10 | 100% | 81.5% | 2 / 1000 | 3 / 1000 | | 11 | 100% | 79.7% | 1 / 1000 | 31 / 1000 | | 12 | 100% | 75.0% | 0 | 57 / 1000 | | 13 | 100% | 69.6% | 0 | 68 / 1000 |

Two findings, both against the default:

  • n=13 leaves 6.8% of the benchmark unscannable, because items shorter than 13 tokens produce no 13-grams and can never match anything. At n=10 that falls to 0.3%. Structural rather than statistical, and the stronger of the two.
  • n=13 recovers 69.6% of lightly edited copies against 81.5% at n=10. At 300 planted items the standard error is about 2.3 points, so the gap is real, and recall falls monotonically with n.

Both find every verbatim copy. n=10 costs two false positives per thousand items against zero at n=13, and both are inspectable, because every hit displays its matching text.

An earlier version of this table claimed 100% against 3.5%. That gap was an artifact of a test fixture that deleted words on a fixed lattice, which decided in advance which values of n could survive. docs/measurements.md has the full account, along with every other defect found by measurement rather than assumed away.

Reports quote n=13 alongside n=10 so findings stay comparable with prior published work.

What was NOT checked

Every report names the benchmark items that produced no surviving n-gram, either because they are shorter than n tokens or because all their grams were filtered as boilerplate. Nothing can ever match those items, so a clean result that stayed silent about them would be hiding the part of the benchmark nobody looked at.

This was a bug first: the validation harness failed on its first run at 95% recall, and three of sixty planted items turned out to be 10 tokens long. Not a scanner defect — a structural limit that was invisible.

The registry

Which public benchmarks appear in which public training corpora, in results/registry.md, reproducible with node scripts/registry-scan.ts. Everything scanned is public, so every number can be re-derived from the same files.

At n=10 the current answer is six flagged items across 15,525, and all six are canonical text rather than leakage. That is published rather than quietly dropped, because it is the finding: a phrase appearing in exactly one benchmark item still carries no evidential weight if the whole world writes it.

The provenance scanner

A second, weaker product: was this batch written by a human or generated? Batch-level, six structural signals, measured against two named reference corpora.

| Signal | What it looks at | Separation on the reference pair | |---|---|---| | Sentence burstiness | how much sentence length varies inside a record | 6.31σ | | Structural repetition | reused openings, closings, markdown scaffolding | 5.18σ | | Lexical cluster tightness | mean pairwise TF-IDF cosine across the batch | 1.24σ | | Lexical variety | type-token ratio and rare-term rate | 0.94σ | | Near-duplicate rate | MinHash 64x over 5-gram shingles | 0.00σ, no information here | | Cross-author style distance | function-word fingerprint per annotator | unavailable, see below |

Signals below 0.25σ are excluded rather than included with a small weight, because a signal that cannot separate the references cannot inform a verdict about a batch.

  contamination    purity (5 draws)      clean control
  ─────────────────────────────────────────────────────
       0%          88.0 ± 4.7            mean 89.9
       5%          90.2 ± 6.6            sd   4.7
      10%          87.6 ± 10.2           n    8 batches
      25%          81.2 ± 5.0
      50%          56.4 ± 7.7

Detection floor: 50%. The first level landing more than two control standard deviations below the clean-human mean. 25% is marginal at about 1.8σ. 5% and 10% are inside the noise and are not detectable with this corpus pair at this batch size. False positive rate: 13% — one of eight clean human batches, at a purity threshold of 85.

That is the honest number, and the four reasons it is not better are known:

  1. The reference pair is not prompt-matched. Dolly and Alpaca answer different questions, so part of the measured separation is task distribution rather than provenance. scripts/generate-paired.ts regenerates the same dolly prompts with a current model, giving a paired design. Needs ANTHROPIC_API_KEY. Highest-value next step.
  2. The machine reference is 2023-era. Treat this curve as an upper bound, never as a claim about 2026 models.
  3. Only 7,208 human records. At batch 1,200 that is roughly six independent batches per half, so split luck alone shifts the reference.
  4. Neither corpus ships annotator ids, so cross-author stylometry never runs. On a real vendor batch it is the strongest available signal.

Contamination leads the product because it is not an inference. A shared n-gram is a fact you can display; authorship is a statistical claim about a distribution.

What it refuses to do

Every one of these was a bug first, found by running the experiment and reading the output instead of trusting it. All are written up in docs/measurements.md and guarded by tests.

  • No per-record verdicts. Batch-level evidence against named references.
  • No extrapolation. A batch outside the interval the references span is unlike both. Clamping such a batch once put a clean human batch at 34/100.
  • No absolute thresholds. Missing baselines means no score, not a fallback guess.
  • No comparing across batch sizes. Near-duplicate rate and type-token ratio move with batch size regardless of provenance.
  • No zero-variance division. A pinned statistic is not a perfect discriminator. Treating it as one produced NaN purity on every row.
  • No silent 0.0. Non-finite scores raise; unparseable lines are counted and listed.

Documentation

  • docs/coverage.md — every way a clean result can mean "we could not have looked" rather than "we looked and found nothing", and what Ingot reports about each one
  • docs/measurements.md — every published number, how it was measured, and every wrong turn taken on the way
  • docs/threat-model.md — what leaves your machine, what the hashes do and do not protect, and where a determined vendor still wins
  • docs/index-format.md — the index format, specified completely enough to implement independently
  • docs/github-action.md — running the scan as a gate in your own CI
  • docs/self-serve-feasibility.md — why the check cannot run entirely in your browser, measured rather than asserted, and what the measurement leaves standing

Licences of the reference data

databricks-dolly-15k is CC-BY-SA-3.0. stanford_alpaca is CC-BY-NC-4.0, so the Alpaca reference is for calibration and research only, not commercial use. Stated here because a verification product that plays loose with data licensing has no standing to audit anyone else.