npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@mailwoman/corpus

v10.0.0

Published

Mailwoman corpus pipeline: BIO-labeled dataset builder for the neural classifier.

Readme

@mailwoman/corpus

BIO-labeled training-corpus pipeline for the Mailwoman address parser.

Generates sequence-labeling training data from reference sources (OpenAddresses, libpostal dictionaries, synthetic training-data subsets) and assembles them into the TSV format consumed by the Modal training pipeline. This package produces the data that trains @mailwoman/neural-weights-*.

// The corpus pipeline is primarily build-time CLI tooling.
// Key entry points:
import { expandGolden } from "@mailwoman/corpus" // Expand reference addresses
import { synthesizeSlice } from "@mailwoman/corpus" // Generate synthetic training rows
import { alignRow } from "@mailwoman/corpus" // Align raw address → BIO tokens
import { validateCorpus } from "@mailwoman/corpus" // Validate corpus integrity

What it produces

The corpus pipeline assembles training data from multiple sources:

| Source | Description | | ------------------ | ------------------------------------------------------------------------ | | OpenAddresses | Real government address point data (US, FR, DE, …) | | NAD | National Address Database (US-specific) | | libpostal | Multilingual street/place name dictionaries | | Synthetic rows | Generated address variations (boundary stress, order variants, all-caps) | | Overture Maps | Address theme ingestion (alpha) | | OpenStreetMap | Pakistan, Bangladesh, Vietnam — ODbL, dropped by --exclude-share-alike |

Output format: TSV rows with raw<TAB>BIO_labels consumed by the Python training pipeline (corpus-python/).

Key modules

| Module | Purpose | | ----------------------- | ---------------------------------------------------------------- | | expand-golden.ts | Expand reference addresses into training rows with alignment | | align.ts | Tokenize raw address → BIO label sequence | | validate.ts | Validate corpus integrity, label coverage, subset balance | | synthesizers/*.ts | Synthetic row generators (boundary stress, order variants, etc.) | | ingest/ | Overture Maps + NAD ingestion | | slice-registry.ts | Subset metadata and composition | | stats.ts | Per-subset and per-tag statistics |

Layout

The top-level ownership boundary is the address system, using ISO 3166-1 alpha-2 country directories. A directory owns its recipes, source adapters, acquisition tools, and tests together: fr/{recipes,adapters,tools}, kr/{adapters,tools}, and us/{adapters,tools} are representative. Regional and global work lives under south-asia/ and international/; cross-locale primitives remain at the recipes/ and synthesizers/ roots.

Sources that deliberately span countries (OpenAddresses, GeoNames, Overture, OSM, and WOF) remain under adapters/, with their common fetching and processing utilities under tools/. Source kind is therefore a nested concern inside a country directory, not the package's organizing principle. Tests live beside the module they cover: fr/adapters/ban/adapter.test.ts belongs with fr/adapters/ban/adapter.ts. Production compilation excludes *.test.ts; the test project includes the same pattern. Public subpaths use this same locale-first layout.

Build-time tooling

The corpus is assembled via scripts in scripts/:

# Validate the corpus
node scripts/validate-corpus.mjs

# Rebuild slices
node scripts/build-boundary-stress-slice.mjs

# Corpus statistics
node scripts/corpus-stats.mjs

Design

  • BIO (Begin/Inside/Outside) labeling over SentencePiece tokens.
  • Character-offset aligned — labels track the raw string, not the normalized form, so the model learns real input distributions.
  • Source-homogeneous subsets — each training-data subset comes from one source, ordered by type, so eval splits are clean (no bleed between train and held-out).

Related

License

AGPL-3.0-only