@runer-hq/peak
v0.1.0
Published
Node.js binding for peak, a multilingual word-to-IPA phonemizer
Readme
peak Node.js Binding
Node.js binding for peak, a multilingual word-to-IPA phonemizer for language learners.
Install
npm install @runer-hq/peakThe package ships prebuilt binaries for macOS (x64/arm64), Linux (x64/arm64), and Windows (x64/arm64). No native compilation required.
Quick start
const { phonemize, Phonemizer } = require('@runer-hq/peak');
// One-off calls (internally cached by lang + dataPath)
phonemize('cat', { lang: 'en-US' }); // ['kæt']
phonemize('bath', { lang: 'en-GB' }); // ['bɑːθ']
phonemize('zapato', { lang: 'es-ES' }); // ['θapato']
phonemize('zapato', { lang: 'es-419' }); // ['sapato']
// For repeated calls, reuse a Phonemizer instance
const us = new Phonemizer('en-US');
us.phonemize('house'); // ['haʊs']
us.phonemize('conduct', { pos: 'noun' }); // ['ˈkɑːndʌkt']
us.phonemize('conduct', { pos: 'verb' }); // ['kənˈdʌkt']TypeScript:
import { phonemize, Phonemizer, PhonemizeOptions, PronunciationEntry } from '@runer-hq/peak';Supported languages
| Code | Language | Status |
|----------|----------------------|--------|
| en-US | English (US) | ✅ |
| en-GB | English (UK) | ✅ |
| es-ES | Spanish (Spain) | ✅ |
| es-419 | Spanish (Latin Am.) | ✅ |
Module-level functions
phonemize(word, options?) → string[]
Phonemize a word, returning IPA strings.
const { phonemize } = require('@runer-hq/peak');
phonemize('cat'); // ['kæt']
phonemize('bath', { lang: 'en-GB' }); // ['bɑːθ']
phonemize('read', { pos: 'VERB', tag: 'VBD' }); // ['rɛd']Creates a Phonemizer internally, cached by lang + dataPath (max 16
entries, LRU-evicted). For high-throughput use, create a Phonemizer
instance instead.
phonemizeEntries(word, options?) → PronunciationEntry[]
Phonemize a word, returning entries with full metadata (accent, variant, source, etc.).
const { phonemizeEntries } = require('@runer-hq/peak');
const entries = phonemizeEntries('her');
// [
// { word: 'her', pos: null, accent: 'us', variant: 'default',
// ipa: 'hɝː', source: 'cambridge-extra' },
// { word: 'her', pos: null, accent: 'us', variant: 'weak',
// ipa: 'ɚ', source: 'cambridge-extra' },
// ...
// ]phonemizeAll(word, options?) → string[]
Like phonemizeEntries but returns deduplicated IPA strings.
const { phonemizeAll } = require('@runer-hq/peak');
phonemizeAll('her'); // ['hɝː', 'ɚ', ...]phonemizeAllEntries(word, options?) → PronunciationEntry[]
Alias for phonemizeEntries. Use when the caller explicitly wants every
entry including duplicates.
Phonemizer class
Create one instance per language and reuse it. Initialization loads dictionaries and rule data, so it is significantly more efficient than module-level functions for repeated calls.
new Phonemizer(lang?, dataPath?)
lang(string, default"en-US") — language codedataPath(string?) — path to a data directory. If omitted, uses data embedded in the binary.
const { Phonemizer } = require('@runer-hq/peak');
const us = new Phonemizer('en-US');
const gb = new Phonemizer('en-GB');
const es = new Phonemizer('es-ES');phonemizer.phonemize(word, options?) → string[]
Phonemize a word, returning IPA strings. lang and dataPath in options
are ignored (set at construction time).
us.phonemize('house'); // ['haʊs']
us.phonemize('read', { pos: 'VERB', tag: 'VBD' }); // ['rɛd']
us.phonemize('house', { profile: 'raw' }); // ["h'aʊs"]phonemizer.phonemizeEntries(word, options?) → PronunciationEntry[]
Phonemize a word, returning entries with full metadata. lang, dataPath,
profile, and preserveStress in options are ignored.
const entries = us.phonemizeEntries('her');
for (const e of entries) {
console.log(e.ipa, e.accent, e.variant, e.source);
}PhonemizeOptions
All properties are optional.
| Property | Type | Default | Description |
|------------------|-----------|---------------|-------------|
| lang | string | "en-US" | Language code. Module-level functions only; ignored on Phonemizer instances. |
| dataPath | string | embedded data | Path to a data directory. Module-level functions only; ignored on Phonemizer instances. |
| profile | string | "cambridge" | Output profile. See Profiles below. |
| pos | string | — | Part of speech filter. See POS tags below. |
| tag | string | — | spaCy fine-grained tag (e.g. "VBD", "VBN"). Used with pos for morphology filtering. |
| morph | string | — | spaCy morph string (e.g. "Tense=Past\|VerbForm=Part"). Alternative to tag. |
| variant | string | "all" | Pronunciation variant filter: "all", "default", "strong", or "weak". |
| preserveStress | boolean | false | Preserve original stress position. Only applies to "default" / "learner" profiles. |
PronunciationEntry
| Property | Type | Description |
|-----------|-------------------|-------------|
| word | string | The lowercased word. |
| pos | string \| null | Part of speech, if known. One of: noun, verb, adjective, adverb, pronoun, determiner, preposition. |
| accent | string | Accent: "us" or "uk". Empty for non-English. |
| variant | string | Pronunciation variant: "default", "strong", or "weak". |
| ipa | string | The IPA pronunciation string. |
| source | string | Where this pronunciation came from. See Sources below. |
Profiles
The profile option controls how IPA is generated:
| Profile | Description | Languages |
|-------------|-------------|-----------|
| "cambridge" | Cambridge Dictionary style IPA. Includes rhotic vowel normalization (ɝː), UK linking-r superscript (ʳ), and Cambridge-specific overrides. Default. | English only; auto-falls back to "default" for other languages. |
| "default" | Learner-friendly IPA. Hides single-syllable stress, moves stress to syllable start. | All languages. |
| "learner" | Same as "default". | All languages. |
| "raw" | Raw eSpeak Kirshenbaum output (ASCII phoneme codes, not IPA). For debugging. | All languages. |
phonemize('house', { profile: 'cambridge' }); // ['haʊs']
phonemize('house', { profile: 'default' }); // ['haʊs']
phonemize('house', { profile: 'raw' }); // ["h'aʊs"]POS tags
pos accepts spaCy coarse (UPOS) tags or Cambridge-style names. Only
meaningful for English Cambridge output — it selects matching dictionary
entries when a word has POS-dependent pronunciations.
| spaCy UPOS | Cambridge name | peak enum |
|------------|----------------|-----------|
| NOUN, PROPN | noun | Noun |
| VERB, AUX | verb | Verb |
| ADJ | adjective | Adjective |
| ADV | adverb | Adverb |
| PRON | pronoun | Pronoun |
| DET | determiner | Determiner |
| ADP | preposition | Preposition |
Unsupported tags (SCONJ, CCONJ, PART, NUM, INTJ, PUNCT, SYM,
X, SPACE) are ignored — the default pronunciation is returned without
error.
us.phonemize('conduct', { pos: 'noun' }); // ['ˈkɑːndʌkt']
us.phonemize('conduct', { pos: 'verb' }); // ['kənˈdʌkt']
us.phonemize('object', { pos: 'NOUN' }); // ['ˈɑːbdʒɛkt']
us.phonemize('object', { pos: 'VERB' }); // ['ɑːbˈdʒɛkt']If the word has no POS-specific entry, POS does not invent a new pronunciation — the default is returned.
Morphology filtering
For verb forms whose pronunciation depends on tense or participle state,
pass tag and/or morph (from spaCy):
us.phonemize('read'); // ['riːd']
us.phonemize('read', { pos: 'VERB', tag: 'VBD' }); // ['rɛd'] (past)
us.phonemize('read', { pos: 'VERB', tag: 'VBN' }); // ['rɛd'] (past participle)
us.phonemize('read', { pos: 'VERB', morph: 'Tense=Past|VerbForm=Part' }); // ['rɛd']Recognized morphology selectors:
| Input | Selects |
|-------|---------|
| tag=VBD | past |
| tag=VBN | past participle |
| morph contains Tense=Past | past |
| morph contains Tense=Past\|VerbForm=Part | past participle |
Variants
Some words have multiple accepted pronunciations tagged as default,
strong, or weak in the dictionary. Use variant to filter:
us.phonemize('from', { variant: 'all' }); // ['frɑːm', 'frʌm']
us.phonemize('from', { variant: 'default' }); // ['frɑːm']
us.phonemize('from', { variant: 'strong' }); // ['frɑːm']
us.phonemize('from', { variant: 'weak' }); // ['frʌm']Sources
The source field in PronunciationEntry indicates where the pronunciation
came from:
| Source | Description |
|--------|-------------|
| cambridge-extra | Hand-curated override from en_cambridge_extra data file. |
| peak-cambridge-profile | Generated by peak's Cambridge display profile (IPA replacements, rhotic normalization, etc.). |
| peak | Generated by peak's default IPA formatter. Used for non-English languages. |
Errors
Errors are thrown as standard Error instances with a descriptive message
prefixed by the error kind:
| Error kind | Cause |
|------------|-------|
| WordNotFound | Empty word, or word not in dictionary and no rules can derive it. |
| UnsupportedLanguage | Invalid or unsupported language code. |
| UnsupportedFormat | Invalid profile, variant, or pos value. |
| DataError | Data file missing or corrupt (only when using dataPath). |
try {
phonemize('');
} catch (e) {
// e.message: "WordNotFound: '' not found"
}
try {
new Phonemizer('invalid-lang');
} catch (e) {
// e.message: "UnsupportedLanguage: ..."
}Spanish
Spanish is rule-based with two regional profiles:
| Code | Region | Differences |
|------|--------|-------------|
| es-ES | Spain | z/ce/ci → θ (theta); ll → ʎ (palatal lateral) |
| es-419 | Latin America | z/ce/ci → s; ll → ʝ (voiced palatal fricative) |
phonemize('zapato', { lang: 'es-ES' }); // ['θapato']
phonemize('zapato', { lang: 'es-419' }); // ['sapato']
phonemize('cielo', { lang: 'es-ES' }); // ['θjelo']
phonemize('cielo', { lang: 'es-419' }); // ['sjelo']
phonemize('lluvia', { lang: 'es-ES' }); // ['ʎubja']
phonemize('lluvia', { lang: 'es-419' }); // ['ʝubja']
phonemize('niño', { lang: 'es-ES' }); // ['niɲo']For Spanish, pos, tag, morph, and variant have no effect (no
POS-specific or variant-specific entries exist). profile auto-falls back
to "default".
Data files
By default, data is embedded in the binary at build time — no external files needed at runtime.
To use external data files (e.g. for updating rules without rebuilding),
pass dataPath:
const us = new Phonemizer('en-US', '/path/to/peak/data');
// or
phonemize('cat', { lang: 'en-US', dataPath: '/path/to/peak/data' });The directory must have this structure:
data/
├── dictsource/ # *_list, *_rules, *_extra files
├── phsource/ # phoneme tables (ph_*, phonemes)
├── lang/ # language config files
└── display/ # display profiles (en_cambridge, en_learner_ipa)Build from source
Requires Rust toolchain and @napi-rs/cli:
cd node
npm install
npm run build # release build
node test.js # run testsTo build for a specific platform (cross-compilation):
rustup target add x86_64-apple-darwin
napi build --release --platform --target x86_64-apple-darwin