npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

uralicnlp

v1.0.0

Published

JavaScript implementation of UralicNLP for server (Node.js) and browser

Readme

UralicNLP-JS

JavaScript implementation of the UralicNLP NLP library for Uralic languages (Finnish, Komi-Zyrian, Erzya, Moksha, Sami, etc.) and others.

Runs universally on Server (Node.js) and Client (modern Web Browsers & Web Workers) using Dr Jack Rueter's hfst-js for pure JavaScript HFST optimized-lookup transducer operations without binary dependencies.


Features

  • Universal Support: Works seamlessly in Node.js (CommonJS & ESM) and in Web Browsers via modern bundlers or a single <script> tag.
  • Pure JavaScript HFST: Uses hfst-js for fast transducer execution.
  • Server vs Client Model Storage:
    • Server (Node.js): Stores language models on disk (following Python UralicNLP logic: checks package models, ~/.uralicnlp, and custom folders).
    • Client (Browser): Stores binary language models persistently in the browser's IndexedDB, eliminating repeated downloads and bypassing localStorage size limits.
  • Built-in Tokenizer: Fast, rule-based sentence splitting and word tokenization bundled with full UralicNLP abbreviation lists.
  • Full API Parity: Supports both JavaScript-friendly option objects and Python-style positional arguments.

Installation

1. NPM Package (Node.js or Frontend Bundlers)

npm install uralicnlp

2. Standalone Browser Script (Non-Node Projects)

For client-side projects without Node.js or bundlers, include the single minified bundle:

<script src="dist/uralicnlp.min.js"></script>

This registers window.uralicNLP globally with zero dependencies.


Quick Start

Node.js (ES Modules)

import {
  lemmatize,
  analyze,
  generate,
  segment,
  get_translation,
  tokenizer,
} from 'uralicnlp';

// 1. Tokenizer (Runs synchronously, 0 network requests)
const sentences = tokenizer.sentences('Tämä on testi. Esim. tämä on toinen lause!');
console.log(sentences);
// ['Tämä on testi.', 'Esim. tämä on toinen lause!']

const tokens = tokenizer.tokenize('Ostin 1.5 kg omenoita.');
console.log(tokens);
// [['Ostin', '1.5', 'kg', 'omenoita', '.']]

// 2. Morphological Analysis
const analyses = await analyze('лыддьыны', 'kpv');
console.log(analyses);
// [['лыддьыны+V+TV+ConNeg+Pl3', 0], ['лыддьыны+V+TV+Inf', 0]]

// 3. Lemmatization
const lemmas = await lemmatize('лыддьыны', 'kpv');
console.log(lemmas);
// ['лыддьыны']

// 4. Generation
const forms = await generate('лыддьыны+V+Ind+Prs+Sg1', 'kpv');
console.log(forms);
// [['лыддя', 0]]

// 5. Segmentation
const morphemes = await segment('лыддьыны', 'kpv');
console.log(morphemes);
// [['лыддьы', 'ны'], ['лыддьы', 'ны']]

Node.js (CommonJS)

const {
  lemmatize,
  analyze,
  generate,
  segment,
  get_translation,
  tokenizer,
} = require('uralicnlp');

Client-Side Browser Usage (Non-Node Projects)

In a web browser, models cannot be written to a local filesystem folder like ~/.uralicnlp. Instead, UralicNLP-JS automatically stores and caches binary transducer models in IndexedDB (uralicnlp_models).

<!DOCTYPE html>
<html>
<head>
  <script src="uralicnlp.min.js"></script>
</head>
<body>
  <script>
    async function run() {
      // 1. Tokenize (instant, synchronous)
      const words = uralicNLP.tokenizer.words('Kissa ja koira!');
      console.log(words); // ['Kissa', 'ja', 'koira', '!']

      // 2. Download language model into browser IndexedDB if needed
      if (!(await uralicNLP.is_language_installed('kpv'))) {
        console.log('Downloading model to browser IndexedDB...');
        await uralicNLP.download('kpv');
      }

      // 3. Analyze wordform
      const result = await uralicNLP.analyze('лыддьыны', 'kpv');
      console.log(result);

      // 4. Lemmatize
      const lemmas = await uralicNLP.lemmatize('лыддьыны', 'kpv');
      console.log(lemmas); // ['лыддьыны']
    }

    run();
  </script>
</body>
</html>

API Reference

Tokenizer

The tokenizer runs synchronously in memory with zero dependencies and zero network requests.

tokenizer.sentences(text: string): string[]

Splits text into sentences, taking abbreviations, numbers, decimals, and punctuation into account.

tokenizer.words(text: string): string[]

Splits a sentence into word and punctuation tokens.

tokenizer.tokenize(text: string): string[][]

Convenience method that splits text into sentences and each sentence into words: sentences(text).map(s => words(s)).


Morphology & Translation

analyze(query, language, options?): Promise<Array<[string, number]>>

Returns morphological analysis for a given word.

  • options.descriptive (boolean, default true): Use descriptive instead of normative analyzer.
  • options.removeSymbols (boolean, default true): Strip internal flag diacritics (@...@).
  • options.languageFlags (boolean, default false): Append +<language> to output strings.
  • options.dictionaryForms (boolean, default false): Use dictionary form analyzer.
  • options.filename (string | Uint8Array | null): Custom transducer path or buffer.
  • options.timeCutoff (number, default 0): Lookup timeout cutoff.

generate(query, language, options?): Promise<Array<[string, number]>>

Generates surface wordforms from morphological analysis tags.

  • options.descriptive (boolean, default false): Use descriptive generator.
  • options.dictionaryForms (boolean, default false): Use dictionary form generator.
  • options.removeSymbols (boolean, default true): Strip internal flag diacritics.

lemmatize(word, language, options?): Promise<string[]>

Extracts base dictionary lemma(s) from a surface wordform.

  • options.wordBoundaries (boolean, default false): Preserve compound word boundaries with |.
  • options.descriptive (boolean, default true): Use descriptive analysis.
  • options.dictionaryForms (boolean, default false): Use dictionary forms.

segment(query, language, options?): Promise<string[][]>

Segments a wordform into morphemes using the morpher transducer.

get_translation(lemma, lang, trans_lang?, options?): Promise<string[] | Record<string, string[]>>

Looks up translations for a lemma from the dictionary transducers.

  • If trans_lang is specified (e.g. 'fin'), returns string[] of translations.
  • If omitted, returns a mapping { [langCode]: string[] }.

Model Management & Storage

download(language: string, options?: DownloadOptions): Promise<void>

Downloads language models from the model server (http://models.uralicnlp.com/nightly/).

  • In Node.js: saves to ~/.uralicnlp/<language>/.
  • In Browser: saves binary Uint8Array files into IndexedDB.
  • options.models: list of specific model types to download (default: all HFST models).
  • options.onProgress: callback ({ language, model, status }) => void.

is_language_installed(language: string): Promise<boolean>

Checks if models for the language are installed locally (on disk or in IndexedDB).

supported_languages(): Promise<string[]>

Queries the model server for the list of supported language codes.

uninstall(language: string): Promise<void>

Removes the language models from storage (disk or IndexedDB) and clears memory cache.


Storage Architecture: Server vs. Client

| Feature | Server (Node.js) | Client (Browser) | | :--- | :--- | :--- | | Storage Engine | File System (node:fs) | IndexedDB (uralicnlp_models) | | Storage Path | ~/.uralicnlp/<language>/ | Object store models | | Binary Model Loading | Direct file stream / buffer | Uint8Array from IndexedDB | | Download Target | Written to local disk | Put into IndexedDB store | | In-Memory Cache | Map<string, Hfst> | Map<string, Hfst> |

Custom storage backends can be registered via setStorage(customStorage).


Building & Testing

# Install dependencies
npm install

# Run build (produces CJS, ESM, and standalone browser bundle in dist/)
npm run build

# Run unit tests
npm test

Cite

If you use UralicNLP in an academic publication, please cite it as follows:

Hämäläinen, Mika. (2019). UralicNLP: An NLP Library for Uralic Languages. Journal of open source software, 4(37), [1345]. https://doi.org/10.21105/joss.01345

@article{uralicnlp_2019, 
    title={{UralicNLP}: An {NLP} Library for {U}ralic Languages},
    DOI={10.21105/joss.01345}, 
    journal={Journal of Open Source Software}, 
    author={Mika Hämäläinen}, 
    year={2019}, 
    volume={4},
    number={37},
    pages={1345}
}