npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

garu-ko

v0.9.16

Published

Ultra-lightweight Korean morphological analyzer for the web (1.2MB model, WASM, F1 96.0% gold testset)

Readme

garu-ko

Browser-native Korean morphological analyzer. No server required.

  • 1.2MB model bundled in npm package (no CDN needed)
  • 412KB WASM engine (176KB gzipped) -- runs in any modern browser
  • F1 96.0% on 9k human-verified gold testset (ep_norm), F1 91.1% on a held-out 2025 spoken-language set
  • ~1ms inference per sentence
  • Offline-ready -- works without network
  • Live Demo -- try it in your browser

Live demos

Try it in the browser — every page below runs the analyzer 100% client-side via WASM.

Quick Start

npm install garu-ko
import { Garu } from 'garu-ko';

const garu = await Garu.load();

// Morphological analysis
const result = garu.analyze('배가 아파서 약을 먹었다');
console.log(result.tokens);
// [
//   { text: '배',   pos: 'NNG', start: 0, end: 2 },
//   { text: '가',   pos: 'JKS', start: 0, end: 2 },
//   { text: '아프', pos: 'VA',  start: 3, end: 6 },
//   { text: '어서', pos: 'EC',  start: 3, end: 6 },
//   { text: '약',   pos: 'NNG', start: 7, end: 9 },
//   { text: '을',   pos: 'JKO', start: 7, end: 9 },
//   { text: '먹',   pos: 'VV',  start: 10, end: 13 },
//   { text: '었',   pos: 'EP',  start: 10, end: 13 },
//   { text: '다',   pos: 'EF',  start: 10, end: 13 },
// ]

// Simple tokenization
const tokens = garu.tokenize('나는 학교에 간다');
// ['나', '는', '학교', '에', '간다']

garu.destroy(); // free WASM memory

Custom Model

// Load from custom URL
const garu = await Garu.load({ modelUrl: '/models/custom.gmdl' });

// Load from ArrayBuffer
const res = await fetch('/models/custom.gmdl');
const garu = await Garu.load({ modelData: await res.arrayBuffer() });

API

Garu.load(options?): Promise<Garu>

Initialize WASM and load model. Uses bundled model by default.

| Option | Type | Description | |---|---|---| | modelData | ArrayBuffer | Provide model bytes directly | | modelUrl | string | Fetch model from URL |

garu.analyze(text, options?): AnalyzeResult

Returns morphological tokens with POS tags (Sejong tagset).

interface Token {
  text: string;   // surface form
  pos: POS;       // POS tag
  start: number;  // eojeol start offset
  end: number;    // eojeol end offset
}

Set options.topN > 1 to get N-best results as an array. Note: topN > 1 is not yet fully supported and may return fewer results.

garu.nouns(text, options?): string[]

Extract nouns (NNG, NNP) from text. Set options.includeSL to also include foreign tokens (SL) like "AI", "BM25".

garu.nouns('인공지능 기술이 발전했다');
// ["인공", "지능", "기술", "발전"]

garu.nouns('AI 기술이 발전했다', { includeSL: true });
// ["AI", "기술", "발전"]

garu.tokenize(text): string[]

Returns surface-form strings only. Lightweight alternative to analyze().

garu.addUserWord(surface, pos, freq?): void

Register a domain word at runtime — no model rebuild required. The word is injected into the lattice and competes with the built-in dictionary on the same cost scale.

garu.nouns('인공지능 기술이 발전했다');
// ["인공", "지능", "기술", "발전"]

garu.addUserWord('인공지능', 'NNG');
garu.nouns('인공지능 기술이 발전했다');
// ["인공지능", "기술", "발전"]

freq uses the same scale as the built-in dictionary — higher values are more likely to beat a split analysis (default 5000). A user word cannot span whitespace.

Useful for product names, organisations and neologisms that a general corpus does not cover, especially when indexing them as single search terms.

garu.addUserWords(words): void

Register several words at once.

garu.addUserWords([
  { surface: '전세사기', pos: 'NNG' },
  { surface: '탄소중립', pos: 'NNG', freq: 20000 },
]);

garu.clearUserWords(): void / garu.userWordCount(): number

Remove every registered user word, or count them. With none registered, analysis is identical to using the built-in dictionary alone.

garu.destroy(): void

Free WASM memory. Instance is unusable after this call.

Integrations

Drop-in tokenizers for popular JS search libraries:

Both solve the same problem: default tokenizers don't handle Korean particles or verb inflections, so "먹다" never matches "먹었다" and "학교" misses "학교에". These adapters run morphological analysis so the inflections fall off before indexing.

import { create, insert, search } from '@orama/orama'
import { createTokenizer } from 'garu-orama-tokenizer'

const db = await create({
  schema: { title: 'string' },
  components: { tokenizer: await createTokenizer() }
})
await insert(db, { title: '학교에서 점심을 먹었다' })

await search(db, { term: '먹다' })  // ← matches

FAQ

What is garu-ko? garu-ko (가루/Garu) is a browser-native Korean morphological analyzer. A 1.2MB model and a 412KB WASM engine run entirely in the browser, so it segments Korean text, tags parts of speech, extracts nouns, and tokenizes with no server and no network.

Does it need a server or API? No. It runs 100% client-side via WebAssembly. After the initial load there is no backend call, so it works offline and inside browser extensions, service-worker PWAs, and intranet apps.

How accurate is it? F1 96.0% on a 9,000-sentence human-verified gold testset (ep_norm), and F1 91.1% on a held-out 2025 spoken-language evaluation set.

Which POS tagset does it use? The Sejong tagset (42 tags) from the National Institute of Korean Language — NNG/NNP for nouns, VV/VA for verbs/adjectives, particles, endings, and symbols.

Does it use a neural network? No. It combines a codebook, an eojeol cache, sentence-level N-best trigram Viterbi, deterministic post-processing rules, and a linear reranking perceptron — not a neural network — so output is deterministic and reproducible.

Where does it run besides the browser? The same npm install garu-ko works in browser ESM, Node.js 18+, Bun, and Deno with one API.

Is it free for commercial use? Yes. garu-ko is open source under the MIT license with no usage fees.

Acknowledgments

The morphological analysis model is trained on the NIKL Morpheme-Tagged Corpus (v1.1) provided by the National Institute of Korean Language (국립국어원). The model contains only derived frequency statistics, not original text.

License

MIT