npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

chunklet

v0.2.0

Published

Token-aware, structure-aware text chunking for RAG - exact source offsets, sentence, markdown, code and HTML modes, pluggable tokenizer, zero dependencies.

Readme

chunklet

npm version CI license: MIT

Token-aware, structure-aware text chunking for RAG pipelines - zero dependencies, runs in Node, browsers, and edge runtimes.

Two ideas define it. Exact source offsets: every chunk guarantees

chunk.text === source.slice(chunk.start, chunk.end)

so you can highlight citations, deep-link retrieval hits, or store embeddings without storing text. And structure awareness: chunks respect the document - markdown and HTML sections with heading breadcrumbs, whole sentences, atomic code fences, whole declarations in source files - and a word is never cut in half unless a single word exceeds the budget.

Install

npm i chunklet

Quick start

import { chunkText } from 'chunklet';

const source = await fs.readFile('handbook.txt', 'utf8');
const chunks = chunkText(source, { maxTokens: 512, overlap: 64 });

for (const chunk of chunks) {
  console.log(chunk.index, chunk.tokens, `[${chunk.start}-${chunk.end}]`);
  console.log(chunk.text.slice(0, 60) + '...');
}

Every chunk is:

{
  text: string;    // exactly source.slice(start, end)
  start: number;   // character offset, inclusive
  end: number;     // character offset, exclusive
  tokens: number;  // per the active tokenizer
  index: number;   // 0-based position
  meta?: { headings?: string[]; language?: string; symbol?: string; source?: { start: number; end: number } }; // markdown / code / html modes
}

How the splitting works

chunkText splits hierarchically: paragraphs (\n\n) first, then lines, then sentences, then words - and only descends a level when a piece is still over the token budget. The resulting pieces are packed back together greedily up to maxTokens. The effect: chunks end at the most natural boundary available, and a mid-word cut can only happen when a single word alone exceeds the budget.

Chunk edges are whitespace-trimmed with the offsets adjusted, so the slice invariant always holds - the text is never normalized, joined with synthetic separators, or otherwise rewritten.

Overlap

overlap repeats trailing context at the start of the next chunk, which softens the "answer was split across two chunks" failure mode of retrieval:

const chunks = chunkText(source, { maxTokens: 512, overlap: 64 });
// chunk 3 begins with the last ~64 tokens of chunk 2

Overlap is honored even when the previous chunk ends in one long piece - chunklet takes a suffix of it rather than silently skipping the overlap.

Sentence mode

import { chunkSentences } from 'chunklet';

const chunks = chunkSentences(article, { maxTokens: 256 });

Whole sentences are packed into the budget, so a chunk never ends mid-sentence (unless a single sentence is itself over budget - then it degrades to word splitting). Boundaries come from Intl.Segmenter, which handles abbreviations like "Mr. Smith" and locale rules correctly; a regex fallback covers runtimes without it.

Use this over chunkText when your chunks are small (embedding models with short context, tweet-sized snippets) and a dangling half-sentence would hurt embedding quality.

Markdown mode

import { chunkMarkdown } from 'chunklet';

const chunks = chunkMarkdown(readme, { maxTokens: 512 });

chunks[4].text;           // "## Install\n\nRun the installer..."
chunks[4].meta?.headings; // ['Guide', 'Install']

What it does differently:

  • Sections follow the headings. A chunk never crosses an ATX heading (# through ######), so retrieval hits map cleanly to document sections.
  • Every chunk knows where it lives. meta.headings is the breadcrumb of enclosing headings, outermost first. Content before the first heading gets [].
  • Code fences are atomic. A fenced block is never merged mid-fence with prose, and only split internally (line by line) when the fence alone exceeds the budget.

Code mode

import { chunkCode } from 'chunklet';

const chunks = chunkCode(source, { maxTokens: 512, language: 'ts' });

chunks[2].text;           // "/** Loads one file. */\nexport async function load(name: string) {..."
chunks[2].meta?.symbol;   // 'export async function load(name: string): Promise<string> {'
chunks[2].meta?.language; // 'ts'

Source files are split at declarations, not at arbitrary lines:

  • Declarations are the unit. A top-level declaration starts at an unindented line that follows a blank line or a closing bracket (or opens the file). Comment lines directly above it belong to it, so a doc comment stays with its function. A declaration that fits the budget is never split, and small ones are packed together.
  • Oversized declarations split at structure. A class or function over the budget starts its own chunks and is cut at blank lines before its least-indented members first (methods before the statements inside them), then after closing brackets and at dedents, then at line breaks. A chunk never starts mid-line unless a single line alone exceeds the budget.
  • Every chunk can say where it came from. meta.symbol is the first line of the enclosing top-level declaration (comments and decorators skipped) whenever the chunk sits inside one declaration with an indented body; imports and one-liners get none. meta.language echoes the hint.
  • Language agnostic. Boundaries come from indentation and blank lines, so anything indented consistently works. The language hint only decides what counts as a comment line (// and /* */ for the C family, # for Python and shells, -- for SQL and Lua, ...); without it every common marker is recognized.

HTML mode

import { chunkHtml } from 'chunklet';

const { text, chunks } = chunkHtml(page, { maxTokens: 512 });

chunks[3].text;           // "Run the installer as shown below."
chunks[3].meta?.headings; // ['Guide', 'Install']
chunks[3].meta?.source;   // { start: 1043, end: 1076 } - range in the original HTML
text.slice(chunks[3].start, chunks[3].end) === chunks[3].text; // true

chunkHtml extracts readable text first and chunks that, so offsets refer to the returned text, not to the HTML. meta.source maps each chunk back to the HTML range that produced it, so highlighting in the page is still a slice: page.slice(source.start, source.end) is the markup between the chunk's first and last character. The extractor is a small tolerant tag scanner, not a full parser, with no dependencies:

  • Structure survives. Block elements (p, div, headings, li, td, ...) become paragraph, line or cell boundaries; inline elements disappear; runs of whitespace collapse to one space, the way a browser renders them.
  • Headings become breadcrumbs. h1 through h6 drive sections exactly like markdown mode, and every chunk carries meta.headings.
  • Code stays whole. <pre> blocks keep their whitespace and are never merged mid-block with prose; a <code> element that stands alone as a block is treated the same way, while inline <code> is plain text. They split internally, line by line, only when one alone exceeds the budget.
  • Noise is dropped. <script>, <style>, <template>, <svg>, <iframe> and comments contribute nothing. Common named entities, &#169; and &#x1F600; are decoded; &nbsp; becomes a plain space; unknown entities stay literal. Stray < in text and unclosed tags are tolerated.

Recipes

RAG ingestion

import { chunkMarkdown } from 'chunklet';

const chunks = chunkMarkdown(doc, { maxTokens: 400, overlap: 40 });

for (const chunk of chunks) {
  // prepend the breadcrumb - cheap and measurably better retrieval
  const context = chunk.meta?.headings?.length
    ? chunk.meta.headings.join(' > ') + '\n\n' + chunk.text
    : chunk.text;

  await db.insert({
    id: `${docId}#${chunk.index}`,
    embedding: await embed(context),
    start: chunk.start, // store offsets, not text
    end: chunk.end,
  });
}

Citation highlighting

Because offsets are exact, mapping a retrieval hit back onto the original document is a slice, not a fuzzy search:

const hit = results[0]; // { start, end } straight from the stored chunk

const before = doc.slice(0, hit.start);
const match = doc.slice(hit.start, hit.end);
const after = doc.slice(hit.end);
render(`${before}<mark>${escape(match)}</mark>${after}`);

Real token counts

The default tokenizer is a fast chars/4 heuristic - fine for packing budgets. For exact counts, plug in any counter:

import { encodingForModel } from 'js-tiktoken';

const enc = encodingForModel('gpt-4o');
const chunks = chunkText(doc, {
  maxTokens: 512,
  tokenizer: (t) => enc.encode(t).length,
});

The tokenizer is called on candidate slices during packing, so a heavyweight tokenizer slows chunking - the heuristic + a safety margin (e.g. budget 480 for a 512 limit) is often the better trade.

API

| Export | Description | |---|---| | chunkText(text, options?) | Hierarchical separator splitting (paragraphs > lines > sentences > words) | | chunkSentences(text, options?) | Sentence-boundary packing via Intl.Segmenter | | chunkMarkdown(text, options?) | Heading-aware sections + breadcrumbs, atomic code fences | | chunkCode(source, options?) | Declaration-aware splitting for source files, meta.symbol + meta.language | | chunkHtml(html, options?) | Readable-text extraction + heading breadcrumbs; returns { text, chunks } with meta.source ranges into the HTML | | estimateTokens(text) | The default chars/4 heuristic |

Options

| Option | Default | Description | |---|---|---| | maxTokens | 512 | Token budget per chunk | | overlap | 0 | Tokens of trailing context repeated at the start of the next chunk (must be < maxTokens) | | tokenizer | chars/4 | (text: string) => number | | language | - | chunkCode only: comment-syntax hint ('ts', 'python', 'sql', ...), echoed as meta.language |

Invalid options throw RangeError. Empty or whitespace-only input returns [].

Alternatives

  • llm-splitter - offset-tracked chunks with a bring-your-own splitter function. chunklet adds the structural layer: markdown sections with heading breadcrumbs, atomic code fences, Intl.Segmenter sentence boundaries, and hierarchical fallback so chunks land on natural boundaries.
  • LangChain / LlamaIndex text splitters - similar strategies inside much larger frameworks; reach for chunklet when you want the splitter without the framework.

License

MIT (c) Muzaffar Qosimov