npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@xberg-io/xberg

v1.0.14

Published

High-performance document intelligence library

Readme

TypeScript (Node.js)

Extract text, tables, images, metadata, and code intelligence from 101 file formats and 371 programming languages including PDF, Office documents, images, and audio/video transcripts where native transcription is available. Native NAPI-RS bindings for Node.js with superior performance, async/await support, and TypeScript type definitions.

What This Package Provides

  • Document intelligence core — extract text, tables, images, metadata, entities, keywords, code intelligence, and transcripts in builds that enable transcription.
  • Format coverage — PDF, Office, images, HTML/XML, email, archives, notebooks, citations, scientific formats, plain text, and audio/video formats in builds that enable transcription.
  • OCR choices — Tesseract, PaddleOCR, Candle where supported, VLM OCR through liter-llm, and plugin hooks for custom backends.
  • Same engine as every binding — Rust, Python, Node.js, Go, Java, PHP, Ruby, .NET, Elixir, WASM, Kotlin Android, Swift, Dart, Zig, and C FFI share the same Rust implementation.
  • Node-first TypeScript API — NAPI-RS package with typed options/results and async extraction.

Installation

Package Installation

pnpm add @xberg-io/xberg

System Requirements

  • Node.js 22+ required (NAPI-RS native bindings)
  • Optional: ONNX Runtime version 1.24+ for ORT-dependent inference features
  • Optional: Tesseract OCR for OCR functionality

Platform Support

Pre-built binaries available for:

  • macOS (arm64, x64)
  • Linux (x64)
  • Windows (x64)

Quick Start

Basic Extraction

Extract text, metadata, and structure from any supported document format:

import { ExtractInputKind, extract } from "@xberg-io/xberg";

const config = {
  useCache: true,
  enableQualityProcessing: true,
};

const output = await extract(
  {
    kind: "uri",
    uri: "document.pdf",
  },
  config,
);

console.log(output.results[0].content);
console.log(`MIME Type: ${output.results[0].mimeType}`);

Common Use Cases

Extract with Custom Configuration

Most use cases benefit from configuration to control extraction behavior:

With OCR (for scanned documents):

import { ExtractInputKind, extract } from "@xberg-io/xberg";

const config = {
  ocr: {
    backend: "tesseract",
    language: ["eng", "fra"],
    tesseractConfig: {
      psm: 3,
    },
  },
};

const output = await extract(
  {
    kind: "uri",
    uri: "document.pdf",
  },
  config,
);

console.log(output.results[0].content);

Table Extraction

import { ExtractInputKind, extract } from "@xberg-io/xberg";

const output = await extract({
  kind: "uri",
  uri: "document.pdf",
});

output.results[0].tables?.forEach((table) => {
  console.log(`Table with ${table.cells?.length ?? 0} rows`);
  console.log(table.markdown);
  table.cells?.forEach((row) => console.log(row.join(" | ")));
});

Processing Multiple Files

import { extractBatch } from "@xberg-io/xberg";

const output = await extractBatch([
  { kind: "uri", uri: "document.pdf" },
  {
    kind: "bytes",
    bytes: Buffer.from("Hello from memory"),
    mimeType: "text/plain",
    filename: "note.txt",
  },
]);

for (const result of output.results) {
  console.log(result.content.slice(0, 200));
}

Async Processing

For non-blocking document processing:

import { ExtractInputKind, extract } from "@xberg-io/xberg";

const output = await extract({
  kind: "uri",
  uri: "document.pdf",
});

console.log(output.results[0].content);
console.log(`Results: ${output.summary.results}`);

Configuration Discovery

import { ExtractInputKind, ExtractionConfig, extract } from "@xberg-io/xberg";

const config = ExtractionConfig.discover();
const input = {
  kind: "uri",
  uri: "document.pdf",
};

if (config) {
  console.log("Found configuration file");
  const output = await extract(input, config);
  console.log(output.results[0].content);
} else {
  console.log("No configuration file found, using defaults");
  const output = await extract(input);
  console.log(output.results[0].content);
}

Next Steps

NAPI-RS Implementation Details

Native Performance

This binding uses NAPI-RS to provide native Node.js bindings with:

  • Zero-copy data transfer between JavaScript and Rust layers
  • Native thread pool for concurrent document processing
  • Direct memory management for efficient large document handling
  • Binary-compatible pre-built native modules across platforms

Threading Model

  • Single documents are processed by Promise-based extraction APIs in the native thread pool
  • Batch operations distribute work across available CPU cores
  • Thread count is configurable but defaults to system CPU count
  • Long-running extractions resolve asynchronously without blocking the JavaScript event loop

Memory Management

  • Large documents (> 100 MB) are streamed to avoid loading entirely into memory
  • Temporary files are created in system temp directory for extraction
  • Memory is automatically released after extraction completion
  • ONNX models are cached in memory for repeated embeddings operations

Features

Supported File Formats (101 formats · 115 file extensions)

101 formats across 115 file extensions in 8 major categories with intelligent format detection and comprehensive metadata extraction.

Office Documents

| Category | Formats | Capabilities | |----------|---------|--------------| | Word Processing | .docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6 | Full text, tables, images, metadata, styles | | Spreadsheets | .xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers | Sheet data, formulas, cell metadata, charts | | Presentations | .pptx, .pptm, .ppt, .ppsx, .potx, .potm, .pot, .odp, .key | Slides, speaker notes, images, metadata | | PDF | .pdf | Text, tables, images, metadata, OCR support | | eBooks | .epub, .fb2 | Chapters, metadata, embedded resources | | Database | .dbf | Table data extraction, field type support | | Hangul | .hwp, .hwpx | Korean document format, text extraction |

Images (OCR-Enabled)

| Category | Formats | Features | |----------|---------|----------| | Raster | .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif | OCR, table detection, EXIF metadata, dimensions, color space | | Advanced | .jp2, .jpx, .jpm, .mj2, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm | OCR via hayro-jpeg2000 (pure Rust decoder), JBIG2 support, table detection, format-specific metadata | | HEIC family | .heic, .heics, .heif, .avif, .avcs | EXIF metadata, optional libheif pixel decoding | | Vector | .svg | DOM parsing, embedded text, graphics metadata |

Audio & Video

| Category | Formats | Features | |----------|---------|----------| | Audio | .mp3, .mpga, .m4a, .wav, .webm | Whisper transcription when native transcription is available | | Video audio track | .mp4, .mpeg, .webm | Audio-track transcription only |

Web & Data

| Category | Formats | Features | |----------|---------|----------| | Markup | .html, .htm, .xhtml, .xml, .svg | DOM parsing, metadata (Open Graph, Twitter Card), link extraction | | Structured Data | .json, .yaml, .yml, .toml, .csv, .tsv | Schema detection, nested structures, validation | | Text & Markdown | .txt, .md, .markdown, .djot, .mdx, .rst, .org, .rtf | CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode |

Email & Archives

| Category | Formats | Features | |----------|---------|----------| | Email | .eml, .msg, .pst | Headers, body (HTML/plain), attachments, threading | | Archives | .zip, .tar, .tgz, .gz, .7z | Recursive extraction of nested archives, file listing, metadata, zip-bomb protection |

Academic & Scientific

| Category | Formats | Features | |----------|---------|----------| | Citations | .bib, .ris, .nbib, .enw | Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX, CSL JSON by MIME type | | Scientific | .tex, .latex, .typ, .typst, .jats, .ipynb | LaTeX, Typst, Jupyter notebooks, PubMed JATS | | Publishing | .fb2, .docbook, .dbk, .docbook4, .docbook5, .opml | FictionBook, DocBook XML, OPML outlines | | Documentation | MIME-only POD, mdoc, troff | Technical documentation formats |

Code Intelligence (371 Languages)

| Feature | Description | |---------|-------------| | Structure Extraction | Functions, classes, methods, structs, interfaces, enums | | Import/Export Analysis | Module dependencies, re-exports, wildcard imports | | Symbol Extraction | Variables, constants, type aliases, properties | | Docstring Parsing | Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats | | Diagnostics | Parse errors with line/column positions | | Syntax-Aware Chunking | Split code by semantic boundaries, not arbitrary byte offsets |

Powered by tree-sitter-language-packdocumentation.

Complete Format Reference

Key Capabilities

  • Text Extraction - Extract all text content with position and formatting information
  • Metadata Extraction - Retrieve document properties, creation date, author, etc.
  • Table Extraction - Parse tables with structure and cell content preservation
  • Image Extraction - Extract embedded images and render page previews
  • Audio/Video Transcription - Extract speech transcripts from MP3, M4A, WAV, WebM, and MP4 inputs when the native transcription feature is available
  • OCR Support - Integrate multiple OCR backends for scanned documents
  • Async/Await - Non-blocking document processing with concurrent operations
  • Plugin System - Extensible post-processing for custom text transformation
  • Embeddings - Generate vector embeddings using ONNX Runtime models or provider-hosted services
  • Batch Processing - Efficiently process multiple documents in parallel
  • Memory Efficient - Stream large files without loading entirely into memory
  • Language Detection - Detect and support multiple languages in documents
  • Code Intelligence - Extract structure, imports, exports, symbols, and docstrings from 371 programming languages via tree-sitter
  • Configuration - Fine-grained control over extraction behavior
  • Six Output Formats - Plain text, Markdown, Djot, HTML, JSON tree structure, or Structured JSON with OCR metadata

OCR Support

Xberg supports multiple OCR backends for extracting text from scanned documents and images:

  • Tesseract

  • Paddleocr

OCR Configuration Example

import { ExtractInputKind, extract } from "@xberg-io/xberg";

const config = {
  ocr: {
    backend: "tesseract",
    language: ["eng", "fra"],
    tesseractConfig: {
      psm: 3,
    },
  },
};

const output = await extract(
  {
    kind: "uri",
    uri: "document.pdf",
  },
  config,
);

console.log(output.results[0].content);

Async Support

This binding provides full async/await support for non-blocking document processing:

import { ExtractInputKind, extract } from "@xberg-io/xberg";

const output = await extract({
  kind: "uri",
  uri: "document.pdf",
});

console.log(output.results[0].content);
console.log(`Results: ${output.summary.results}`);

Plugin System

Xberg supports extensible post-processing plugins for custom text transformation and filtering.

For detailed plugin documentation, visit Plugin System Guide.

Embeddings Support

Generate vector embeddings for extracted text using the built-in ONNX Runtime support. Requires ONNX Runtime installation.

Embeddings Guide

Batch Processing

Process multiple documents efficiently:

import { extractBatch } from "@xberg-io/xberg";

const output = await extractBatch([
  { kind: "uri", uri: "document.pdf" },
  {
    kind: "bytes",
    bytes: Buffer.from("Hello from memory"),
    mimeType: "text/plain",
    filename: "note.txt",
  },
]);

for (const result of output.results) {
  console.log(result.content.slice(0, 200));
}

Configuration

For advanced configuration options including language detection, table extraction, OCR settings, and more:

Configuration Guide

Documentation

Contributing

Contributions are welcome! See Contributing Guide.

Part of Xberg.dev

  • crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
  • html-to-markdown — fast, lossless HTML→Markdown engine.
  • liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
  • tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
  • alef — the polyglot binding generator that produces this README and all per-language bindings.
  • Discord — community, roadmap, announcements.

License

MIT License — see LICENSE for details.

Support