@xberg-io/xberg
v1.0.14
Published
High-performance document intelligence library
Readme
TypeScript (Node.js)
Extract text, tables, images, metadata, and code intelligence from 101 file formats and 371 programming languages including PDF, Office documents, images, and audio/video transcripts where native transcription is available. Native NAPI-RS bindings for Node.js with superior performance, async/await support, and TypeScript type definitions.
What This Package Provides
- Document intelligence core — extract text, tables, images, metadata, entities, keywords, code intelligence, and transcripts in builds that enable transcription.
- Format coverage — PDF, Office, images, HTML/XML, email, archives, notebooks, citations, scientific formats, plain text, and audio/video formats in builds that enable transcription.
- OCR choices — Tesseract, PaddleOCR, Candle where supported, VLM OCR through liter-llm, and plugin hooks for custom backends.
- Same engine as every binding — Rust, Python, Node.js, Go, Java, PHP, Ruby, .NET, Elixir, WASM, Kotlin Android, Swift, Dart, Zig, and C FFI share the same Rust implementation.
- Node-first TypeScript API — NAPI-RS package with typed options/results and async extraction.
Installation
Package Installation
pnpm add @xberg-io/xbergSystem Requirements
- Node.js 22+ required (NAPI-RS native bindings)
- Optional: ONNX Runtime version 1.24+ for ORT-dependent inference features
- Optional: Tesseract OCR for OCR functionality
Platform Support
Pre-built binaries available for:
- macOS (arm64, x64)
- Linux (x64)
- Windows (x64)
Quick Start
Basic Extraction
Extract text, metadata, and structure from any supported document format:
import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = {
useCache: true,
enableQualityProcessing: true,
};
const output = await extract(
{
kind: "uri",
uri: "document.pdf",
},
config,
);
console.log(output.results[0].content);
console.log(`MIME Type: ${output.results[0].mimeType}`);Common Use Cases
Extract with Custom Configuration
Most use cases benefit from configuration to control extraction behavior:
With OCR (for scanned documents):
import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = {
ocr: {
backend: "tesseract",
language: ["eng", "fra"],
tesseractConfig: {
psm: 3,
},
},
};
const output = await extract(
{
kind: "uri",
uri: "document.pdf",
},
config,
);
console.log(output.results[0].content);Table Extraction
import { ExtractInputKind, extract } from "@xberg-io/xberg";
const output = await extract({
kind: "uri",
uri: "document.pdf",
});
output.results[0].tables?.forEach((table) => {
console.log(`Table with ${table.cells?.length ?? 0} rows`);
console.log(table.markdown);
table.cells?.forEach((row) => console.log(row.join(" | ")));
});Processing Multiple Files
import { extractBatch } from "@xberg-io/xberg";
const output = await extractBatch([
{ kind: "uri", uri: "document.pdf" },
{
kind: "bytes",
bytes: Buffer.from("Hello from memory"),
mimeType: "text/plain",
filename: "note.txt",
},
]);
for (const result of output.results) {
console.log(result.content.slice(0, 200));
}Async Processing
For non-blocking document processing:
import { ExtractInputKind, extract } from "@xberg-io/xberg";
const output = await extract({
kind: "uri",
uri: "document.pdf",
});
console.log(output.results[0].content);
console.log(`Results: ${output.summary.results}`);Configuration Discovery
import { ExtractInputKind, ExtractionConfig, extract } from "@xberg-io/xberg";
const config = ExtractionConfig.discover();
const input = {
kind: "uri",
uri: "document.pdf",
};
if (config) {
console.log("Found configuration file");
const output = await extract(input, config);
console.log(output.results[0].content);
} else {
console.log("No configuration file found, using defaults");
const output = await extract(input);
console.log(output.results[0].content);
}Next Steps
- Installation Guide - Platform-specific setup
- API Documentation - Complete API reference
- Examples & Guides - Full code examples and usage guides
- Configuration Guide - Advanced configuration options
NAPI-RS Implementation Details
Native Performance
This binding uses NAPI-RS to provide native Node.js bindings with:
- Zero-copy data transfer between JavaScript and Rust layers
- Native thread pool for concurrent document processing
- Direct memory management for efficient large document handling
- Binary-compatible pre-built native modules across platforms
Threading Model
- Single documents are processed by Promise-based extraction APIs in the native thread pool
- Batch operations distribute work across available CPU cores
- Thread count is configurable but defaults to system CPU count
- Long-running extractions resolve asynchronously without blocking the JavaScript event loop
Memory Management
- Large documents (> 100 MB) are streamed to avoid loading entirely into memory
- Temporary files are created in system temp directory for extraction
- Memory is automatically released after extraction completion
- ONNX models are cached in memory for repeated embeddings operations
Features
Supported File Formats (101 formats · 115 file extensions)
101 formats across 115 file extensions in 8 major categories with intelligent format detection and comprehensive metadata extraction.
Office Documents
| Category | Formats | Capabilities |
|----------|---------|--------------|
| Word Processing | .docx, .docm, .doc, .dotx, .dotm, .dot, .odt, .pages, .wpd, .wp, .wp5, .wp6 | Full text, tables, images, metadata, styles |
| Spreadsheets | .xlsx, .xlsm, .xlsb, .xls, .xla, .xlam, .xltm, .xltx, .xlt, .ods, .numbers | Sheet data, formulas, cell metadata, charts |
| Presentations | .pptx, .pptm, .ppt, .ppsx, .potx, .potm, .pot, .odp, .key | Slides, speaker notes, images, metadata |
| PDF | .pdf | Text, tables, images, metadata, OCR support |
| eBooks | .epub, .fb2 | Chapters, metadata, embedded resources |
| Database | .dbf | Table data extraction, field type support |
| Hangul | .hwp, .hwpx | Korean document format, text extraction |
Images (OCR-Enabled)
| Category | Formats | Features |
|----------|---------|----------|
| Raster | .png, .jpg, .jpeg, .gif, .webp, .bmp, .tiff, .tif | OCR, table detection, EXIF metadata, dimensions, color space |
| Advanced | .jp2, .jpx, .jpm, .mj2, .jbig2, .jb2, .pnm, .pbm, .pgm, .ppm | OCR via hayro-jpeg2000 (pure Rust decoder), JBIG2 support, table detection, format-specific metadata |
| HEIC family | .heic, .heics, .heif, .avif, .avcs | EXIF metadata, optional libheif pixel decoding |
| Vector | .svg | DOM parsing, embedded text, graphics metadata |
Audio & Video
| Category | Formats | Features |
|----------|---------|----------|
| Audio | .mp3, .mpga, .m4a, .wav, .webm | Whisper transcription when native transcription is available |
| Video audio track | .mp4, .mpeg, .webm | Audio-track transcription only |
Web & Data
| Category | Formats | Features |
|----------|---------|----------|
| Markup | .html, .htm, .xhtml, .xml, .svg | DOM parsing, metadata (Open Graph, Twitter Card), link extraction |
| Structured Data | .json, .yaml, .yml, .toml, .csv, .tsv | Schema detection, nested structures, validation |
| Text & Markdown | .txt, .md, .markdown, .djot, .mdx, .rst, .org, .rtf | CommonMark, GFM, Djot, MDX, reStructuredText, Org Mode |
Email & Archives
| Category | Formats | Features |
|----------|---------|----------|
| Email | .eml, .msg, .pst | Headers, body (HTML/plain), attachments, threading |
| Archives | .zip, .tar, .tgz, .gz, .7z | Recursive extraction of nested archives, file listing, metadata, zip-bomb protection |
Academic & Scientific
| Category | Formats | Features |
|----------|---------|----------|
| Citations | .bib, .ris, .nbib, .enw | Structured parsing: RIS, PubMed/MEDLINE, EndNote XML, BibTeX/BibLaTeX, CSL JSON by MIME type |
| Scientific | .tex, .latex, .typ, .typst, .jats, .ipynb | LaTeX, Typst, Jupyter notebooks, PubMed JATS |
| Publishing | .fb2, .docbook, .dbk, .docbook4, .docbook5, .opml | FictionBook, DocBook XML, OPML outlines |
| Documentation | MIME-only POD, mdoc, troff | Technical documentation formats |
Code Intelligence (371 Languages)
| Feature | Description | |---------|-------------| | Structure Extraction | Functions, classes, methods, structs, interfaces, enums | | Import/Export Analysis | Module dependencies, re-exports, wildcard imports | | Symbol Extraction | Variables, constants, type aliases, properties | | Docstring Parsing | Google, NumPy, Sphinx, JSDoc, RustDoc, and 10+ formats | | Diagnostics | Parse errors with line/column positions | | Syntax-Aware Chunking | Split code by semantic boundaries, not arbitrary byte offsets |
Powered by tree-sitter-language-pack — documentation.
Key Capabilities
- Text Extraction - Extract all text content with position and formatting information
- Metadata Extraction - Retrieve document properties, creation date, author, etc.
- Table Extraction - Parse tables with structure and cell content preservation
- Image Extraction - Extract embedded images and render page previews
- Audio/Video Transcription - Extract speech transcripts from MP3, M4A, WAV, WebM, and MP4 inputs when the native transcription feature is available
- OCR Support - Integrate multiple OCR backends for scanned documents
- Async/Await - Non-blocking document processing with concurrent operations
- Plugin System - Extensible post-processing for custom text transformation
- Embeddings - Generate vector embeddings using ONNX Runtime models or provider-hosted services
- Batch Processing - Efficiently process multiple documents in parallel
- Memory Efficient - Stream large files without loading entirely into memory
- Language Detection - Detect and support multiple languages in documents
- Code Intelligence - Extract structure, imports, exports, symbols, and docstrings from 371 programming languages via tree-sitter
- Configuration - Fine-grained control over extraction behavior
- Six Output Formats - Plain text, Markdown, Djot, HTML, JSON tree structure, or Structured JSON with OCR metadata
OCR Support
Xberg supports multiple OCR backends for extracting text from scanned documents and images:
Tesseract
Paddleocr
OCR Configuration Example
import { ExtractInputKind, extract } from "@xberg-io/xberg";
const config = {
ocr: {
backend: "tesseract",
language: ["eng", "fra"],
tesseractConfig: {
psm: 3,
},
},
};
const output = await extract(
{
kind: "uri",
uri: "document.pdf",
},
config,
);
console.log(output.results[0].content);Async Support
This binding provides full async/await support for non-blocking document processing:
import { ExtractInputKind, extract } from "@xberg-io/xberg";
const output = await extract({
kind: "uri",
uri: "document.pdf",
});
console.log(output.results[0].content);
console.log(`Results: ${output.summary.results}`);Plugin System
Xberg supports extensible post-processing plugins for custom text transformation and filtering.
For detailed plugin documentation, visit Plugin System Guide.
Embeddings Support
Generate vector embeddings for extracted text using the built-in ONNX Runtime support. Requires ONNX Runtime installation.
Batch Processing
Process multiple documents efficiently:
import { extractBatch } from "@xberg-io/xberg";
const output = await extractBatch([
{ kind: "uri", uri: "document.pdf" },
{
kind: "bytes",
bytes: Buffer.from("Hello from memory"),
mimeType: "text/plain",
filename: "note.txt",
},
]);
for (const result of output.results) {
console.log(result.content.slice(0, 200));
}Configuration
For advanced configuration options including language detection, table extraction, OCR settings, and more:
Documentation
Contributing
Contributions are welcome! See Contributing Guide.
Part of Xberg.dev
- crawlberg — web crawling and scraping with HTML→Markdown and headless-Chrome fallback.
- html-to-markdown — fast, lossless HTML→Markdown engine.
- liter-llm — universal LLM API client with native bindings for 14 languages and 165 providers.
- tree-sitter-language-pack — tree-sitter grammars and code-intelligence primitives.
- alef — the polyglot binding generator that produces this README and all per-language bindings.
- Discord — community, roadmap, announcements.
License
MIT License — see LICENSE for details.
Support
- Discord Community: Join our Discord
- GitHub Issues: Report bugs
- Discussions: Ask questions
