@damorris25/content-extractor
v0.1.1
Published
Extract text content and metadata from unstructured file formats (Office, ODF, PDF, images, audio, CAD)
Maintainers
Readme
@damorris25/content-extractor
Extract text content and metadata from unstructured file formats — Office
documents, OpenDocument, PDF, images, audio, email, CAD, and more. Inspired by
Apache Tika, built browser-first with a single
runtime dependency (fflate, MIT).
This is the TypeScript implementation of content-extractor. A Go implementation lives in the same repository; both produce identical output, enforced by golden-file parity tests.
Install
npm install @damorris25/content-extractorOptional peer dependencies:
npm install pdfjs-dist # PDF text extraction
npm install tesseract.js # OCR for imagesQuick Start
import { extract } from '@damorris25/content-extractor';
const bytes = new Uint8Array(await file.arrayBuffer()); // or fs.readFileSync
const result = await extract(bytes, { filename: 'report.docx' });
console.log(result.text); // extracted text content
console.log(result.contentType); // detected MIME type
console.log(result.metadata); // Record<string, string[]> — title, creator, dates, ...Works in browsers and Node.js. Input is always a Uint8Array; extraction is
async and accepts an optional AbortSignal:
const controller = new AbortController();
const result = await extract(bytes, {
filename: 'big.xlsx',
signal: controller.signal,
});Supported Formats
| Family | Formats | Extracted |
|--------|---------|-----------|
| Microsoft Office | DOCX, XLSX, PPTX | Text + document metadata |
| OpenDocument | ODT, ODS, ODP | Text + document metadata |
| PDF | PDF (text layer) | Text (via pdfjs-dist peer) |
| Web/Markup | XML, HTML, SVG | Visible text, titles |
| Structured data | JSON, CSV, TSV, YAML, Markdown | Pass-through + detection |
| Rich text | RTF | Text with Unicode support |
| Email | EML (RFC 2822) | Body text + headers |
| Images | JPEG, PNG, TIFF/GeoTIFF, BMP, GIF, WebP, HEIC, HEIF, AVIF | EXIF/metadata, optional OCR |
| Audio | MP3, OGG, FLAC, WAV | ID3v2 / Vorbis / RIFF metadata |
| CAD | DXF | Text entities, layers, metadata |
Security
Built for untrusted input: ZIP bomb protection (compression ratio + size caps), integer overflow checks, NUL byte stripping, and bounded parsing for ISOBMFF containers. Report vulnerabilities via GitHub Security Advisories.
Documentation
Full documentation, architecture contracts, benchmarks, and the Go implementation: github.com/damorris25/content-extractor
License
MIT © 2026 Dana Morris
