officeparser
v8.1.1
Published
A robust, strictly-typed Node.js and Browser library for parsing office files (.docx, .pptx, .xlsx, .odt, .odp, .ods, .odg, .pdf, .rtf, .csv, .md, .html, .epub, .tex) and generating high-fidelity outputs in DOCX, ODT, LaTeX, Markdown, HTML, CSV, RTF, PDF,
Maintainers
Readme
officeParser: Universal Office Document Parser & Generator
A robust, strictly-typed Node.js and Browser library for parsing office files into a rich Abstract Syntax Tree (AST) and generating high-fidelity output in multiple formats.
Parses: docx · pptx · xlsx · odt · odp · ods · odg · pdf · rtf · csv · md · html · epub · tex (LaTeX, including Overleaf project zips)
Generates: DOCX · ODT · LaTeX · Markdown · HTML · CSV · RTF · PDF · EPUB · Plain Text · RAG Chunks
🌟 Live Interactive AST Visualizer & Documentation 🌟
Upload any office file in your browser: inspect the AST, tweak config, and preview generated output in real-time.
- AST Visualizer: Inspect the hierarchical node tree, metadata, and raw content
- Config Configurator: Tweak options (
ignoreNotes,ocr,newlineDelimiter) and see results instantly - Debugging: Identify exactly how nodes are interpreted
- Format Specs: Read detailed specs for the AST structure and all config options
📝 Changelog
What's New in v8
- Rebuilt PDF text extraction. PDF is no longer treated as a page of flat lines. Tagged PDFs now yield real
heading(with correct levels),table/row/cell,listand footnote/endnotenotenodes, and oneparagraphper paragraph. Untagged PDFs recover the same structure geometrically. Multi-column and float-beside-text pages are read in the correct order (recursive XY-cut), broken and glued words are fixed from inter-fragment spacing, super/subscripts and hyphenated line-breaks are rejoined, rotated text is recovered, and internal links resolve to the target section..to('text')is layout-faithful by default, rendering each page as a spatial grid so columns and tables line up like the source. Per-run color/highlight extraction (on by default; setpdfParserConfig.extractTextColor: falseto skip it) and merged-cell (colSpan/rowSpan) recovery round it out, and every node carries page geometry (bounds). - Password-protected documents. Encrypted PDF, OOXML (
docx/xlsx/pptx) and ODF (odt/ods/odp/odg) open through one unifiedpassword/onPasswordoption, across parsing, conversion and templating. - Native DOCX & ODT generation, plus a native PDF engine (
pdfConfig.engine: 'native', built onpdf-lib) that produces real PDF bytes with no headless browser, in Node and the browser alike. - Templates / mail-merge via
OfficeTemplate.render(fill a DOCX template's{{placeholders}}, single or batch), and ODG parsing (LibreOffice Draw). - LaTeX in both directions (8.1):
.texfiles and Overleaf project zips parse into the same AST as every other format (sections, lists, tables with merged cells, figures, math, footnotes, citations, cross-references, user macros,beamerslides), so LaTeX converts to DOCX, ODT, HTML, Markdown and the rest. Andto('tex')turns any parsed document into LaTeX source that compiles unmodified with pdfLaTeX, XeLaTeX, LuaLaTeX, upLaTeX, pLaTeX andlatex(the last three through dvipdfmx), presentations included (asbeamerframes), carrying its images inside the one.texfile (or, in bundle mode, as files in a zip). See LaTeX Support.
See the full changelog for the complete list, including breaking changes.
Table of Contents
- What's New in v8
- Install
- Command Line Usage
- Quick Decision Guide
- Library Usage: Parsing
- OfficeGenerator
- OfficeConverter: One-Step API
- OfficeTemplate: Mail-Merge / Document Generation
- Native RAG Chunking
- The AST Structure
- Deep Dive: Document Components
- Markdown Dialect Support
- EPUB Support
- LaTeX Support
- Performance Highlights
- Advanced AST Usage
- Configuration Reference
- OCR Scheduler & Resource Management
- Browser Usage
- Troubleshooting & Common Issues
- Known Limitations
- Security & Trust Boundary
- Contributing
Install via npm
npm i officeparser[!NOTE] Requires Node.js >= 22.13.
Command Line Usage
# Full AST as JSON (default)
npx officeparser /path/to/file.docx
# Plain text output
npx officeparser /path/to/file.docx --to=text
# Convert DOCX to Markdown and save
npx officeparser report.docx --to=md --output=report.md
# Convert PPTX to HTML with OCR (OCR runs over extracted images, so --extractAttachments is required)
npx officeparser presentation.pptx --to=html --output=preview.html --ocr --extractAttachments
# Convert XLSX to CSV with a custom delimiter
npx officeparser data.xlsx --to=csv --csvDelimiter=";"
# Generate RAG chunks
npx officeparser document.pdf --to=chunks
# Convert DOCX to EPUB (--extractAttachments is required to embed images)
npx officeparser book.docx --extractAttachments --to=epub --output=book.epub
# Convert Markdown (or any source) to a Word document
npx officeparser notes.md --extractAttachments --to=docx --output=notes.docx
# Convert a Word document (or any source) to OpenDocument Text
npx officeparser report.docx --extractAttachments --to=odt --output=report.odt
# Convert any source to LaTeX: a .tex file, or a zip of main.tex plus its images
npx officeparser paper.docx --extractAttachments --to=tex --output=paper.tex
npx officeparser paper.docx --extractAttachments --to=tex --texConfig.bundle --output=paper.zip
# Convert LaTeX to Word: a single .tex, or an Overleaf project zip (which brings its \input files and images)
npx officeparser paper.tex --to=docx --output=paper.docx
npx officeparser overleaf-project.zip --extractAttachments --to=docx --output=paper.docx
# Overriding file extension mapping
npx officeparser my_document --fileType=docx --to=jsonCLI Syntax
- Values: Flags can be passed as
--flag=valueor--flag value. - Booleans: Bare flags imply
true(e.g.--ocris equivalent to--ocr=true). Negation flags start withno-(e.g.--no-ocris equivalent to--ocr=false). - Nested Objects: You can pass nested properties directly using JSON dot-notation (e.g.
--ocrConfig.language=fraor--htmlConfig.containerWidth=900px). - Images: the CLI parses directly, so add
--extractAttachmentsfor images to reach any output (HTML/EPUB embed them, DOCX/ODT/Markdown/native-PDF include them, and LaTeX carries PNG and JPEG images inside the.tex, or with--texConfig.bundlepackages them beside it). Without it, an image node has no bytes and HTML/Markdown emit a name-only<img src="image1.png">reference. (TheOfficeConverter/convert()API auto-enables this; the CLI does not.)
CLI Options
| Flag | Values | Default | Description |
|------|--------|---------|-------------|
| --to | json\|text\|md\|html\|csv\|rtf\|pdf\|docx\|odt\|tex\|epub\|chunks | json | Output format (latex is accepted as an alias of tex) |
| --output | path | (none) | Write output to a file |
| --fileType | docx\|xlsx\|pptx\|odt\|odp\|ods\|odg\|pdf\|rtf\|csv\|md\|html\|epub\|tex | (none) | Explicitly override input file type detection. Also accepts latex/ltx, the ODF template names ott/ots/otp/otg, and zip (parsed as whatever the archive holds) |
| --ocr | boolean | false | Enable OCR for images (also requires --extractAttachments; OCR runs over extracted images) |
| --ocrConfig.language | string | eng | Tesseract language(s), e.g. deu or eng+fra |
| --ocrConfig.preserveLayout | boolean | true | Keep the line layout of recognized text |
| --password | string | (none) | Password for an encrypted document (PDF, OOXML, or ODF) |
| --extractAttachments | boolean | false | Extract images/charts as Base64 |
| --ignoreNotes | boolean | false | Ignore footnotes/endnotes/speaker notes |
| --ignoreComments | boolean | false | Ignore inline comments |
| --ignoreHeadersAndFooters | boolean | false | Ignore headers and footers |
| --ignoreSlideMasters | boolean | false | Ignore slide masters |
| --ignoreInternalLinks | boolean | false | Ignore internal links |
| --newlineDelimiter | string | \n | Delimiter between lines/blocks in plaintext outputs |
| --csvDelimiter | string | , | Custom delimiter for CSV files |
| --includeRawContent | boolean | false | Include raw XML/RTF in nodes |
| --serializeRawContent | boolean | true | Include stringified XML in metadata |
| --preserveXmlWhitespace | boolean | false | Keep raw formatting space |
| --includeBreakNodes | boolean | false | Include layout break nodes (DOCX, ODF and LaTeX page and column breaks; a typed line break is always kept) |
| --ignorePageGeometry | boolean | false | Omit per-node bounding boxes and page dimensions |
| --pdfParserConfig.useTags | boolean | true | Use the PDF tag tree; false forces geometry-only structure |
| --pdfParserConfig.detectColumns | boolean | true | Multi-column reading-order detection |
| --pdfParserConfig.pageRange | string | all | Parse only the given pages, e.g. 1-3,7 |
| --htmlParserConfig.preserveComments | boolean | false | Keep HTML/EPUB <!-- --> comments as comment nodes |
| --texParserConfig.today | string | the date of the parse | What \today prints in LaTeX input, e.g. --texParserConfig.today="May 1, 2024" |
| --pdfParserConfig.headingDetection | auto\|font-size\|off | auto | How headings are inferred on the geometry path |
| --pdfParserConfig.mergeHyphenatedWords | boolean | true | Rejoin words hyphenated across line breaks |
| --pdfParserConfig.normalizeText | boolean | true | Unicode/ligature normalization of extracted text |
| --pdfParserConfig.extractTextColor | boolean | true | Record each run's fill colour in formatting.color (set false to skip for speed) |
| --pdfParserConfig.maxTextItems | number | 20000 | Base of the text items a PDF may yield (plus one per byte of the file) |
| --pdfParserConfig.maxOperators | number | 250000 | Base of the drawing operators kept from a PDF (plus four per byte of the file) |
| --pdfParserConfig.maxAnnotations | number | 10000 | Base of the annotations read from a PDF (plus one per 32 bytes of the file) |
| --pdfParserConfig.maxTimeMs | number | 5000 | Base of the CPU time the separate pdf.js process may spend on a PDF (plus 20 ms per KB of the file) |
| --pdfParserConfig.separateProcess | boolean | true | Run pdf.js in a separate process under a memory limit (Node) |
| --pdfParserConfig.processMemoryMb | number | 1024 | Heap the separate pdf.js process may use |
| --verbose | boolean | false | Show full error stack traces and warning logs |
| --includeFormatting | boolean | true | Include formatting style map matching |
| --renderMetadata | boolean | false | Render metadata as visible content in the generated output |
| --includeImages | image-only\|image+ocr-text\|ocr-text-only\|none | image-only | How image nodes render. Works as --includeImages=<mode> or --includeImages <mode>; a bare --includeImages means image-only |
| --maxInlineImageBytes | number | 1500000 | Largest image HTML/Markdown inlines as a data: URI (0 never inlines) |
| --htmlConfig.containerWidth | string | number | auto | HTML output container width (e.g. 900px, 100%) |
| --textConfig.pageSeparator | string | \n | Separator written between pages in text output |
| --pdfConfig.engine | html\|native | html | PDF engine: Puppeteer (html) or pdf-lib (native, no browser) |
| --texConfig.bundle | boolean | false | LaTeX: write a zip of main.tex plus its images/ instead of the .tex alone |
| --texConfig.embedImages | boolean | true | LaTeX: false refers to images/ files instead of carrying PNG and JPEG images inside the .tex |
| --texConfig.documentClass | auto\|article\|report\|book\|beamer | auto | LaTeX document class (auto = beamer for presentations, article otherwise) |
| --texConfig.standalone | boolean | true | LaTeX: false writes the body only, for pasting into an existing document |
| ~~--format~~ | json\|text\|md\|html\|csv\|rtf\|pdf\|docx\|odt\|tex\|epub\|chunks | json | Deprecated. Use --to |
| ~~--toText~~ | | | Removed in v8. Use --to=text. |
| ~~--ocrLanguage~~ | | | Removed in v8. Use --ocrConfig.language. |
| ~~--putNotesAtLast~~ | | | Removed in v8. Notes are attached structurally via node.notes. |
| ~~--outputErrorToConsole~~ | | | Removed in v8. Use --verbose. |
Every removed flag above exits with status 1 and prints its replacement, rather than being accepted
and ignored. An unrecognized or renamed config key (say --ocrConfig.autoTerminateTimeout) is not fatal, but the
CLI always prints the warning naming its replacement, with or without --verbose.
Quick Decision Guide
| Goal | API to use |
|------|-----------|
| Extract text / AST from a file | OfficeParser.parseOffice(file) |
| Convert directly to another format | OfficeConverter.convert(file, 'md') |
| Parse first, then generate | parseOffice() → OfficeGenerator.generate(ast, 'html') |
| Convert on the AST itself (shorthand) | ast.to('md') |
| RAG pipeline chunking | OfficeConverter.convert(file, 'chunks', {...}) |
Library Usage: Parsing
Async/Await
const officeParser = require('officeparser');
const ast = await officeParser.parseOffice('/path/to/file.docx');
console.log(ast.type); // 'docx'
console.log(ast.metadata); // { author, title, created, ... }
console.log(ast.content); // Array of hierarchical nodes
console.log(ast.attachments);// Images/charts (if extractAttachments: true)
console.log(ast.warnings); // Non-fatal issues from parsing phaseTypeScript (named import):
import { OfficeParser } from 'officeparser';
const ast = await OfficeParser.parseOffice('report.docx', {
extractAttachments: true,
ocr: true,
});Callback (Backward Compat)
officeParser.parseOffice('/path/to/file.docx', async function(ast, err) {
if (err) { console.error(err); return; }
console.log((await ast.to('text')).value);
});File Buffers, ArrayBuffers & Blobs
Pass a Buffer, ArrayBuffer, Uint8Array, or a web Blob/File instead of a file path:
const fs = require('fs');
const buffer = fs.readFileSync('/path/to/file.pdf');
const ast = await officeParser.parseOffice(buffer);In the browser you can hand a File/Blob straight from an <input type="file">, with no need to
read it into a buffer first. A File's name drives type detection, so no fileType hint is
needed when the name has a recognizable extension:
// input.files[0] is a File (e.g. "report.docx")
const ast = await officeParser.parseOffice(input.files[0]);[!IMPORTANT] Text-based formats from buffers need a
fileTypehint. Formats likemd,html,csvandtexhave no magic bytes, so the parser cannot auto-detect them from a buffer. You must providefileTypein that case:const ast = await officeParser.parseOffice(markdownBuffer, { fileType: 'md' });
[!NOTE] ZIP-backed formats are identified from inside the archive. DOCX, XLSX, PPTX, ODT, ODS, ODP and EPUB are all ZIP files, and telling them apart from the first bytes alone is unreliable for archives written by streaming producers or holding very many parts. When the byte signature is inconclusive, the archive is opened and the format is read from its own declaration (
[Content_Types].xml, or themimetypeentry), so these parse from a buffer without a hint. A LaTeX project zip is recognized the same way, by a.texfile with a\documentclassnear the archive root, and so is a file named.zip: that extension names no format, so the archive's contents decide which parser runs. SupplyingfileTyperemains the fastest and most certain route: it decides which parser runs, and for these formats no archive inspection is done at all.
Cancellation with AbortSignal
You can pass a standard AbortSignal (e.g. from an AbortController) to cancel an active parse operation. This is especially useful for setting request-level timeouts or canceling long-running parses (like large PDFs with OCR).
const controller = new AbortController();
// Cancel parsing if it takes longer than 5 seconds
setTimeout(() => controller.abort(), 5000);
try {
const ast = await officeParser.parseOffice('large_scanned_file.pdf', {
abortSignal: controller.signal,
ocr: true,
extractAttachments: true // page-image OCR needs this; ocr alone does nothing
});
} catch (err) {
if (err.name === 'AbortError') {
console.log('Parsing was cancelled.');
} else {
console.error('Parsing failed:', err);
}
}The signal is checked between steps, and it stops work that waits: OCR, PDF pages, reading an
archive. Reading a document's own markup (its XML, HTML, Markdown or LaTeX) is synchronous work, and
a timer cannot fire while it runs, so an abort requested during it takes effect only once it ends.
To bound the time an untrusted document can take, parse it in a worker thread or child process and
end that when your time limit passes; the signal alone cannot. PDFs are the exception in Node: pdf.js
runs in a separate process (pdfParserConfig.separateProcess), which the signal ends mid-stream, and
pdf.js in the host stops at its next time slice.
[!IMPORTANT] AbortError Propagation When parsing is cancelled via
AbortSignal, the parser rejects with a standardAbortError(aDOMExceptionor an Error withname: 'AbortError'). This error is not wrapped in standard OfficeParser error types so that you can reliably detect cancellation usingerror.name === 'AbortError'.
[!NOTE] Cancellation is cooperative The signal is checked between steps: an already-aborted signal rejects before any work, and an abort is seen at the next check (between archive reads, pages, OCR jobs, or batches of parsed tokens). A step that runs synchronously, such as parsing a
.texfile or one large XML part, finishes before a timer's abort can run; those steps are bounded in size instead.
[!NOTE] Worker Cleanup on Abort If an OCR job is actively running in the background when the signal is aborted,
officeParserautomatically terminates the Tesseract worker process immediately and removes it from the pool to prevent thread/memory leaks.
Custom OCR Timeouts
To prevent the parser from hanging indefinitely due to slow network connections (when downloading Tesseract language datasets) or complex image processing, you can configure granular timeouts under ocrConfig.timeout.
const ast = await officeParser.parseOffice('scanned_document.pdf', {
ocr: true,
extractAttachments: true, // required: OCR of a PDF's page images runs through the attachment path
ocrConfig: {
timeout: {
workerLoad: 30000, // 30s max to load worker & download language training files
recognition: 15000, // 15s max per image text recognition
autoTerminate: 10000 // 10s of inactivity before terminating idle workers
}
}
});[!TIP] Non-Fatal Timeout Recovery If
workerLoadorrecognitiontimeouts are exceeded, the parser will log a warning inast.warningsand continue parsing the rest of the document. The overall promise resolves successfully with the text extracted from the document layers (rather than failing the entire parse).
OCR Layout Reconstruction
By default (ocrConfig.preserveLayout: true) the recognized text keeps its two-dimensional page layout, rebuilt from Tesseract's per-word bounding boxes: a scanned table, form or multi-column page keeps its columns (right-hand text stays on the right, labels and values line up) instead of collapsing to a flat reading-order string. It is the OCR analogue of textConfig.preserveLayout for born-digital PDFs. Set it false for the plain, linearized text.
const ast = await officeParser.parseOffice('scanned_invoice.pdf', {
ocr: true,
extractAttachments: true, // required: OCR of a PDF's page images runs through the attachment path
ocrConfig: { preserveLayout: true } // default; false = flat reading-order text
});ast.to(): Generate from AST
The preferred way to convert a parsed AST to another format. Returns a ConversionResult.
// ConversionResult shape:
// { value: string | Uint8Array | OfficeChunk[], messages: OfficeIssue[] }
const { value: markdown, messages } = await ast.to('md');
const { value: html } = await ast.to('html', { includeFormatting: false });
const { value: chunks } = await ast.to('chunks', { chunksConfig: { strategy: 'fixed-size', chunkSize: 800 } });
const { value: pdfBytes } = await ast.to('pdf'); // Uint8Array.to('text'): Plain Text Extraction
Plain text comes from .to('text'), which is asynchronous and configurable. Its defaults render
tables as aligned grids, lists with markers and indentation, and include notes and image
placeholders:
// Default: aligned tables, list markers, notes, image placeholders, layout-faithful PDF pages
const { value } = await ast.to('text');
// Flat stream of text, no grid alignment or markers
const { value } = await ast.to('text', {
includeImages: false,
textConfig: { preserveLayout: false, renderNotes: false },
});| Feature | default | flat (preserveLayout: false) | governed by |
|---|---|---|---|
| Tables | aligned grid | one cell per line, tab-separated | textConfig.preserveLayout (default true) |
| Lists | markers + indentation | plain text | textConfig.preserveLayout (default true) |
| PDF pages | spatial monospace grid (columns/tables aligned like the page) | flowing text | textConfig.preserveLayout + geometry |
| Footnotes/endnotes | emitted | emitted | textConfig.renderNotes (default true) |
| Image placeholders | emitted | emitted | includeImages (default true) |
For PDFs with page geometry (the default, unless ignorePageGeometry is set), preserveLayout renders
each page as a spatial monospace grid so multi-column text and tables line up much like the original
page, similar to pdftotext -layout. Use textConfig.pageSeparator (default '\n', or '\f' for a
form feed) to control what goes between pages.
Spreadsheets (CSV/ODS/XLSX) are unaffected by preserveLayout: it governs table/list nodes,
while spreadsheet content is sheet/row/cell. There the default aligned grid is the most
faithful rendering.
[!NOTE] The synchronous
ast.toText()method was removed in v8. Use(await ast.to('text')).value, which produces the same content at its defaults and adds the configuration above.
OfficeGenerator
Use OfficeGenerator.generate(ast, format, config?) when you need to produce output from an already-parsed AST:
import { OfficeParser, OfficeGenerator } from 'officeparser';
const ast = await OfficeParser.parseOffice('report.docx');
// Convert to Markdown
const { value: md } = await OfficeGenerator.generate(ast, 'md');
// Convert to HTML with style mapping
const { value: html } = await OfficeGenerator.generate(ast, 'html', {
includeFormatting: true,
styleMap: [
{
selector: { nodeType: 'paragraph', attributes: { style: 'Heading 1' } },
output: { tag: 'h1', classes: ['main-title'] }
}
]
});
// Convert to CSV (spreadsheets)
const { value: csv } = await OfficeGenerator.generate(ast, 'csv');Supported destinations: 'text' · 'md' · 'html' · 'csv' · 'rtf' · 'pdf' · 'docx' · 'odt' · 'tex' (alias 'latex') · 'epub' · 'chunks'
[!NOTE] PDF generation uses a headless browser by default (
pdfConfig.engine: 'html'), which needs the optionalpuppeteerpeer dependency:npm install puppeteerOr choose
pdfConfig.engine: 'native'to lay the document out directly withpdf-lib(npm install pdf-lib): no browser, and the only engine that produces a real PDF in the browser (import fromofficeparser/browser-native-pdffor the client-side path). See PdfGeneratorConfig.EPUB generation with images requires
extractAttachments: trueon the parse step that produced the AST. See EPUB Support.
OfficeConverter: One-Step API
OfficeConverter.convert() combines parsing and generation in a single call. It automatically syncs parser options from the generator config: unless you set parseConfig.extractAttachments explicitly, it is enabled when the output will render images or charts, or when you enable parseConfig.ocr. An explicit parseConfig.extractAttachments (including false) always wins, and parseConfig.ocr is honored (so { parseConfig: { ocr: true } } produces OCR text through the converter, given an image-or-OCR output mode).
import { OfficeConverter } from 'officeparser';
// Minimal usage
const { value: markdown } = await OfficeConverter.convert('report.docx', 'md');
// With config
const { value: html, messages } = await OfficeConverter.convert('data.xlsx', 'html', {
parseConfig: {
ignoreNotes: true,
newlineDelimiter: '\n\n',
},
generatorConfig: {
includeFormatting: true,
styleMap: [
{
selector: { attributes: { style: { value: 'Header', operator: '~=' } } },
output: { tag: 'h2', classes: ['data-header'] }
}
]
},
onWarning: (issue) => console.warn(`[${issue.code}] ${issue.message}`)
});[!IMPORTANT] The
OfficeConverterConfigshape uses nestedparseConfigandgeneratorConfigsub-objects. Do not put parser or generator options at the top level; onlyonWarninglives there. An option placed there has no effect, and is reported asUNRECOGNIZED_CONFIG_OPTIONnaming where it belongs (texConfigundergeneratorConfig,ocrunderparseConfig).
OfficeTemplate: Mail-Merge / Document Generation
OfficeTemplate.render() (alias renderTemplate) fills a DOCX template's {{placeholder}} tags from your data and returns a new .docx. It is not parsing or conversion: the template is copied and only the placeholders are substituted, so all of the template's formatting, layout and structure are preserved. Give it one data object for one document, or an array for a batch (one document per entry, a classic mail-merge). Think of it as a zero-dependency take on Adobe's Document Generation API.
import { OfficeTemplate } from 'officeparser';
import { writeFileSync } from 'fs';
// One document.
const bytes = await OfficeTemplate.render('invoice-template.docx', {
data: { name: 'Acme Corp', amount: '$1,250.00', due: '2026-10-01' },
});
writeFileSync('invoice-acme.docx', bytes); // Uint8Array
// A batch: one .docx per row.
const docs = await OfficeTemplate.render('invoice-template.docx', {
data: [
{ name: 'Acme Corp', amount: '$1,250.00' },
{ name: 'Globex', amount: '$980.00' },
],
});
docs.forEach((d, i) => writeFileSync(`invoice-${i}.docx`, d));- Run-aware. Word often splits a typed
{{name}}across several runs ({{,na,me}}); it is filled anyway, and a value takes the formatting of the run its placeholder sat in (a bold{{amount}}renders bold). - Everywhere text lives. Placeholders in the body, headers, footers, footnotes/endnotes and comments are all filled. Values may contain
\n(rendered as line breaks). - Placeholder names are Unicode letters and digits plus
_,.,-(e.g.{{invoice.total}},{{customer-name}}). A name containing a space or other punctuation is not recognized and is left as literal text, so surrounding prose between the delimiters is never mistaken for a field. - Deterministic output (pinned zip timestamps): the same template + data always renders byte-identical bytes.
| Option | Type | Default | Description |
|--------|------|---------|-------------|
| data | TemplateData \| TemplateData[] | (required) | Field values. One object is one document; an array is one document per entry |
| delimiters | { start: string; end: string } | {{ }} | Placeholder delimiters |
| onMissing | 'keep' \| 'empty' \| 'error' | 'keep' | A placeholder with no matching field: leave it, blank it, or reject with TEMPLATE_FIELD_MISSING. A field present but null/undefined always renders empty |
| password | string | (none) | Decrypt the template first, if it is itself password-protected |
Only DOCX is supported today (other OOXML/ODF formats will follow); a non-DOCX template rejects with TEMPLATE_UNSUPPORTED_FORMAT.
Native RAG Chunking
officeParser provides native document chunking for Retrieval-Augmented Generation (RAG) pipelines with three strategies:
Strategy 1: Document Structure (Default)
Splits at natural AST boundaries (paragraphs, headings, pages, slides, sheets). Preserves logical flow.
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
generatorConfig: {
chunksConfig: {
strategy: 'document-structure',
splitBy: 'heading', // 'paragraph' | 'heading' | 'page' | 'slide' | 'sheet'
maxChunkSize: 1500,
tableSplitStrategy: 'row', // repeats header row in every chunk, ideal for RAG
}
}
});Strategy 2: Fixed-Size (Recursive)
Splits by character count with overlap. Equivalent to LangChain's RecursiveCharacterTextSplitter.
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
generatorConfig: {
chunksConfig: {
strategy: 'fixed-size',
chunkSize: 1000,
chunkOverlap: 200,
}
}
});
console.log(`Generated ${chunks.length} chunks`);Strategy 3: Semantic
Uses cosine similarity between sentence embeddings to find topic boundaries. Requires you to provide an embeddingFunction.
import OpenAI from 'openai';
const openai = new OpenAI();
const { value: chunks } = await OfficeConverter.convert('report.docx', 'chunks', {
generatorConfig: {
chunksConfig: {
strategy: 'semantic',
embeddingFunction: async (text) => {
const res = await openai.embeddings.create({
input: text, model: 'text-embedding-3-small'
});
return res.data[0].embedding;
},
similarityThreshold: 0.8,
maxChunkSize: 2000,
}
}
});The OfficeChunk Object
generate(ast, 'chunks') (and ast.to('chunks')) resolves to a real OfficeChunk[] array, not a JSON string - serialize it to JSON/JSONL yourself if your pipeline needs that.
Every chunk contains text and rich metadata for citations and filtered retrieval:
interface OfficeChunk {
text: string;
/** Rich metadata for filtered retrieval */
metadata: {
sourceType: string; // e.g., 'docx', 'pdf'
pageNumber?: number; // (PDF only)
slideNumber?: number; // (PPTX only)
sheetName?: string; // (XLSX only)
closestHeading?: string; // Nearest heading above this chunk
isTableChunk?: boolean; // True if part of a split table
};
startIndex?: number; // Character offset (if addStartIndex: true)
endIndex?: number; // End character offset (if addStartIndex: true)
}The AST Structure
OfficeParserAST is a format-agnostic document representation:
OfficeParserAST
├── type: 'docx' | 'pdf' | 'xlsx' | 'csv' | 'md' | 'epub' | 'tex' | ... (14 formats)
├── metadata: { author, title, created, modified, keywords, customProperties, nativeProperties, styleMap, ... }
├── content: [ OfficeContentNode ]
│ ├── type: 'paragraph' | 'heading' | 'table' | 'list' | 'image' | 'chart' | 'comment' | 'admonition' | 'embed' | 'definitionList' | ...
│ ├── text: string (concatenated text of node + all descendants)
│ ├── children: [ OfficeContentNode ] (recursive structural children)
│ ├── notes: [ OfficeContentNode ] (footnotes/endnotes/slide notes attached to this node; see below)
│ ├── comments: [ OfficeContentNode ] (inline comments attached to this node)
│ ├── formatting: { bold, italic, underline, color, size, font, alignment, ... }
│ └── metadata: { level, listId, row, col, rowSpan, colSpan, backgroundColor, style, ... }
├── auxiliary?: OfficeAuxiliaryContent (out-of-band layout elements)
│ ├── headers?: OfficeContentNode[] (DOCX, PDF top band, ODT master pages, LaTeX fancyhdr)
│ ├── footers?: OfficeContentNode[] (DOCX, PDF bottom band, ODT master pages, LaTeX fancyhdr)
│ ├── slideMasters?: OfficeContentNode[] (PPTX slide masters)
│ └── outline?: OfficeContentNode[] (PDF bookmark outline)
├── attachments: [ OfficeAttachment ] (populated when extractAttachments: true)
│ ├── type: 'image' | 'chart'
│ ├── name: string
│ ├── mimeType: string
│ ├── data: string (Base64)
│ ├── ocrText?: string (if ocr: true AND extractAttachments: true)
│ └── chartData?: { title, dataSets, labels }
├── warnings: OfficeIssue[] (non-fatal issues from the parsing phase)
├── config: OfficeParserConfig (the resolved parse config; `.to()` inherits newlineDelimiter/onWarning from it)
└── to(format, config?) (format: 'html'|'md'|'text'|'csv'|'rtf'|'pdf'|'docx'|'odt'|'tex'|'epub'|'chunks', returns { value, messages })A note or comment referred to from several places is one node that each reference's notes (or comments) array holds. Test for a note you have already seen by identity before handling it again, and do not mutate it expecting only one reference to change. (JSON.stringify writes it at every reference.)
OfficeIssue: Warning / Error Object
All warnings and errors (from both parsing and generation) use this shape:
interface OfficeIssue {
type: 'warning' | 'info' | 'error';
code: OfficeWarningType | OfficeErrorType; // typed enum, e.g. 'OCR_FAILED'
message: string;
node?: OfficeContentNode; // the node that triggered the issue, if any
details?: any; // original error or extra context
}Thrown errors carry the same object on error.officeIssue, so a failed parse is identified by
the same stable code you would branch on for a warning, rather than by matching message text:
try {
const ast = await officeParser.parseOffice(buffer, { fileType: 'docx' });
} catch (err) {
switch (err.officeIssue?.code) {
case 'ZIP_NO_ENTRIES_FOUND': // not a ZIP archive at all
case 'ZIP_TRUNCATED': // cut off in transfer, entries incomplete
case 'REQUIRED_PART_MISSING': // readable ZIP, but not the format it claims
console.error('Unusable file:', err.officeIssue.message);
break;
default:
throw err;
}
}[!IMPORTANT] A corrupt file throws; it does not parse as an empty document. If an archive is not readable, is truncated, or is missing the part its format requires (
word/document.xml,xl/workbook.xml,ppt/presentation.xml, ODFcontent.xml, the EPUB OPF), parsing rejects with one of the codes above. An empty result therefore means the document really is empty. Files that are legitimately empty still parse, and say so throughonWarning/ast.warnings(NO_WORKSHEETS_FOUNDfor a chartsheet-only workbook,NO_SLIDES_FOUNDfor a presentation with no slides).
Warning codes (type: 'warning' | 'info', delivered to onWarning and collected in ast.warnings)
These never throw; they report a degraded-but-successful outcome you may branch on by code.
| Code | Phase | Meaning / what to do |
|---|---|---|
| OCR_REQUIRES_ATTACHMENTS | parse | ocr: true without extractAttachments: true; no OCR ran. Set both. |
| PDF_NO_TEXT_EXTRACTED | parse | A PDF yielded ~no text (likely scanned). Set ocr: true + extractAttachments: true. |
| PDF_STRUCT_TREE_UNRELIABLE | parse | PDF tag tree absent/incomplete; structure recovered geometrically. |
| PDF_TEXT_ENCODING_SUSPECT | parse | A fifth or more of a PDF's characters (of at least 50) are unmappable glyphs (broken ToUnicode); text may be garbage. Consider OCR. |
| PDF_OUTLINE_TRUNCATED | parse | Bookmark outline hit the depth/size cap; ast.auxiliary.outline is partial. |
| PDF_WORKER_MISSING / PDF_WORKER_FALLBACK | parse | The pdf.js worker could not be loaded / a fallback was used (set pdfWorkerSrc). |
| NO_WORKSHEETS_FOUND / NO_SLIDES_FOUND | parse | A legitimately empty workbook/presentation. |
| TABLE_CELL_LIMIT_EXCEEDED | parse | The cells an ODF or XLSX document yields passed decompressionLimits.maxTableCells plus one per byte of the document; the rest were not read. |
| REPEATED_CONTENT_LIMIT_EXCEEDED | parse | The document repeated decompressionLimits.maxRepeatedContent (plus 16 characters per byte of the document) of content by reference (ODF repeated cells and chart values, XLSX shared strings, style values, link targets, chart text per frame, LaTeX titles per reference); later repeats were not made, shortened or went without the value. |
| PDF_SEPARATE_PROCESS_UNAVAILABLE | parse | pdf.js could not start in a separate process (Node), so it runs in the host; a hostile PDF can then exhaust its memory, and pdf.js's time is not bounded (bound it with abortSignal). |
| PDF_CONTENT_LIMIT_EXCEEDED | parse | A PDF passed one of its content limits (plus an allowance per byte), which the message names: past pdfParserConfig.maxTextItems or maxTimeMs the rest of it was not read; past maxOperators its images, text colours and font styles from that page on were not (its text was). |
| ALT_CHUNK_NOT_READ | parse | A DOCX alternative-format chunk (w:altChunk) was not read: its part is missing, is of a format other than HTML, MHT, RTF, plain text or DOCX, is a DOCX inside a DOCX chunk, or could not be read (not a ZIP, no document part, nested too deep). The rest of the document is read; a chunk past the document's limits (maxXmlElements, maxUncompressedBytes) still fails the parse. Saving the document again in Word merges chunks into it. |
| CONTENT_PART_NOT_READ | parse | A part of the document was not read. A chapter an EPUB's spine lists: its file is missing from the archive (or has an extension other than .xhtml, .html, .htm, .xht or .xml), or it is encrypted (DRM, listed in META-INF/encryption.xml); a book whose chapters are all encrypted throws DOCUMENT_DECRYPTION_FAILED instead. In a PPTX, a part that is not XML and holds none of the slides' text: the presentation's slide list or relationships (the slides are then read in the order of their file numbers), or the relationships of a notes page or a slide master (its links and pictures are then not resolved). |
| RAW_CONTENT_LIMIT_EXCEEDED | parse | With includeRawContent, the document's nodes reached decompressionLimits.maxRawContentLength of raw content; the remaining nodes carry none. |
| IMAGE_EXTRACTION_FAILED / IMAGE_PROCESSING_FAILED / ATTACHMENT_EXTRACTION_FAILED | parse | An image/attachment could not be extracted or decoded; it was skipped or degraded. |
| ANNOTATION_EXTRACTION_FAILED / CHART_DATA_EXTRACTION_FAILED | parse | A PDF annotation / a chart's data could not be read. |
| OCR_FAILED | parse | OCR ran but failed for an image (see details). |
| LATEX_CONSTRUCT_NOT_INTERPRETED | parse | The LaTeX input used commands or environments the parser does not interpret (the message names them). Text inside them was kept; drawings such as TikZ pictures were omitted. |
| LATEX_EXPANSION_LIMIT_REACHED | parse | A LaTeX document hit a bound on macro expansion, file inclusion or nesting depth (the guard against expansion bombs, include cycles and runaway nesting); macros or files past it were not expanded, and content nested past it is kept as plain text. |
| LATEX_FILE_NOT_FOUND | parse | The LaTeX input includes files or images the parser could not read (a .tex holds only the files it carries in filecontents blocks). Parse the project as a .zip to include them; images are kept as path references. |
| FILE_TYPE_DETECTION_FAILED / BUFFER_TYPE_MISMATCH | parse | Type could not be sniffed / disagreed with the fileType hint. |
| PASSWORD_REQUIRED / PASSWORD_INCORRECT | parse | Encrypted input; supply password/onPassword (these also throw when parsing cannot continue). |
| UNRECOGNIZED_CONFIG_OPTION | config | A config key this version does not know (often a typo or a removed/renamed option); it had no effect. Raised for parser and generator configs, and for a convert() option placed at the top level instead of under parseConfig/generatorConfig (the message says where it belongs). |
| INVALID_CONFIG_VALUE | config | A generator option got a value it does not accept (a documentClass, paper format or margin, pdfConfig.engine, htmlConfig.standalone.styles, a Markdown dialect preset, a chunking strategy/splitBy/tableSplitStrategy) or texParserConfig.today is not a string; the option's default was used, and the message names the option, the value and what it accepts. |
| CONTENT_NOT_REPRESENTABLE | generate | A node type has no faithful form in the target format and was downgraded or omitted (e.g. math/embeds in DOCX/ODT, a table-less document to CSV). |
| METADATA_NOT_REPRESENTABLE | generate | A metadata field could not be represented in the target format. |
| TABLE_GRID_LIMIT_EXCEEDED | generate | The document's tables needed more empty grid positions than one output fills (a million, plus 16 per byte of the document): a table's cells were laid out closer, short rows were not padded to the table's width, or a sparse sheet's empty rows were not written; or plain text stopped lining columns up (16 million spaces, plus 16 per byte). |
| IMAGE_NOT_INLINED | generate | An image was referenced by name instead of inlined: it is over maxInlineImageBytes (Markdown / fragment HTML), or the pictures inlined in the document reached 128 MB in all (Markdown, HTML of either kind, RTF; a picture is inlined at every place that shows it). |
| IMAGES_NOT_BUNDLED | generate | LaTeX output references image files the .tex does not carry (with texConfig.embedImages: false, an image other than a readable PNG or JPEG, or one past the decoding limits); the message names them. Ship them alongside, or set texConfig.bundle: true. It also names images the source referred to only by a relative path, with no image data (a .tex without its project, an HTML page's <img src="pics/a.png">): supply those at that path yourself, since not even a bundle can contain them. |
| MATH_WRITTEN_AS_TEXT | generate | A math expression used an unsafe LaTeX command (file access, shell, redefinition) or was malformed, so LaTeX output shows it as literal text instead of typesetting it. |
| CITATIONS_NOT_RESOLVED | generate | LaTeX output cites keys (\cite{key}) that have no entry in a bibliography it holds, so LaTeX prints [?] for them until one is added (a \bibliography{file} with a .bib file, or a thebibliography list); the message names the keys. |
| PDF_GENERATION_FAILED | generate | PDF generation failed (e.g. Puppeteer missing for engine: 'html'). |
| INVALID_STYLE_MAPPING / INVALID_STYLE_MAP_TAG | generate | A styleMap entry/tag was invalid and ignored. |
| TEMPLATE_UNSUPPORTED_FORMAT / TEMPLATE_FIELD_MISSING | template | The template format is unsupported / a {{field}} had no value under onMissing: 'error'. |
| PAGE_LOAD_FAILED | parse | A PDF page could not be processed and was skipped (partial content). |
| SHEET_RANGE_NOT_FOUND | generate | A csvConfig.sheets range matched no sheet, so CSV output is empty. |
| EMPTY_CHUNK_GENERATED / WHITESPACE_NODE_SKIPPED / BROWSER_GENERATION_LIMITATION / PERFORMANCE_TIP / DEPENDENCY_LOAD_FAILED | generate | Diagnostic/informational notes from the chunking and PDF generators. |
Generating throws OUTPUT_TOO_LARGE when the output would grow past what a string or array holds, or when an AST built in code shares its nodes, records, lists or long strings along more paths than a writer follows (a shared value is written at every use, so a few KB of AST could ask for gigabytes). Parsed documents stay inside the bound: a record many nodes share (a spreadsheet's style, given to every cell) weighs at each later holder only what it weighs past what a node's own record may, and the bound grows with what the AST holds and with the document it was parsed from (its own config.decompressionLimits.maxRepeatedContent, up to 1 GiB, plus 16 characters per byte of that document). An AST nested past what the stack holds throws MAX_NESTING_DEPTH_EXCEEDED.
The full enum lives in OfficeWarningType / OfficeErrorType (src/types.ts); the error codes used in the catch above are the OfficeErrorType members.
Per-Format Capability Matrix
What each parser extracts differs by format, because the source formats themselves differ. This table
is the authoritative reference; the option docs point back to it. Y = extracted by default (subject to
the relevant ignore*/extractAttachments flag), – = the format has no such construct or it is not
extracted (the matching ignore* flag is then a no-op).
| Input | Comments (node.comments) | Notes (node.notes) | Headers/footers (ast.auxiliary) | Slide masters | Images (needs extractAttachments) | Charts | Tables (colSpan/rowSpan) |
|---|---|---|---|---|---|---|---|
| DOCX | Y | footnotes/endnotes | Y | – | Y | Y | Y |
| XLSX | Y | – | – | – | Y | Y | grid |
| PPTX | Y | speaker notes | – | Y | Y | Y | Y |
| ODT | Y (in text) | footnotes/endnotes | Y (master pages) | – | Y | Y | Y |
| ODS | Y (cell notes) | – | – | – | Y | Y | grid |
| ODP | Y (page) | speaker notes | – | – (ODP masters not extracted) | Y | Y | Y |
| ODG | Y (page) | – | – | – | Y | – | Y |
| PDF | – | footnotes/endnotes (tagged) | Y (top/bottom bands) | – | Y | – | Y (spans: tagged only) |
| RTF | Y (annotations) | footnotes/endnotes | Y | – | Y | – | Y |
| HTML | <!-- --> become comment nodes* (opt-in: preserveComments) | footnotes/endnotes | – | – | Y (data: only) | – | Y |
| MD | <!-- --> become comment nodes* | footnotes/endnotes | – | – | Y (data: only) | – | Y (HTML-table fallback) |
| CSV | #-rows become comment nodes* | – | – | – | – | – | rows |
| EPUB | <!-- --> become comment nodes* (opt-in: preserveComments) | footnotes/endnotes | – | – | Y | – | Y |
| TEX | Y (% Comment (Author, date): lines); % <!-- --> lines become comment nodes* | footnotes/endnotes | Y (fancyhdr) | – | Y (from a project zip) | – | Y |
Notes: comments land on node.comments[] (with author/date, and a PowerPoint reply's parentId) except source-level comments, which
are not governed by ignoreComments: CSV's leading-# rows become top-level comment nodes, and
Markdown/HTML <!-- ... --> (and LaTeX % <!-- ... --> lines) become comment nodes (block or
inline) marked metadata.sourceSyntax: 'html' whose text is the raw comment body. The Markdown,
HTML and LaTeX generators keep those as comments; every other output format omits them (a hidden note
stays hidden). ignoreNotes /
ignoreComments / ignoreHeadersAndFooters / ignoreSlideMasters each remove the corresponding
column and are a no-op wherever it shows –. OCR (ocr: true) recognizes text from any extracted
image and therefore also needs extractAttachments: true.
Content kept in parts of its own is read where it stands: SmartArt text (as a nested bulleted list, in
DOCX after the paragraph drawing it), DOCX content controls, text boxes (their paragraphs, lists and
tables, after the paragraph drawing them) and alternative-format chunks (w:altChunk: HTML, MHT, RTF,
plain text or a DOCX, in the body, headers, footers, notes and comments), and RTF shape text boxes (as
blocks right after the paragraph the shape is anchored in). Text a tracked
change deleted or moved away (DOCX, ODT) is not read; insertions are.
A PPTX's slides are read in the order its slide list (p:sldIdLst) shows them, not the order of their
file names, and slideNumber is a slide's place in that order. Each slide has the notes page its own
relationships name. A slide part the list no longer names (a deleted slide an editor left in the
package) is not read. A slide list or presentation relationships that are not XML leave the slides in
the order of their file numbers, and relationships of a notes page or a slide master that are not XML
leave its links and pictures unresolved; each is reported with a CONTENT_PART_NOT_READ warning.
Deep Dive: Document Components
1. Lists
List Node
├── type: 'list'
├── metadata: {
│ listId: '1', // items with the same listId belong to one logical list
│ listType: 'ordered' | 'unordered',
│ indentation: 0, // nesting level (0-based)
│ itemIndex: 0, // sequential position within the list level
│ paragraphIndentation: { left, hanging, right, firstLine }
│ }
└── children: [ Text content ][!TIP] Even if a list is interrupted by a regular paragraph,
itemIndexkeeps incrementing for the samelistId, so numbering stays correct.
2. Tables
Tables follow a strict table → row → cell hierarchy:
Table Node (type: 'table')
└── children: Row Nodes (type: 'row')
└── children: Cell Nodes (type: 'cell')
├── metadata: { row, col, rowSpan?, colSpan? }
└── children: [ Paragraph | List | Table | ... ]row/col: zero-based grid position (a cell without them takes the next place in reading order). Generators fill the grid between cells, so the grids of a document's tables and sheets may hold 1,000,000 empty positions in all, plus 16 per byte of the document parsed (spans count); one holding more than remain (a cell far from the rest) is laid out closer when written, with aTABLE_GRID_LIMIT_EXCEEDEDwarning: the rows and columns no cell starts or ends in are left out, and if that is not enough, each row's cells follow one anotherrowSpan/colSpan: merged cells (DOCX, ODF, HTML, Markdown HTML-tables, and tagged PDF)- Cells can contain nested tables
[!NOTE] Header rows. A header row is flagged on its cells' metadata (
style: 'header', orisHeader), set by the parsers that mark one (DOCXw:tblHeader, ODFtable:table-header-rows, HTML<th>/<thead>, tagged-PDFTH); generators read it through one shared heuristic (a marked row, or an all-bold first row). One format-imposed asymmetry: a Markdown table always renders a header row (the GFM| --- |separator is mandatory syntax), whereas HTML emits<thead>only for a detected header. So a table with no real header prints a header in Markdown output but not in HTML.
3. Images & OCR
Image Node (type: 'image')
├── metadata: { attachmentName: 'img1.png', altText: '...' }
└── → Attachment: { data: 'base64...', ocrText: '...' }- Set
extractAttachments: trueto populateattachment.data - Set
ocr: true(requiresextractAttachments: true) to populateocrText
4. Charts
Chart Node (type: 'chart')
├── metadata: { attachmentName: 'chart1.xml' }
└── → Attachment: { chartData: { title, dataSets, labels } }5. Text Formatting
formatting: {
bold?: boolean
italic?: boolean
underline?: boolean
strikethrough?: boolean
color?: string // '#RRGGBB'
backgroundColor?: string
size?: string // e.g. '12pt'
font?: string
subscript?: boolean
superscript?: boolean
alignment?: 'left' | 'center' | 'right' | 'justify'
}[!NOTE] On a content node, an absent flag and
falsemean the same thing: the flag is simply not applied. Onast.metadata.styleMap, they differ: an absent flag means the style says nothing about that property (so it inherits), whilefalsemeans the style explicitly turns it off (ODF'sfo:font-weight="normal", DOCX's<w:b w:val="0"/>). Code resolving inheritance itself must test=== undefined, not truthiness, or it will treat "explicitly off" as "unspecified".
6. Break Nodes (DOCX, ODF and LaTeX)
When includeBreakNodes: true, break elements appear as nodes:
Break Node (type: 'break')
└── metadata: {
breakType: 'textWrapping' | 'page' | 'column' | 'lastRenderedPage' | 'carriageReturn' | 'thematic',
clear?: 'all' | 'left' | 'none' | 'right'
}[!NOTE] Break nodes have no
textproperty, butast.to('text')automatically converts them to the configured newline delimiter.
[!NOTE]
includeBreakNodesgates layout breaks in DOCX, ODF and LaTeX only (page and column breaks, otherwise invisible). A line break the author typed (DOCXw:br/w:cr, ODFtext:line-break, LaTeX\\) is content, and is abreaknode whatever the flag says. HTML and Markdown always emit break nodes regardless of the flag, because a break is explicit content there: a<br>/hard line break becomes acarriageReturnbreak, and<hr>/---athematicbreak (a Markdown\f-style page break maps topage).
[!NOTE] DOCX writes breaks inline (
w:br/w:cr), so they land as children of the paragraph. ODF instead carries page and column breaks on the paragraph style (fo:break-before/fo:break-after), so those are emitted as siblings around the paragraph rather than inside it.<text:soft-page-break/>maps ontolastRenderedPage, the same type as DOCX'sw:lastRenderedPageBreak.
6b. Equations
Equations are extracted from every format that can carry them and normalized to LaTeX, so a formula means the same thing whichever format it arrived in:
| Source format | Markup in the file |
|---|---|
| DOCX, PPTX | OOXML <m:oMath> / <m:oMathPara> |
| ODT, ODP, ODS | MathML inside the embedded formula object |
| HTML, EPUB | native MathML <math> |
| Markdown | $inline$ / $$block$$ |
They all land as the same node:
Code Node (type: 'code')
├── text: '\\frac{1}{2}' // LaTeX, whatever the source markup was
└── metadata: { math: 'inline' | 'block' }Fractions, sub/superscripts, radicals, delimiters, n-ary operators (sums, integrals), named
functions, accents, bars, matrices and math alphabets (ℝ, 𝒜, …) are all preserved. When a
document supplies its own TeX source in an <annotation encoding="application/x-tex">, that is
used verbatim in preference to anything reconstructed from the presentation markup.
[!NOTE] Equation text is structure, not prose: a fraction whose numerator and denominator are simply concatenated reads as a different number rather than as obviously-missing content. Consumers that index document text should treat
codenodes carryingmathas opaque LaTeX rather than splitting them as words.
On generation, an equation's fate depends on the target: LaTeX output typesets it as real math
($…$, \[…\], or a bare align-style environment), after a safety check that refuses any command
reaching outside the formula (see TexGeneratorConfig); HTML and Markdown keep it
as LaTeX (a $…$/$$…$$ delimited block or a data-math attribute); DOCX and ODT downgrade it to its
LaTeX text and emit a CONTENT_NOT_REPRESENTABLE warning (no native OMML/ODF-math is written); plain
text, RTF and the PDF engines render the LaTeX string as-is without a warning. So the LaTeX always
survives, and LaTeX, HTML and Markdown keep it as math.
7. Document Metadata
ast.metadata = {
author?: string
title?: string
created?: Date
modified?: Date
description?: string
keywords?: string // NEW: Keywords from document properties
customProperties?: Record<string, any> // User-defined metadata from the document
nativeProperties?: Record<string, any> // NEW: All format-specific raw metadata
styleMap?: Record<string, TextFormatting> // Named styles → formatting definitions
formatting?: TextFormatting // Document-wide defaults
}Accessing native properties (format-specific metadata):
const ast = await officeParser.parseOffice('contract.docx');
console.log(ast.metadata.nativeProperties);
// DOCX: { Pages: 5, Application: 'Microsoft Word' }
// HTML: { description: 'My page', 'og:title': 'Title' }
// PDF: { Title: 'Report', XMP: { ... } }8. Admonitions, Embeds & Definition Lists
Admonition Node (type: 'admonition')
├── metadata: { admonitionType: 'note' | 'tip' | 'important' | 'warning' | 'caution', title?: string }
└── children: [ Paragraph | List | ... ] (block content)
Embed Node (type: 'embed')
└── metadata: { embedType: 'youtube' | 'iframe', videoId?: string, url?: string, width?: string, height?: string, align?: string }
Definition List Node (type: 'definitionList')
└── children:
├── Definition Term (type: 'definitionTerm')
└── Definition Description (type: 'definitionDescription')admonitionround-trips through both Markdown (> [!NOTE]/:::note ... :::/ Pandoc's::: {.note} ... :::) and HTML (<div class="admonition admonition-note" data-type="note">)embedmodels YouTube videos and generic iframes. Markdown form is selected bymdConfig.dialect.embeds:'html'(default; the<div data-youtube-video>/<iframe>block),'directive'(a::youtube[…]{…}/::embed[…]{…}leaf directive),'link', or'thumbnail'(YouTube-only clickable preview). A generic iframe is captured only underhtmlParserConfig.preserveIframes(the trust input) and can be emitted as an inert click-to-load placeholder viahtmlConfig.gatedEmbeds. The'directive'form is an editor round-trip format, not GitHub-rendered- Abbreviations (
*[HTML]: Hypertext Markup Language) are stored asTextMetadata.abbreviationTitleon the abbreviated text node rather than as a separate node type
Markdown Dialect Support
Beyond CommonMark/GFM basics, MarkdownParser/MarkdownGenerator support an extended dialect aimed at
full-fidelity round-tripping with rich Markdown editors. Every construct below parses to a first-class
AST node/metadata field and regenerates back to the canonical syntax shown, so .md → AST → .md is
idempotent and .md → AST → HTML → AST → .md survives unchanged.
Markdown-input parsing options that are not dialect toggles live on
htmlParserConfig(Markdown shares the HTML parser for embeds):preserveIframesandembedFolkFormsgovern raw<iframe>blocks and folk embed forms encountered in.md. There is no separatemdParserConfig. (preserveCommentsgoverns HTML and EPUB input only: comments in Markdown are always kept, see the table below.)
| Feature | Markdown syntax | AST representation |
|---|---|---|
| Task lists (GFM) | - [x] Done / - [ ] Todo | ListMetadata.isTask / .checked |
| Admonitions | > [!NOTE] (also accepts GLFM :::note ... ::: and Pandoc ::: {.note} ... ::: on import) | type: 'admonition', AdmonitionMetadata |
| Footnotes | Text[^1] + [^1]: Definition | type: 'note', keyed by footnote id |
| Definition lists | Term\n: Definition | type: 'definitionList' / 'definitionTerm' / 'definitionDescription' |
| Abbreviations | *[HTML]: Hypertext Markup Language | TextMetadata.abbreviationTitle |
| HTML comments | <!-- note --> on its own lines (may span lines, blank ones included) or inline in a run | type: 'comment', CommentMetadata.sourceSyntax: 'html', raw body in text; re-emitted byte-for-byte by the Markdown and HTML generators, kept by the LaTeX generator as % <!-- ... --> lines (which the LaTeX parser reads back), omitted by every other generator. <!--> and <!---> are empty comments; write \<!-- for literal text |
| Attribute lists | {width=50% .centered} | ImageMetadata.width / .align, TableMetadata.align |
| Citations | [@smith2024] | TextMetadata.citationKey |
| Wikilinks | [[Page]] / [[Page\|Alias]] | TextMetadata.wikilink, .link, .linkType |
| Highlight | ==text== | TextMetadata.backgroundColor |
| Link/image titles | [text](url "Title") /  | TextMetadata.title / ImageMetadata.title |
| Linked images (badges) | [](url "Link title") | ImageMetadata.link / .linkType / .linkTitle |
| Inline/block math | $E=mc^2$ / $$...$$ | type: 'code', CodeMetadata.math ('inline' \| 'block'). $$...$$ is display math wherever it is written: inside a paragraph, the paragraph is split around it; in a heading, list item, table cell, quote or note, where a block cannot go, it is inline math |
| Embeds | ::youtube[Label]{id=… width=… align=…} / ::embed[Label]{src=… …} (leaf directive; see mdConfig.dialect.embeds) | type: 'embed', EmbedMetadata |
| Frontmatter arrays | tags: [a, b] or tags: ["a","b"] | Real array in metadata.customProperties/nativeProperties |
| MDX components (import-only) | <Component prop="x">...</Component> | Stripped; inner Markdown is kept. Never generated back. |
Text is written so that it reads back as itself, in CommonMark renderers and in the parser alike: a
character that would be markup gets a backslash (\*, \`, \[, \$, and \_ except inside a
word), a block marker gets one only where it starts a line (\# , \- , 1\. ), and the & of a
character reference is written &. The parser reads emphasis, code spans, links and fences as
CommonMark does: ***text*** is bold and italic, an underscore inside a word is text, link text may
hold brackets and a title parentheses, a hard break is two trailing spaces or a backslash, and a fence
may be indented under a list item.
[!NOTE] MDX/JSX stripping is one-directional (parse-only): officeParser never authors JSX back into Markdown, and a component inside code is code, left as written. Wikilink enable/disable and citekey→bibliography resolution are application-level concerns; officeParser always parses/generates the syntax itself.
The same round-trip fidelity extends to HTML, so content saved from a rich-text editor survives a save→reload cycle:
| HTML attribute | AST
