npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@opencraw/office-reader

v0.0.13

Published

Reads Office files (.xlsx workbooks, .pptx presentations, .docx documents) into plain objects: cells with merged ranges and hidden rows, slides with positioned text, tables and chart data, document paragraphs with headings, lists and tables. No Node built

Readme

@opencraw/office-reader

Reads Office files into plain objects:

  • workbooks (.xlsx, .xlsm) as sheets of cells, with their merged ranges and hidden rows;

  • presentations (.pptx, .pptm, .ppsx) as slides of positioned text boxes, tables, chart data and speaker notes;

  • documents (.docx, .docm, .dotx) as paragraphs with their heading and list levels and links, tables with their merged cells, headers, footers and notes.

  • What the file holds, faithfully. Formulas give their cached value, and nothing is evaluated. Dates are dates, in both the 1900 and 1904 systems. Merged ranges and hidden rows and sheets are reported, not guessed away.

  • Runs anywhere. Two small dependencies (fflate and htmlparser2), and no Node built-in except when you pass a path. It works in Node, browsers, workers and edge runtimes.

  • Safe on hostile files. Zip entries are capped by size. XML entities a file declares are never expanded (no "billion laughs"), and no external entity is ever fetched (no XXE). Nothing in the file runs.

It is part of OpenCraw, which uses it to crawl price lists and incentive sheets, but it depends on nothing from it.

Install

npm install @opencraw/office-reader

Read a workbook

import { readXlsx } from '@opencraw/office-reader/xlsx'

const book = await readXlsx('./listino.xlsx')
for (const sheet of book.sheets) {
  console.log(sheet.name, sheet.rows.length)
}
// { date1904: false, sheets: [
//   { name: 'Incentivi giugno', hidden: false,
//     rows: [
//       ['Incentivi concessionari – giugno 2026'],
//       [],
//       ['Marca', 'Modello', 'Prezzo', null, 'Sconto', 'Valido dal', 'Attivo', 'Nota'],
//       [null, null, 'Listino', 'Netto'],
//       ['Fiat', 'Pandina', 15950, 13955.625, 0.125, Date(2026-06-01), true, 'Solo rottamazione'],
//       [null, 'Pandina Cross', 17950, 15706.25, 0.125, Date(2026-06-01T09:30), false, { error: '#DIV/0!' }],
//       …
//     ],
//     hiddenRows: [6],
//     merges: ['A1:H1', 'A3:A4', 'B3:B4', 'C3:D3', 'E3:E4', 'F3:F4', 'A5:A6'] },
//   { name: 'Archivio', hidden: true, … } ] }

Sources

readXlsx(source, options?) takes any of these:

| Source | Notes | |---|---| | a path (string), or a file: URL | Node only. A string is always a path. | | Uint8Array, Buffer, ArrayBuffer, any ArrayBufferView | | | Blob, File | An upload in a browser or a server framework. | | ReadableStream<Uint8Array> | response.body from fetch, for example. | | any AsyncIterable<Uint8Array> | Node streams: fs.createReadStream(path), a request body. |

It never fetches: pass await (await fetch(url)).arrayBuffer() or response.body.

Options

| Option | Default | | |---|---|---| | sheets | all | A name, a RegExp, or ({ name, hidden }) => boolean. Unselected sheets are never inflated. | | values | 'typed' | 'typed' or 'text' (below). | | limits | 256 MiB per entry, 512 MiB per file | { entryBytes, totalBytes }: what the zip entries may declare. A part read twice counts once. |

Values

| Cell | values: 'typed' | values: 'text' | |---|---|---| | text, rich text | 'Solo rottamazione' | same | | number | 13955.625 | '13955.625' (shortest round-trip form: 78.6, not 78.599999999999994) | | percentage | 0.125 (display formats are not applied) | '0.125' | | date / date-time | Date holding the wall-clock time as UTC | '2026-06-01' / '2026-06-01T09:30:00' | | time of day (a serial under 1) | Date on the date system's day zero (1899-12-30T12:00Z, or 1904-01-01T12:00Z in a 1904 workbook) | '12:00:00' | | serial 60 in the 1900 system | '1900-02-29', the day Excel shows, as text (below) | same | | boolean | true | 'true' | | error | { error: '#DIV/0!' } | '#DIV/0!' | | formula | its cached value, as above | same | | empty | null | '' |

Spreadsheets have no time zones, so a Date carries the wall-clock time in its UTC fields. Read it with getUTCHours() or toISOString(), not getHours().

The 1900 date system counts a 29 February 1900 that never was (Lotus 1-2-3's bug, which Excel kept). Serials before it read a day earlier than their count, as Excel shows them, and serial 60 itself, which Excel shows as 29/02/1900, reads as the text '1900-02-29' ('1900-02-29T18:00:00' with a time): no Date can hold it, and 28 February is serial 59's.

Sheets

Each sheet is { name, hidden, rows, hiddenRows, merges }:

  • rows: top to bottom from row 1. Each row runs to its last stored cell, so rows can be ragged and empty rows are [].
  • hidden: the sheet is hidden or very hidden in the workbook.
  • hiddenRows: hidden rows, 0-based.
  • merges: merged ranges as A1 references. A merged range's value sits in its top-left cell only, as the file stores it. To read the table a person sees, copy that value into every cell the range covers.
  • Chart sheets hold no cells and are left out.

Errors

A file that cannot be read throws OfficeReadError. Its code is one of the following, and its message says what to do:

| code | The file is | |---|---| | legacy-format | a legacy binary .xls, .ppt or .doc: save it as .xlsx / .pptx / .docx, or export it as PDF | | encrypted | password-protected | | unsupported-format | an OpenDocument .ods / .odp / .odt | | not-xlsx / not-pptx / not-docx | a zip package of another kind (a .pptx given to readXlsx, say) | | not-zip | not a zip at all | | too-large | past the limits | | malformed | damaged: a part does not inflate | | bad-source | not something to read bytes from (an http: URL, a text stream) |

Read a presentation

import { readPptx } from '@opencraw/office-reader/pptx'

const deck = await readPptx('./incentivi.pptx')
// { width: 960, height: 540, slides: [
//   { number: 1, title: 'Incentivi giugno', hidden: false,
//     shapes: [{ x: 60, y: 30, width: 840, height: 60, text: 'Incentivi giugno', placeholder: 'title' }],
//     tables: [{ name: 'table 1', hidden: false,
//                rows: [['Incentivi giugno 2026', '', '', ''], ['Modello', 'Prezzo', '', 'Sconto'], ['', 'Listino', 'Netto', ''], …],
//                hiddenRows: [], merges: ['A1:D1', 'A2:A3', 'B2:C2', 'D2:D3'] }],
//     charts: [], notes: 'Prezzi IVA inclusa.\nValidi fino al 30 giugno.' },
//   { number: 3, title: 'Vendite', …,
//     charts: [{ type: 'bar', title: 'Immatricolazioni',
//                series: [{ name: 'Pandina', categories: ['Aprile', 'Maggio', 'Giugno'], values: [1200, 1350.5, 1410] }, …] }] },
//   …] }

readPptx(source, options?) takes the same sources as readXlsx. Its options:

| Option | Default | | |---|---|---| | slides | all | Numbers from 1 ([2, 5]), a RegExp on titles, or ({ number, title, hidden }) => boolean. | | values | 'typed' | Chart values: numbers (null where a point is missing), or 'text'. | | notes, charts | true | false skips reading them. | | limits | as above | |

What each slide holds:

  • shapes: its text boxes in reading order (top to bottom, left to right), in points from the top-left corner.
    • A title or body placeholder with no position of its own takes the one its layout gives it, else the one its master gives it, the way PowerPoint draws it.
    • Boxes inside a group are placed through the group's scaling.
    • Slide numbers, dates and footers are left out.
  • tables: its native tables as sheets, with their merged cells (gridSpan, rowSpan) as ranges. A cell inside a merge is ''.
  • charts: the type, the title and each series' name, categories and values, from the data the chart caches next to its formulas. The embedded workbook is not needed.
    • The title is its text once: its rich text runs, or, for a title linked to a cell, the cell's cached text (never the reference). A title Office generates (no text of its own) is left out, and so are axis titles.
    • The chart types Office added in 2016 (chartEx parts: waterfall, treemap, sunburst, histogram and Pareto, box and whisker, funnel, region map) are read too. Their type is the series' layout (waterfall, treemap, sunburst, clusteredColumn for a histogram, boxWhisker, funnel, regionMap), their categories are the leaves of a hierarchy, and their values the val points (a region map's colorVal, a treemap's size).
  • notes: the speaker notes.
  • hidden: hidden in a slideshow.

Slides come in presentation order, which is not always the order of the files inside the zip.

Tested against the 100 presentations of Apache POI's test corpus. It reads 90 of them (562 slides, 41 tables, 28 charts), and every placeholder gets a position. The other 10 are fuzzer cases and a truncated zip, all refused with an OfficeReadError.

Read a Word document

import { readDocx } from '@opencraw/office-reader/docx'

const document = await readDocx('./circolare.docx')
// { title: 'Circolare giugno',
//   body: [
//     { kind: 'paragraph', text: 'Circolare incentivi giugno 2026', style: 'Title', heading: 1 },
//     { kind: 'paragraph', text: 'Condizioni', style: 'Heading1', heading: 1 },
//     { kind: 'paragraph', text: 'Solo rottamazione', style: 'ListBullet', list: { level: 0, ordered: false } },
//     { kind: 'paragraph', text: 'Prezzi', style: 'Titolo2', heading: 2 },
//     { kind: 'table', name: 'table 1', hidden: false, hiddenRows: [],
//       rows: [['Modello', 'Prezzo', '', 'Sconto'], ['', 'Listino', 'Netto', ''], ['Pandina', '15.950', '13.955', '12,5%'], …],
//       merges: ['B1:C1', 'A1:A2', 'D1:D2'] },
//     { kind: 'paragraph', text: 'Listino completo: listino giugno', links: [{ text: 'listino giugno', href: 'https://…/listino.pdf' }] },
//     … ],
//   headers: [[{ kind: 'paragraph', text: 'Stellantis Italia – riservato' }]],
//   footers: [[…]],
//   notes: [{ kind: 'footnote', id: '1', text: 'Prezzi chiavi in mano, IPT esclusa.' }] }

readDocx(source, options?) takes the same sources as readXlsx. Options: extras (default true; false skips headers, footers and notes) and limits.

  • Headings: a paragraph whose style is named heading N (the built-in names stay English in every language: an Italian Titolo2 is still named heading 2), has an outline level (its own or its style's, through basedOn), or is Title (level 1).
  • Lists: numbering from the paragraph or its style (List Bullet); ordered says whether the level is numbered (1., a), i.) or bulleted.
  • Tables: grids like a workbook's sheets. A cell spanning columns (gridSpan) fills the columns after it with ''; a cell merged down (vMerge) leaves '' below it; both are listed in merges. A table inside a cell is a block of its own, and its text is in the cell too. Tables are named table 1, table 2… once over the whole document, headers first, then the body, the footers and the notes, so no two share a name.
  • Text: tabs as \t, line breaks as \n. Tracked insertions read as text, deletions do not. Field codes are left out, their results kept. A text box's paragraphs follow the paragraph that holds it, once (Word also writes a fallback copy for old readers, which is skipped).
  • Links: external links by URL, internal ones as #bookmark.

Tested against the 130 documents of Apache POI's test corpus. It reads 115 of them (1,500 paragraphs, 5,115 tables of which 5,000 are one stress-test file, 352 links) and refuses the other 15, fuzzer cases and truncated zips, with an OfficeReadError.

How it compares

| | office-reader | SheetJS (xlsx on npm) | ExcelJS | read-excel-file | |---|---|---|---|---| | Merged ranges | yes | yes | yes | no | | Hidden rows and sheets | yes | yes | yes | no | | Error cells | { error } | yes | yes | null | | Dependencies | 2 | 7 | 9, including archiver and tmp | 4 | | Browser, workers, edge | yes | yes | Node-first | yes | | Writes files | no | yes | yes | no |

For presentations, no JavaScript library we found reads positions, tables and chart data. officeparser returns flattened text, and the others are text-only or browser renderers.

SheetJS's npm copy (0.18.5) carries two high-severity advisories (prototype pollution, ReDoS); the fixed versions are published on its own CDN only. Choose SheetJS or ExcelJS when you need to write workbooks, evaluate formulas, or read .xls and .ods. This package only reads.

Tested against the 367 workbooks of Apache POI's test corpus, real files and fuzzer cases alike. It reads 350 of them and refuses the other 17 with an OfficeReadError: encrypted, truncated or corrupted files.

Not supported

  • Writing files.
  • Evaluating formulas, and applying display formats. A percentage stays 0.125 and a price stays 15950.
  • Legacy .xls, .ppt, .doc, OpenDocument .ods / .odp / .odt, and binary .xlsb. These are refused with a code.
  • In workbooks: charts, pivot tables, images, comments and data validation.
  • In presentations: SmartArt text, text inside images (no OCR), animations and themes.
  • In documents: comments, formatting (bold, colours, fonts), images, equations, and the text of charts.