npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@s8fy/pptx-parser

v3.1.0

Published

Browser-first JavaScript and WebAssembly parser for PPTX files

Downloads

2,015

Readme

@s8fy/pptx-parser

A browser-first TypeScript library for parsing .pptx files into structured data with JavaScript or WebAssembly, and converting original PPTX bytes to PDF. Node.js 22.12+ is also supported through ESM and CommonJS entrypoints.

Architecture

  • Structured text output — text content outputs TextParagraph[] JSON arrays with supported formatting metadata instead of raw HTML strings
  • Pure data output — parser only outputs raw parsed data; model adaptation lives in @s8fy/pptx-core and DOM/SVG rendering lives in @s8fy/pptx-renderer
  • Typed parsing API — strict TypeScript parsing code and declarations, with bundled MoonBit WASM parsing/PDF engines

Capabilities

  • Embedded font extraction — extracts font binary data (EOT/ODTTF/TTF/OTF/WOFF) from PPTX embedded font list; opt-in via extractFonts option (default false to avoid loading large CJK fonts)
  • Arrow endpoints — parses a:headEnd/a:tailEnd on lines with support for arrow, stealth, diamond, dot marker types and sm/med/lg sizing
  • Chart models — ordinary and extended chart types, inherited styling, missing data slots, bubble sizing, and data-label deletion/reset metadata
  • SmartArt — cached drawing extraction and native projection for supported uncached linear-process and organization-hierarchy layouts, with local resources and deferred text-fitting metadata
  • DrawingML effects and media — glyph outlines, glow/reflections, shape inner shadows/reflections, 3D metadata, and authored audio/video posters
  • Animation metadata — supported chart/SmartArt subtargets, Box effects, and slide transition facts including dissolve; playback belongs to the renderer

Animation diagnostics in the planned 3.1 release

This release intentionally includes an incompatible diagnostic type change in a minor version while the SDK has no external consumers. No compatibility variant is retained.

ParsedAnimationDegraded no longer includes the effect-fallback variant or its fromEffect / toEffect fields. Both parsers retain effect.kind: 'dissolve' and stop emitting the dissolve-to-fade diagnostic; the Renderer now uses particle masks. Update exhaustive diagnostic handlers to the remaining text-granularity variant and reparse original PPTX bytes when migrating stored parser output. Upgrade Parser, Core, Renderer and Viewer together. PDF output remains static.

Install

pnpm add @s8fy/pptx-parser
# or
npm install @s8fy/pptx-parser

Usage

Browser

<input
  type="file"
  accept="application/vnd.openxmlformats-officedocument.presentationml.presentation"
/>
import { parsePptx, revokeBlobUrls } from '@s8fy/pptx-parser'

const input = document.querySelector('input[type="file"]')
if (!(input instanceof HTMLInputElement)) throw new Error('Missing file input')
input.addEventListener('change', async () => {
  const file = input.files?.[0]
  if (!file) return
  try {
    const result = await parsePptx(await file.arrayBuffer(), { parser: 'wasm' })
    console.log(result)
    // Keep result while its media is used; see cleanup below.
    revokeBlobUrls(result)
  } catch (error) {
    console.error(error)
  }
})

Node.js

const { parsePptx, revokeBlobUrls } = require('@s8fy/pptx-parser')
const { readFileSync } = require('node:fs')

async function main() {
  const bytes = readFileSync('test.pptx')
  const buffer = bytes.buffer.slice(bytes.byteOffset, bytes.byteOffset + bytes.byteLength)
  const result = await parsePptx(buffer)
  try {
    console.log(result)
  } finally {
    revokeBlobUrls(result)
  }
}

main().catch(console.error)

API

parsePptx(file: ArrayBuffer, options?: ParseOptions): Promise<ParseResult>

Options

| Option | Type | Default | Description | | -------------- | ---------------- | ---------------- | --------------------------------------------------------------- | | extractFonts | boolean | false | Whether to extract embedded font binary data | | parser | 'js' \| 'wasm' | 'js' | Parser engine for this low-level package | | unzipMode | 'js' \| 'wasm' | 'wasm' | ZIP engine when parser: 'wasm' | | wasmUrl | string \| URL | Package-relative | Override the parser WASM URL for CDN or self-hosted deployments | | timing | boolean | false | Include per-phase timing in the optional _timing result field | | signal | AbortSignal | None | Cancel this parse request without cancelling other callers |

@s8fy/pptx-core, the renderer, and the viewer facade default to the WASM parser. The low-level parser defaults to JavaScript for backward compatibility.

The package ships package-relative JS/WASM/PDF workers and both WASM binaries. Modern bundlers such as Vite copy these assets from the static runtime references. The UMD entry resolves adjacent assets from the current <script> URL, so keep the published dist/ files together when loading it directly. CommonJS is intended for Node.js/SSR and can load the packaged WASM files directly when parser: 'wasm' is requested. Browser applications should prefer the ESM entry. Both the parser and PDF APIs accept explicit wasmUrl overrides; the PDF API additionally accepts workerUrl.

Parser WASM deployment

Automatic package-relative resolution is the default. For a CDN, offline bundle, or private deployment, copy node_modules/@s8fy/pptx-parser/dist/mbt/main.wasm into the application's static assets and pass its deployed URL:

import { parsePptx } from '@s8fy/pptx-parser'

await parsePptx(buffer, {
  parser: 'wasm',
  wasmUrl: '/vendor/s8fy/main.wasm',
})
  • The package-relative Worker/WASM paths are covered by the repository's Vite consumer smoke. For other bundlers, verify the deployed assets; use the bundler's public/copy-asset facility when supplying an override.
  • Next.js and Nuxt serve the copied file from public/; call the SDK from client-side code and pass a root-relative URL.
  • SvelteKit serves it from static/; pass the matching root-relative URL.
  • Angular CLI can add the source file as an assets glob in angular.json, then pass the configured output URL.

The same-origin server should return the file without HTML fallback content and preferably with Content-Type: application/wasm. Cross-origin URLs must allow the application origin through CORS.

PDF and resource helpers

  • pptxToPdf(file, options?) returns PDF bytes as Uint8Array.
  • pptxToPdfBlob(file, options?) returns a browser Blob.
  • revokeBlobUrls(data) releases Blob URLs created while parsing media.
  • transferBlobUrls(source, target) moves cleanup ownership to a custom adapted object, without copying bytes. After transfer, clean up target instead of source.
  • ResourcePolicyError and RESOURCE_POLICY_CONTRACT expose the parser's deterministic input/resource-limit contract.

PDF conversion accepts explicit wasmUrl, wasmBytes, and workerUrl overrides, font/emoji fallback sources, watermarks, and signed-license runtime data. Prefer the higher-level @s8fy/pptx-core facade unless the raw parser result is required.

import { pptxToPdf } from '@s8fy/pptx-parser'

export async function convertFile(file: File): Promise<Uint8Array> {
  return pptxToPdf(await file.arrayBuffer(), {
    unicodeFallbackFont: { url: '/fonts/NotoSansSC-Regular.ttf' },
  })
}

Serve that font URL yourself or omit it for documents that do not need a Unicode fallback. The PDF API always uses its PDF WASM engine; choosing JS for parsing does not make PDF conversion JavaScript-only. PDF wasmUrl points to pdf-converter.wasm, not parser main.wasm. Browser-window conversion requires a working PDF Worker and does not silently retry on the main thread if the Worker fails. Node conversion runs on the current thread.

The bundled PDF engine renders supported native SmartArt layouts, glyph glow/text reflections, shape inner shadows, solid 3D text, and static audio posters. Shape-frame reflections currently require zero blur; this restriction is separate from glyph reflections. Unsupported or excessive effect branches retain foreground fallbacks. Reflections do not duplicate foreground text in PDF text extraction. PDF output is static and does not play animations or audio. Missing fonts and unsupported Office features can reduce fidelity.

Both PDF functions accept signal for cancellation. Font options include regularFonts (with regular/bold/italic/boldItalic styles), unicodeFallbackFont, mathFallbackFont, emojiFallbackFont, and emojiBitmapFallback. Browser font CSS does not automatically supply PDF font bytes. The low-level watermarks/licenseRuntime options differ from core/viewer's watermark/license facade; do not interchange their option objects.

Cancellation, errors, and cleanup

Pass new AbortController().signal to parsing or PDF conversion, then call that controller's abort() when the request is obsolete. Cancellation rejects the request; synchronous work on the same JavaScript thread cannot be interrupted until control returns.

When a parsed result is no longer displayed or otherwise used, call revokeBlobUrls(result) on that same object. Media URLs may be blob: URLs rather than base64 strings and are not portable across sessions. A JSON clone does not retain the original cleanup ownership. Use transferBlobUrls only when intentionally replacing the owner with an adapted model.

ResourcePolicyError.details contains reason, stage, subject, limit, observed, unit, and policyVersion. These engineering limits are independent of commercial licenses (commercialUpgradeApplicable is false); catch and inspect the error rather than retrying the same oversized input. Other parse/conversion failures can be ordinary Error instances.

ParseResult

import type { EmbeddedFont, Slide } from '@s8fy/pptx-parser'

interface ParseResult {
  slides: Slide[]
  themeColors: string[]
  hlinkColor: string | null
  fonts: EmbeddedFont[]
  size: {
    width: number // px
    height: number // px
  }
}

Upgrading to 3.0

ParsedAnimationDegraded is now a discriminated union rather than an interface with text-granularity fields on every entry. Existing diagnostic objects must include kind: 'text-granularity'; consumers must narrow on kind before reading variant-specific fields. Interface extension and declaration merging against the old interface must also be replaced with application-owned types.

import type { ParsedAnimationDegraded } from '@s8fy/pptx-parser'

export function describeDegradation(diagnostic: ParsedAnimationDegraded): string {
  if (diagnostic.kind === 'text-granularity') {
    return `${diagnostic.fromGranularity} -> ${diagnostic.toGranularity}`
  }
  return `${diagnostic.fromEffect} -> ${diagnostic.toEffect}`
}

An exact dissolve entrance/exit element effect remains kind: 'dissolve' in the parsed source model, without an effect-fallback diagnostic. The Renderer plays scattered-cell masking; source timing and phase are preserved. Unknown filter variants remain unsupported. Object and slide-transition dissolve have separate lifecycles.

Output Example

{
  slides: [
    {
      fill: { type: 'color', value: '#FFFFFF' },
      elements: [
        {
          type: 'text',
          left: 100,
          top: 50,
          width: 600,
          height: 80,
          content: [
            {
              align: 'center',
              lineHeight: 1.2,
              runs: [
                {
                  text: 'Hello World',
                  fontSize: '24pt',
                  fontFamily: 'Calibri',
                  fontWeight: 'bold',
                  color: '#333333'
                }
              ]
            }
          ],
          vAlign: 'mid',
          name: 'Title 1',
          order: 0,
          // ...
        },
        // more elements...
      ],
      layoutElements: [/* master/layout elements */],
      note: 'Speaker notes...',
      transition: { type: 'fade', duration: 500, direction: null }
    },
    // more slides...
  ],
  themeColors: ['#4472C4', '#ED7D31', '#A5A5A5', '#FFC000', '#5B9BD5', '#70AD47'],
  hlinkColor: null,
  fonts: [/* embedded fonts if extractFonts: true */],
  size: { width: 960, height: 540 }
}

Element Types

Text (type: 'text')

Text box with structured paragraphs. content is a TextParagraph[] array containing runs with font/color/decoration info, list/bullet info, and paragraph-level alignment/spacing.

Shape (type: 'shape')

Predefined or custom shapes. Also uses TextParagraph[] for content. Includes shapType, optional path (SVG), keypoints, and arrow endpoints (headEnd/tailEnd).

Image (type: 'image')

Image element with src (data URI or an owned Blob URL), optional rect (crop), geom (clip shape), and filters (sharpen, brightness, contrast, saturation, color temperature).

Table (type: 'table')

Table with data (2D TableCell[][]), rowHeights, colWidths. Cells support rowSpan/colSpan merging, per-cell borders and fill.

Chart (type: 'chart')

Chart data with chartType (bar, column, line, pie, ring, area, radar, scatter, bubble, stock, surface, waterfall, sunburst, treemap, histogram, boxWhisker, or funnel), rawChartType (original OOXML type), and type-specific data/style fields. Combo charts retain ordered chartGroups.

Missing ChartValue.y and scatter coordinates can be null; JavaScript cache gaps can also be NaN. Preserve point indexes and check values before arithmetic rather than replacing gaps with zero. Bubble bubbleSizeRepresents is 'area' or 'w' (diameter). An omitted data-label deleted flag means inherit, true hides the label, and false explicitly keeps it; textTransform: 'none' resets inherited capitalization.

Video (type: 'video')

Video element with blob (Blob URL) or src (external URL), optional poster (thumbnail image from slide), ext (file extension), and rotate.

Audio (type: 'audio')

Audio element with optional blob (playable Blob URL), poster (authored audio-object image), ext (file extension), isFlipH/isFlipV, and rotate. The poster is independent of the playable resource and can remain available when audio cannot be played.

Diagram (type: 'diagram')

SmartArt with elements (Shape/Text sub-elements) and textList (fallback text content), plus optional point/connection identities and textFit metadata. Cached drawings take precedence. Without a cache, supported linear-process and organization-hierarchy layout profiles can be projected into native shapes; unsupported or ambiguous layouts retain the existing fallback. This does not implement every SmartArt layout.

Math (type: 'math')

Math formula with latex (LaTeX expression), picBase64 (fallback image), and optional text (mixed text/formula content).

Group (type: 'group')

Group container with nested elements array.

Structured Text Format

Text content uses structured TextParagraph[] instead of HTML strings:

import type {
  BulletInfo,
  TextColor,
  TextEffects,
  TextMathRun,
  TextOutline,
} from '@s8fy/pptx-parser'

// Abridged shapes of the exported types; use the package's declarations in applications.
interface TextParagraph {
  align: string // 'left' | 'center' | 'right' | 'justify'
  lineHeight?: number // CSS line-height value
  spaceBefore?: string // e.g. '12pt'
  spaceAfter?: string // e.g. '6pt'
  list?: {
    type: 'ul' | 'ol'
    level: number
    bullet?: BulletInfo
  }
  runs: (TextRun | TextLink | TextMathRun | TextBreak)[]
}

interface TextRun {
  text: string
  color?: TextColor
  fontSize?: string // e.g. '18pt'
  fontFamily?: string
  fontWeight?: string // e.g. 'bold'
  fontStyle?: string // e.g. 'italic'
  textDecoration?: string // e.g. 'underline'
  textDecorationLine?: string // e.g. 'line-through'
  letterSpacing?: string
  kern?: string // inherited pair-kerning threshold in pt (e.g. '12pt'); '0pt' disables kerning
  verticalAlign?: string // 'super' | 'sub'
  textShadow?: string
  textEffects?: TextEffects
  textOutline?: TextOutline
  textCaps?: 'none' | 'small' | 'all'
  highlightColor?: string
}

interface TextLink extends TextRun {
  linkURL: string
  linkColor?: string
}

interface TextBreak {
  type: 'br'
}

textEffects carries glyph glow/reflection; an explicitly empty object cancels inherited run effects. textOutline is a separate solid glyph stroke in pixels, with zero width disabling it; gradient/dashed/compound outlines use a solid fallback. textCaps changes display capitalization without changing source text. Reflection distances/blur use pixels, angles use degrees, and alpha/position/scale values use ratios. Parsed metadata does not guarantee that every consumer paints every effect.

Fill Types

All fill-capable elements support four fill types:

  • Color (type: 'color') — solid color value
  • Image (type: 'image') — picture source with opacity, stretch/tile mapping, and optional crop metadata
  • Gradient (type: 'gradient') — linear/circle/rect/shape gradient with color stops
  • Pattern (type: 'pattern') — pattern type with foreground/background colors

Embedded Fonts

When extractFonts: true is passed:

interface EmbeddedFont {
  typeface: string
  styles: {
    regular?: { path: string; data: ArrayBuffer }
    bold?: { path: string; data: ArrayBuffer }
    italic?: { path: string; data: ArrayBuffer }
    boldItalic?: { path: string; data: ArrayBuffer }
  }
}

Build Outputs

| Format | File | Usage | | ------ | ------------------- | --------------------------------------------- | | ESM | dist/index.js | Modern bundlers (Vite, Webpack 5+) | | CJS | dist/index.cjs | Node.js require() | | UMD | dist/index.umd.js | Browser <script> tag (global: pptxParser) |

Type Definitions

The package root exports the complete bundled TypeScript declarations through its types and conditional exports fields. No deep import is required.

WASM parsing/PDF conversion requires WebAssembly GC support. Workers may be unavailable in some hosts: parsing has a main-thread fallback, while browser-window PDF conversion requires its Worker. The parsed model is a subset of PowerPoint behavior; output presence does not imply pixel-identical rendering or support for every animation/chart effect.

Credits

This library is based on pptxtojson by pipipi-pikachu.

License

Default license for original contributions: PolyForm-Noncommercial-1.0.0

Commercial use of those contributions requires a separate commercial license.