npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@pptx-studio/xml

v0.1.2

Published

Byte-preserving XML tokenizer and serializer for OOXML

Readme

@pptx-studio/xml

A byte-preserving XML tokenizer, node model, serializer and editor for OOXML.

Pre-alpha. Sub-phase 0.6 has landed schema-ordered insertion, the Markup Compatibility walker, extLst as an opaque list, and edits that each carry their exact inverse, on top of 0.4's tokenizer and tree and 0.5's serializer. Sub-phase 1.4 adds canonicalXml, the normal form two parts are compared in when the question is whether they say the same thing.

Part of PPTX Studio. Runs in the browser and in a Web Worker; there is no Node build and there never will be.

What it does today

import { parseXml, descendantElements, attributeValue, textContent } from '@pptx-studio/xml';

const doc = parseXml(store.read('/ppt/slides/slide1.xml'));

for (const el of descendantElements(doc.root)) {
  if (el.qname === 'a:t') console.log(textContent(el));
}

// Every node knows where it came from. While `dirty` is false, this *is* the node.
const shape = doc.root.children[0];
doc.source.slice(shape.start, shape.end); // byte-for-byte what was in the file

And back again:

import { parseXml, serializeXml, markAttributeDirty } from '@pptx-studio/xml';

const doc = parseXml(bytes);
serializeXml(doc); // === bytes, for every part of every deck we have tried

const name = doc.root.attributes.find((a) => a.qname === 'name')!;
name.value = 'Rectangle 4';
markAttributeDirty(doc.root, name);

serializeXml(doc); // the same bytes, except those eleven characters

Why not DOMParser

Because the DOM specification permits a serializer to rewrite namespace prefixes, and Markup Compatibility attributes hold prefixes, not URIs. mc:Ignorable="a14 p14" and mc:Choice Requires="a14" name prefixes declared elsewhere in the document. Rename one and ignorable extension markup silently becomes a hard error in PowerPoint — with no diagnostic, because PowerPoint does not emit one.

@xmldom/xmldom loses namespaces on createElementNS beneath a prefixed parent. fast-xml-parser makes no guarantee about whitespace, self-closing form or quote style. And DOMParser is not reliably available in a Web Worker, which is where our parsing runs.

The bar is set by PowerPoint, and PowerPoint is lossy

Hand PowerPoint a slide part and ask it to save. Measured, by doing exactly that:

| | PowerPoint's own resave | | --------------------------------- | ---------------------------- | | comments, processing instructions | discarded | | <![CDATA[x]]> | rewritten as plain text | | &#72; | resolved to H | | <p:spPr></p:spPr> | collapsed to <p:spPr/> | | single-quoted attributes | rewritten as double | | byte order mark | stripped | | xml:space="preserve" | removed | | an unused xmlns:zz | dropped |

Every row is something this package hands back untouched. That is the whole product: preservation is the default state, not a feature.

A node's source and a node's value are different things

XML 1.0 requires the parser to transform what it reads — line endings (§2.11), attribute values (§3.3.3), references (§4.6) — and a tokenizer that stores only the transformed form can never reproduce its input. So spans are the truth and values are derived.

The trap worth knowing, because the obvious implementation gets it backwards:

attr('a\tb'); // "a b"   — a literal tab is normalized to a space
attr('a&#9;b'); // "a\tb" — a reference to a tab is not

Expand references first and squash whitespace second and those two collapse into one value. The mistake passes every test written against a corpus with no character references in it — such as ours, which has zero across 2834 parts.

Things that are easy to get wrong

  • TextDecoder's ignoreBOM is named backwards. The default, false, means "delete the BOM". 1037 of the 2834 XML parts in our corpus carry one, so the default corrupts 37% of them — and invisibly, until an export is diffed against its input.
  • String offsets are not byte offsets. Decode to a string, scan it, and report byte offsets and they are wrong for every part containing non-ASCII. We work in string space throughout and encode once at the edge; encode(decode(bytes)) is byte-identical for all 2834 parts, and for any valid UTF-8, by construction.
  • A prefix does not have one meaning. p14 is bound to two different URIs in different parts of our corpus, 213 namespace declarations sit on non-root elements, and a template Microsoft ships contains <p14:discardImageEditData xmlns="" xmlns:p14="…" val="0"/> nested mid-document. So resolution walks up the tree; there is no table.
  • An unprefixed attribute is in no namespace, not the default one (Namespaces in XML §6.2).
  • <x/> and <x></x> are different, and so is <x />. 70 822 of the 98 777 self-closing tags in our corpus have a space before the slash.
  • xml:space is not needed in PresentationML. There is none in the corpus, 120 <a:t> elements carry edge whitespace without it, and PowerPoint round-trips them intact. This package reads the attribute and never writes one.

The gate is coverage, not "no errors"

An off-by-one in a span raises no error. It parses every part of every deck happily and loses a character the first time anything re-serializes. So checkSpanCoverage and checkTreeCoverage assert that spans tile the source exactly — no gaps, no overlaps, and inside a start tag the name, each attribute and the closing delimiter account for every character too. Every test that parses anything runs through them.

The tag interior was added after a review: the first version tiled the nodes but never read trailingSpaceStart or any attribute offset, so an off-by-one there passed the whole suite and would have passed 0.5's byte-identical gate, surfacing for the first time in 0.6 on an edited deck.

Across the corpus: 2834 parts, 314 814 tokens, 194 148 elements, 170 019 attributes, zero gaps, overlaps or byte mismatches.

What a round trip actually guarantees

Two different claims, and only one of them is about bytes.

| | | | --------------------------------------------------------------- | ------------------ | | parse → serialize, nothing edited → byte-identical | 2834 / 2834 (100%) | | every node forced dirty → serialize → parse → same document | 2834 / 2834 (100%) | | every node forced dirty → serialize → byte-identical | 2178 / 2834 (77%) |

The first row is the headline and it is exact. The third is measured rather than promised, and there is exactly one cause: CRLF and LF are the same value after §2.11, so a text node rebuilt from its value cannot know which it came from, and 656 parts have a CRLF between the declaration and the root. Every other part of a full rebuild comes back character for character — every entity reference, every quote character, every <x /> with its space, every <x></x> left long.

That third row costs nothing in practice, because dirty travels up and never down: a text node is rebuilt only when someone edits that node, so the whitespace between siblings in a part you did not touch is never rebuilt at all.

The canonical form, for when the two parts came from different writers

canonicalXml(doc) returns one string per document, and two documents that mean the same thing return the same string. It is what @pptx-studio/writer compares parts in, and it exists because "the same bytes" is the wrong question for a part somebody else wrote.

canonicalXml(parseXmlString('<a:off x="1" y="2"/>')) ===
  canonicalXml(parseXmlString(`<a:off y='2' x="1" />`)); // true

Absorbed: attribute order, quote style, <a/> against <a></a>, a character reference against the character it names, CDATA against the text it holds, the encoding pseudo-attribute, a byte order mark.

Not absorbed, and this is the half that matters: whitespace anywhere, a namespace prefix, a namespace declaration nothing appears to use, a comment, the presence of the XML declaration.

The name is borrowed from W3C Canonical XML and two of C14N's rules are wrong here. It prunes unused namespace declarations — but PowerPoint writes xmlns:a14="…" alongside mc:Ignorable="a14" with no a14: element in the part, so pruning it leaves a document naming an undeclared prefix. And it is free to rename prefixes, which is the exact rewrite this whole package exists to prevent.

Writing a value back is not the same rules read backwards

Four characters are lost outright by a serializer that writes values as it finds them:

| a value containing | written literally, reparses as | by | | ------------------ | ------------------------------ | ------ | | text \r | \n | §2.11 | | attribute \r | a space | §3.3.3 | | attribute \n | a space | §3.3.3 | | attribute \t | a space | §3.3.3 |

So each is written as a character reference. > is escaped everywhere although almost nothing requires it — measured, the corpus contains zero literal > inside a value and five &gt;, so escaping it always re-spells nothing while the narrow rule re-spelled all five. Only the delimiting quote is escaped, because all 170 019 attributes are double-quoted and 26 carry a &quot;.

Editing, and the promise that undo is exact

import { applyEdits, insertInOrder, newElement, newAttribute } from '@pptx-studio/xml';

const spPr = firstChild(shape, 'p:spPr')!;

// Where does <a:ln> go inside <a:spPr>? Not the caller's problem.
const undo = applyEdits([
  insertInOrder(spPr, newElement('a:ln', [newAttribute('w', '9525')])),
  { kind: 'setAttribute', element: off, qname: 'x', value: '914400' },
]);

serializeXml(doc); // the deck, with those two changes and nothing else

applyEdits(undo);
serializeXml(doc); // === the original bytes. Not "an equivalent document" — the bytes.

That last line is the whole of sub-phase 0.6, and it is harder than it looks. dirty === false is an invariant, not a hint: it asserts that source.slice(start, end) is this node's serialization. An inverse that restores the value but leaves the flag set gives back the same document and different bytes — the subtree is rebuilt rather than sliced, so a <a:off … /> loses its space, a &#62; comes back as >, and a CRLF between two elements comes back as LF, which is 656 of our 2834 parts. So every edit records the flags it set and its inverse clears exactly those, never one it did not set.

Measured over the corpus: 194 244 edits applied and undone across 2834 parts, every part byte-identical afterwards, no dirty flag left behind.

Two properties worth knowing before you plan a gesture:

  • A batch is all of it or none of it. If any edit is refused, the ones that already landed are undone before the error is rethrown. Some refusals are only knowable at apply time, and without this a caller would be left with a half-applied document and no way back.
  • Removal is addressed by node, insertion by index. An index is a position at the moment the edit is applied, so two insertions planned against one parent can both compute the same one. Removal has a stable address and uses it.

Where a new child goes

OOXML complex types are xsd:sequence, and PowerPoint answers a misplaced child with "found a problem" and nothing more. insertInOrder is the only sanctioned way to add one; the table behind it is generated from the ECMA-376 Transitional schemas by tools/schema-codegen.

It is a rank, not a list, and that distinction is load-bearing: p:spTree admits sp, grpSp, graphicFrame, cxnSp, pic and contentPart under one repeating xsd:choice, and their freedom to interleave is the z-order of the slide. A new shape lands on top of it; nothing already there moves, ever. An mc:AlternateContent or a p14: element among the children is stepped over rather than ranked, so it stays exactly where its producer put it.

It refuses rather than guesses. a:ext inside an extLst, and a:graphicData where a chart lives, are xsd:any in the schema and get no table at all. Neither does anything in the chart or diagram namespaces — deliberately, because those parts are copied byte-for-byte and never rewritten.

Checked against 194 148 real elements written by PowerPoint: zero out of order.

Markup Compatibility, read but never rewritten

const branch = selectAlternateContent(alternate, new Set([NS.p, NS.a]));
for (const child of effectiveChildren(shape, supported)) { … }

mc:Choice/@Requires holds prefixes, and it is written unprefixed — <mc:Choice Requires="a14"> — which by Namespaces in XML §6.2 puts it in no namespace at all. Look for mc:Requires and you find nothing in any real file, and a walker that then treats the Choice as requiring nothing selects the first branch of every switch in the document.

A prefix that resolves to nothing is a hard error here rather than a namespace we happen not to support, because treating it as unsupported would silently select a different branch. An element in a namespace nobody declared ignorable is kept, not dropped: the specification says the producer should have annotated it, and losing content a producer forgot to annotate is the worse failure.

Nothing in this module mutates anything.

Hostile input

DOCTYPE is rejected outright, with its own error code. That removes XXE, parameter entities and billion-laughs from the attack surface entirely rather than defending against them — and it costs nothing, because PowerPoint refuses a DOCTYPE too.

The builder never recurses, because PowerPoint opens a part nested 5000 elements deep and a recursive-descent parser would answer that with a RangeError. Depth, attribute count and node count all have configurable ceilings. Every failure throws an XmlError with a machine-readable code.

Known limitation

UTF-8 only. PowerPoint honours the encoding pseudo-attribute and will open a part written in windows-1252 or UTF-16; we refuse it with ERR_UNSUPPORTED_ENCODING, naming the encoding. Zero of 2834 parts in our corpus are affected. A part labelled windows-1252 whose content is ASCII parses fine — the refusal is about bytes we cannot reproduce, not about labels.

Licence

Apache-2.0