@pptx-studio/xml
v0.1.2
Published
Byte-preserving XML tokenizer and serializer for OOXML
Maintainers
Readme
@pptx-studio/xml
A byte-preserving XML tokenizer, node model, serializer and editor for OOXML.
Pre-alpha. Sub-phase 0.6 has landed schema-ordered insertion, the Markup Compatibility walker,
extLstas an opaque list, and edits that each carry their exact inverse, on top of 0.4's tokenizer and tree and 0.5's serializer. Sub-phase 1.4 addscanonicalXml, the normal form two parts are compared in when the question is whether they say the same thing.
Part of PPTX Studio. Runs in the browser and in a Web Worker; there is no Node build and there never will be.
What it does today
import { parseXml, descendantElements, attributeValue, textContent } from '@pptx-studio/xml';
const doc = parseXml(store.read('/ppt/slides/slide1.xml'));
for (const el of descendantElements(doc.root)) {
if (el.qname === 'a:t') console.log(textContent(el));
}
// Every node knows where it came from. While `dirty` is false, this *is* the node.
const shape = doc.root.children[0];
doc.source.slice(shape.start, shape.end); // byte-for-byte what was in the fileAnd back again:
import { parseXml, serializeXml, markAttributeDirty } from '@pptx-studio/xml';
const doc = parseXml(bytes);
serializeXml(doc); // === bytes, for every part of every deck we have tried
const name = doc.root.attributes.find((a) => a.qname === 'name')!;
name.value = 'Rectangle 4';
markAttributeDirty(doc.root, name);
serializeXml(doc); // the same bytes, except those eleven charactersWhy not DOMParser
Because the DOM specification permits a serializer to rewrite namespace prefixes, and Markup
Compatibility attributes hold prefixes, not URIs. mc:Ignorable="a14 p14" and
mc:Choice Requires="a14" name prefixes declared elsewhere in the document. Rename one and ignorable
extension markup silently becomes a hard error in PowerPoint — with no diagnostic, because PowerPoint
does not emit one.
@xmldom/xmldom loses namespaces on createElementNS beneath a prefixed parent.
fast-xml-parser makes no guarantee about whitespace, self-closing form or quote style. And
DOMParser is not reliably available in a Web Worker, which is where our parsing runs.
The bar is set by PowerPoint, and PowerPoint is lossy
Hand PowerPoint a slide part and ask it to save. Measured, by doing exactly that:
| | PowerPoint's own resave |
| --------------------------------- | ---------------------------- |
| comments, processing instructions | discarded |
| <![CDATA[x]]> | rewritten as plain text |
| H | resolved to H |
| <p:spPr></p:spPr> | collapsed to <p:spPr/> |
| single-quoted attributes | rewritten as double |
| byte order mark | stripped |
| xml:space="preserve" | removed |
| an unused xmlns:zz | dropped |
Every row is something this package hands back untouched. That is the whole product: preservation is the default state, not a feature.
A node's source and a node's value are different things
XML 1.0 requires the parser to transform what it reads — line endings (§2.11), attribute values (§3.3.3), references (§4.6) — and a tokenizer that stores only the transformed form can never reproduce its input. So spans are the truth and values are derived.
The trap worth knowing, because the obvious implementation gets it backwards:
attr('a\tb'); // "a b" — a literal tab is normalized to a space
attr('a	b'); // "a\tb" — a reference to a tab is notExpand references first and squash whitespace second and those two collapse into one value. The mistake passes every test written against a corpus with no character references in it — such as ours, which has zero across 2834 parts.
Things that are easy to get wrong
TextDecoder'signoreBOMis named backwards. The default,false, means "delete the BOM". 1037 of the 2834 XML parts in our corpus carry one, so the default corrupts 37% of them — and invisibly, until an export is diffed against its input.- String offsets are not byte offsets. Decode to a string, scan it, and report byte offsets and
they are wrong for every part containing non-ASCII. We work in string space throughout and encode
once at the edge;
encode(decode(bytes))is byte-identical for all 2834 parts, and for any valid UTF-8, by construction. - A prefix does not have one meaning.
p14is bound to two different URIs in different parts of our corpus, 213 namespace declarations sit on non-root elements, and a template Microsoft ships contains<p14:discardImageEditData xmlns="" xmlns:p14="…" val="0"/>nested mid-document. So resolution walks up the tree; there is no table. - An unprefixed attribute is in no namespace, not the default one (Namespaces in XML §6.2).
<x/>and<x></x>are different, and so is<x />. 70 822 of the 98 777 self-closing tags in our corpus have a space before the slash.xml:spaceis not needed in PresentationML. There is none in the corpus, 120<a:t>elements carry edge whitespace without it, and PowerPoint round-trips them intact. This package reads the attribute and never writes one.
The gate is coverage, not "no errors"
An off-by-one in a span raises no error. It parses every part of every deck happily and loses a
character the first time anything re-serializes. So checkSpanCoverage and checkTreeCoverage assert
that spans tile the source exactly — no gaps, no overlaps, and inside a start tag the name, each
attribute and the closing delimiter account for every character too. Every test that parses anything
runs through them.
The tag interior was added after a review: the first version tiled the nodes but never read
trailingSpaceStart or any attribute offset, so an off-by-one there passed the whole suite and
would have passed 0.5's byte-identical gate, surfacing for the first time in 0.6 on an edited deck.
Across the corpus: 2834 parts, 314 814 tokens, 194 148 elements, 170 019 attributes, zero gaps, overlaps or byte mismatches.
What a round trip actually guarantees
Two different claims, and only one of them is about bytes.
| | | | --------------------------------------------------------------- | ------------------ | | parse → serialize, nothing edited → byte-identical | 2834 / 2834 (100%) | | every node forced dirty → serialize → parse → same document | 2834 / 2834 (100%) | | every node forced dirty → serialize → byte-identical | 2178 / 2834 (77%) |
The first row is the headline and it is exact. The third is measured rather than promised, and there
is exactly one cause: CRLF and LF are the same value after §2.11, so a text node rebuilt from its
value cannot know which it came from, and 656 parts have a CRLF between the declaration and the root.
Every other part of a full rebuild comes back character for character — every entity reference,
every quote character, every <x /> with its space, every <x></x> left long.
That third row costs nothing in practice, because dirty travels up and never down: a text node is
rebuilt only when someone edits that node, so the whitespace between siblings in a part you did not
touch is never rebuilt at all.
The canonical form, for when the two parts came from different writers
canonicalXml(doc) returns one string per document, and two documents that mean
the same thing return the same string. It is what @pptx-studio/writer compares
parts in, and it exists because "the same bytes" is the wrong question for a part
somebody else wrote.
canonicalXml(parseXmlString('<a:off x="1" y="2"/>')) ===
canonicalXml(parseXmlString(`<a:off y='2' x="1" />`)); // trueAbsorbed: attribute order, quote style, <a/> against <a></a>, a character
reference against the character it names, CDATA against the text it holds, the
encoding pseudo-attribute, a byte order mark.
Not absorbed, and this is the half that matters: whitespace anywhere, a namespace prefix, a namespace declaration nothing appears to use, a comment, the presence of the XML declaration.
The name is borrowed from W3C Canonical XML and two of C14N's rules are wrong
here. It prunes unused namespace declarations — but PowerPoint writes
xmlns:a14="…" alongside mc:Ignorable="a14" with no a14: element in the
part, so pruning it leaves a document naming an undeclared prefix. And it is
free to rename prefixes, which is the exact rewrite this whole package exists to
prevent.
Writing a value back is not the same rules read backwards
Four characters are lost outright by a serializer that writes values as it finds them:
| a value containing | written literally, reparses as | by |
| ------------------ | ------------------------------ | ------ |
| text \r | \n | §2.11 |
| attribute \r | a space | §3.3.3 |
| attribute \n | a space | §3.3.3 |
| attribute \t | a space | §3.3.3 |
So each is written as a character reference. > is escaped everywhere although almost nothing
requires it — measured, the corpus contains zero literal > inside a value and five
>, so escaping it always re-spells nothing while the narrow rule re-spelled all five. Only the
delimiting quote is escaped, because all 170 019 attributes are double-quoted and 26 carry a
".
Editing, and the promise that undo is exact
import { applyEdits, insertInOrder, newElement, newAttribute } from '@pptx-studio/xml';
const spPr = firstChild(shape, 'p:spPr')!;
// Where does <a:ln> go inside <a:spPr>? Not the caller's problem.
const undo = applyEdits([
insertInOrder(spPr, newElement('a:ln', [newAttribute('w', '9525')])),
{ kind: 'setAttribute', element: off, qname: 'x', value: '914400' },
]);
serializeXml(doc); // the deck, with those two changes and nothing else
applyEdits(undo);
serializeXml(doc); // === the original bytes. Not "an equivalent document" — the bytes.That last line is the whole of sub-phase 0.6, and it is harder than it looks. dirty === false is
an invariant, not a hint: it asserts that source.slice(start, end) is this node's
serialization. An inverse that restores the value but leaves the flag set gives back the same
document and different bytes — the subtree is rebuilt rather than sliced, so a <a:off … />
loses its space, a > comes back as >, and a CRLF between two elements comes back as LF, which
is 656 of our 2834 parts. So every edit records the flags it set and its inverse clears exactly
those, never one it did not set.
Measured over the corpus: 194 244 edits applied and undone across 2834 parts, every part byte-identical afterwards, no dirty flag left behind.
Two properties worth knowing before you plan a gesture:
- A batch is all of it or none of it. If any edit is refused, the ones that already landed are undone before the error is rethrown. Some refusals are only knowable at apply time, and without this a caller would be left with a half-applied document and no way back.
- Removal is addressed by node, insertion by index. An index is a position at the moment the edit is applied, so two insertions planned against one parent can both compute the same one. Removal has a stable address and uses it.
Where a new child goes
OOXML complex types are xsd:sequence, and PowerPoint answers a misplaced child with "found a
problem" and nothing more. insertInOrder is the only sanctioned way to add one; the table behind
it is generated from the ECMA-376 Transitional schemas by tools/schema-codegen.
It is a rank, not a list, and that distinction is load-bearing: p:spTree admits sp, grpSp,
graphicFrame, cxnSp, pic and contentPart under one repeating xsd:choice, and their freedom
to interleave is the z-order of the slide. A new shape lands on top of it; nothing already there
moves, ever. An mc:AlternateContent or a p14: element among the children is stepped over rather
than ranked, so it stays exactly where its producer put it.
It refuses rather than guesses. a:ext inside an extLst, and a:graphicData where a chart lives,
are xsd:any in the schema and get no table at all. Neither does anything in the chart or diagram
namespaces — deliberately, because those parts are copied byte-for-byte and never rewritten.
Checked against 194 148 real elements written by PowerPoint: zero out of order.
Markup Compatibility, read but never rewritten
const branch = selectAlternateContent(alternate, new Set([NS.p, NS.a]));
for (const child of effectiveChildren(shape, supported)) { … }mc:Choice/@Requires holds prefixes, and it is written unprefixed — <mc:Choice Requires="a14">
— which by Namespaces in XML §6.2 puts it in no namespace at all. Look for mc:Requires and you
find nothing in any real file, and a walker that then treats the Choice as requiring nothing selects
the first branch of every switch in the document.
A prefix that resolves to nothing is a hard error here rather than a namespace we happen not to support, because treating it as unsupported would silently select a different branch. An element in a namespace nobody declared ignorable is kept, not dropped: the specification says the producer should have annotated it, and losing content a producer forgot to annotate is the worse failure.
Nothing in this module mutates anything.
Hostile input
DOCTYPE is rejected outright, with its own error code. That removes XXE, parameter entities and
billion-laughs from the attack surface entirely rather than defending against them — and it costs
nothing, because PowerPoint refuses a DOCTYPE too.
The builder never recurses, because PowerPoint opens a part nested 5000 elements deep and a
recursive-descent parser would answer that with a RangeError. Depth, attribute count and node count
all have configurable ceilings. Every failure throws an XmlError with a machine-readable code.
Known limitation
UTF-8 only. PowerPoint honours the encoding pseudo-attribute and will open a part written in
windows-1252 or UTF-16; we refuse it with ERR_UNSUPPORTED_ENCODING, naming the encoding. Zero of
2834 parts in our corpus are affected. A part labelled windows-1252 whose content is ASCII parses
fine — the refusal is about bytes we cannot reproduce, not about labels.
Licence
Apache-2.0
