mausmark
v0.2.0
Published
A fast, dependency-free Markdown parser compiled to WebAssembly
Maintainers
Readme
Mausmark
A fast, renderer-agnostic Markdown parser.
It aims to support and extend Markdown. It is also practical, so it parses CommonMark too.
Pre-1.0: minor versions may change the API.
Install
npm install mausmarkUsage
import {createParser, Dialect, NodeType} from "mausmark";
const parser = await createParser();
const document = parser.parse("# Hello\n\nThis is **Mausmark**.");
// ...or parser.parse(source, Dialect.COMMONMARK) for CommonMark input.
for (let index = 0; index < document.nodeCount; index += 1) {
if (document.type(index) === NodeType.HEADING) {
console.log(document.joinedText(index)); // "Hello"
}
}The document is a flat array of nodes, so finding every node of one type
is a single loop; Writing a renderer walks it as a
tree. Initialize the parser once and reuse it. Initialization is
asynchronous; parse() is synchronous. Returned documents are
independent snapshots and remain valid after later parses.
AST
The root is node 0. Nodes use integer indexes into a compact preorder
table. TypeScript declarations and runtime constants are included.
type(index)returns aNodeType.children(index)iterates direct content children.text(index)decodes one node’s source slice.joinedText(index)joins a node’s text chunks back into one string.link(index)returns raw URL, title, label, andReferenceStyle.level(index)carries heading levels and other type-specific values.flags(index)carriesListFlags.ORDEREDandListFlags.LOOSEon lists.offset(index)andlength(index)are UTF-8 byte offsets.subtreeEnd(index)supports lower-level preorder traversal.
Text arrives in chunks, one per source line: a soft line break is not a
byte inside a slice but the next chunk’s join code, so line structure is
read off the tree rather than by scanning for newlines. joinedText()
reassembles a run when you just want the string; children() walks the
chunks when you care about source lines.
In the CommonMark dialect, an escaped ampersand is
NodeType.LITERAL_TEXT rather than NodeType.TEXT: its text is literal
and must not be entity-decoded together with the following chunk. Both
types contain source-backed text slices.
The tree also holds internal URL, title, and label nodes; children()
skips them. Reference links remain unresolved so renderers can choose
their own lookup approach. The tree is meant to carry enough info about
the document so that you could reproduce it exactly (minus whitespace).
Raw HTML is returned as NodeType.HTML_BLOCK or NodeType.HTML_INLINE;
sanitize it when parsing untrusted input.
Compare node types against the exported constants: their numeric values can change between versions.
Writing a renderer
Walk from node 0 with children(index), which yields direct content
children and skips a link’s leading URL/title/label fields. Do not
iterate every node index as though each were a sibling:
subtreeEnd(index) is the next sibling in the preorder table. For
example, this deliberately simple plain-text renderer omits URLs, raw
HTML, custom nodes, list markers, and entity decoding:
import {createParser, Join, NodeType} from "mausmark"
import type {MausmarkDocument} from "mausmark"
function plainText(doc: MausmarkDocument, index = 0): string {
const type = doc.type(index)
switch (type) {
case NodeType.TEXT: {
const join = doc.level(index)
return (join === Join.SPACE ? " " : join === Join.NL ? "\n" : "") + doc.text(index)
}
case NodeType.LITERAL_TEXT:
return doc.text(index)
case NodeType.HARD_BREAK:
return "\n"
case NodeType.CODE_SPAN:
case NodeType.CODE_BLOCK:
return doc.joinedText(index)
case NodeType.HTML_INLINE:
case NodeType.HTML_BLOCK:
case NodeType.CUSTOM:
case NodeType.REF_DEFS:
case NodeType.REF_DEF:
return ""
}
let text = ""
for (const child of doc.children(index)) text += plainText(doc, child)
if (type === NodeType.PARAGRAPH || type === NodeType.HEADING || type === NodeType.LIST_ITEM) text += "\n"
return text
}
const parser = await createParser()
console.log(plainText(parser.parse("# Hello\n\nA **wrapped** paragraph.")))For a complete renderer, the package includes dom-example.js. It
builds DOM nodes directly: text goes in through text nodes and
attributes through setAttribute, so it never has to escape HTML by
hand. It covers every node type, including entity decoding, reference
lookup, tight lists, fence info strings and image alt text, and it
leaves raw HTML inert. It is an example, not an export: copy it and
adapt it.
Things to decide when building a richer renderer:
- Text is source-backed and split into chunks, often by line or by the
65,535-byte slice limit. For
TEXTchunks,level()is aJoincode to emit beforetext()(Join.NLis a soft break).joinedText()assembles a node’s immediate text chunks, useful for code spans/blocks and metadata fields; it does not recursively render arbitrary inline markup. In CommonMark, aLITERAL_TEXTleaf must not be entity-decoded with its neighboringTEXT. link(index)works onLINK,IMAGE, andREF_DEFand returns raw URL, title, label, andReferenceStyle. Inline links and autolinks already have destinations;FULL,COLLAPSED, andSHORTCUTneed a definition lookup. Scan all node indexes forREF_DEF(including definitions in containers), store the first definition for each normalized label, and decide how to render unresolved references. CommonMark label matching uses Unicode case folding and whitespace collapse, not justtoLowerCase(). AnAUTOLINK_EMAILURL is the bare address; addmailto:when emitting it.title === ""alone does not distinguish an absent title from an explicit empty title; inspect the leadingTITLEfield if that matters to your output.- For lists,
flags(index)hasListFlags.ORDEREDandListFlags.LOOSEonLIST/LIST_ITEM.level()on a list records its marker character;text()on a CommonMark ordered list holds its starting digits. Tight list items normally omit paragraph wrappers. A fenced code block’s info string is its leadingLABELfield: when nodeindex + 1is inside the block and is aLABEL, read it withjoinedText(index + 1).link(index)covers link-like nodes only. HTML_BLOCKandHTML_INLINEare raw input, not sanitized HTML. Escape output for your target format and apply an explicit policy for untrusted input.
Notes
- ESM-only.
- Node.js 22.14+ or a modern browser with WebAssembly SIMD.
- Browser bundlers must emit the package-relative
.wasmasset. - Input is limited to 64 MiB; parser arena growth is limited to 512 MiB.
- Two input dialects, parsed at the same speed.
Dialect.MARKDOWN(the default) is original Markdown;Dialect.COMMONMARKis CommonMark 0.31.2, currently 649/652 (99.5%) of the 0.31.2 spec suite byte-exact. Three definition-dependent reference-link cases remain unresolved. - No GFM yet: no tables, strikethrough, task lists or autolink literals.
The npm package contains ESM bindings, TypeScript declarations, precompiled WebAssembly, and the renderer example. The implementation source is currently private; the distributed package is licensed under MIT.
Project page and demo: https://myshkin.eu/mausmark
License
MIT. See LICENSE.
