@dornick/parsers
v0.1.0
Published
File parser registry for Dornick. Lazy VFS sub-mounts under /user/files/<id>/parsed/*. Built-in CSV/JSON/text; companion packages add PDF, DOCX, XLSX, OCR, ML, audio.
Readme
@dornick/parsers
File parser registry for Dornick — lazy VFS sub-mounts under /user/files/<id>/parsed/*.
Each parser materialises text or structured artifacts from an uploaded file. They register with a ParserRegistry; reads against /user/files/<id>/parsed/<sub> look up the matching parser and run it on demand, caching the result. WASM dependencies inside companion packages (@dornick/parsers-pdf, @dornick/parsers-office, @dornick/parsers-ocr, @dornick/parsers-ml, @dornick/parsers-audio) load only when first accessed.
Install
npm i @dornick/parsersCapability-aware via a peer of @dornick/capabilities.
Built-in parsers
Three batteries-included parsers ship in this package:
| Parser | Matches | Produces |
|---|---|---|
| textParser | .txt, .md, .log, .yaml, anything text/* | text.txt |
| jsonParser | .json, application/json | pretty.json, summary.txt |
| csvParser | .csv, .tsv, text/csv | rows.jsonl, summary.txt |
The CSV parser handles RFC 4180-style quoting + escaped quotes + multi-line cells, without any runtime dependency.
Use
import { Dornick } from "@dornick/core";
import { ParserRegistry, BUILTIN_PARSERS } from "@dornick/parsers";
const parsers = new ParserRegistry(dornick.capabilities);
for (const p of BUILTIN_PARSERS) parsers.add(p);
// Pass via DornickOptions (planned wiring in @dornick/core)
new Dornick({ config, transport, parsers });The model then reads /user/files/<id>/parsed/rows.jsonl (CSV), /parsed/text.txt (plain text), etc. Companion parser packages slot in identically:
import { pdfTextParser } from "@dornick/parsers-pdf"; // future package
import { mammothParser, xlsxParser } from "@dornick/parsers-office";
parsers.add(pdfTextParser);
parsers.add(mammothParser);
parsers.add(xlsxParser);Define your own
import { defineSkill } from "@dornick/skills"; // unrelated — for context
import type { FileParser } from "@dornick/parsers";
export const exifParser: FileParser = {
name: "exif",
description: "Read EXIF metadata from JPEG/PNG uploads.",
requires: ["wasm"], // capability gate
matches: (file) => file.mimeType.startsWith("image/"),
paths: () => ["exif.json"],
parse: async (file, sub) => {
if (sub !== "exif.json") throw new Error(`exif does not produce ${sub}`);
const { extractExif } = await import("./_lazy/exif.js");
return JSON.stringify(await extractExif(file.base64), null, 2);
},
};Two rules:
- Be lazy — dynamic-import any heavy dependency inside
parse(), not at module load. - Declare your gate — list every required
BrowserCapabilitiesdot-path inrequires. The registry filters you out on browsers that can't run you.
Capability gating
ParserRegistry.lookup and ParserRegistry.paths consult the registry's CapabilityRegistry. A parser declaring requires: ["webgpu", "crossOriginIsolated"] simply doesn't appear in paths() output or read() resolution on a browser without WebGPU. The model never sees an affordance that can't run.
Links
License
MIT — © Dipankar Sarkar
