npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@jamiedavenport/htomd

v0.1.1

Published

Focused Markdown and metadata from messy HTML, with no runtime dependencies

Readme

htomd

Extract Markdown and metadata from decoded HTML, with no runtime dependencies or network access. Requires Node 24 or later; the package is ESM.

npm install @jamiedavenport/htomd
import { convert, extract, type Document } from "@jamiedavenport/htomd";

const html = "<article><h1>Hello</h1><p>Readable text.</p></article>";
console.log(convert(html));

const document: Document = extract(html, {
  url: "https://example.com/article",
});
console.log(document.metadata.title);

extract(html, options?) returns { markdown, metadata, diagnostics }. convert(html, options?) returns its Markdown string. Both functions are synchronous. HTML must be a primitive string; invalid HTML or URL argument types throw TypeError. The optional url accepts a string, null, or undefined. It provides source context and resolves relative references without fetching. An empty URL remains an empty string in metadata.

Metadata fields are title, author, description, language, publishedTime, url, and canonicalUrl. Missing values are null. Diagnostics contain a strategy (semantic, scored, fallback, or none) and a notes array. Results, nested metadata, diagnostics, and notes are frozen at runtime and readonly in TypeScript. Nonempty Markdown ends with a newline; empty content produces "".

CLI

cat page.html | npx @jamiedavenport/htomd convert > page.md
cat page.html | npx @jamiedavenport/htomd extract > page.json
cat page.html | npx @jamiedavenport/htomd convert --url https://example.com/article

Both commands read strict UTF-8 from stdin, stripping an initial BOM. convert writes Markdown and extract writes indented JSON. JSON retains the Python wire keys published_time and canonical_url. Use htomd --help, htomd help extract, or htomd --version after installing the CLI. Help and version do not read stdin. Invalid arguments exit with status 2; input/output errors exit with status 1. A broken output pipe exits quietly with status 1.

Behavior and limitations

The maintained implementation follows the Python reference: an ordered HTML tree, explicit metadata extraction, visibility and clutter filtering, content scoring and sibling recovery, then iterative Markdown serialization. Selected headings and local author/date information refine the metadata. Deep trees do not depend on the JavaScript call stack.

Extraction is best-effort, intended for articles and documentation. It does not execute JavaScript or implement browser layout. Simple tables use GFM; complex tables become row/cell text. Active and unknown URL schemes are omitted.

The shared conformance cases record runtime differences explicitly:

  • Whitespace normalization uses Unicode White_Space, with JavaScript trimming. Python-only control separators are not treated as ordinary whitespace; BOM characters at string boundaries are trimmed.
  • List counters use bigint, preserving large decimal values. List attributes accept ASCII decimal digits; Python's Unicode digits and underscore syntax are not accepted.
  • Relative references use Node's WHATWG URL, which may normalize host spelling, default ports, and paths differently from Python's URL joining. Absolute references retain their supplied spelling after validation.
  • JSON-LD uses JSON.parse, so nonstandard NaN and Infinity tokens are rejected. These words inside valid JSON strings are preserved.

HTML entity names come from the WHATWG table; see NOTICE for its source and attribution. Package code uses the repository's MIT license.

Development

From the repository root, mise run setup, mise run build, and mise run check install dependencies, build both packages, and validate the existing artifacts. Within this directory:

bun install --frozen-lockfile
bun run build
bun run lint
bun run format:check
bun run typecheck
bun run test

The compiler builds JavaScript and declarations into dist/ and checks source types. The separate typecheck command checks tests against those declarations. Tests run on Node and read the root fixtures without making package-local copies. bun pm pack builds once through prepack; test and package-check commands never rebuild. Root validation checks the tarball in an isolated consumer and compares both installed CLIs on synthetic, edge, and saved-page fixtures.