npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@shiftlabs/markit

v0.6.1

Published

Convert anything to markdown. PDF, DOCX, PPTX, XLSX, HTML, EPUB, Jupyter, RSS, images, audio, URLs, and more. Works as a CLI and as a library.

Readme

markit

Convert anything to markdown. PDF, DOCX, PPTX, XLSX, HTML, EPUB, Jupyter, RSS, images, audio, URLs, and more. Works as a CLI and as a library.

npm install -g @shiftlabs/markit

Quick Start

# Documents
markit report.pdf
markit document.docx
markit slides.pptx

# Data
markit data.csv
markit config.json
markit schema.yaml

# Web
markit https://example.com/article
markit https://en.wikipedia.org/wiki/Markdown

# Media
markit photo.jpg                          # EXIF metadata
markit recording.mp3                      # Audio metadata

# Text only, skipping image extraction
markit report.pdf --no-images

# Write to file
markit report.pdf -o report.md

# Pipe it
markit report.pdf | pbcopy
markit data.xlsx -q | napkin create "Imported Data"

Supported Formats

| Format | Extensions | How | |--------|-----------|-----| | PDF | .pdf | Built-in Rust engine | | Word | .docx | Built-in Rust engine; headings, images, and tables | | PowerPoint | .pptx | Built-in Rust engine; slides, notes, and tables | | Excel | .xlsx | Built-in Rust engine; sheets become markdown tables | | HTML | .html .htm | Built-in Rust engine; scripts/styles stripped | | EPUB | .epub | Built-in Rust engine; spine-ordered chapters | | Jupyter | .ipynb | Built-in Rust engine; cells and outputs | | RSS/Atom | .rss .atom .xml | Built-in Rust engine; dated feed items | | CSV/TSV | .csv .tsv | Built-in Rust engine; markdown tables | | JSON | .json | Built-in Rust engine; formatted code block | | YAML | .yaml .yml | Built-in Rust engine; code block | | XML/SVG | .xml .svg | Built-in Rust engine; code block | | Images | .jpg .png .gif .webp | Built-in Rust metadata extraction | | Audio | .mp3 .wav .m4a .flac | Built-in Rust metadata extraction | | ZIP | .zip | Built-in Rust recursive conversion | | URLs | http:// https:// | Rust fetcher with markdown negotiation | | Wikipedia | *.wikipedia.org | Built-in Rust main-content extraction | | Code | .py .ts .go .rs ... | Built-in Rust fenced code block | | Plain text | .txt .md .rst .log | Built-in Rust pass-through |


Benchmarks

Non-OCR PDF to Markdown: parsers that read the text layer already inside a PDF, with no OCR and no model calls. Measured on one machine, with image extraction off for every parser.

| | markit | liteparse | anydoc | |---|---:|---:|---:| | olmOCR-bench score | 45.0% | 38.8% | 31.2% | | shitty-pdf-bench text recovered | 100.0% | 83.2% | 91.0% | | shitty-pdf-bench conversion time | 23.2s | 55.4s | 322.2s |

olmOCR-bench is AllenAI's public benchmark. shitty-pdf-bench is ours: 40 hash-pinned public PDFs, 50,884 pages of semiconductor manuals, standards, RTL documents, scans, and malformed files.

Method, provenance, and per-category results are in rust/BENCHMARKS.md; the harness that reproduces them is in benchmark/.


For Agents

Every command supports --json. Raw markdown with -q.

markit report.pdf --json       # Structured output for parsing
markit report.pdf -q           # Raw markdown, nothing else
markit onboard                 # Add instructions to CLAUDE.md

SDK

markit is also a Node library. Every format uses the same bundled Rust engine as the CLI on supported macOS and Linux platforms—there is no second fallback implementation:

import { Markit } from "@shiftlabs/markit";

const markit = new Markit();
const { markdown } = await markit.convertFile("report.pdf");
const { markdown } = await markit.convertUrl("https://example.com");
const { markdown } = await markit.convert(buffer, { extension: ".docx" });

CLI Reference

markit <source>                          # Convert file or URL
markit <source> -o output.md             # Write to file
markit <source> --json                   # JSON output
markit <source> -q                       # Raw markdown only
markit <source> --no-images              # Skip image extraction
markit <source> -i ./images              # Extract images to a directory
cat file.pdf | markit -                  # Read from stdin
markit formats                           # List supported formats
markit onboard                           # Add to CLAUDE.md

Development

bun install
bun run build:native
bun run dev -- report.pdf
bun test
bun run check
cd rust && cargo test

Distribution

Releases publish @shiftlabs/markit plus six scoped native binary packages for macOS and Linux (x64/ARM64, glibc/musl). npm selects the matching binary through optionalDependencies. The former markit-ai package is a deprecated compatibility shim that re-exports this SDK.

# 1. Update the same version in package.json, rust/Cargo.toml,
#    and npm/*/package.json.
# 2. Verify locally.
bun run verify

# 3. Tag and push. GitHub Actions builds and publishes all seven packages.
git tag v0.6.0
git push origin v0.6.0

The release workflow requires an npm automation token in the repository secret NPM_TOKEN. It publishes platform packages first through napi prepublish, then @shiftlabs/markit, then the deprecated markit-ai shim, and finally creates the GitHub release.

License

MIT