npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

files-to-markdown

v0.1.1

Published

Convert uploaded files (PDF, Word, Excel, HTML, CSV, images) to Markdown to cut LLM token usage. Node-native, no Python.

Readme

files-to-markdown

Convert uploaded files — PDF, Word, Excel, CSV, HTML, JSON, text, and images — into clean Markdown so you can feed text to an LLM instead of raw files. That cuts token usage dramatically (and lets you cache the markdown, e.g. in S3, so repeat turns never re-upload the file).

Node-native. No Python, no markitdown runtime, no extra service.

Install

npm install files-to-markdown

# Optional — only if you need .xlsx/.xls Excel support:
npm install xlsx

Excel uses SheetJS (xlsx), kept as an optional peer dependency because its npm build carries an unfixed advisory. Everything else (PDF, Word, CSV, HTML, JSON, text, images) works without it.

Usage

import { fileToMarkdown } from 'files-to-markdown';

// PDF / Word / Excel / CSV / HTML / JSON / text — auto-detected:
const { markdown, kind, truncated } = await fileToMarkdown({
  buffer,                 // Buffer of the uploaded file
  filename: 'report.pdf', // or pass mimeType
  maxChars: 200_000,      // optional safety cap
});

Images need a vision function

The library is model-agnostic: images are turned into markdown by your OCR / vision call (Gemini, Claude, Tesseract — your choice). Without it, image inputs throw, so the token-free types still work on their own.

import { fileToMarkdown, VisionInput } from 'files-to-markdown';

async function geminiOcr({ buffer, mimeType }: VisionInput): Promise<string> {
  // call your vision model, return markdown (e.g. transcribed text / a table)
  return '# Invoice\n\n| Item | Qty | Price |\n| --- | --- | --- |\n| Pen | 2 | 20 |';
}

const { markdown } = await fileToMarkdown({
  buffer,
  mimeType: 'image/png',
  vision: geminiOcr,
});

Suggested flow (token-saving)

  1. On upload, fileToMarkdown(...) → markdown.
  2. Store the markdown in S3 (keyed by file hash).
  3. On every chat turn, pass the markdown text to the model — never the raw file again.

API

function fileToMarkdown(input: ConvertInput): Promise<ConvertResult>;
function detectKind(input: { mimeType?: string; filename?: string }): ConvertKind;

interface ConvertInput {
  buffer: Buffer;
  filename?: string;
  mimeType?: string;
  vision?: (input: VisionInput) => Promise<string>; // images only
  maxChars?: number;
}

interface ConvertResult {
  markdown: string;
  kind: 'pdf' | 'docx' | 'xlsx' | 'csv' | 'html' | 'image' | 'json' | 'text';
  bytes: number;
  truncated: boolean;
}

License

MIT