npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

mineru-mcp

v1.2.0

Published

MinerU document-parsing cloud API as an MCP server and a CLI (mineru-cloud) — batch by URL, long documents in ≤200-page slices, 73% token reduction

Readme

mineru-mcp

MCP server for MinerU document parsing API — extract text, tables, and formulas from PDFs, DOCs, and images.

Features

  • VLM model — 90%+ accuracy for complex documents
  • Pipeline model — Fast processing for simple documents
  • Local file upload — Upload files from disk for batch parsing
  • Batch processing — Parse up to 200 documents at once
  • Download & rename — Extract markdown with original filenames
  • Page ranges — Extract specific pages only
  • Long documents — MinerU caps files at 200 pages; mineru_parse_long slices and mineru_merge_slices stitches
  • CLI twinmineru-cloud runs the same tools from a shell (no MCP context cost)
  • 109 language OCR support
  • Optimized for Claude Code — 73% token reduction vs alternatives

Tools

| Tool | Description | |------|-------------| | mineru_parse | Parse a document URL | | mineru_status | Check task progress, get download URL | | mineru_batch | Parse multiple URLs (max 200) | | mineru_batch_status | Get batch results with pagination | | mineru_upload_batch | Upload local files for batch parsing | | mineru_download_results | Download results as named markdown files | | mineru_parse_long | Document >200 pages: one batch of ≤200-page page_ranges slices | | mineru_merge_slices | Stitch a sliced batch into one {name}.md + {name}_content.json (page_idx re-based) + images/ |

Installation

Requires Node.js 18+ and a MinerU API key.

CLI Install (one-liner)

# Claude Code
claude mcp add mineru-mcp -e MINERU_API_KEY=your-api-key -- npx -y mineru-mcp

# Codex CLI (OpenAI)
codex mcp add mineru --env MINERU_API_KEY=your-api-key -- npx -y mineru-mcp

# Gemini CLI (Google)
gemini mcp add -e MINERU_API_KEY=your-api-key mineru npx -y mineru-mcp

Claude Desktop

Add to your claude_desktop_config.json:

| OS | Config path | |----|-------------| | macOS | ~/Library/Application Support/Claude/claude_desktop_config.json | | Windows | %APPDATA%\Claude\claude_desktop_config.json | | Linux | ~/.config/Claude/claude_desktop_config.json |

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

VS Code

Add to .vscode/mcp.json (workspace) or open Command Palette > MCP: Open User Configuration (global):

{
  "servers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Note: VS Code uses "servers" as the top-level key, not "mcpServers". Other VS Code forks (Trae, Void, PearAI, etc.) typically use this same format.

Cursor

Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (project):

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Windsurf

Add to ~/.codeium/windsurf/mcp_config.json (Windows: %USERPROFILE%\.codeium\windsurf\mcp_config.json):

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Cline

Open MCP Servers icon in Cline panel > Configure > Advanced MCP Settings, then add:

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Cherry Studio

In Settings > MCP Servers > Add Server, set Type to STDIO, Command to npx, Args to -y mineru-mcp, and add environment variable MINERU_API_KEY. Or paste in JSON/Code mode:

{
  "mineru": {
    "name": "MinerU",
    "command": "npx",
    "args": ["-y", "mineru-mcp"],
    "env": {
      "MINERU_API_KEY": "your-api-key"
    },
    "isActive": true
  }
}

Witsy

In Settings > MCP Servers, add a new server with Type: stdio, Command: npx, Args: -y mineru-mcp, and set environment variable MINERU_API_KEY to your API key.

Codex CLI (TOML config)

Alternatively, edit ~/.codex/config.toml directly:

[mcp_servers.mineru]
command = "npx"
args = ["-y", "mineru-mcp"]

[mcp_servers.mineru.env]
MINERU_API_KEY = "your-api-key"

Gemini CLI (JSON config)

Alternatively, edit ~/.gemini/settings.json directly:

{
  "mcpServers": {
    "mineru": {
      "command": "npx",
      "args": ["-y", "mineru-mcp"],
      "env": {
        "MINERU_API_KEY": "your-api-key"
      }
    }
  }
}

Windows

On Windows, npx requires a shell wrapper. Replace "command": "npx" with:

{
  "command": "cmd",
  "args": ["/c", "npx", "-y", "mineru-mcp"],
  "env": {
    "MINERU_API_KEY": "your-api-key"
  }
}

For CLI tools on Windows:

claude mcp add mineru-mcp -e MINERU_API_KEY=your-api-key -- cmd /c npx -y mineru-mcp
codex mcp add mineru --env MINERU_API_KEY=your-api-key -- cmd /c npx -y mineru-mcp

ChatGPT

ChatGPT only supports remote MCP servers over HTTPS — local stdio servers like this one are not directly supported. You would need to deploy behind a public URL with HTTP transport.

CLI: mineru-cloud

Every tool is also a shell command — the CLI runs the MCP server in-process over an in-memory transport, so the two can't drift. Same env vars (MINERU_API_KEY, MINERU_BASE_URL, MINERU_DEFAULT_MODEL).

mineru-cloud list                                    # commands + options (from the tool schemas)
mineru-cloud parse --url https://arxiv.org/pdf/2303.08774 --pages 1-10
mineru-cloud status --task-id <id> --wait            # --wait polls every 10s until done/failed
mineru-cloud batch --urls '["https://…/a.pdf","https://…/b.pdf"]'
mineru-cloud download-results --batch-id <id> --output-dir ./papers --wait

# > 200 pages: slice, then stitch
mineru-cloud parse-long --url https://…/book.pdf --total-pages 520 --name book
mineru-cloud merge-slices --batch-id <id> --output-dir ./books --wait

Options mirror the tool parameters with _- (--total-pages, --output-dir); numbers, true/false and JSON arrays are coerced. Install: bun add -g mineru-mcp (or npm i -g).

Configuration

| Environment Variable | Default | Description | |---------------------|---------|-------------| | MINERU_API_KEY | (required) | Your MinerU API Bearer token | | MINERU_BASE_URL | https://mineru.net/api/v4 | API base URL | | MINERU_DEFAULT_MODEL | pipeline | Default model: pipeline or vlm |

Get your API key at mineru.net

Usage

Parse a single URL

mineru_parse({
  url: "https://example.com/document.pdf",
  model: "vlm",        // optional: "pipeline" (default) or "vlm" (90% accuracy)
  pages: "1-10,15",    // optional: page ranges
  ocr: true,           // optional: enable OCR (pipeline only)
  formula: true,       // optional: formula recognition
  table: true,         // optional: table recognition
  language: "en",      // optional: language code
  formats: ["html"]    // optional: extra export formats
})

Check task progress

mineru_status({
  task_id: "abc-123",
  format: "concise"    // optional: "concise" (default) or "detailed"
})

Concise output: done | abc-123 | https://cdn-mineru.../result.zip

Batch parse URLs

mineru_batch({
  urls: ["https://example.com/doc1.pdf", "https://example.com/doc2.pdf"],
  model: "vlm"
})

Check batch progress

mineru_batch_status({
  batch_id: "batch-123",
  limit: 10,           // optional: max results (default: 10)
  offset: 0,           // optional: skip first N results
  format: "concise"    // optional: "concise" or "detailed"
})

Upload local files

mineru_upload_batch({
  directory: "/path/to/pdfs",  // scan directory for supported files
  // OR
  files: ["/path/to/doc1.pdf", "/path/to/doc2.pdf"],  // explicit file list
  model: "vlm",        // optional
  formula: true,       // optional
  table: true,         // optional
  language: "en",      // optional
  formats: ["html"]    // optional
})

Returns batch_id for tracking. Each file's original name is preserved via data_id (spaces become underscores).

Download results as markdown

mineru_download_results({
  batch_id: "batch-123",       // from mineru_upload_batch or mineru_batch
  output_dir: "/path/to/output",
  overwrite: false             // optional: overwrite existing files
})

Output filenames are derived from data_id (e.g., my_paper_title.md). Spaces in original filenames become underscores.

Typical local file workflow

mineru_upload_batch → mineru_batch_status (poll) → mineru_download_results

Supported Formats

  • PDF, DOC, DOCX, PPT, PPTX
  • PNG, JPG, JPEG

Limits

  • Single file: 200MB max, 200 pages max (use pages to parse a longer file in ≤200-page slices — verified 2026-09-16)
  • Daily quota: 1000 pages at high priority (excess is deprioritized, not rejected)
  • Batch: max 200 files per request

Release 1.1.6

Restores Node.js 18 HTTP compatibility for fresh installs by retaining MCP SDK 1.29.x and its Node 18-compatible Hono adapter. SDK 1.30 permits an adapter that requires Node.js 20. Version 1.1.5 passed the locked dependency checks but the published-package check exposed an HTTP initialization failure on a fresh install. CI now installs the packed package without the repository lock and exercises both transports on Node.js 18. The SDK compatibility bound is intentional; revisit it with this consumer-install gate before adopting a newer SDK.

Release 1.1.5

Maintenance release: audited dependency updates, Express 5 and Zod 4 compatibility, and regression coverage for both transports. The MCP handshake and HTTP startup message now report the package version instead of the stale 1.0.2 value. Tool inputs and document-processing behavior are unchanged.

Development

Use Bun 1.4.2 and Node.js 24 for the build and CI checks:

bun install --frozen-lockfile
bun audit
bun run build
bun run test
bun run test:package

The runtime tests exercise the built stdio and HTTP servers against a local MinerU API double. They check tool schemas, request mapping, pagination defaults, provider errors, malformed HTTP requests, and session termination without real credentials or API calls. They do not verify live parsing or file extraction. Dependabot updates the Bun manifest and lockfile together. CI audits dependencies and runs the build and runtime tests before publishing on version tags.

Publishing

Bump package.json and both version fields in server.json, complete the checks above, merge, then push the matching vX.Y.Z tag. CI publishes to npm, waits for the exact package version to become available, then registers it with the MCP Registry. If registry registration fails after npm succeeds, retry only registration using the existing immutable tag:

gh workflow run publish-mcp.yml --ref main -f registry_tag=v1.1.6

License

MIT

Links