npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@tokcalc/mcp-server

v0.2.8

Published

tokcalc MCP server — open-source LLM serving capacity planner for AI agents. stdio + stateless HTTP transport with bearer API key auth + KV-backed rate limiting + public hosted endpoint. 11 read-only tools: estimate_capacity, compare_gpus, recommend_topol

Readme

@tokcalc/mcp-server

LLM serving capacity planner for AI agents.

Open-source MCP (Model Context Protocol) server that lets AI agents (Cursor, Claude Desktop, Cline) estimate LLM serving capacity — model fit, KV cache, throughput, latency, multi-GPU topology, and cost.

Tools

| Tool | What it does | |---|---| | estimate_capacity | VRAM/KV/throughput/latency/cost for one config | | compare_gpus | Ranked GPU comparison for one workload | | recommend_topology | TP/CP topology recommendation | | estimate_api_vs_self_host | Break-even analysis | | list_models | Discover supported model IDs (39 models) | | list_gpus | Discover supported GPU IDs (30 GPUs) | | get_mlperf_benchmarks | Curated MLPerf v4.1 reference configs | | find_config_for_slo | Inverse planner — SLOs + traffic → feasible configs ranked by cost/throughput/value | | plan_deployment | One-call decision brief — memory + perf + build-vs-buy + risks + next steps | | fetch_model_spec | Diff the catalog against live HuggingFace config.json (24h cache) | | record_measured | Store real tok/s measurements; future estimates self-calibrate |

All tools are read-only — no side effects, no cloud credentials, no deployments.

Install

Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "tokcalc": {
      "command": "npx",
      "args": ["-y", "@tokcalc/mcp-server"]
    }
  }
}

Restart Claude Desktop. The tokcalc server will be available as an MCP tool source.

Cursor

Add to .cursor/mcp.json in your project:

{
  "mcpServers": {
    "tokcalc": {
      "command": "npx",
      "args": ["-y", "@tokcalc/mcp-server"]
    }
  }
}

Cline (VS Code)

Add the same config to Cline's MCP settings.

Example prompts

Ask your AI agent:

"I need to serve Llama 3.3 70B at 32K context for 50 concurrent users. What GPU topology do you recommend, and how much will it cost per month?"

"Compare H100 vs H200 for serving Qwen 2.5 72B in FP8 with continuous batching."

"At what daily request volume does self-hosting Llama 70B on H200 beat the GPT-4o API?"

The agent calls list_models → list_gpus → recommend_topology → estimate_capacity and returns a structured plan with throughput ranges, latency, VRAM, cost, and confidence levels.

Supported models (35)

Llama 3/3.1/3.3, Llama 4 Scout/Maverick, Mistral 7B, Mixtral 8x7B/8x22B, Mistral Large 3, Pixtral 12B, Codestral, Qwen 2/2.5/3 (incl. MoE + VL), DeepSeek V3/R1/Coder V2, Gemma 2, Phi-3/4, SmolLM2, Falcon 3, OLMo 2, BGE-M3, E5, GTE.

Supported GPUs (30)

NVIDIA H100/H200/B200/B300, A100, L40S, L4, T4, V100, RTX 4090/3090/5090, RTX PRO 6000 Blackwell, AMD MI300X/MI325X, Intel Gaudi 3, Google TPU v5p/Trillium, Groq LPU, Cerebras CS-3, Apple M2/M3/M4 Ultra/Max.

License

Apache 2.0 — same as the main tokcalc project.

Links