npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@paxalabs/mcp

v0.1.0-beta.12

Published

Official MCP server for the Paxa Labs API: Thai, English, and Mandarin text to speech with local playback, speech to text with subtitles, translation to Thai, document OCR, and field extraction.

Readme

@paxalabs/mcp

npm CI node license Claude Desktop bundle

Add to Cursor Install in VS Code

Official MCP server for the Paxa Labs API: Thai, English, and Mandarin Chinese speech AI for your agent, including local audio playback.

An agent connected to this server can speak out loud through your machine's speakers, read long content aloud as a managed playback queue, save speech to audio files, listen for a spoken reply, transcribe recordings with subtitles, translate any language into Thai, read PDFs and images with OCR, and extract typed fields from documents.

Beta. The tool set is complete and tested end to end, but tool names and behavior may still change before 1.0 as feedback comes in. Report problems at https://github.com/paxalabs/mcp/issues.

Quick start

You need a Paxa API key from paxalabs.com. New accounts include free credits.

Claude Code

claude mcp add paxa -e PAXA_API_KEY=pxa_your_key_here -- npx -y @paxalabs/mcp

Claude Desktop, Cursor, and other MCP clients

Add to your client's MCP configuration (for Claude Desktop: claude_desktop_config.json):

{
  "mcpServers": {
    "paxa": {
      "command": "npx",
      "args": ["-y", "@paxalabs/mcp"],
      "env": {
        "PAXA_API_KEY": "pxa_your_key_here"
      }
    }
  }
}

One click for Cursor or VS Code: the buttons install the same entry into ~/.cursor/mcp.json or VS Code's MCP settings. Then replace pxa_your_key_here in the paxa entry with your key.

Add to Cursor Install in VS Code

Tools

| Tool | What it does | Credits | | --- | --- | --- | | speak | Synthesize a short line and play it through the speakers, blocking until done | 10 per 1000 chars | | queue_speech | Read long content aloud: auto-chunks, synthesizes ahead while playing, returns immediately | 10 per 1000 chars | | control_playback | Control the shared audio queue: status, pause, resume, skip, clear | free | | play_audio | Play a local audio file through the speakers | free | | text_to_speech | Synthesize speech to an audio file (mp3, opus, wav) without playing it | 10 per 1000 chars | | translate_to_thai | Translate any language into Thai, with formality, glossary, and context controls | 25 per 1000 chars | | ocr_document | OCR a local PDF, PNG, JPEG, or WebP into Markdown or structured blocks | 6.5 per page | | extract_fields | Fill a schema of typed fields (Thai IDs, dates, amounts, banks, line items) from a local PDF or image; every value is printed in the document or null with the reason | 13 per page (19.5 for schemas over 50 fields) | | listen | Record one spoken turn from the microphone and return its transcript; the API detects when the user stops talking | 12.5 per minute listened | | send_feedback | Rate an output or report a problem to the Paxa team, tied to the request ids the tools print | free | | transcribe_audio | Transcribe a local recording (Thai, English, mixed) to text, with optional speaker labels, word timings, and srt or vtt subtitles saved next to it | 8.33 per minute | | list_voices | The TTS voice roster with character notes | free | | list_models | Available models, limits, and pricing | free | | get_account | Credit balance, plan, and rate limits | free |

All audio flows through one ordered queue, so sounds never overlap: speak lines slip in ahead of queued long-form segments, and queue_speech keeps a book or article flowing gap-free by synthesizing the next segment while the current one plays.

With a streaming-capable player installed (see below), speech starts on the first bytes from the API instead of after the full download: about 0.3 s to the first word regardless of length, versus 0.7 s for a short line and 2.5 s for a long paragraph when buffered.

Voice mode for Claude Code

Three pieces turn Claude Code into something you can walk away from: it talks when it has news, and it calls you when it needs you.

1. Install the server (Quick start above).

2. Tell Claude when to talk. Add this to ~/.claude/CLAUDE.md, or to one project's CLAUDE.md:

## Voice

I have the Paxa MCP server (tools: speak, queue_speech, control_playback).
I am often away from the screen, so use voice like this:

- At the end of a turn where you did real work, call speak with a one or
  two sentence summary before writing the final message: what you did,
  what is next, and anything you need from me.
- When you need a decision from me, speak the question too.
- Keep it short and conversational. Never read code, file paths, logs, or
  long lists aloud. Those stay in text.
- Do not speak for quick back-and-forth or trivial answers.
- If I ask to hear something long, use queue_speech.
- Speak in the language I write in.

3. Get told when Claude needs you. When Claude Code waits for a permission or an answer, the model is not running, so it cannot call speak. Claude Code fires a hook at those moments instead, and paxa say turns the hook into a spoken phrase such as "Permission needed." Put the paxa command on your PATH:

npm install -g @paxalabs/mcp

Then add to ~/.claude/settings.json:

{
  "hooks": {
    "Notification": [
      {
        "matcher": "permission_prompt|idle_prompt|agent_needs_input",
        "hooks": [{ "type": "command", "command": "paxa say" }]
      }
    ]
  }
}

paxa say takes the key from PAXA_API_KEY, or from the paxa entry in ~/.claude.json when that is unset, so step 1 is all the setup it needs. A Stop hook configured the same way speaks "Done." at the end of every turn.

The built-in phrases are synthesized once per voice and kept in your user cache directory (~/Library/Caches/paxa/say on macOS, ~/.cache/paxa/say on Linux, %LOCALAPPDATA%\paxa\cache\say on Windows). After that first play, which costs well under one credit, a notification plays from disk: no network round trip and no credits.

To change the words, write your own text in the hook command and add --cache so it gets the same treatment. One entry per event, since the matcher picks the event:

{
  "hooks": {
    "Notification": [
      {
        "matcher": "permission_prompt",
        "hooks": [{ "type": "command", "command": "paxa say --cache \"Hey, need your OK\"" }]
      },
      {
        "matcher": "idle_prompt|agent_needs_input",
        "hooks": [{ "type": "command", "command": "paxa say --cache --voice cookie \"Your turn\"" }]
      }
    ]
  }
}

Without --cache, nothing you type or pipe into paxa say is written to disk, and messages carried inside a hook payload never are.

On macOS the built-in afplay needs about half a second just to start and stop, which is most of the delay you hear on a short phrase. With brew install mpg123 (or ffmpeg) installed, cached phrases play through that instead: mpg123 starts in about 50 ms, ffplay in about 300 ms.

paxa say also works on its own:

paxa say "Build finished"
paxa say --voice cookie "Deploy is live"

If your editor or desktop app was not launched from a terminal, its PATH may not include your node bin directory, and the hook will fail silently. Use the absolute path to paxa in the hook command if that happens.

Listening

listen turns the microphone into an answer. Ask a question with speak, call listen, and the user's spoken reply comes back as text as soon as they stop talking: Paxa's realtime transcription detects the end of the turn, so nothing has to guess at silence. The audio goes straight to the API and is never written to disk. A few seconds of listening costs well under one credit (12.5 credits per minute, silence included).

It needs a command-line recorder, which stock macOS and Windows do not ship: brew install ffmpeg or brew install sox on macOS; sox on Windows (or ffmpeg with PAXA_MIC set to the DirectShow device name); ffmpeg, sox, parecord, or arecord on Linux. It also needs Node 22 or newer, and the app that launched the server (Terminal, VS Code, Claude Desktop) must have microphone permission. PAXA_MIC picks a device when the system default is not the one you want. To check recognition without a microphone, node scripts/listen-file.mjs recording.mp3 feeds a file through the same session.

Environment variables

| Variable | Required | Default | Purpose | | --- | --- | --- | --- | | PAXA_API_KEY | yes | | Your API key. The server starts without it, but every tool that calls the API then fails with setup instructions the agent can relay | | PAXA_OUTPUT_DIR | no | working directory | Where text_to_speech saves files | | PAXA_DEFAULT_VOICE | no | nomyen | Voice used when a tool call does not pick one. Every voice is designed around one language. English text usually sounds best with an English voice (donut, cookie, toast, latte, espresso, mocha), Mandarin with taohuay or oolong | | PAXA_VOCABULARY | no | | Keyword pinning for transcribe_audio: comma-separated names and terms the transcript should spell as written (product names, people, jargon) | | PAXA_MIC | no | system default | Microphone device for listen: an AVFoundation device name on macOS, a PulseAudio or ALSA device on Linux, a DirectShow device name on Windows | | PAXA_VOCABULARY_FILE | no | | A text file with one term per line (# starts a comment), also pinned on every transcription. Read at call time, so edits apply without a restart | | PAXA_BASE_URL | no | https://api.paxalabs.com | API origin override |

To pin a different list per project, set the vocabulary variables in a project-scope server entry (Claude Code's project scope, .cursor/mcp.json, .vscode/mcp.json). On each call the tool takes the call's own terms first, then the configured ones, deduplicated and cut at the API's limit of 50, and reports how many it pinned.

Playback support

| Platform | File player | Streaming player | Pause/resume | | --- | --- | --- | --- | | macOS | afplay (built in) | mpg123, ffplay, or mpv if installed | yes | | Linux | ffplay, mpv, mpg123, paplay, or aplay | mpg123, ffplay, or mpv | yes | | Windows | ffplay if installed, else PowerShell (wav) | mpg123, ffplay, or mpv if installed | no |

Streaming needs a player that reads from stdin. On macOS, brew install mpg123 (the quickest to start) or ffmpeg enables it; without one, speech still plays through afplay after the download completes. If no player is found at all, speech tools report it clearly and text_to_speech still works.

Claude Desktop extension

Each release on GitHub ships a .mcpb bundle. Download it, open it with Claude Desktop, and enter your API key in the extension settings. The bundle carries its own copy of the server and its dependencies, so it works without Node.js or npm on the machine.

Development

pnpm install
pnpm build        # compile to dist/
pnpm typecheck
pnpm mcpb         # build release/paxalabs-mcp-<version>.mcpb for Claude Desktop

# live smoke test (spends a few credits, plays audio out loud)
PAXA_API_KEY=pxa_... TEST_OUT_DIR=/tmp/paxa-out node scripts/e2e.mjs