npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@shen866/vision-bridge-mcp

v0.5.0

Published

Vision Bridge MCP server: describe, OCR, and ask about images via multi-provider LLM APIs.

Readme

Vision Bridge MCP

功能

提供 3 个 MCP 工具:

  • describe_image:详细描述图片内容(对象、场景、文字、人物动作、颜色、布局等)。
  • extract_image_text:提取图片中的所有文字(OCR)。
  • ask_about_image:针对图片回答具体问题。

支持四种图片输入方式:

  • image_url:图片 URL
  • image_path:本地图片路径
  • image_base64:Base64 编码的图片数据
  • imageMessages API 格式的图片内容块,例如:
{
  "type": "image",
  "source": {
    "type": "base64",
    "media_type": "image/png",
    "data": "iVBORw0KGgo..."
  }
}

支持的 API 格式(Provider)

| Provider | 说明 | 默认 Endpoint | |----------|------|---------------| | anthropic | Anthropic Messages API | /v1/messages | | openai | OpenAI Chat Completions API | /chat/completions | | gemini | Gemini Native generateContent API | /v1beta/models/{model}:generateContent |

通过 PROVIDER 环境变量切换 API 格式。

快速开始

全局安装

npm install -g @shen866/vision-bridge-mcp

本地开发

git clone https://github.com/shen866/vision-bridge-mcp.git
cd vision-bridge-mcp
npm install
npm run build

配置 Claude Code

API_KEYBASE_URLMODEL 为必填项。

Anthropic Messages

{
  "mcpServers": {
    "vision_bridge": {
      "command": "npx",
      "args": ["-y", "@shen866/vision-bridge-mcp"],
      "env": {
        "API_KEY": "your-api-key",
        "BASE_URL": "https://api.anthropic.com",
        "MODEL": "claude-3-5-sonnet-20241022"
      }
    }
  }
}

如果使用本地构建版本,将 command 改为 node 并把 args 改为绝对路径:

{
  "mcpServers": {
    "vision_bridge": {
      "command": "node",
      "args": ["/Users/shen/workspace/kimi-vision-mcp/dist/index.js"],
      "env": {
        "API_KEY": "your-api-key",
        "BASE_URL": "https://api.anthropic.com",
        "MODEL": "claude-3-5-sonnet-20241022"
      }
    }
  }
}

OpenAI Chat Completions

{
  "mcpServers": {
    "vision_bridge": {
      "command": "npx",
      "args": ["-y", "@shen866/vision-bridge-mcp"],
      "env": {
        "PROVIDER": "openai",
        "API_KEY": "your-api-key",
        "BASE_URL": "https://api.openai.com/v1",
        "MODEL": "gpt-4o"
      }
    }
  }
}

Gemini Native generateContent

{
  "mcpServers": {
    "vision_bridge": {
      "command": "npx",
      "args": ["-y", "@shen866/vision-bridge-mcp"],
      "env": {
        "PROVIDER": "gemini",
        "API_KEY": "your-api-key",
        "BASE_URL": "https://generativelanguage.googleapis.com",
        "MODEL": "gemini-1.5-pro-latest"
      }
    }
  }
}

使用

在 Claude Code / Claude Desktop 中发送图片或引用图片路径后,Claude 会自动调用 describe_image 等工具获取多模态模型对图片的文字描述,然后将描述交给纯文本模型继续处理。

如果 Claude Code 经常直接把图片发给模型导致报错,可以在项目根目录创建 .claude/CLAUDE.md

当前模型不支持图片输入。当用户发送图片时,必须调用 vision_bridge MCP 工具处理,不要直接传给模型。

环境变量

| 变量 | 说明 | 默认值 | |------|------|--------| | PROVIDER | API 格式:anthropicopenaigemini | anthropic | | API_KEY | API 密钥 | 必填 | | BASE_URL | API 基础地址 | 必填 | | ENDPOINT | API 端点路径 | 取决于 provider | | MODEL | 模型名称 | 必填 | | MAX_TOKENS | 最大输出 token 数 | 4096 | | API_VERSION | Messages API 版本 | 2023-06-01 | | AUTH_HEADER | Messages API 认证方式:x-api-keybearer | x-api-key |

协议

基于 MCP(Model Context Protocol)JSON-RPC 2.0,通过 stdio 进行通信。

401 排查

如果返回 401,通常是认证头不对。各 provider 默认认证方式如下:

  • anthropicx-api-key + anthropic-version
  • openaiAuthorization: Bearer
  • gemini:URL query parameter ?key=

如果 Messages API 网关需要 Bearer Token,可设置:

"AUTH_HEADER": "bearer"