npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@watermelonpm/greennode-rag-mcp

v0.1.3

Published

MCP server exposing GreenNode RAG APIs (knowledge bases, documents, search, ingest) over stdio or streamable HTTP.

Readme

greennode-rag-mcp

An MCP server that exposes the GreenNode RAG REST APIs (knowledge bases, documents, search, ingest) as 19 tools. It proxies agent-platform-api via its public gateway with pass-through OAuth bearer auth and optional engine (agent name) scoping. Runs locally over stdio (default) or remotely over streamable HTTP, with any MCP-speaking client.

Table of contents

Quick start

Connect a local MCP client (Claude Code, Cursor, Windsurf, …) to the server over stdio in under a minute.

Prerequisites

  • Node.js ≥ 20 (see package.json engines)
  • An OAuth bearer token valid against the platform. BACKEND_URL is optional — it defaults to prod.

1. Install — from npm (no clone needed):

npm install -g @watermelonpm/greennode-rag-mcp

…or run one-off with npx -y @watermelonpm/greennode-rag-mcp.

2. Run (stdio is the default transport — no need to set TRANSPORT)

GREENNODE_RAG_TOKEN=<your-token> \
greennode-rag-mcp          # global install; or: npx -y @watermelonpm/greennode-rag-mcp

Uses the prod backend by default. Set BACKEND_URL=https://aiplatform.console-dev.vngcloud.tech/agent-api to use dev.

3. Wire up your client. Claude Code — .mcp.json:

{
  "mcpServers": {
    "greennode-rag": {
      "command": "npx",
      "args": ["-y", "@watermelonpm/greennode-rag-mcp"],
      "env": {
        "GREENNODE_RAG_TOKEN": "<your-token>"
      }
    }
  }
}

Add "BACKEND_URL": "https://aiplatform.console-dev.vngcloud.tech/agent-api" to env to use the dev environment; it defaults to prod.

From source (development): clone the repo, npm ci, then use "args": ["tsx", "src/index.ts"] with "cwd": "<repo path>" in the .mcp.json above, or npm run dev.

Add "ENGINE": "<agent name>" to env to scope search and list_knowledge_bases to that engine's attached knowledge bases; omit it to use every KB in the account. For Cursor, Windsurf, Cline, Roo Code, Claude Desktop, and other clients, use the same command + env under each client's own config key.

First call flow: list_knowledge_bases → list_documents → search (see How it works).

How it works

The server exposes 19 tools that map onto the agent-platform-api RAG endpoints. Auth is pass-through: the MCP server forwards the caller's OAuth bearer to the gateway and never handles portal-user-id — the gateway validates the token and injects ownership. When an engine (agent name) is set, the server resolves it to KB ids via GET /agents?searchName= and scopes search / list_knowledge_bases to those KBs.

┌───────────────────────────────────────────────────────────────┐
│  MCP client (Claude Code, Cursor, …)                          │
│    stdio JSON-RPC  ·or·  streamable HTTP (POST /mcp)          │
└───────────────────────────────────────────────────────────────┘
                          ▼
┌───────────────────────────────────────────────────────────────┐
│  greennode-rag-mcp  (hand-written TypeScript)                 │
│    • 19 tools: search, ingest_*, documents, knowledge_bases   │
│    • inbound auth: env token (stdio) / Authorization header   │
│    • engine scoping: ENGINE env / X-Engine header → KB ids    │
│    • list-response truncation (MAX_RESPONSE_BYTES)            │
└───────────────────────────────────────────────────────────────┘
                          ▼  pass-through OAuth bearer
┌───────────────────────────────────────────────────────────────┐
│  agent-platform-api gateway  (BACKEND_URL)                    │
│    validates bearer, injects ownership — MCP never sees it    │
└───────────────────────────────────────────────────────────────┘

Tools (19)

| Tool | Key args | Notes | |---|---|---| | search | question, filters? | Semantic search over in-scope KB(s); returns chunks {content, documentId, similarity} | | ingest_document | kbId, filename, content ∣ data (base64), mimeType? | One file; async — poll get_ingest_status | | ingest_batch | kbId, documents[] | Multiple files in one call | | ingest_file | kbId, path, filename?, mimeType? | Read one local file by path → multipart (no base64); stdio local | | ingest_files | kbId, files[] | Multiple files by path in one call | | get_ingest_status | kbId, documentId? | Poll KB + document ingest status | | list_documents | kbId, page?, size? | Paginated | | get_document | kbId, documentId, maxPages? | Lists client-side; bounded by maxPages | | restart_document | kbId, documentId | Restart (re-parse) a document; returns {jobIds} | | cancel_document | kbId, documentId | Cancel an in-flight document parse | | download_document | kbId, documentId | Transport-aware: stdio writes to disk, http returns base64 | | update_document_metadata | kbId, documentId, metadata[] | Update a document's metadata ({key, value, type?}) | | delete_document | kbId, documentIds[] | Batch delete | | list_knowledge_bases | page?, size?, searchName? | When engine set, only the engine's KBs | | create_knowledge_base | name, description, embeddingModel, llmModel?, … | llmModel optional; valid values from list_models(type=chat) | | update_knowledge_base | kbId, description? | Update a knowledge base's description | | get_knowledge_base | kbId | — | | delete_knowledge_base | kbId | Fails if agents still use it | | list_models | type? | List active embedding/chat models (valid embeddingModel/llmModel for create_knowledge_base) |

// 1) orient on the available knowledge bases
list_knowledge_bases()
// → [{ "id": "kb-1", "name": "docs", … }, …]

// 2) see what's inside one
list_documents({ kbId: "kb-1" })
// → [{ "id": "d-1", "name": "handbook.pdf", "status": "ACTIVE", … }, …]

// 3) ask a question over the in-scope KB(s)
search({ question: "how do I rotate a token?" })
// → [{ "content": "…", "documentId": "d-1", "similarity": 0.83 }, …]

Ingest is async — ingest_document / ingest_batch return immediately; pair them with get_ingest_status to poll until ACTIVE.

How to upload a file is spelled out in the tool description itself, and it differs by transport so the agent never has to guess:

  • stdio, path-based (preferred for local files): call ingest_file({ kbId, path }) — the server reads the file from disk and uploads it as multipart. No base64, no inlining; works for large files up to MAX_INGEST_FILE_BYTES.
  • stdio, inline: read the file from disk → base64-encode → pass as data with mimeType (or content for plain text).
  • streamable HTTP (server is remote, can't read your disk): you must read and inline the file yourself. Check the size first — base64 is ~33% larger than the file. If it's large (roughly ≥ 50 KB), stop and don't inline it; tell the user to run the MCP locally over stdio instead.

The filename / content / data / mimeType fields are documented inline in the tool's inputSchema.

Transports: stdio vs. streamable HTTP

| | stdio | streamable HTTP | |---|---|---| | Use case | local, any MCP client | deployed runtime / remote clients | | Default | yes (TRANSPORT=stdio) | opt-in (TRANSPORT=http) | | Lifecycle | one server for the process lifetime | fresh server + transport per request (stateless) | | Token source | env var named by TOKEN_ENV (default GREENNODE_RAG_TOKEN) | Authorization: Bearer header, per request | | Engine source | ENGINE env var | X-Engine header, per request | | Endpoint | stdin/stdout (JSON-RPC) | POST /mcp | | Health | — | GET /healthz, GET /health |

stdio (default)

The server reads JSON-RPC from stdin and writes responses to stdout. stdout is the protocol — all diagnostics and the one-line startup banner go to stderr, so they never corrupt the stream.

BACKEND_URL=https://aiplatform.console-dev.vngcloud.tech/agent-api \
GREENNODE_RAG_TOKEN=<your-token> \
npm start                     # TRANSPORT=stdio is the default

The token is read once at startup from the env var named by TOKEN_ENV (default GREENNODE_RAG_TOKEN). The server runs for the process lifetime and exits when the client closes stdin. See Quick start for the client-wiring snippet.

Streamable HTTP

For a deployed runtime or remote clients. Each POST /mcp builds a fresh server + StreamableHTTPServerTransport for that request (stateless) and authenticates from the Authorization header. The token is not read from the environment in this mode.

TRANSPORT=http \
BACKEND_URL=https://agent-rag.api.vngcloud.vn \
npm start                     # listens on :8080 (PORT); pass the token per request, not via env

Smoke-test it:

curl http://localhost:8080/healthz          # → {"ok":true}

# a raw initialize request to /mcp (clients normally build this JSON-RPC envelope for you)
curl -X POST http://localhost:8080/mcp \
  -H "Authorization: Bearer <your-token>" \
  -H "Content-Type: application/json" \
  -H "Accept: application/json, text/event-stream" \
  -d '{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-06-18","capabilities":{},"clientInfo":{"name":"curl","version":"0"}}}'

To scope to an engine in HTTP mode, add -H "X-Engine: <agent name>".

Configuration

All config is via environment variables, read once at startup by loadEnvConfig (src/config/env.ts). No dotenv — set vars in the shell, or for dev use Node 20's built-in --env-file: node --env-file=.env --import tsx src/index.ts.

| Var | Default | Notes | |---|---|---| | BACKEND_URL | https://agent-rag.api.vngcloud.vn | Backend base URL. Optional — defaults to prod. Set to https://aiplatform.console-dev.vngcloud.tech/agent-api for dev. | | TRANSPORT | stdio | stdio or http. Any other value throws at boot — the process exits non-zero, nothing listens. | | GREENNODE_RAG_TOKEN | — | Upstream OAuth bearer, stdio only. Forwarded to the gateway on every call. | | TOKEN_ENV | GREENNODE_RAG_TOKEN | Name of the env var that holds the token, stdio only. Set this to read the token from a differently-named var. | | ENGINE | — | Optional RAG engine / agent name, stdio only. Scopes search + list_knowledge_bases to that engine's KBs. | | PORT | 8080 | HTTP transport listen port. | | MAX_RESPONSE_BYTES | 25000 | Hard cap on list responses; over-cap responses are truncated with a notice. | | DEFAULT_PAGE_SIZE | 10 | Default size for list_documents / list_knowledge_bases. | | MAX_GET_DOCUMENT_PAGES | 10 | Max pages get_document will scan before giving up. | | LOG_LEVEL | info | debug / info / warn / error. Logs go to stderr — never stdout, so stdio JSON-RPC is never corrupted. debug adds per-call backend traces. | | BACKEND_TIMEOUT_MS | 300000 | Hard timeout (ms) for each backend call (search, list, upload, …). On timeout the tool returns a 504-style error instead of hanging forever. 0 disables. | | DOWNLOAD_DIR | system temp dir (os.tmpdir()) | Where download_document writes files over stdio. |

In streamable HTTP mode the token and engine are not read from env at all — clients supply them per request via Authorization: Bearer and X-Engine. GREENNODE_RAG_TOKEN / TOKEN_ENV / ENGINE apply only to stdio.

Development & operations

Scripts (package.json):

| Script | What it does | |---|---| | npm start | Run the compiled server (node dist/index.js) — run npm run compile first | | npm run dev | Run from source with reload (tsx watch src/index.ts) | | npm run build | Typecheck only (tsc --noEmit) | | npm run compile | Compile to dist/ (tsc -p tsconfig.build.json); the published form runs via node dist/index.js / npx greennode-rag-mcp | | npm test / npm run test:watch | Vitest |

Logs & timeouts. Every backend call logs backend → (method, path, url, body size, timeout) on start and backend ← (status, ms, bytes) on completion to stderr — so a hang shows up as a backend → with no matching backend ←. ingest_document / ingest_batch also log ingest start (kbId, per-file filename/mimeType/size) and ingest done / ingest failed. Set LOG_LEVEL=debug for more detail. BACKEND_TIMEOUT_MS (default 300000 ms) bounds every upstream call; on timeout the tool returns a 504-style error instead of hanging. For a deployed HTTP server, stderr is wherever the runtime collects it (e.g. docker logs).

Docker:

docker build -t greennode-rag-mcp .
docker run -p 8080:8080 greennode-rag-mcp

The shipped Dockerfile bakes ENV TRANSPORT=http and exposes 8080, so the container runs in streamable-HTTP mode by default. The bearer token is supplied per request via the Authorization header (same as HTTP mode) — not via env.

Further reading