mcp-tdqs
v0.2.0
Published
Reference implementation of the Tool Definition Quality Score (TDQS): score how well an MCP tool definition communicates to an AI agent.
Downloads
850
Maintainers
Readme
mcp-tdqs
The reference implementation of the Tool Definition Quality Score (TDQS): a CLI and a library that score how well an MCP tool definition communicates to an AI agent, exactly as the specification defines it.
TDQS scores a definition, not behaviour. The inputs are what an MCP client sees from tools/list — name, title, description, input schema, output schema, annotations — and the output is a score from 1.0 to 5.0 with a letter tier, per tool and per server, with a justification for every dimension. The evaluator reads the full output schema; documented return fields reduce what the description must explain, while a bare object does not.
Install
npm install --global mcp-tdqs
# or run it without installing
npx mcp-tdqs --helpThe package is mcp-tdqs; the command it installs is tdqs. Node 22 or newer. The same implementation is on PyPI as tdqs for Python, with the same command; the two are held to the same fixtures.
Lint: deterministic, no model, no key
tdqs lint --file tools.json
tdqs lint --command "npx -y @scope/mcp-server"
tdqs lint --url https://mcp.example.com/mcp --header "Authorization: Bearer …"lint runs the stages of the pipeline that need no model: the context signals (parameter counts, schema description coverage, annotation values, invocation cost, the definition's hash and byte size), the hard gates (no description, tautological description), the shadow prefilter across the tool set, and the checklist the specification ranks highest. It exits 1 on an error-level finding, which makes it a pull request check:
tdqs lint --file tools.json --fail-on warning --format markdown --output tdqs-lint.mdA lint finding names a fix. It is not a score, and it never pretends to be one.
Score: the full rubric
export TDQS_BASE_URL=https://api.openai.com/v1 # any OpenAI-compatible endpoint
export TDQS_API_KEY=…
export TDQS_MODEL=…
tdqs score --file tools.json
tdqs score --command "npx -y @scope/mcp-server" --fail-under B --format markdownscore sends every tool through the rubric (six dimensions, 1–5 each, with the specification's system prompt verbatim), runs the server coherence evaluation (four dimensions plus shadowing-risk confirmation), and rolls both up into the server score with integer arithmetic. The report is stamped with the specification version and the model, because a score is calibrated to a rubric+model pair and is not comparable to anything without both.
Turn extended reasoning off. The reference model reasons before it answers unless told not to, which makes a call take a minute instead of seconds — and the specification's calibration examples reproduce with reasoning off. How to say so is provider-specific, so it is an opaque JSON object merged into every request:
tdqs score --file tools.json --request-overrides '{"reasoning":{"enabled":false}}' # OpenRouter
# or TDQS_REQUEST_OVERRIDES in the environment; DeepSeek directly takes {"thinking":{"type":"disabled"}}--hosted https://tdqs.example scores through a hosted TDQS site instead of a model key of your own, and prints the report's URL. It takes that site's API key as --api-key or TDQS_API_KEY; the site's account page is where keys come from.
Input is exactly one of --file (a tools/list result, an array of tools, or a single tool; - reads stdin), --command (a stdio server) or --url (a Streamable HTTP server).
| Exit code | Meaning |
| --------- | -------------------------------------------------------------------------- |
| 0 | done |
| 1 | the threshold was not met (--fail-on for lint, --fail-under for score) |
| 2 | usage error, unreadable input, unreachable server, or a model failure |
--format is text (default), markdown or json. The JSON formats are described by schemas/score-report.json and schemas/lint-report.json.
Library
import { createLlmClient, lintServer, parseToolDefinitions, scoreServer } from 'mcp-tdqs';
const { serverName, tools } = parseToolDefinitions(await response.json());
// No model involved.
const lint = lintServer({ serverName: serverName ?? 'my-server', tools });
// The full pipeline.
const report = await scoreServer({
llm: createLlmClient({
apiKey,
baseUrl,
model,
requestOverrides: { reasoning: { enabled: false } },
}),
serverName: serverName ?? 'my-server',
tools,
});
report.serverScore.overallTier; // 'A' | 'B' | 'C' | 'D' | 'F'
report.tools[0].justifications.usage_guidelines; // { score, justification }Every stage is exported on its own — computeContextSignals, evaluateHardGates, computeTdqs, findShadowCandidates, buildToolScoringPrompt, scoreToolDefinition, scoreServerCoherence, rollupServerScore — along with the specification's metadata (TOOL_DIMENSIONS, COHERENCE_DIMENSIONS, FLAGS, TIERS, LINT_RULES, SPEC_VERSION) and the two system prompts, so a site or a second implementation can build on the same pieces.
What is deterministic and what is not
Stages 1, 2 and 4 of the pipeline, the shadow prefilter, and every rollup are deterministic and reproducible from the definitions alone; inputHash is computed the same way the Glama registry computes it, so a hash here matches the one on a server's public score page. Stage 3 — the rubric — and the coherence evaluation are model calls. The specification pins the prompts, the output contract and the calibration examples; the model is the remaining variable, which is why every report names it. Swap models and expect to re-score.
Specification
This package follows TDQS 1.3. The prompts are compared byte for byte against the specification in the test suite, so the two cannot drift apart silently.
