@zespan/cli
v0.2.0
Published
CI quality gate for prompt versions. `zespan-gate` calls the same [quality gate](https://docs.zespan.com/dashboard/prompts#the-quality-gate) the Zespan dashboard's **Versions** tab uses, and turns the pass/fail verdict into a process exit code — so a CI j
Readme
@zespan/cli
CI quality gate for prompt versions. zespan-gate calls the same quality gate the Zespan dashboard's Versions tab uses, and turns the pass/fail verdict into a process exit code — so a CI job can block a merge or deploy on prompt quality.
For full documentation visit docs.zespan.com/sdk/cli
Install
npm install --save-dev @zespan/cli
# or run without installing:
npx @zespan/cli ...Before you gate: link a dataset run
The gate scores a dataset run of the candidate prompt version against a baseline run — it doesn't run your prompt for you. Produce that run first, either from the dashboard's Run over dataset button or from your own pipeline using DatasetsClient (see Dataset runs). Once the candidate run's items are linked, zespan-gate triggers scoring automatically if it isn't scored yet.
Usage
zespan-gate \
--name customer-support-agent \
--version 3 \
--dataset-run-id run_abc123 \
--evaluator-id eval_def456 \
--api-key $ZESPAN_API_KEY| Flag | Required | Description |
|---|---|---|
| --name | yes | The prompt's name |
| --version | yes | The candidate prompt version number to gate |
| --dataset-run-id | yes | Dataset run (with the candidate version's traces linked) to score and compare |
| --evaluator-id | yes | Evaluator to use as judge for both candidate and baseline runs |
| --api-key | yes* | Falls back to ZESPAN_API_KEY env var |
| --baseline-run-id | no | Compare against a specific run; defaults to whatever holds the production label |
| --baseline-label | no | Compare against a specific label instead of production |
| --connection-id | no | Pin scoring to a specific project LLM connection (BYOK) |
| --regression-run-id | no | Opt into the regression-resolution-rate check |
| --api-url | no | Defaults to https://api.zespan.com/v1; override for self-hosted |
| --poll-interval-ms | no | Default 3000 |
| --timeout-ms | no | Default 120000 |
Exit codes
| Code | Meaning |
|---|---|
| 0 | Passed — safe to proceed |
| 1 | Failed — a real quality regression (avg score drop, regression count, or tool-call accuracy) |
| 2 | Usage/config error or transport failure — missing flag, unreachable API, timeout, or no dataset-run items linked yet |
Treat 1 and 2 differently in your pipeline: 1 is a genuine regression worth surfacing on the PR; 2 usually means the job is misconfigured or the run wasn't prepared yet.
Example GitHub Actions step
- name: Gate prompt quality
run: npx @zespan/cli --name customer-support-agent --version ${{ steps.publish-prompt.outputs.version }} --dataset-run-id ${{ steps.link-run.outputs.run_id }} --evaluator-id ${{ vars.ZESPAN_EVALUATOR_ID }}
env:
ZESPAN_API_KEY: ${{ secrets.ZESPAN_API_KEY }}License
Apache-2.0
