@plune-ai/cli
v0.13.0
Published
AI-powered assertion testing for LLM apps — write assertions, run them against any provider, get a pass/fail report (local, CI, or diff).
Maintainers
Readme
@plune-ai/cli
AI-powered assertion testing for LLM apps — a test runner for model behaviour.
Plune runs an assertion suite against an LLM provider and gives you a pass/fail report —
locally, in CI, or as a regression diff between two runs. You describe the checks in one
plune.yaml; Plune calls the model, evaluates each assertion, caches results, and reports
token cost. Ten built-in assertion types cover plain text, JSON-schema, LLM-as-judge, and
RAG metrics (faithfulness, answer-relevance, context-precision).
Links: plune.ai | npm | GitHub Action on Marketplace
Install
npm install -g @plune-ai/cli # or: pnpm add -g @plune-ai/cli
plune --versionOr run it without installing:
npx -y @plune-ai/cli runRequires Node.js ≥ 20.
Quickstart
# 1. Scaffold plune.yaml, an example dataset, and .env.example
plune init
# 2. Add your provider key (read from the environment / .env — never written to disk)
echo 'ANTHROPIC_API_KEY=sk-ant-...' >> .env
# 3. Run the assertions
plune run
# → 1/1 passed · 0 failed · 0 errored · $0.0008
# 4. Re-render the last run, or diff two runs to catch regressions
plune report --format markdown
plune diff baseline.json current.json --fail-on-regressionEach run writes its full result to .plune/last-run.json.
Configuration
Plune reads a single plune.yaml, discovered by walking up from the working directory (or
passed with -c <path>). A minimal example:
version: 1
provider:
type: anthropic # anthropic | openai | openrouter
model: claude-3-5-sonnet-latest
evals:
- id: example
prompt: "Answer concisely. {{question}}" # {{vars}} come from each dataset row
dataset: datasets/example.jsonl # a file path, or an inline `examples:` list
assertions:
- type: contains
value: "Paris"Datasets are JSONL, one row per line, shaped { "vars": { ... }, "expected"?: "..." }. The
provider API key is read from the environment based on provider.type:
| Provider | provider.type | Environment variable |
| ---------- | --------------- | -------------------- |
| Anthropic | anthropic | ANTHROPIC_API_KEY |
| OpenAI | openai | OPENAI_API_KEY |
| OpenRouter | openrouter | OPENROUTER_API_KEY |
Assertion types
| Type | Passes when… |
| --------------------- | ---------------------------------------------------------------- |
| exact-match | output equals value (optional trim, ignore_case) |
| contains | output contains value |
| contains-any | output contains at least one of values |
| contains-all | output contains every one of values |
| json-schema | output validates against the JSON schema |
| llm-judge | an LLM grades the output against criteria (≥ pass_threshold) |
| semantic-similarity | embedding similarity to reference ≥ threshold |
| faithfulness | output is grounded in context (RAG) |
| answer-relevance | output actually answers the question (RAG) |
| context-precision | context is relevant to the question (RAG) |
Commands
| Command | Summary |
| ------- | ------- |
| plune run | Run the suite. Flags: --dry-run, --only <id\|tag> (repeatable), --bail, --no-cache, --concurrency <n>, --format console\|json\|markdown, -o, --output <file>. |
| plune report | Re-render the most recent run. Flags: --format, -o. |
| plune diff <baseline> <current> | Compare two plune run --format json outputs and report pass→fail regressions. Flags: --fail-on-regression, --format, -o. |
| plune init | Scaffold plune.yaml, a sample dataset, and .env.example. Flags: --yes (non-interactive), --force. |
| plune login | Save a Plune platform API token so sync and ingest can reach it. The token is checked against the API before it is saved, so a wrong one fails here rather than two commands later. Get one at https://beta.plune.ai → Settings → API tokens. Flags: --token <token> (omit to paste it or pipe it via stdin), --skip-verify (save without checking, for offline setup). In CI, skip this step: every platform command reads PLUNE_TOKEN from the environment first, and the saved login only when it is not set. |
| plune logout | Remove the saved token. |
| plune sync | Upload the latest local run to the platform. Flags: --file <path> to send a specific run JSON. |
| plune run import <file> | Turn a JUnit XML or Playwright JSON report into a run in Plune. Needs no provider key — nothing is generated. Flags: --format junit\|playwright-json (detected from the file when omitted), --key <externalKey> to land several reports in one run (with PLUNE_SHARED_RUN=1 — see below), --create to offer unmatched tests to the review queue. |
| plune run start | Open a platform run — or join the one already carrying --key — and print its id. Flags: --key <externalKey> (generated when absent), --json for one machine-readable line. |
| plune run finish <id> | Close a platform run. This is what the reporter tells you to do for a run it had to leave open. Flags: --terminate to record it as cut short, --reason <text>. |
| plune run exec -- <command> | Open a run, run the command inside it, close the run — and exit with whatever the command returned. Sets PLUNE_SHARED_RUN for you, so anything reporting inside joins that run. Flags: --key <externalKey>. |
| plune run delete <id> | Delete a run and everything it produced — its results and the review-queue entries it raised. Approved test cases and the audit log stay. Recoverable for six months (ask Plune to put it back), then gone for good. No prompt: it is your data, and this command belongs in scripts. |
| plune run report | Replay .plune/pending-results.jsonl — send what the reporter could not. Flags: --file <path>. |
| plune ingest [dir] | Record a Cairn run in Plune. Omit [dir] for the newest run under ./runs, or name the directory holding report.json. Generated cases arrive as review proposals — nothing is created until a person approves it. |
| plune pull [file] | Write the project's test cases as one Markdown document — Testomat's classical format, plus Plune's own columns — to plune/cases.md (or [file]). Refuses to overwrite a file git sees as modified; --force overrides, --suite <id> takes one suite or folder and what is under it. |
| plune push [file] [--dry-run] | Send the document back. Cases are matched by id; a block without one becomes a draft; an unknown id is refused by line; nothing is deleted. Prints the report — created, updated, unchanged, refused, warnings — and with --dry-run writes nothing. Exit 4 when anything was refused. |
| plune plan grep <id> | Print a test plan as one --grep pattern — npx playwright test --grep "$(plune plan grep <id>)" or npx vitest run -t "$(plune plan grep <id>)" runs exactly what the plan collects. Each case contributes its title path from the path-title key the reporter wrote (the file left out, [ >#]+ between the segments — a space for Playwright, jest, mocha and vitest ≤ 4, > for vitest 5), escaped, joined with \|; a case without one goes in by title. The pattern is all that reaches stdout; the summary — how many cases, how many keyed — goes to stderr. An empty plan prints (?!), which matches nothing, because an empty --grep would run everything. |
Global flags: -c, --config <path> · -v, --verbose · --no-color.
Exit codes: 0 everything passed · 1 an assertion failed · 2 configuration or execution error · 3 pull refused to overwrite uncommitted changes · 4 push had refusals.
Already running tests? Bring the results in
Everything above generates checks, which is why it needs a provider key. If you already have a suite, there is nothing to generate — the results exist and only have to arrive, and that route costs nothing beyond a Plune token.
plune login # once, on your machine; in CI export PLUNE_TOKEN instead
plune run import ./junit.xml # Jest, Vitest, pytest, PHPUnit, Surefire, Cypress, …
plune run import ./playwright.json # or Playwright's own JSON reportThe format is read from the file, not from its name. Statuses go over as the report wrote them and
Plune maps them per project, so error can mean something different to your team than to ours.
A test Plune has no case for is counted, and --create offers it to the review queue instead —
nothing becomes a test case until a person approves it.
For a Playwright suite there is also @plune-ai/playwright,
which reports as the run happens and needs no second step:
// playwright.config.ts
reporter: [['list'], ['@plune-ai/playwright']],Several jobs, one run
Two suites, or a sharded matrix, report into a single run when every job shares a key and knows it is not the last one:
PLUNE_SHARED_RUN=1 plune run import ./junit.xml --key "$GITHUB_RUN_ID"
PLUNE_SHARED_RUN=1 plune run import ./app/junit.xml --key "$GITHUB_RUN_ID"
plune run finish "$RUN_ID" # once, when they are all donePLUNE_SHARED_RUN=1 is what stops a job closing a run its siblings are still reporting into; the
key alone says where the results go, not who ends the run. Leave it unset in the last job and that
job closes the run instead of the explicit finish.
plune run exec sets it for you — anything reporting inside it, this command included, joins
without closing.
When a test is deleted
A test that vanishes from the code stays in Plune as an active case until something says otherwise. The run that covers the whole suite is what says it:
PLUNE_FULL_RUN=1 plune run import ./junit.xml # or the same variable on the Playwright reporter's jobWhen that run finishes having reported everything it expected, every case its source used to
report and did not this time is marked detached — not deleted, not retired: it keeps its history
and its place, and the dashboard shows it with the run as the reason. The next result on it brings
it back by itself. Set the variable only on the job that runs everything: a run of one file or a
--grep also reports everything it expected, and with the flag it would detach the rest. A run that
lost a shard (notRun is not empty) never detaches anything. plune run start has no list of
tests to compare against, so the flag means nothing there.
What the run is called
A run nobody named is called after the directory, the minute it started and where it ran —
plune · 2026-09-15 15:26 · ci — so a row in the list reads without being opened. The clock is the
reporting machine's own; ci comes from the CI variable every hosted runner sets. To call it
something else, set PLUNE_RUN_TITLE; PLUNE_ENV and PLUNE_LABELS mark where it ran and how, and
those stay marks rather than becoming part of the name.
Optional: keep a history
Everything above works with no account, no network, and no token — that does not change. run,
report, diff and init never open a socket, and that is the half this promise is about. If you
also want run history, trends, and a shared dashboard, the commands below push your local runs to
the Plune platform:
plune login # paste the API token from your platform settings page
plune run # exactly as before — the run is saved locally
plune sync # upload .plune/last-run.json, print the read-back URL
plune pull # the test cases as plune/cases.md — edit them anywhere
plune push --dry-run # what a push would change, before it changes anythingThe token is stored at ~/.config/plune/credentials.json (mode 0600, honours XDG_CONFIG_HOME)
and is never printed or logged. PLUNE_API_URL points sync, pull and push at a different server —
useful for a self-hosted backend. The Markdown format pull writes and push reads is documented at
docs.plune.ai/platform/cases/markdown.
Programmatic API
The same engine that powers plune run is exported for use from your own code. Unlike the
CLI, the library does not parse argv or auto-load .env — set the provider key in
process.env yourself.
import { run } from '@plune-ai/cli';
import type { RunResult } from '@plune-ai/cli';
const result: RunResult = await run({ dryRun: false, configPath: 'plune.yaml' });
console.log(result.summary); // { total, passed, failed, errored, ... }Use in CI
Run Plune on every pull request and post a regression diff as a sticky comment with the companion GitHub Action, plune-ai/eval-action:
- uses: plune-ai/eval-action@v1
with:
config: plune.yaml
fail-on-regression: trueContributing
Bug reports and pull requests are welcome — see CONTRIBUTING.md. For security issues, see SECURITY.md.
License
MIT © Plune Contributors
