@kindlm/cli
v2.3.2
Published
CLI for KindLM — behavioral regression testing for AI agents
Downloads
35
Maintainers
Readme
KindLM
Behavioral regression testing for AI agents. Test what your agents do — not just what they say.
Why KindLM?
LLM evals measure text quality. KindLM tests behavior — the tool calls your agent makes, the decisions it takes, and whether it leaks PII or violates compliance rules. It runs in CI so regressions never ship.
Features
- Tool call assertions — verify agents call the right tools with the right arguments, in the right order
- Schema validation — structured output checked against JSON Schema (AJV)
- PII detection — a guardrail (not a full DLP system). Default detection covers SSN, credit card, and email; phone, IBAN, IP, JWT, and API-key detectors are available via opt-in
detectors: [...] - LLM-as-judge — score responses against natural-language criteria (0.0–1.0). Results are probabilistic; use repeated runs for critical gates
- Drift detection — semantic + field-level comparison against saved baselines
- Keyword guards — require or forbid specific phrases in output
- Latency & cost budgets — fail tests that exceed time thresholds, or cost thresholds for priced models (OpenAI, Anthropic, Gemini)
- EU AI Act documentation draft — map test/gate results to selected EU AI Act articles. A starting-point document, not legal advice and not a conformity assessment
- CI-native — exit code 0/1, JUnit XML reporter, GitHub Actions ready
Supported Providers
Providers are configured under providers: and referenced by key from each entry in models: (see the Quick Start below). The provider: key is one of:
| Provider | provider: key | Example model | Notes |
|----------|-----------------|---------------|-------|
| OpenAI | openai | gpt-4o | Azure OpenAI works via openai + a custom baseUrl |
| Anthropic | anthropic | claude-sonnet-4-5-20250929 | |
| Google Gemini | gemini | gemini-2.0-flash | key is gemini, not google |
| Mistral | mistral | mistral-large-latest | no cost estimation yet |
| Cohere | cohere | command-r-plus | no cost estimation yet |
| Ollama | ollama | llama3 | local; no cost |
| HTTP | http | any | generic OpenAI-compatible endpoint |
| MCP | mcp | — | passthrough HTTP POST to an MCP-style tool server |
Not yet supported: AWS Bedrock, and a first-class Azure adapter (use openai + baseUrl for Azure).
Quick Start
Try it instantly:
npx @kindlm/cli initOr install globally:
npm install -g @kindlm/cli
kindlm initEdit the generated kindlm.yaml:
kindlm: 1
project: "my-agent"
suite:
name: "refund-agent"
providers:
openai:
apiKeyEnv: "OPENAI_API_KEY"
models:
- id: "gpt-4o"
provider: "openai"
model: "gpt-4o"
params:
temperature: 0
prompts:
refund:
system: "You are a refund support agent. Use lookup_order(order_id) to find orders."
user: "{{message}}"
tests:
- name: "looks-up-order"
prompt: "refund"
vars:
message: "I want to return order #12345"
tools:
- name: "lookup_order"
responses:
- when: { order_id: "12345" }
then: { order_id: "12345", status: "eligible" }
expect:
toolCalls:
- tool: "lookup_order"
argsMatch: { order_id: "12345" }
guardrails:
pii:
enabled: true
judge:
- criteria: "Response is empathetic and professional"
minScore: 0.8Run your tests:
kindlm testOutput:
refund-agent / looks-up-order
gpt-4o
✓ looks-up-order (1.3s)
✓ tool_called: lookup_order
✓ pii: no PII detected
✓ judge: 0.92 ≥ 0.80
1 passed, 0 failed
Gates: ✓ PASSEDCLI Flags
| Flag | Description | Added |
|------|-------------|-------|
| --reporter <type> | Output format: pretty (default), json, junit. Report is written to stdout — redirect to a file with > report.xml | v1.0.0 |
| --compliance | Generate the EU AI Act documentation draft (not legal advice) | v1.0.0 |
| --pdf <path> | Export the compliance draft as PDF (requires --compliance) | v1.0.0 |
| -s, --suite <name> | Assert the configured suite matches this name (a config has one suite) | v1.0.0 |
| --runs <count> | Override the repeat count from config | v1.0.0 |
| --gate <percent> | Fail if suite pass rate falls below threshold (0–100) | v1.0.0 |
| -c, --config <path> | Path to config file (default kindlm.yaml) | v2.1.0 |
| --dry-run | Validate config and print the test plan without executing | v2.1.0 |
| --watch | Re-run tests when kindlm.yaml changes | v2.1.0 |
| --no-cache | Disable response caching | v2.1.0 |
| --isolate | Copy the config + referenced schema files into a detached-HEAD git worktree and run there (requires git). This isolates the config, not your agent code or node_modules. | v2.0.0 |
| --concurrency <n> | Override the concurrency setting from config (≥ 1) | v2.1.0 |
| --timeout <ms> | Override the per-test timeout in ms (does not affect provider HTTP timeout) | v2.1.0 |
Reports go to stdout; redirect with > to write a file (there is no --output flag).
Other Commands
| Command | Description |
|---------|-------------|
| kindlm init | Scaffold a kindlm.yaml |
| kindlm validate | Validate config without running |
| kindlm baseline set\|compare\|list | Manage drift baselines |
| kindlm trace | Ingest OTLP traces and run assertions |
| kindlm cache | Inspect or clear the response cache |
| kindlm redteam | Run adversarial red-team suites (experimental) |
| kindlm login / kindlm upload | KindLM Cloud auth and run upload |
CI Integration
# .github/workflows/test.yml
- run: npm install -g @kindlm/cli
- run: kindlm test --reporter junit > results.xmlRepository Layout
packages/
core/ @kindlm/core — Business logic, zero I/O dependencies
cli/ @kindlm/cli — CLI entry point
cloud/ @kindlm/cloud — Cloudflare Workers API + D1 database
docs/ Technical specs and documentation
site/ Documentation website (Next.js)Documentation
Full docs: kindlm.dev | Source: docs/
License
MIT (core + CLI) | AGPL (cloud)
