npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@evals-ax/evx

v0.7.3

Published

Test changes to agent environments with independently verified outcomes

Downloads

338

Readme

evx

The command-line interface for Evals AX. Investigate context and run predeclared experiments on your host. Sign in to use hosted audits and save evidence in the same account as the web app.

The canonical package is @evals-ax/evx; its executable is evx. Previous package names are retained for compatibility, not promoted as alternatives. Registry publication is a separate release step. The unrelated bare npm package evx is not this project. Node.js 22.18 or newer is required.

What is distributed

This package includes the MIT-licensed portable framework: local inspection, experiment execution and evidence comparison run on your host. It is a bundled runtime, not only an HTTP client. The two synthetic experiment init templates, external observer helpers and opt-in agent integrations are included because commands use them. The hosted website auditor, account-service implementation, private research snapshots and customer data are not in this package.

Source maps and unused example plans are excluded to reduce distributed material. The JavaScript and Python implementations remain readable public code under their existing licenses; this packaging change does not conceal the framework or retract earlier releases. Local agents use your own runtime/provider credentials and may incur your provider's charges. Hosted requests and saved reports use Evals AX's account and admission controls.

Installation

From a repository checkout, install workspace dependencies and build the framework, then invoke the CLI directly:

pnpm install
pnpm --filter @evals-ax/framework build
node packages/cli/src/main.js --help

The commands below target the 0.7.3 archive shipped with this README. The archive and its install instructions are deployed together; before that release is available, use the source-checkout command above.

Install the versioned early-access archive from Evals.AX:

bun add --global @evals-ax/evx@https://evals.ax/downloads/evx-0.7.3.tgz
evx --version

Bun is the default package runner: bunx --package @evals-ax/evx@https://evals.ax/downloads/evx-0.7.3.tgz evx --help. Alternatives are npx --yes https://evals.ax/downloads/evx-0.7.3.tgz --help, pnpm dlx https://evals.ax/downloads/evx-0.7.3.tgz --help, and yarn dlx -p @evals-ax/evx@https://evals.ax/downloads/evx-0.7.3.tgz evx --help. These hosted commands work once this release is deployed; they do not depend on npm registry publication.

First audit

evx auth login
evx vibecheck https://example.com --wait
evx reports get REPORT_ID

Website audit reports have public-link visibility. Anyone given that URL can read the report. Network eligibility and admission limits are enforced by the shared API. Static findings describe observed interface properties; they do not establish that a change improves agent outcomes.

Website histories and interface reports

Each exact normalized URL has one private history. Paths and query strings remain separate properties; a fragment does not create another property.

evx api GET /websites
evx api GET /websites/PROPERTY_KEY
evx auth login
evx inspect ./inspection.json --output ./interface-report.json --title "Production API"
evx interfaces list
evx interfaces get REPORT_ID --output ./saved-report.json

inspect reads an explicit API, CLI or MCP plan and runs on the host without needing an Evals AX account. With a CLI login, it saves the report privately to that account automatically after writing and flushing the local file. Reports can include interface names and diagnostic evidence. Use --local-only to keep an authenticated run on the host; unauthenticated runs also stay local. Write the explicit plan using the interface report documentation. API probes are declared GET/HEAD requests. CLI and stdio MCP plans execute the supplied argument arrays; review them before running. MCP discovery does not call tools. Plans bound repetitions, time and output; interrupted or partial collection remains visible in the result. HTTP credentials are named through environment references and are not copied into the report.

These reports describe collected observations, not agent task performance. Account saves preserve private, immutable evidence. The returned raw report can be downloaded and validated through the same framework contract.

Successful inspect JSON retains output and report, adding cloud: { status: "saved", reportId, reportUrl } after account confirmation. A local run returns cloud: { status: "local-only", reason: "requested" | "not-authenticated", reportId: null, reportUrl: null }. --title sets the saved title; the default is the plan name.

A failed account save exits 1 with INSPECTION_SAVE_FAILED on stderr and keeps the local file. Its structured details include output, reportSaved: true, reason, cloud delivery status and retry: { command: "evx", args: [...] }. Run that argument array to retry the same retained report and title without rerunning collection. A timeout or lost response means delivery is unconfirmed; the shared API deduplicates exact report-ID/content/title retries. No write is retried automatically. Expired login, missing write scope or required terms need action before retrying; terms acceptance stays in the browser. Cancellation exits 130 and preserves the local report. Authenticated experiments follow the automatic saving flow below.

Account access on the host

evx auth login
evx auth status
evx auth tokens
evx reports list
evx auth logout

Login opens the Evals.AX authorization page. Enter the displayed code, check the account shown, and approve access. Use --no-browser when a browser cannot be opened automatically. Authorization does not complete until the browser approves it. auth tokens lists account token metadata, never bearer tokens. auth tokens revoke TOKEN_ID revokes a selected token.

To provision a separate environment without putting a token in terminal output, use evx auth tokens create --name NAME --output /absolute/private/config-directory. The new directory becomes that environment’s EVX_CONFIG_HOME; add --read-only to limit the token. This command refuses to overwrite your current login or an existing credential. Make the directory available only to the intended user or sandbox.

Credentials are saved atomically under ~/.config/evx (or $XDG_CONFIG_HOME/evx), in a private directory with permissions 0700 and origin-specific files with permissions 0600 on POSIX. They persist across CLI processes. Expired credentials require a fresh login. Logout revokes the server token before removing the saved credential; logout --local only removes the local file, leaving its server token valid until expiry or revocation elsewhere. On Windows, use a private user-profile directory with appropriate account ACLs.

Logout refuses to proceed while EVX_TOKEN is set, including with --local, so an environment credential cannot cause the separate host login to be deleted. Unset the variable to log out the host. To explicitly revoke the active environment token without touching the host login, use evx api DELETE /auth/token and then remove that variable from the parent environment.

A sandboxed agent can reuse the login when its host exposes the same config directory. Set EVX_CONFIG_HOME to that absolute directory when necessary. If the sandbox cannot access it, use an explicitly provided EVX_TOKEN with EVX_API_ORIGIN; the CLI does not bypass the sandbox to discover host secrets. Avoid putting a token in prompts, command arguments, or source files. Environment tokens override saved credentials and are bound to EVX_API_ORIGIN (https://evals.ax by default). Changing --api-origin cannot forward an environment token to another origin. Credentials and authentication endpoints always require HTTPS, including local development. Use a local HTTPS server with a certificate trusted by that CLI process; the test suite uses temporary certificates through NODE_EXTRA_CA_CERTS without changing the host trust store or disabling certificate verification. Anonymous HTTP loopback requests remain available. Redirects are refused.

Local experiments

evx experiment init ./context-study
evx experiment validate ./context-study/manifest.json
evx experiment run ./context-study/manifest.json --output ./study-results
evx experiment compare ./study-results/result.json
evx context scan .

With a CLI login, experiment run checks account write access and current Terms acceptance before executing commands. It registers the declared manifest, runs on the host, then saves the immutable result and bounded trace attachments privately. Manifest uploads contain the task prompt and configured paths; captured streams can include source content and agent answers. Add --project UUID to choose a private project; otherwise the CLI uses an account-and-origin-bound “CLI experiments” project across your hosts. Use --local-only to avoid account access and saving. Without a login, runs stay local; an authenticated preflight error stops execution instead of silently downgrading.

To run a declaration already registered in the web account, supply the matching local manifest:

evx experiment run ./context-study/manifest.json --output ./study-results --experiment EXPERIMENT_UUID

This starts its trials; it does not resume a partial run. A declaration with an attached result cannot be rerun. The registration timestamp is a service receipt, not independent proof of execution timing. Concurrent registration retries use the same client ID; the service checks owner and content before reusing it.

Results remain in ./study-results/result.json, with local trace artifacts beside them. Standard output adds a cloud receipt to the result fields; use the retained result.json for validation and saving. On cancellation or failed saving, the error includes a save-only retry command:

evx experiment upload ./study-results/result.json --experiment EXPERIMENT_UUID

That command attaches the same immutable evidence without rerunning any trial. Keep the returned experiment ID and API origin; saving under another account cannot access the original declaration. A failed registration provides the original run command for a deduplicated retry. Existing output directories are never overwritten, so inspect their retained evidence before choosing whether to save it or run a new experiment elsewhere.

Automatic trace attachment requires evx 0.7.2 or later. The CLI collects only primary trial and verifier-control trace files listed in that result. It verifies their original bytes and hashes, keeps the originals unchanged locally, and saves a compressed copy within the service's per-file and total limits. No directory crawl occurs. The cloud.traces receipt lists each saved or omitted file, explicit redaction provenance, and missing coverage. External observer artifacts are not retained by this attachment version. Saving a result from standard input cannot locate its trace directory; use the local result file to save attachments.

Before saving, selected credential syntax, structured credential string fields, and every nonempty secret-named environment value are redacted from captured streams. This is limited detection, not a guarantee that all secrets were removed. A changed copy gets a different hash and a redaction manifest; a copy with no detected changes retains its exact original bytes. Structural collisions with a known secret or nonempty credential containers that cannot be safely inspected cause omission instead of changing protocol keys or nonstring values. Missing, unsafe, oversized or undecodable artifacts remain explicit omissions. If attachment delivery fails, the result is already saved and the error supplies the same upload-only retry; inspect cloud.traces for what was confirmed. Identical retries reuse retained attachments; changed redacted copies can conflict with an earlier saved copy.

Inspect the initialized manifest, agent command, and verifier before running it. experiment run is an explicit local execution boundary: it runs argv commands defined by that file. Downloading or retrieving an experiment never executes it. The bundled example calls your locally installed Codex using the declared model and can incur provider usage charges. Its small task qualifies the runner and illustrates an experiment; it cannot establish a universal design rule.

Experiment execution supports macOS and Linux, including Linux under WSL. This initial release rejects native Windows execution because it cannot yet guarantee cleanup of descendant processes. Context inspection, result validation and comparison, and account commands remain separate from this execution restriction.

Predeclared paired trials preserve failures and trials that could not run. A comparison summarizes one paired result using the framework’s uncertainty method; it does not pool unrelated experiments. Results contain the manifest, including the user-authored task prompt; treat that output as potentially sensitive. Saved traces represent captured provider and verifier streams, not complete model-visible context. Context scanning reports inspectable facts and hypotheses, with coverage limits.

experiment validate checks a result’s manifest and retained trace hashes by default. Use --schema-only for a result downloaded without its trace files; the response explicitly reports that local evidence was not checked. Hash verification establishes file integrity, not the scientific validity of the hypothesis or verifier.

To save a validated result:

evx projects create --name "Context discovery"
evx experiment upload ./study-results/result.json --project PROJECT_ID
evx experiment list
evx experiment get EXPERIMENT_ID

API parity and telemetry

Resource commands use the same versioned API as the web application. evx api exposes a resource that has not yet received a dedicated command:

evx api GET /account
evx projects update PROJECT_ID --file project.json
evx telemetry ingest --file event.json
evx telemetry list
evx api GET /experiments

--file - reads JSON from stdin. List commands accept --cursor and return {items,nextCursor?}; they do not silently fetch every page. Telemetry must distinguish measured values from unknown values and identify its evidence source. The server validates submissions and preserves account ownership.

Responses are JSON on stdout. Help is plain text; progress and structured errors use stderr. Exit status is 0 for successful commands, 1 for failures, and 130 for cancellation. An HTTP success can contain a failed audit or an inconclusive experiment: inspect the returned domain status. Request errors include stable codes and retry timing when provided. Mutating requests are never automatically retried because a timeout does not establish that a write failed. --timeout bounds an individual API request (30 seconds by default); vibecheck --wait waits up to five minutes for completion, while experiment limits are declared in its manifest.

Agent hooks and plugins

The package includes integrations/evals-ax, with one shared skill and Codex/Claude plugin manifests. Its README documents manual loading and inert hook examples; installation never enables telemetry or edits host configuration. Run evx integrations path to find the installed directory independently of your package manager.

evx hooks adapt --format claude-result --file claude-result.json
evx hooks ingest --format claude-result --file claude-result.json --upload

adapt is local. ingest requires explicit --upload and authentication. Supported formats are codex-jsonl, codex-notify, claude-result, and claude-hook. Codex JSONL and Claude hook captures require a --run-id UUID retained for retries. Only allowlisted provider measurements are emitted; raw text and paths are discarded. Lifecycle observations contain no usage metrics. Outcomes remain unknown without an independent verifier; Claude cost estimates include subagents while its result token counts cover the main agent. See the bundled integration guide for coverage and opt-in hook configuration.

Package checks

node --test packages/cli/test/*.test.js

The tests exercise credential permissions and origin binding, redirect refusal, real CLI process output, anonymous submission, and device authorization. HTTPS fixture tests require OpenSSL and generate short-lived local certificates. Provider email delivery and real model behavior require separate live verification.

External trial observation (0.1.1)

An experiment manifest can opt into execution.kind: "linux-observed". This requires a dedicated Linux host with bubblewrap, strace and Python 3; it never falls back to ordinary host execution. Private syscall logs and model-request bodies stay alongside the result, and evx experiment validate RESULT verifies the receipt and its artifacts. See the observer setup and coverage limits before running it. Codex uses an explicit API gateway with a supervisor-only credential; existing ChatGPT session files are not copied.

Website report handoff

evx reports fixes REPORT_ID fetches the complete public action plan: every recorded observation, editorial priority and rationale, evidence, candidate change and limitation. No inspection or model is rerun. Priorities are not predicted benefits or causal effects. The returned agentBrief can be pasted into an existing Codex or Claude Code session in the relevant project. The CLI does not launch an agent or install host integrations. report_not_completed means a completed report is not available; deletion disables this endpoint along with the original report.

Safe Vibecheck retries

Vibecheck creates a private audit request UUID before submission. After an interrupted response, reuse it with evx vibecheck URL --request-id UUID and the same comparison. Errors contain the ID and structured recovery arguments. An existing request cannot bind to another account or different input. After admission, evx reports get REPORT_ID reads its outcome; stopping CLI waiting does not cancel admitted server work.

The configured operator can explicitly activate a seven-day noncash pricing trial in the website. That trial reserves its accepted $0.01 per single-page audit and settles only a completed report. The response includes a private chargeUrl; use evx api GET /billing/audits/REPORT_ID for its financial outcome. This is not a payment or paid Pro subscription. Other accounts retain ordinary free admission and all service caps still apply.