npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

chaosbringer

v0.11.2

Published

Playwright-based chaos testing tool for web applications

Readme

chaosbringer

Playwright-based chaos testing for web apps. Crawls the pages you point it at, performs weighted random actions, injects network faults, evaluates invariants, and reports what broke — with a seed you can replay.

Where chaosbringer fits in the chaos layering

chaosbringer only injects faults that a browser-driven test can reach. Server-internal failure modes need a sibling library. Use this table to pick the layer before reaching for a fault provider:

| Layer | Library | What it touches | When to use | |---|---|---|---| | Application state | (your test setup; see #60 for a built-in hook) | Backend rows, storage state, fixtures | "Crawler needs N todos to navigate" | | Network | chaosbringer faults.* | HTTP between browser and server (Playwright route()) | "What does the UI do when /api/x is 500 / slow / aborted" | | Page lifecycle / runtime | chaosbringer lifecycleFaults / runtimeFaults | Browser DOM, storage wipe, CPU throttle, fetch / clock monkey-patches | "Does the SPA recover when localStorage gets wiped mid-action" | | Server-side | @mizchi/server-faults | Inside the server process, before the handler runs | "Do the server's own OTel traces / metrics show the fault, and does the handler degrade gracefully" | | Cloudflare bindings | (proposed #61 @mizchi/cf-faults) | KV / Service Binding / D1 / Cache wrappers | "How does the Worker behave when its KV throws" |

Common confusion: faults.status(500, ...) from chaosbringer does not produce server-side telemetry — the route is intercepted in the browser, the server is never called. To see a fault inside the server's OTel trace, mount @mizchi/server-faults (or run both layers together).

Features

  • Weighted random actions targeted by ARIA role and visible text (nav links > buttons > inputs > scroll).
  • Thorough link extraction — <a>, <area>, <iframe>, <link rel="canonical"/"alternate">, <meta http-equiv="refresh">, and SPA history.pushState / replaceState navigations all feed the queue, so React Router / Vue Router / SvelteKit / Next.js client-side routes get discovered without static <a href>.
  • Seeded reproducibility — same seed, same action order. Every report prints a Repro: line you can paste into CI logs.
  • Network fault injection via Playwright's route API: serve a 500, abort, or add latency to any URL pattern.
  • Lifecycle fault injection — CDP CPU throttling, storage wipe (localStorage / sessionStorage / cookies / IndexedDB), Service Worker cache eviction, and key/value tampering, applied at named stages of every page visit (beforeNavigation / afterLoad / beforeActions / betweenActions).
  • Runtime fault injection — persistent in-page monkey-patches installed via addInitScript on every navigation; subverts JS APIs that no network mock can reach: reject-fetch (TypeError or AbortError), never-settle-fetch, reject-body (res.json() rejects after the fetch resolved), resolve-rejected-thenable, clock-skew.
  • Deterministic fault schedules — schedule: { decisions: ["inject", "pass"] } on any fault layer replaces the probability roll with a per-occurrence decision table, so "fail the first call, pass the retry" is a test rather than a lucky run. Consumes no RNG, so seeds stay stable.
  • Model-driven fault coverage — a temporal-logic model (Quint / ITF) enumerates the failure space, chaosbringer model compile turns each witness into a committed FaultPlan, model run replays every state with the model as the oracle, and model shrink minimises a failing plan to the smallest schedule that still fails the same way. Reports which states are reachable, which are unreachable within the bound, and which plans the app never actually exercised.
  • Coverage-guided action selection — opt-in V8 precise coverage feedback (CDP Profiler.takePreciseCoverage) attributes per-action coverage deltas to the target that fired them and biases subsequent action weights toward targets that historically delivered new code paths.
  • Declarative invariants evaluated on every page. A violation fails the run regardless of --strict. Trans-page state — e.g. state-machine transitions — is supported via a run-scoped ctx.state Map and an invariants.stateMachine() helper.
  • Accessibility checks via an invariants.axe() preset — axe-core is an optional peer dep.
  • Performance budgets per TTFB / FCP / LCP / TBT — budget breaches fail the run.
  • Visual regression via pixelmatch — compare per-page screenshots against baselines, fail on diff.
  • Error detection: console errors, failed requests, JS exceptions, unhandled rejections, invariant violations.
  • Recovery from 404 / 5xx — records what actions preceded the failure.
  • HAR record / replay + trace record / replay / minimize for fully deterministic runs and delta-debugged repros.
  • Failure artifact bundles — every failing page dumps a directory with screenshot, HTML, errors, trace prefix, and a repro.sh to replay it.
  • Baseline diff — surface new clusters / newly failing pages vs a previous run.
  • Flake detection — rerun the same crawl N times and flag clusters / pages whose outcome varies.
  • Action heatmap — pure aggregation of report.actions[] exposing the most-hit and most-failed targets.
  • JUnit XML output — Surefire-style junit.xml so any CI dashboard (Jenkins, CircleCI, GitLab, GitHub Actions test summary, Allure) can ingest the run.
  • Authenticated crawls via Playwright storageState, device emulation (iPhone, Pixel, …), network throttling (slow-3g, fast-3g, offline).
  • Sitemap seeding — prepend every URL in a sitemap.xml (or sitemap index) to the queue.
  • Parallel sharding — split a crawl across N processes with --shard i/N, then merge via the shard subcommand.
  • GitHub Actions annotations — emit ::error / ::warning lines so failures show up on the run summary.
  • Playwright Test integration for when you'd rather run chaos inside an existing test file.
  • CLI for running from a shell or CI, with minimize / flake / shard / diff / parity / cluster-artifacts / recipes / load subcommands.

Install

pnpm add chaosbringer playwright @playwright/test
npx playwright install chromium

chaosbringer targets ESM. Programmatic consumers need "type": "module" (or .mts files) and playwright as a peer dependency.

Installing from a git ref

chaosbringer is currently distributed via GitHub (no npm package yet). Both forms work:

# pin a commit SHA
pnpm add chaosbringer@github:mizchi/chaosbringer#<sha>

# track main
pnpm add chaosbringer@github:mizchi/chaosbringer

The package's prepare script runs tsc on install, so pnpm install / npm install builds dist/ automatically — no manual step needed.

Quick start — CLI

# Crawl, then exit 0 / 1 based on navigation outcomes
chaosbringer --url http://localhost:3000

# Dev mode: ignore third-party analytics noise
chaosbringer --url http://localhost:3000 --ignore-analytics

# CI mode: console errors and JS exceptions also fail the run
chaosbringer --url http://localhost:3000 --strict --compact --ignore-analytics

# Reproduce a failing run by pasting its Repro: line
chaosbringer --url http://localhost:3000 --seed 1234567 --max-pages 20

# Sweep an unknown site: clean crawl + API-failure crawls + ranked findings
chaosbringer scan --url http://localhost:3000 --max-pages 30   # → chaosbringer-scan/scan-report.md

scan is described in docs/recipes/scan.md.

Preview a crawl in an existing Chromium tab

Start your app server first. In one terminal tab, open the site in terminal-browser and leave it running. In another terminal tab, run chaosbringer. The terminal-browser command must be on PATH:

# Terminal tab 1: keep this open so the crawl stays visible
terminal-browser open http://localhost:3000

# Terminal tab 2: crawl the visible tab
chaosbringer --url http://localhost:3000 --terminal-browser

--terminal-browser runs terminal-browser ls --all --json and selects a tab with the same URL as --url, or an active tab on the same origin. The URLs must use the same hostname, such as localhost in both commands. If more than one tab matches, select one explicitly with --cdp <port> --cdp-target <targetId>:

terminal-browser ls --all --json
chaosbringer --url http://localhost:3000 --cdp 9222 --cdp-target TARGET_ID

Use the cdpPort and targetId values from ls in place of 9222 and TARGET_ID. The crawl navigates the selected visible tab, including subsequent pages, and leaves the tab and browser open when it finishes. --cdp also accepts an HTTP or WebSocket CDP endpoint. For a programmatic crawl, set terminalBrowser: true in CrawlerOptions.

This mode uses the tab's existing browser context. launchOptions, device/viewport/user-agent emulation, storage state, and HAR record/replay cannot be used with an attached browser.

Quick start — programmatic

The shortest path, using the chaos() convenience and the faults helpers:

// chaos-test.ts
import { chaos, faults } from "chaosbringer";

async function main() {
  const { report, passed } = await chaos({
    baseUrl: "http://localhost:3000",
    seed: 42,
    maxPages: 20,
    strict: true,
    faultInjection: [
      faults.status(500, { urlPattern: /\/api\// }),
      faults.delay(2000, { urlPattern: /\/slow\// }),
    ],
    invariants: [
      {
        name: "cart-count-non-negative",
        when: "afterActions",
        async check({ page }) {
          const n = Number(await page.locator("[data-cart-count]").textContent());
          return n >= 0 || `cart count was ${n}`;
        },
      },
    ],
  });

  console.log(report.reproCommand);
  process.exit(passed ? 0 : 1);
}
main();

The examples wrap in async function main() because a plain .ts file under an unconfigured project can't use top-level await.

Lower-level API

If you need more control than chaos() exposes:

import { ChaosCrawler, getExitCode } from "chaosbringer";

async function main() {
  const crawler = new ChaosCrawler({
    baseUrl: "http://localhost:3000",
    seed: 42,
  });
  const report = await crawler.start();
  process.exit(getExitCode(report, /* strict */ true));
}
main();

Reproducible runs

Every report includes:

  • report.seed — the seed actually used (random if you didn't pass one).
  • report.reproCommand — a shell-safe invocation that rebuilds the same run.

Both are printed in the compact header ([PASS] … (seed=42)) and the full report (Seed: 42 / Repro: chaosbringer --url … --seed 42).

To rerun an exact failure locally, copy the Repro: line from CI.

Fault injection

Fault rules let you force specific network requests to fail, delay, or return a canned response. Use the faults helpers to build rules without the discriminated-union noise:

import { chaos, faults } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  faultInjection: [
    // Always 500 on /api/*
    faults.status(500, { urlPattern: /\/api\// }),

    // 30% of the time, return a 429 with a Retry-After body
    faults.status(429, {
      urlPattern: /\/api\/orders$/,
      methods: ["POST"],
      probability: 0.3,
      body: "Retry-After: 5",
      contentType: "text/plain",
    }),

    // Abort tracking pixels
    faults.abort({ urlPattern: /tracking/ }),

    // Add 2s of latency to one endpoint
    faults.delay(2000, { urlPattern: /\/api\/search/ }),
  ],
});

probability is evaluated against the seeded RNG — same seed, same pattern of injections.

resourceTypes limits a rule to Playwright resource types (request.resourceType()): faults.abort({ urlPattern: /\/docs/, resourceTypes: ["fetch", "xhr"] }) fails what the page's code requests and never a document the browser navigates to. That matters when a framework fetches a route's data from the route's own URL (Next.js App Router's /docs?_rsc=…): a URL pattern alone cannot tell the two apart. Without it, a rule matches every type.

Behaviour change: probability: 0 no longer draws from the RNG. It is a rule that can never fire, so rolling for it was a wasted draw — but the draw was part of the sequence, so any existing seeded config containing a probability: 0 fault rule now produces a different action sequence than it did before. Parking a rule with probability: 0 instead of deleting it is the ordinary way to do that, so if you have a pinned seed and a pinned expected action order, re-record it. probability: 1 and values in (0, 1) are unaffected, and the lifecycle layer never drew for 0 in the first place. Nothing else about seed stability changed: RNG is consumed only for a probability strictly inside (0, 1).

Behaviour change, the second one: every rule's urlPattern is now test()ed against every request, where before the handler returned as soon as a rule injected. This is what lets two rules on the same URL agree about occurrence numbers, and it is invisible for a stateless pattern — but a RegExp carrying g or y has a lastIndex that test() writes, so an extra test used to renumber it. Those flags are now stripped when a matcher is compiled (your own RegExp object is left alone), which means a /g pattern matches what it reads as matching rather than firing on alternating requests. If you have a pinned expectation recorded against the old alternating behaviour, it will change — in the direction of what the pattern says.

Deterministic schedules

probability cannot say "fail the first call, let the retry through". schedule can: a decision table indexed by how many times the rule has already matched.

faultInjection: [
  faults.status(500, {
    urlPattern: /\/api\/cart$/,
    schedule: { decisions: ["inject", "pass"] }, // 1st call 500s, retry works
  }),
],
  • afterEnd decides what happens past the table: "pass" (default — spent), "inject" (keep firing), "repeat" (cycle it).
  • Available on all four layers (faultInjection, lifecycleFaults, runtimeFaults, iframeFaults). probability + schedule together is a validation error.
  • A schedule consumes no RNG, so adding one leaves the seed sequence — and therefore chaos action selection — untouched.
  • Faults watching the same URL (or the same iframe) share occurrence numbering on every layer: each evaluates every matching rule, so occurrence 0 can get one fault kind and occurrence 2 another, and a rule that decided inject and lost the race is counted in suppressed rather than dropped. Don't split one endpoint across the network and runtime layers, though: a client-side rejection issues no request, so the network counter never advances.

To enumerate every combination rather than the ones you thought of, see model-driven faults.

Per-rule matched / injected counters end up in report.faultInjections. At the end of a run, chaosbringer warns on the logger about every configured fault that did not take effect — on all four layers, not just the network one, and naming which of three things happened: fault_rule_unmatched (nothing matched the pattern), fault_rule_unfired (it matched and the firing policy said no), fault_rule_uncounted (the row carried no usable counter, so nothing measured it). Each carries rule and layer. Useful for catching typo'd urlPattern regexes, rules shadowed by an earlier catch-all, and schedules that never reach the occurrence you meant.

Two exported readers give you the same thing without scraping the log: faultWarnings(report) returns those events for a report you already hold, and unfiredFaults(report) returns the three diagnoses as strings — an empty array is the assertion you want in a test, since report.faultInjections[0].injected > 0 only ever covers one layer (and reads undefined > 0 on the other three).

A third counter, suppressed, appears on a row only when it is non-zero. Rules are first-match-wins, but a scheduled rule advances its occurrence whenever its pattern matches — that is what lets two rules on one URL agree about what "occurrence 1" means. So a scheduled rule can decide inject and still not act, because a rule ahead of it answered the request. Without suppressed that reads as matched: 3, injected: 0, which is exactly what an all-pass schedule reports: a planned fault that did not happen, indistinguishable from one that was never planned. RuntimeFaultStats carries the same field for the same reason, and there fired counts effects, not decisions.

Rule order: first match wins

Rules are evaluated top-to-bottom in the order you pass them, and the first match wins. This is the opposite of Playwright's raw page.route(...) API, where later registrations override earlier ones (LIFO). Put specific rules first and broad catch-alls last:

faultInjection: [
  // ✅ specific overrides first
  faults.status(200, {
    urlPattern: /^https:\/\/api\.example\.com\/p\//,
    body: '{"foo":"bar"}',
    name: "fulfill-api",
  }),
  // ✅ catch-all last
  faults.abort({ urlPattern: /^https?:\/\/(?!127\.0\.0\.1)/, name: "block-external" }),
],

If you reverse the order, the catch-all swallows every request and fulfill-api will show matched: 0 in the report (and trigger the unmatched-rule warning described above). Chaosbringer also runs a best-effort static check at crawl start and emits one fault_rule_shadowed warning per shadowed pair — catching the common pathological case (broad regex before specific one) before the crawl has consumed a single request.

Fault profiles

Hand-authored probabilities are easy to start with but hard to share. Profiles wrap operator knowledge — "S3 503 burst", "flaky third-party CDN", "regional degradation" — into a single function that returns a ready-made array of fault rules:

import { chaos, profiles } from "chaosbringer";

await chaos({
  baseUrl,
  faultInjection: [
    ...profiles.flakyThirdPartyCdn(/cdn\.example\.com/),
    ...profiles.s3FivexxBurst(/s3\.amazonaws\.com/),
    ...profiles.regionalDegradation({ urlPattern: /\/api\//, severity: 0.3 }),
    ...profiles.slowAuthService(/\/auth\//),
    ...profiles.partialDataLoss(/\/api\/feed/),
  ],
});

Available profiles (all return FaultRule[]):

| Profile | What it models | |---|---| | flakyThirdPartyCdn(urlPattern) | Slow + occasional drops on a third-party CDN | | s3FivexxBurst(urlPattern) | Mostly 503 with a 500 sprinkle — retry-storm provocation | | regionalDegradation({ urlPattern, severity }) | Severity-scaled mix of slow / 5xx / drop. severity is clamped to [0, 1] | | slowAuthService(urlPattern, { ms?, rate? }) | One slow dependency. Defaults: 3000 ms, rate 0.5 | | partialDataLoss(urlPattern, { rate? }) | Empty 200 body + occasional 5xx — catches JSON.parse(\"\") bugs |

Each rule is named (profile:behavior), so per-profile counters land in report.faultInjections without extra wiring. Override knobs by passing options or compose faults.* directly when a profile doesn't fit.

Trace correlation (W3C traceparent)

When the server is OTel-instrumented, it's useful to find the server-side trace that corresponds to a specific browser-driven action. Enable traceparent: true and chaosbringer will inject a fresh W3C traceparent header onto every request the browser sends:

await chaos({
  baseUrl: "http://localhost:3000",
  traceparent: true,
});

To capture the generated trace IDs in your own report, pass an onInject hook:

const traceIds: Array<{ url: string; traceId: string }> = [];

await chaos({
  baseUrl: "http://localhost:3000",
  traceparent: {
    onInject: ({ url, traceId, existing }) => {
      // `existing` is true when the request already carried a traceparent
      // (e.g. set by an outer middleware) — chaosbringer never overwrites it.
      traceIds.push({ url, traceId });
    },
  },
});

The injected header is the standard 00-{trace-id}-{span-id}-01 format. Existing traceparent headers are honoured (passed through unchanged), so explicit upstream propagation always wins. The fault-injection layer also keeps the header attached, so a fault response and the matching server-side trace share the same correlation id.

Pre-run setup hook

State-driven apps (CRUD, anything with a list) often start empty — the BFS frontier dries up at pages=2 and maxPages becomes meaningless. chaos({ setup }) runs before the crawler starts, in a disposable browser context, and gives you a page to seed backend state.

await chaos({
  baseUrl: "http://localhost:3000",
  setup: async ({ page, baseUrl }) => {
    for (let i = 0; i < 5; i++) {
      await page.request.post(`${baseUrl}/api/todos`, {
        data: { title: `seed-${i}` },
        // pair with @mizchi/server-faults' bypassHeader to keep seeds out of the chaos surface
        headers: { "x-chaos-bypass": "1" },
      });
    }
  },
  // ... maxPages, faultInjection, etc.
});

The setup browser is closed before the crawler starts; carry shared state through the server (REST seed) or by saving storageState to a file and pointing options.storageState at it.

Lifecycle faults (client-side)

faultInjection is request-scoped; lifecycle faults are page-scoped client-side perturbations that fire at well-defined stages of every page visit. Use them to simulate slow CPUs, stale auth tokens, evicted Service Worker caches, and other browser-side conditions that aren't expressible at the network layer.

import { chaos, faults } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  lifecycleFaults: [
    // Throttle the CPU 4× before navigation, so the load itself is slow.
    faults.cpu(4),

    // Wipe localStorage + cookies right after the page loads.
    faults.clearStorage({ scopes: ["localStorage", "cookies"] }),

    // Drop every Service Worker cache before chaos clicks fire — only on /app/*.
    faults.evictCache({ urlPattern: /\/app\// }),

    // Replace the auth token with an expired value on the dashboard, with a
    // 50% probability per visit.
    faults.tamperStorage({
      scope: "localStorage",
      key: "auth_token",
      value: "expired",
      urlPattern: /\/dashboard/,
      probability: 0.5,
    }),
  ],
});

Stages

Each lifecycle fault declares a when stage:

| Stage | Fires | Typical use | | --- | --- | --- | | beforeNavigation | Before page.goto. | CDP-level conditions that need to apply during the load (CPU throttle). | | afterLoad | Right after navigation, before afterLoad invariants. | In-page mutations (storage wipes / tamper). | | beforeActions | After afterLoad invariants, before chaos clicks. | One-shot evictions that should not affect invariants but should precede user simulation (Service Worker cache). | | betweenActions | After every chaos action. | Sustained-pressure faults that need re-application across the action loop. |

Helpers default to a sensible stage per action kind (cpu → beforeNavigation, clearStorage / tamperStorage → afterLoad, evictCache → beforeActions); pass when to override.

Action kinds

  • faults.cpu(rate, opts?) — rate ≥ 1 multiplier (1 = no throttle, 4 ≈ 4× slower) applied via CDP Emulation.setCPUThrottlingRate.
  • faults.clearStorage({ scopes, ... }) — wipes one or more of localStorage, sessionStorage, cookies, indexedDB. Cookies are cleared at the BrowserContext level; the rest run in-page via page.evaluate.
  • faults.evictCache(opts?) — drops entries from the Service Worker caches API. With no cacheNames, every cache is dropped.
  • faults.tamperStorage({ scope, key, value, ... }) — sets a single key in localStorage or sessionStorage. Useful for forcing logged-in apps into "stale auth token" / "corrupted client state" scenarios without touching the rest of storage.

Common options

Every lifecycle helper accepts the same overrides:

| Option | Description | | --- | --- | | when | Override the helper's default stage. | | urlPattern | Restrict the fault to URLs matching this regex / regex string. Omit to apply on every page. | | probability | 0..1, default 1. Uses the crawler's seeded RNG so the firing pattern is reproducible. RNG is consumed only when probability is in (0, 1) — adding a probability-1 (or probability-0) fault doesn't shift the seed sequence for chaos action selection. Note that probability: 0 used to consume a draw on the network layer; see the behaviour change above. | | name | Override the auto-derived stats label (e.g. cpu-throttle:4x). |

Stats

Every fault gets one row in report.lifecycleFaults with matched (URL-pattern matches), fired (post-probability), and errored (executor threw — e.g. SecurityError on opaque origins). Misbehaving faults are caught and counted; they never abort the rest of the crawl.

{
  "lifecycleFaults": [
    { "name": "cpu-throttle:4x", "matched": 12, "fired": 12, "errored": 0 },
    { "name": "clear-storage:localStorage", "matched": 12, "fired": 6, "errored": 0 },
    { "name": "tamper-storage:localStorage.auth_token", "matched": 3, "fired": 1, "errored": 0 }
  ]
}

Like network-side fault injection, lifecycle faults are programmatic-only — they're not expressible as flat shell flags and so are absent from the CLI.

Runtime faults (in-page monkey-patches)

runtimeFaults is a third fault layer, distinct from request-scoped faultInjection and stage-scoped lifecycleFaults. Each entry is a persistent monkey-patch installed via addInitScript on every page navigation, subverting in-page JS APIs so the app sees client-side failures that no network mock would expose.

import { chaos, faults } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  seed: 42,
  runtimeFaults: [
    // 30% of fetch() calls reject with "Failed to fetch" before any
    // network round-trip — exposes Service Worker fallbacks, retry
    // logic, and "offline indicator" code paths.
    faults.flakyFetch({
      urlPattern: /\/api\//,
      probability: 0.3,
      rejectionMessage: "simulated network failure",
    }),
    // Skew the clock 25 minutes forward on /dashboard pages — surfaces
    // token-expiry and cache-bust bugs without waiting real time.
    faults.clockSkew(25 * 60_000, { urlPattern: /\/dashboard/ }),
  ],
});

Promise-shaped kinds, for the failure modes a network mock cannot express:

| Helper | What the app sees | Bug it exposes | | --- | --- | --- | | faults.rejectFetch({ rejectAs }) | fetch rejects with a TypeError (default) or a DOMException named AbortError | Handlers that branch on instanceof TypeError; a retry banner shown on a user cancel | | faults.rejectBody() | fetch resolves, then res.json() rejects | The classic missed catch: guarded fetch, unguarded await res.json() | | faults.neverSettleFetch() | the promise never settles, no request is issued | Missing timeout. Because nothing is in flight, networkidle still fires — the UI simply never leaves loading | | faults.rejectedThenable() | same rejection, one microtask later, via thenable assimilation | Handlers attached too late |

faults.flakyFetch() still works; it is rejectFetch({ rejectAs: "TypeError" }). One thing to know if you migrate: the stats label changes with it. An unnamed flakyFetch reported rule: "flaky-fetch" and the rejectFetch form reports rule: "reject-fetch:TypeError", so anything matching on report.runtimeFaults[].rule needs updating — or pass an explicit name and stop depending on the derived one.

faults.status(500, { urlPattern }) with no body does not send an empty body — the default is {"error":500} with content-type: application/json. That default decides which app bug a 500 finds: a client that skips res.ok and calls res.json() renders junk out of it and reports success, where an HTML or empty body makes res.json() reject and the client takes its error path (or leaks an unhandled rejection). Two different defects behind one status code, so both are worth testing; pass body: "" or body: "<html>…" explicitly for the second.

(The default was originally justified by a spurious ERR_ABORTED Chromium was said to emit alongside an empty intercepted body. That does not reproduce on Chromium 147 — no requestfailed, no ERR_ABORTED on any channel, the same single console line either way — so treat the body choice as being about your client's parsing path, not about browser noise.)

The network layer gains the matching faults.hang({ urlPattern, releaseAfterMs }): the request is held open and never answered. Without releaseAfterMs the route is parked and counted in report.heldRequests, then aborted when the run is done with the page: at teardown for a page the crawler owns, and before testPage() returns for one it does not (a parked route left behind would make the caller's next action on that page wait on a request nothing will ever answer). If you drive the page yourself across several steps, crawler.release() drains on demand.

Since the crawler navigates with waitUntil: "networkidle", prefer hanging what an action fires after load, or set the bound.

Know what that costs before you read the report: a hang on a load-time request means page.goto spends its whole timeout and then throws, and that throw is recorded as a page error of type exception — so summary.jsExceptions reads 1 and an error cluster appears carrying Playwright's own timeout message. The page threw nothing. The classification is the crawler's, not the app's, and it is the expected outcome of the fault rather than a finding. report.heldRequests (now printed in the text report too) is the number that tells you which one you are looking at.

never-settle-fetch honours init.signal, which is what makes it a fair test of a bounded request rather than a way to fail every client. Note what the caller sees: under AbortSignal.timeout(ms) the rejection is a TimeoutError, not an AbortError — a catch branching only on err.name === "AbortError" misses it. An explicit AbortController.abort() still gives AbortError.

The probability roll is deterministic given the same (seed, runtimeFaults) pair — the in-page LCG is seeded from seed so two runs roll identically.

urlPattern means different things per kind: fetch-scoped kinds (flaky-fetch, reject-fetch, never-settle-fetch, reject-body, resolve-rejected-thenable) match the request URL passed to fetch(), per call; page-scoped kinds (clock-skew) match location.href once, when the init script installs.

Layer comparison:

| Layer | Where it runs | Targets | Example | | --- | --- | --- | --- | | faultInjection | Playwright route() (Node side) | individual network requests | serve 500 on /api/* | | lifecycleFaults | per-page hook (CDP / page eval) | one-shot at named stages | wipe localStorage afterLoad | | runtimeFaults | addInitScript (in-page) | persistent JS API patches | reject fetch(), skew Date.now | | iframeFaults | addInitScript (in-page) | HTMLIFrameElement.prototype.src per matching iframe | delay / starve / mid-load-remove an iframe load |

Stats land in report.runtimeFaults — one row per fault with matched (URL filtered ok, probability about to roll) and fired (actually triggered) counts.

Like the other fault layers, runtimeFaults is programmatic-only.

Iframe-load faults

Some classes of bugs only surface when a third-party library injects an iframe and the host page reacts to that iframe's load lifecycle — ad SDKs, embeddable widgets, checkout iframes, social plugins, video players. faultInjection operates on the request layer (requests inside the iframe) and lifecycleFaults operates on the host page lifecycle; neither can perturb the iframe element's own load as observed by the parent.

iframeFaults is a fourth fault layer that monkey-patches HTMLIFrameElement.prototype.src (and setAttribute("src", ...)) so faults fire the moment the host page assigns the iframe's URL — three primitives that no other layer can express:

import { chaos, faults } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  iframeFaults: [
    // Delay every ad iframe's load by 3s so the host library races a
    // visibility timer against the contained document arriving.
    faults.iframeLoadDelay(3000, { selector: "iframe.ad-slot" }),

    // 20% of the time, never fire `load` on the player iframe — swap to
    // about:blank so the host's onload-driven impression event is starved.
    faults.iframeNeverLoad({
      selector: "iframe[data-widget='player']",
      probability: 0.2,
    }),

    // 10% of the time, remove the iframe from the DOM 500ms after src
    // is set — exposes listener-teardown and pending-callback races.
    faults.iframeRemoveMidLoad({
      selector: "iframe",
      atMs: 500,
      probability: 0.1,
    }),
  ],
});

Selectors are matched via iframe.matches(selector) at the moment src is set. If the library does iframe.src = "..."; container.appendChild(iframe); (i.e. assigns src before attaching to the DOM), ancestor combinators like #container iframe won't match — prefer attribute / class selectors on the iframe itself (iframe[data-widget], iframe.ad-slot).

Why not just faults.delay({ urlPattern })?

  • faults.delay slows the request inside the iframe, which does delay the parent's iframe.onload — that case (iframe loads slowly) overlaps.
  • But there's no way to express "iframe never fires load" via route() (the response body has to actually arrive).
  • And there's no way to express "iframe removed mid-load" via route() at all — it's a DOM-side operation.

Stats land in report.iframeFaults — one row per fault with selector, action, matched (iframes whose selector matched), and fired (post-probability) counts. Like the other fault layers, iframeFaults is programmatic-only.

{
  "iframeFaults": [
    { "rule": "iframe-load-delay:3000ms", "selector": "iframe.ad-slot", "action": "load-delay", "matched": 4, "fired": 4 },
    { "rule": "iframe-never-load", "selector": "iframe[data-widget='player']", "action": "never-load", "matched": 2, "fired": 1 }
  ]
}

Invariants

Invariants are assertions that must hold on every page. They run either afterLoad (right after navigation) or afterActions (default — after chaos clicks/inputs). Returning false, throwing, or returning a string all count as a failure; returning true or void means the invariant held.

import { chaos, type Invariant } from "chaosbringer";

const invariants: Invariant[] = [
  {
    name: "has-h1",
    when: "afterLoad",
    async check({ page }) {
      return (await page.locator("h1").count()) > 0 || "no <h1>";
    },
  },
  {
    name: "no-loading-spinner-after-actions",
    urlPattern: /\/spa\//,
    async check({ page }) {
      const t = (await page.locator("#app").textContent()) ?? "";
      return !/loading/i.test(t) || `app still shows loading: "${t}"`;
    },
  },
];

const { passed } = await chaos({ baseUrl: "http://localhost:3000", invariants });

Violations always fail the run (exit 1), whether or not strict is set — a declared invariant is a stronger signal than console noise.

Trans-page state — ctx.state

Each invariant's check() receives a ctx.state: Map<string, unknown> shared with every other invariant on every page. The same instance is passed for the lifetime of one crawler.start() call and reset on the next, so invariants can carry data across pages and flag regressions that need history (monotonic counters, set-membership, ordered events).

const cartCountMonotonic: Invariant = {
  name: "cart-monotonic-after-add",
  when: "afterActions",
  async check({ page, state }) {
    const n = Number((await page.locator("[data-cart-count]").textContent()) ?? "0");
    const prev = (state.get("cart:max") as number | undefined) ?? 0;
    if (n + 1 < prev) {
      // Allow one decrement to model a removed item; flag larger drops.
      return `cart count went from ${prev} to ${n}`;
    }
    state.set("cart:max", Math.max(prev, n));
  },
};

Use state.set / state.get directly, or build on top via stateMachine() below.

State-machine invariants

For discrete app modes (anonymous → logged-in → in-checkout → purchased), invariants.stateMachine() compiles down to a regular Invariant that detects illegal transitions across pages.

import { chaos, invariants } from "chaosbringer";

type Auth = "anonymous" | "logged-in" | "in-checkout" | "purchased";

const auth = invariants.stateMachine<Auth>({
  name: "auth-flow",
  initial: "anonymous",
  // Self-loops are legal automatically. Terminal states have no outgoing edges.
  transitions: {
    anonymous: ["logged-in"],
    "logged-in": ["anonymous", "in-checkout"],
    "in-checkout": ["logged-in", "purchased"],
    // `purchased` left out → terminal: leaving it is illegal.
  },
  // Run after chaos clicks so post-action page state is reflected.
  when: "afterActions",
  async derive({ page }) {
    if (await page.locator("[data-receipt]").count() > 0) return "purchased";
    if (await page.locator("[data-checkout-step]").count() > 0) return "in-checkout";
    if (await page.locator("[data-user-id]").count() > 0) return "logged-in";
    return "anonymous";
  },
});

await chaos({ baseUrl: "http://localhost:3000", invariants: [auth] });

When derive() returns a label that the previous label's transition list doesn't allow, the invariant fails with illegal transition "<prev>" → "<next>" (allowed: …) — surfaced as a regular invariant-violation PageError, clustered like any other.

derive() receives { page, url, prev, errors } so the caller can branch on the previous label or the current URL when classifying the page.

The state-machine helper is one preset on top of ctx.state; for non-discrete properties (counters, set membership, ordered event log), drop down to a plain Invariant and use state.set / state.get directly.

Coverage-guided action selection

coverageFeedback opts the crawler into AFL-style feedback: V8 precise coverage (CDP Profiler.startPreciseCoverage / takePreciseCoverage) is collected per page, the coverage delta of every chaos action is attributed to the action target that fired it, and on subsequent visits each target's weight is multiplied by 1 + boost · log1p(score). Targets that have historically delivered new V8 functions get picked more often; dead-end targets fade.

import { chaos } from "chaosbringer";

const { report } = await chaos({
  baseUrl: "http://localhost:3000",
  seed: 42,
  maxPages: 50,
  maxActionsPerPage: 5,
  coverageFeedback: { enabled: true, boost: 2 },
});

console.log(report.coverage);
// {
//   totalFunctions: 312,
//   pagesWithNewCoverage: 18,
//   topNovelTargets: [
//     { url: "http://localhost:3000/cart", selector: "button:has-text(\"Checkout\")", score: 47 },
//     { url: "http://localhost:3000/", selector: "[role=link]:has-text(\"Sign in\")", score: 23 },
//     ...
//   ],
// }

| Option | Description | Default | | --- | --- | --- | | enabled | Master switch — attaches the coverage collector and biases weights. | required (false if omitted entirely) | | boost | Multiplier applied via 1 + boost · log1p(score). 0 keeps coverage tracked but disables the weight bias. 2 is moderate. 4+ aggressively concentrates picks. | 2 | | topN | Cap top-N novel targets emitted in report.coverage. | 20 |

Reproducibility

The collector never consumes the seeded RNG, so the seed sequence is unchanged. Action selection still differs vs a no-feedback run because the weight inputs to weightedPick are different — reproducibility is now (seed, coverageFeedback) rather than seed alone. The Repro: line emitted by the report only encodes flags expressible on the CLI, so a coverage-feedback run is reproducible programmatically (same chaos({...}) config) but not from the CLI alone.

Cost

Each chaos action incurs one Profiler.takePreciseCoverage CDP roundtrip. Chromium-only (the API is Chrome-DevTools-Protocol-specific). For typical chaos runs (≤100 pages, ≤5 actions per page) the overhead is below 5%; expect a noticeable slowdown on heavy SPAs with hundreds of scripts.

Device emulation & network throttling

Emulate mobile devices or throttle the network to catch bugs that only surface on slow connections or small viewports.

chaosbringer --url http://localhost:3000 --device "iPhone 14" --network slow-3g
  • --device <name> — any Playwright device descriptor (iPhone 14, Pixel 7, iPad Pro 11, Desktop Chrome, …). Sets viewport, user-agent, device pixel ratio, mobile / touch flags via newContext({ ...devices[name] }). Unknown names fail validation up-front.
  • --network <profile> — slow-3g, fast-3g, or offline. Attaches a CDP session per page and calls Network.emulateNetworkConditions with the same values Chrome DevTools' presets use.

Combining the two lets you measure perf budgets under realistic conditions: chaosbringer --url … --device "Pixel 7" --network slow-3g --budget lcp=4000.

Sitemap seeding

Prepend every URL in a sitemap.xml (or sitemap index) to the crawl queue — essential for sites whose nav is JS-rendered and so gets missed by DOM link extraction.

chaosbringer --url https://docs.example.com --seed-from-sitemap https://docs.example.com/sitemap.xml

Accepts a URL or a local path. Sitemap indexes are followed breadth-first; referenced URLs outside the baseUrl origin are dropped to avoid wasting visit budget. A runaway index (suspected cycle) fails fast.

import { fetchSitemapUrls } from "chaosbringer";
const urls = await fetchSitemapUrls("https://docs.example.com/sitemap.xml");

Authenticated crawls (storage state)

To crawl pages behind a login, point chaosbringer at a Playwright storageState file — the JSON containing cookies + localStorage that a logged-in browser context produces. Run a one-off login script once, save the state, then reuse it for every chaos run.

// auth-setup.ts — run once, or as a Playwright global setup
import { chromium } from "playwright";

const browser = await chromium.launch();
const context = await browser.newContext();
const page = await context.newPage();
await page.goto("http://localhost:3000/login");
await page.getByLabel("Email").fill("[email protected]");
await page.getByLabel("Password").fill(process.env.TEST_PASSWORD!);
await page.getByRole("button", { name: "Sign in" }).click();
await page.waitForURL("**/dashboard");
await context.storageState({ path: "auth.json" });
await browser.close();
# Chaos-test the authenticated surface
chaosbringer --url http://localhost:3000/dashboard --storage-state auth.json
await chaos({
  baseUrl: "http://localhost:3000/dashboard",
  storageState: "auth.json",
});

The file is read by Playwright and not modified by the crawl. If the session expires mid-run, you'll see auth-redirect pages surface as errors — regenerate the state file and rerun.

HAR record / replay

Chaosbringer can capture network traffic to a HAR file on one run and replay it on the next. A replay run is deterministic even if the backend is flaky — every request that was in the HAR gets served from the HAR, not the network.

# First run: capture responses
chaosbringer --url http://localhost:3000 --seed 42 --har-record chaos.har

# Later: replay without the server running
chaosbringer --url http://localhost:3000 --seed 42 --har-replay chaos.har

Programmatic:

await chaos({
  baseUrl: "http://localhost:3000",
  seed: 42,
  har: { path: "chaos.har", mode: "record" },
});

// Replay
await chaos({
  baseUrl: "http://localhost:3000",
  seed: 42,
  har: { path: "chaos.har", mode: "replay", notFound: "abort" },
});
  • notFound: "fallback" (default) lets unmatched URLs fall through to the real network.
  • notFound: "abort" fails them — useful when you want to prove a run is fully deterministic.
  • Fault injection rules still apply in replay mode and take precedence over HAR responses.

Heads up — notFound: "abort" with traceparent injection: when traceparent is enabled, every outgoing request carries a freshly-generated traceparent header that wasn't in the recorded HAR. Depending on your Playwright version and HAR matcher, this can cause notFound: "abort" to fail on every request during replay. If you hit this, either record the HAR with traceparent also enabled (so the matcher sees consistent headers), set traceparent: false for replay-mode runs, or use notFound: "fallback".

Accessibility (axe-core)

Install axe-core as a peer and opt in with either the invariants.axe() preset or the --axe flag. Each visited page is scanned; violations are reported as invariant failures (name: a11y-axe), which always fail the run.

pnpm add axe-core
chaosbringer --url http://localhost:3000 --axe
chaosbringer --url http://localhost:3000 --axe --axe-tags wcag2aa,best-practice
import { chaos, invariants } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  invariants: [
    invariants.axe({
      tags: ["wcag2aa"],
      exclude: [".third-party-widget"],
      disableRules: ["color-contrast"],
    }),
  ],
});

axe-core is an optional peer dependency — the preset fails with a clear install hint if it isn't present. The preset is thin; drop to a custom invariant if you need multiple axe runs per page, per-URL rule overrides, or full-result capture (passes / incomplete).

A failing scan is rendered on one line: [a11y-axe] 3 a11y violations: color-contrast(×5, serious), image-alt(×2, critical), region(×1). Because violations cluster by their fingerprint, a11y regressions show up in the baseline diff just like any other invariant.

Action heatmap

Aggregate report.actions[] into per-target stats — count, success rate, blocked-external count, shard-skipped count — sorted by frequency. Useful when you want to know which targets the chaos driver is hitting most and which ones disproportionately fail.

chaosbringer --url http://localhost:3000 --heatmap --heatmap-top 30
chaosbringer --url http://localhost:3000 --heatmap-out heatmap.json
import { buildActionHeatmap, formatHeatmap, chaos } from "chaosbringer";

const { report } = await chaos({ baseUrl: "http://localhost:3000" });
const entries = buildActionHeatmap(report.actions);
console.log(formatHeatmap(entries, 20));
// entries is sorted by count desc, then failureCount desc, then key asc.

It's pure aggregation over the existing actions array — works on any report (current run, baseline, or one loaded from disk). Action types remain distinct, so click Search and input Search count separately.

JUnit XML output

Render the report as Surefire-style junit.xml so existing CI dashboards (Jenkins, CircleCI, GitLab CI, GitHub Actions test summaries, Allure) ingest chaosbringer runs without bespoke parsing.

chaosbringer --url http://localhost:3000 --junit junit.xml
import { buildJunitXml, chaos } from "chaosbringer";
import { writeFileSync } from "node:fs";

const { report } = await chaos({ baseUrl: "http://localhost:3000" });
writeFileSync("junit.xml", buildJunitXml(report, { suiteName: "smoke" }));

Mapping is one <testcase> per visited page:

  • status="error" / "timeout" → <error> (HTTP code or timeout in type)
  • status="success" with errors[].length > 0 → <failure> (concatenates all PageError entries)
  • otherwise → passing testcase, no children

Test names strip the baseUrl prefix so /docs/intro shows up rather than the full URL. Special XML chars (< > & " ') in messages and URLs are escaped.

Visual regression

Compare each page's screenshot against a baseline PNG on disk. Differences beyond the configured budget are recorded as invariant violations (visual-regression), which fail the run.

# First run: baselines don't exist yet — chaosbringer records them and passes.
chaosbringer --url http://localhost:3000 --visual-baseline ./__snapshots__

# Subsequent runs: compare against the recorded baselines.
chaosbringer --url http://localhost:3000 --visual-baseline ./__snapshots__ \
  --visual-max-diff-pixels 100 \
  --visual-diff-dir ./__diffs__

# After an intentional UI change, overwrite the baselines.
chaosbringer --url http://localhost:3000 --visual-baseline ./__snapshots__ --visual-update
import { chaos, invariants } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  invariants: [
    invariants.visualRegression({
      baselineDir: "./__snapshots__",
      threshold: 0.1,          // pixelmatch color distance (0..1)
      maxDiffPixels: 100,      // absolute tolerance
      maxDiffRatio: 0.001,     // or proportional tolerance
      diffDir: "./__diffs__",
    }),
  ],
});
  • Baseline filenames are derived from each page's URL (path + query, sanitized) so different routes don't collide.
  • pixelmatch and pngjs are optional peer deps — install them explicitly (pnpm add pixelmatch pngjs). The invariant fails with a clear install hint when they're missing.
  • Dimension mismatches between baseline and current are treated as full-diff failures — resize or re-record the baseline intentionally rather than auto-accepting.
  • Takes fullPage: true screenshots by default. Flip to viewport-only via fullPage: false in the programmatic API if your layout is sensitive to scroll position.
  • Pair with --device iPhone 14 to record device-specific baselines; the baseline dir is per-crawl so split baselines across devices by using different dirs.

Failure artifact bundles

When a page errors, times out, recovers from a 4xx/5xx, or surfaces an invariant violation, the crawler can dump a self-contained bundle so the failure is reproducible without re-running the whole crawl.

chaosbringer --url http://localhost:3000 --failure-artifacts ./failures
import { chaos } from "chaosbringer";

await chaos({
  baseUrl: "http://localhost:3000",
  failureArtifacts: { dir: "./failures", maxArtifacts: 50 },
});

Each failing page becomes a numbered subdirectory under --failure-artifacts <dir>:

failures/0000__checkout_review__a91c2f0e/
├── screenshot.png   # full-page PNG at the moment of failure
├── page.html        # `await page.content()` snapshot
├── errors.json      # full PageError[] (console / exception / network / invariant)
├── trace.jsonl      # meta + visits + actions, sliced up to and including this page
├── repro.sh         # `chaosbringer --url <base> --trace-replay ./trace.jsonl`
└── info.json        # URL, status, sourceUrl, recovery, seed, timestamps

repro.sh is executable — cd into the bundle and run it to replay the same sequence locally. Combine with --strict or --baseline to gate CI on the same shape of failure.

maxArtifacts caps the bundle count per run for runs that produce many failures (default: unlimited). Per-artifact opt-outs are available programmatically (saveScreenshot: false, saveHtml: false, saveTrace: false) when bundle size matters more than completeness.

Performance budget

Declare a per-metric budget (in ms). Any page whose measured metric exceeds its limit is recorded as an invariant violation (perf-budget.<metric>), which fails the run just like any other invariant.

# CLI — comma-separated pairs, or repeat the flag
chaosbringer --url http://localhost:3000 --budget ttfb=200,fcp=1800,lcp=2500
await chaos({
  baseUrl: "http://localhost:3000",
  performanceBudget: { ttfb: 200, fcp: 1800, lcp: 2500 },
});

Supported keys: ttfb, fcp, lcp, tbt, domContentLoaded, load. Omitted keys are not enforced. Metrics that weren't captured (e.g. lcp on a page that didn't render anything large) don't produce violations — only observed-and-over-limit cases do.

Where the numbers come from, measured once the page load has settled:

  • ttfb, fcp, domContentLoaded, load — Navigation and Paint Timing.
  • lcp — the latest Largest Contentful Paint candidate reported by web-vitals, through the lightbringer collector the crawler installs on every page.
  • tbt — the sum of duration − 50 ms over long tasks that started at or after FCP. This is an approximation: real Total Blocking Time runs from FCP to Time to Interactive, and the crawler has no TTI, so tbt covers FCP to the end of the load instead.

The collector is installed before every other init script, so a clock-skew runtime fault cannot shift lcp or tbt. On a page the crawler does not own (testPage() on a page opened elsewhere), lcp and tbt are absent rather than 0, and their budgets are not enforced.

Budget violations are clustered by metric name, so perf-budget.lcp firing on 20 pages shows up as one cluster with count: 20 in the report and the baseline diff.

Per-step performance (--perf)

performanceBudget sees one number per page. --perf measures every step: each page load and each chaos action becomes a span, with its own network, main-thread blocking, render, memory and interaction-latency breakdown. The measuring engine is lightbringer, running on the page's shared CDP session.

chaosbringer --url http://localhost:3000 --seed 42 --perf
chaosbringer --url http://localhost:3000 --perf-out perf/          # + a full JSON report per page
chaosbringer --url http://localhost:3000 --perf-trace --perf-out perf/   # + a Chrome trace per page (large)
await chaos({
  baseUrl: "http://localhost:3000",
  perf: true, // ≡ { level: "light" }
  // perf: { level: "trace", memory: { forceGc: true }, coverage: true, outDir: "perf", actions: true },
});

The text report gains one line per page and the five slowest actions:

PER-STEP PERFORMANCE
  /items/:id  load 412ms  blocking 0ms  6 req / 38.2KB  LCP 180ms
Slowest actions:
    2104ms  blocking 132ms  interaction 148ms  /items/:id :: click #buy

Spans (PageResult.perf, ActionResult.perf), their stable perfKeys, levels, overhead and artefacts are in docs/recipes/perf.md.

Before you trust a number:

  • Under the default networkidle settle, a span's durationMs is mostly the crawler's wait. Rank hotspots by busyMs, blockingMs and scriptMs, or crawl with --settle adaptive, where durationMs follows the app. The settle mode is recorded in the report, and perf gate / regress refuse (exit 2) to compare runs crawled under different modes.
  • --settle adaptive does not wait for timers, so errors a page throws a few hundred ms after it goes quiet are missed. Hunt late async errors under networkidle or a long fixed settle.
  • Leak trends need --perf-mem. Without the forced GC, a step's uncollected garbage reads as growth, and non-leaking pages get flagged.
  • A probabilistic fault shares the crawler's RNG, so the same seed with and without it is a different crawl. Compare within one run (perf.degradation) or use a schedule.
  • Gate CPU regressions on scriptMs, which scales with the work; blockingMs reads 0 until a task crosses 50 ms. For hand-set budgets on a shared runner, allow +20–25% or +20 ms, whichever is larger, over the median of 3–5 runs.
  • A click that navigates has no interactionMs, and work a harness runs through page.evaluate does not show in blockingMs.

The details, with the measurements behind them, are in perf.md's settling, noise and caveats sections.

Trace record / replay / minimize

For failures that are hard to diagnose from a seed alone, record the exact sequence of visits + actions to a JSONL file, then replay or minimize that sequence.

# Record
chaosbringer --url http://localhost:3000 --seed 42 --trace-out chaos.trace.jsonl

# Replay the exact sequence (no RNG, no discovery)
chaosbringer --url http://localhost:3000 --trace-replay chaos.trace.jsonl

# Shrink the trace to the minimum subsequence that still reproduces a failure
chaosbringer minimize --url http://localhost:3000 \
  --trace chaos.trace.jsonl \
  --match "Cannot read properties of undefined" \
  --trace-out min.trace.jsonl

A trace is line-delimited JSON: a leading meta entry with the seed + baseUrl, then alternating visit and action lines. Each action carries the selector that was clicked (or the scroll amount, or the input target), so replay can locate the same element in a fresh page. The format version is tracked — parsing refuses traces written by incompatible future versions rather than silently misinterpreting them.

Replay skips link discovery and the RNG entirely: only URLs listed as visit entries are loaded, and only the recorded actions are performed. Missing selectors are logged as failed actions and the run continues.

minimize drives repeated replays via delta debugging (ddmin) — it keeps removing action entries and re-running as long as --match still fires against an error cluster. Output goes to --trace-out (defaults to min.trace.jsonl).

Baseline diff (regression detection)

Pass a previous report to --baseline and the current run is diffed against it — new error clusters and newly failing pages are surfaced separately from ones that were already broken.

# First run: writes chaos-report.json as usual (no baseline yet, warns and continues)
chaosbringer --url http://localhost:3000 --baseline chaos-report.json

# Subsequent runs: compare against the prior report
chaosbringer --url http://localhost:3000 --baseline chaos-report.json --baseline-strict
  • --baseline <path> — diff against this report. A missing file produces a warning, not an error (the run still writes its own report so a later invocation has a baseline to compare against).
  • --baseline-strict — exit 1 when the diff contains new clusters or newly failing pages. Resolved / unchanged entries never fail the run.

Programmatic:

import { chaos } from "chaosbringer";

const { report, passed } = await chaos({
  baseUrl: "http://localhost:3000",
  baseline: "chaos-report.json",
  baselineStrict: true,
});

for (const c of report.diff?.newClusters ?? []) {
  console.log(`NEW [${c.type}]×${c.after}: ${c.fingerprint}`);
}

Clusters are matched by the same fingerprint used for errorClusters (URL / line:col / long numeric ids stripped), so HTTP 500 on /api/users/42 and HTTP 500 on /api/users/99 collapse to the same entry. Pages are matched by URL.

GitHub Actions annotations

Opt in with --github-annotations and chaosbringer prints a workflow command for every error cluster and dead link. GitHub surfaces these on the Checks tab alongside test output.

chaosbringer --url http://localhost:3000 --strict --github-annotations

Severity maps from cluster type: invariants / exceptions / network errors / crashes are ::error, console errors and unhandled rejections are ::warning (upgraded to error under --strict). Dead links always annotate as error with the source page in the message.

Model-driven fault coverage

Probability sampling tells you what fired; it cannot tell you what was never attempted. A temporal-logic model (Quint, or anything emitting ITF) enumerates the failure space instead, and each enumerated state replays as one deterministic run with the model's prediction as the oracle:

# dev-time: ITF witnesses -> committed plan files (pure Node)
chaosbringer model compile --traces model/traces --out model/plans

# CI: replay every plan, check every oracle (no Quint, no JVM)
chaosbringer model run --plans model/plans --url http://localhost:3000 \
  --config model/bridge.mjs
=== MODEL COVERAGE ===
States: 16/18 reachable (depth <= 4), 2 unreachable
Plans run: 16
Mismatches: 13
  [unhandledRejection] cart-fulfilled__shipping-rejected: a rejection escaped every handler, which the model's contract forbids
  [ui] cart-hung__shipping-fulfilled: model predicted ui="error", page reported "stuck"

Programmatic equivalent: compilePlan / runPlans / aggregateCoverage, exported from the package root. The runner checks three things per plan — the UI label via your uiProbe, whether a rejection escaped, and whether the planned faults actually fired (a plan whose request the app never issues is reported, not counted as a pass).

A checker returns the first counterexample it reaches, not the smallest, so a failing plan is routinely longer and harsher than the bug. model shrink minimises one:

chaosbringer model shrink --plan model/plans/refresh-storm.plan.json \
  --url http://localhost:3000 --config model/bridge.mjs --out min.plan.json
2 step(s) -> 1 over 5 run(s), preserving unhandledRejection
1-minimal: every remaining edit was tried and none of them still fails.

It drops steps that do not matter, weakens outcomes that need not be that strong (hang → status), and lowers occurrences that need not be that late — every candidate a real run judged by the same oracle, so the minimum provably still fails, and the same way: a candidate that breaks differently is a different finding and is rejected.

Only contract findings can be shrunk — an escaping rejection, a uiInvariant violation. expect.ui and expect.state are what the model predicted for that exact schedule, and a smaller schedule has no recomputed prediction, so shrinking on them would "minimise" a plan to one that injects nothing and still call it a reproduction. Those exit 1 with schedule-relative rather than being answered wrongly. Likewise the search exits 0 only when it actually finished: running out of runs, or hitting a candidate the oracle could not judge, exits 1 and says which, because a minimum nobody established is not a minimum. shrinkPlan is the exported equivalent.

Full walkthrough: model-driven faults. Runnable: examples/model-faults/.

Error clustering

CrawlReport.errorClusters collapses repeated errors so a run with 100 identical console.error("Failed to load X") calls surfaces as one cluster line with count: 100. Each cluster is keyed by type + a normalised fingerprint (URLs, line:col, and long numeric ids stripped).

ERROR CLUSTERS
  [console]×42 [5 urls] Failed to load resource: the server responded with a status of <n> (Not Found)
  [exception]×3 fixture: boom

Use it to triage noisy fuzz runs — high-count clusters are the first thing to look at.

Exit codes

| Condition | Exit | | --- | --- | | No navigation errors, no invariant violations | 0 | | At least one page with status: "error" or "timeout" | 1 | | At least one invariant violation (any mode) | 1 | | --strict and any console error / JS exception / unhandled rejection | 1 |

chaos() returns { passed, exitCode }; the CLI applies the same rule via getExitCode.

Behaviour change: --strict now also fails on an unhandled rejection. It used to ignore them, so a run whose entire finding was "the app left a rejection unhandled" exited 0 while a single console.error exited 1 — and an escaping rejection is the failure mode this library's Promise fault kinds exist to produce. If you were running --strict over a page with a known unhandled rejection, that run now fails; add the pattern to ignoreErrorPatterns, or fix it.

Playwright Test integration

Use the pre-configured chaosTest:

import { chaosTest, chaosExpect } from "chaosbringer";

chaosTest("chaos-test homepage", async ({ page, chaos }) => {
  await page.goto("http://localhost:3000");
  const result = await chaos.testPage(page, page.url());
  chaosExpect.toHaveNoExceptions(result);
  chaosExpect.toLoadWithin(result, 3000);
});

Or extend your existing test:

import { test as base } from "@playwright/test";
import { withChaos, type ChaosFixtures } from "cha