redstorm
v2.1.0
Published
REDSTORM — distributed HTTP load testing engine with a realtime web dashboard
Maintainers
Readme
REDSTORM
Built by N4OD.
A distributed HTTP load-testing engine with a realtime web dashboard — a from-scratch
alternative to hey and boom.
Multi-core request generation on undici, HDR-histogram latency percentiles, 10 Hz telemetry over WebSocket, and a React dashboard you can pause, ramp and compare from.
# straight from a clone — run the CLI through npm
npm install && npm run d:install && npm run d:build
npm start -- -z 30s -c 100 https://api.example.com/health
npm run web # dashboard on http://127.0.0.1:3000
# or install the command globally, then drop the `npm start --` prefix
npm install -g redstorm # from a clone: npm install -g . (or npm link)
redstorm -z 30s -c 100 https://api.example.com/healthContents
- Why it exists
- Install
- CLI
- Dashboard
- Engine architecture
- HTTP API
- Reports and export
- Traffic profiles
- Dynamic request templating
- Testing
- Publishing
- Design notes
Why it exists
hey and boom are single-process and give you a number at the end. REDSTORM is built around
the assumption that you want to watch a system degrade in real time, intervene while it is
happening, and then compare the run against a baseline.
| | hey / boom | REDSTORM | |---|---|---| | Concurrency | single event loop | one worker thread per core, shared-memory counters | | Latency | mean + stddev | HDR histogram, p50→p99.9, sub-µs resolution | | Feedback | none until exit | 10 Hz WebSocket stream + SSE fallback | | Control | Ctrl-C only | pause, resume, live ramp, abort | | Failure modes | refused/socket errors lumped in | HTTP status vs transport error separated and classified | | Comparison | n/a | 50-run history with a side-by-side diff | | Protocol | HTTP/1.1 | HTTP/1.1, HTTP/2, optional HTTP/3 |
Install
Requires Node.js 20 or newer.
npm install # engine + server
npm run d:install # dashboard dependencies (React, Vite, Tailwind)
npm run d:build # bundle the dashboard into src/dashboard/distThe dashboard is optional. Without dist/, redstorm --web still serves the API and realtime
stream and shows a placeholder page telling you how to build it.
These install the project locally; they do not put a redstorm command on your PATH. From a
clone, invoke the CLI through npm — npm start -- -z 30s -c 100 <url> — or run it directly
with node bin/redstorm.js. To get the bare redstorm command everywhere:
npm install -g redstorm # from the registry
npm install -g . # from a clone
# or, to symlink the clone into your global bin: npm linkCLI
redstorm -z 30s -c 100 https://api.example.com/health| Flag | Description |
|---|---|
| -z, --duration <d> | run duration (30s, 2m, 1h) |
| -n, --requests <n> | total request budget; stops exactly on the last one |
| -c, --concurrency <n> | virtual users |
| --connections <n> | max TCP connections per worker |
| --threads <n> | worker threads (default: all cores) |
| -q, --rps <n> | fleet-wide rate limit |
| -m, --method <m> | HTTP method |
| --http-version <v> | 1.1, 2, or 3 |
| --pipelining <n> | HTTP/1.1 requests in flight per socket |
| --timeout <ms> | per-request timeout |
| -H, --header <k:v> | extra header (repeatable) |
| --header-template <k:v> | header with {{dynamic}} values (repeatable) |
| -d, --body <s> | request body with placeholders |
| --json <s> / --json-template <s> | JSON body, static or templated |
| --body-file <p> | read the body from a file |
| --mock-file <p> | JSON dataset exposed to templates as {{mock}} |
| --token <t> | auth token for {{token}} rotation (repeatable) |
| --expect <code> | treat any other status as an error |
| --error-body-bytes <n> | bytes of each distinct error body to keep in the report (0 disables, default 512) |
| --assert-p95/p99/mean <ms> | fail the run if that latency exceeds the value (bare number = ms) |
| --assert-errors <pct> | fail the run if the error rate exceeds this percent (5 or 5%) |
| --assert-rps <n> | fail the run if average throughput drops below this |
| --correct-omission | inject the requests a stall hid when grading latency (needs -q) |
| --profile <t> | static, ramp, wave |
| --ca/--cert/--key | CA, client cert, client key for mTLS |
| --insecure | skip TLS verification |
| --label <text> | name the run; shown in the report and in the dashboard history |
| -o, --output <p> | write the JSON report to a file, or to stdout with - |
| --web, -p, --port | start the dashboard and open it in the browser |
| --host, --strict-port | bind address; fail instead of falling back to an ephemeral port |
| --no-open | don't open the dashboard in a browser |
| --quiet, --no-color | output control |
The live renderer and the report are suppressed automatically when stdout is not a
terminal, so redstorm ... > report.txt produces clean text. Unicode meter glyphs fall back to
ASCII for the same reason — the report is meant to be pasted into a ticket.
In a terminal the live block is redrawn in place rather than appended, so a long run keeps its scrollback and the report stays visible at the end. Where colour is unavailable, frames are appended instead, at a slower rate.
Exit codes
| Code | Meaning |
|---|---|
| 0 | success |
| 1 | runtime failure (the dashboard could not bind, or an unexpected error) |
| 2 | usage error — unknown flag, missing URL, unparseable threshold |
| 3 | an --assert-* threshold was breached |
| 4 | the run recorded errors and no error budget was declared |
| 130 | interrupted at the keyboard |
CI can tell a regression from a typo:
redstorm -z 30s --assert-p99 250 --expect 200 https://api/health
case $? in
0) echo "healthy" ;;
3) echo "SLO regressed" ;;
4) echo "request failures" ;;
*) echo "could not run the test" ;;
esac--expect 200 makes any non-200 an error; drop it if the endpoint legitimately returns
other 2xx codes.
Machine-readable output
-o - writes the report to stdout and nothing else — no live renderer, no human report — so
it can be piped:
redstorm -z 30s https://api.example.com -o - | jq '.latency.p99, .totals.errorRate'Thresholds
--assert-* turns REDSTORM into a pass/fail gate, which is the point of running a load
test in CI. Declared thresholds are graded against the final report, printed as a
Thresholds section, and reported as exit code 3:
# Fail unless p99 stays under 250ms, errors under 1%, and throughput holds 500 rps
redstorm -z 30s -c 100 https://api/health \
--assert-p99 250 --assert-errors 1 --assert-rps 500 Thresholds
x p99 latency at most 250ms observed 431.2ms
+ error rate at most 1% observed 0.04%
+ throughput at least 500 observed 1187Two rules are worth knowing, because they exist to avoid false failures:
- A declared error budget replaces the strict default. Without
--assert-errors, any recorded error exits4. Passing--assert-errors 5means "up to 5% is acceptable" and that budget is used instead — otherwise the flag could never pass. - Aborted runs are skipped, not failed. If you Ctrl+C a run, a p99 gate over a
partial sample is noise, so thresholds report as
skippedand the run is not failed. The same applies to latency gates when the target returned nothing at all to measure.
Latency values accept 250, 250ms, or 0.25s; a bare number is milliseconds,
matching --timeout. Error rates accept 5 or 5%, and a bare fraction like .5
is rejected rather than guessed at, since 0.5% and 50% are very different gates.
Coordinated omission
redstorm -z 30s -c 50 -q 500 --correct-omission https://api/healthA closed-loop client blocked on a slow response does not send the requests it meant to
send while it waits, so their latencies are never measured and the tail is understated —
the coordinated-omission bias. --correct-omission applies the HdrHistogram remedy: each
observed latency also records what the requests skipped during that window would have
seen. It needs -q/--rps, because the target rate is the schedule that gets omitted;
without a rate the flag does nothing and says so on stderr. The report records whether it
was on as config.correctOmission.
Dashboard
npm run webnpm run web opens the dashboard in your default browser and prints the URL it bound
to. Pass --no-open, or run where stdout is not a terminal, to skip the browser launch.
The dashboard tries 3000 by default. If that port is unavailable — taken by another
process, or inside a range Windows has reserved — it falls back to an ephemeral port and
prints the URL it actually bound to. This is common: Hyper-V and WinNAT reserve large
blocks of the low port range when Docker or WSL2 is installed, which frequently swallows
3000. To see what is reserved on your machine:
netsh interface ipv4 show excludedportrange protocol=tcpPin a port when something depends on it, or pass --strict-port to fail instead of
falling back:
npm run web -- --port 8080
npm run web -- --port 8080 --strict-portLive tiles — requests, error rate, p95, p99, throughput gauge with peak marker.
Charts — requests/sec and p50/p95 over a 12 s window.
Status + error breakdown — status classes, and transport failures split into timeout / refused / reset / DNS / TLS.
Control — start a run, then pause, resume, ramp or abort it live.
History — the last 50 runs; click one to diff it against the most recent.
Export — JSON, CSV, or a print-optimised page for saving as PDF.
Run form — target, concurrency, duration or request budget, target rate, traffic profile, headers/body, and a collapsible transport/TLS section (timeout, keep-alive, pipelining, expected status, skip-verification). SLO thresholds and coordinated-omission correction live under the same form. HTTP/3 shows up in the version picker only when the optional
httpxQUIC binding is installed.
Development
npm run d:dev # Vite on :5173, proxying the API to a REDSTORM server on :3000Run npm run web in another terminal so the proxy has something to talk to.
Engine architecture
┌──────────────────────────────┐
bin/redstorm.js ───► │ LoadEngine │
Fastify ──────► │ (main thread, no I/O hot) │
WS/SSE clients └──────┬──────────────┬────────┘
│ │
SharedArrayBuffer │ │ control Int32Array
(Float64 counters) │ │ (state, inflight)
┌────────▼───┐ ┌────▼────────────┐
│ worker 0 │ ... │ worker N-1 │
│ undici │ │ undici │
│ Agent / │ │ Agent / │
│ Client │ │ Client │
└─────┬──────┘ └──────┬──────────┘
│ histogram delta │
└────────► engine ◄─┘ (merge → global HDR histogram)- Counters are shared memory. The main thread never polls a worker; it sums a
Float64Arrayin aSharedArrayBuffer. The 10 Hz tick is a straight memory read, so the orchestration loop never competes with request I/O. - Control is Atomics. Pause/resume/ramp land in an
Int32Arrayin the same buffer, so a running request loop notices without messaging. - Counters are plain increments, not Atomics.
Atomics.addrejectsFloat64Array, and each worker's slice is single-threaded anyway. - Histogram deltas are transferred, not re-read. Each worker encodes its 100 ms interval snapshot; the engine merges it into one global HDR histogram. Percentiles are always global and exact, never an average of per-worker averages.
- Latency includes failures. A timeout is the user's latency, so it is recorded.
Responses and transport errors are counted separately, and
attemptsis the denominator for error rate — so a run where every request times out reports 100%, not 0%. - Termination is worker-authoritative. Workers report
donewhen their VUs exit; the engine finalises when the whole fleet is down. A success counter is not a completion signal when nothing succeeded.
HTTP API
| Method | Path | Purpose |
|---|---|---|
| GET | /api/health | status, cores, client count |
| GET | /api/config | engine defaults and limits |
| POST | /api/run | start a run |
| POST | /api/pause · /api/resume · /api/abort | live controls |
| POST | /api/ramp | ramp concurrency without restarting |
| GET | /api/stats | current snapshot + series |
| GET | /api/report | latest full report |
| GET | /api/history | completed runs (summaries) |
| GET | /api/export?format=json\|csv&runId= | download a report |
| GET | /report?runId= | printable report (PDF via the browser) |
| GET | /ws | WebSocket realtime stream |
| GET | /sse | Server-Sent Events fallback |
POST /api/run
Durations are explicitly unit-suffixed because the two forms are easy to confuse:
| Field | Type | Notes |
|---|---|---|
| durationMs | number | milliseconds — preferred for API clients |
| duration | string | human form, e.g. "30s", "2m" |
| requests | number | request budget; alternative to a duration |
| profile | object or string | { type: "ramp", from, to, durationMs } or "ramp" |
| concurrency | number | virtual users |
| connections | number | sockets per worker |
| httpVersion | "1.1" | "2" | "3" | |
| headers | array of [k, v] pairs or object | |
| body, contentType | string | body templates see the same {{placeholders}} as the CLI |
| expect | number | status that counts as success |
| errorBodyBytes | number | bytes of each distinct error body to capture in the report (0 disables, default 512) |
| rps | number | target rate across the fleet; 0 (default) saturates |
| timeout / keepAlive | number | per-request timeout and socket keep-alive, ms |
| pipelining | number | HTTP/1.1 requests in flight per socket (1-128) |
| tls | object | e.g. { "rejectUnauthorized": false } for self-signed targets |
| slo | object | { p95, p99, mean, errors, rps } pass/fail gates (see Thresholds) |
A run must have some bound — a duration or a request budget. Unbounded runs are rejected
with 400, because at full tilt they take the machine down before you notice.
curl -X POST localhost:3000/api/run -H 'content-type: application/json' -d '{
"url": "http://127.0.0.1:8080/api/orders",
"durationMs": 30000,
"concurrency": 200,
"connections": 64,
"profile": { "type": "ramp", "from": 20, "to": 200, "durationMs": 15000 },
"headers": { "Authorization": "Bearer {{token}}" }
}'Stream frames
hello (full state + series) → start → tick* → end (report + history). state and
warning frames interleave. The client never has to poll; hello is sent once on connect
so a late joiner is immediately current.
WebSocket is the primary transport. The dashboard falls back to /sse automatically when an
upgrade is blocked, because the frame shapes are identical.
Reports and export
Every run produces a report with global percentiles, the status distribution, the transport error taxonomy, and the full 100 ms series.
When a run records errors, the report also carries an exact status-code breakdown and a ranked
list of error signatures: HTTP failures show the code with their first sample request
(method, path, and a truncated response body), and transport failures show the err.code and
message. This is what answers "what went wrong" without re-running. The same signatures drive
the What went wrong section of the terminal report, the printable page, and the dashboard
report card; --error-body-bytes 0 turns body capture off.
- JSON — everything, for re-processing.
- CSV — one row per 100 ms interval, for spreadsheets. No comment preamble, so it imports cleanly.
- PDF —
GET /reportrenders a print-optimised page with inline SVG charts. "Save as PDF" from the browser gives better output than embedding a font-heavy PDF library, and keeps the dependency count at zero. - History diff — select any retained run to see per-metric deltas against the newest.
- Thresholds — when
--assert-*is declared, the JSON report and the printable page also carry the gradedsloverdict.
Traffic profiles
# linear ramp
redstorm --profile ramp --ramp-from 5 --ramp-to 500 --ramp-duration 2m -z 3m http://host/
# oscillating load
redstorm --profile wave --wave-base 100 --wave-amplitude 50 --wave-period 15s -z 5m http://host/ramp and wave are also available live via the dashboard's Ramp… button, which
re-profiles a running load without dropping in-flight requests.
Dynamic request templating
Headers and bodies are compiled once per worker and can interpolate per request, which is what stops a load test from being rejected by anything that expects unique nonces.
| Placeholder | Expands to |
|---|---|
| {{uuid}} | random v4 UUID |
| {{id}} | incrementing integer, unique per worker |
| {{ts}} | unix seconds |
| {{now}} | ISO timestamp |
| {{token}} | rotates through --token values |
| {{mock}} / {{mockKey:field}} | lookup in a --mock-file dataset |
redstorm -m POST --json-template '{"id":"{{uuid}}","ts":"{{now}}"}' http://host/apiTesting
npm test # engine + slo + server + CLI
npm run test:e2e # dashboard in a real browsernpm test runs four suites against real HTTP servers — no mocks of the network:
test/engine.test.js— duration and budget termination, 5xx handling, connection refusal, pause/resume/abort, wave profiles, paced (-q) rate accuracy, and worker-failure containment.test/slo.test.js— threshold parsing and grading, including the strict-default override and the aborted-run and no-sample skip rules.test/server.test.js— REST contract and validation, WebSocket frame sequence, SSE fallback, exports, the printable report, SLO passthrough, and run-scoped lookups.test/cli.test.js— the real binary, spawned as a subprocess: help, the full exit-code contract, exact request budgets, 100% error-rate runs, profiles, JSON to a file and to stdout,--label, and the in-place live renderer driven through a fake TTY.
npm run test:e2e drives Chromium through the built dashboard: it starts a run from the
form, asserts the tiles update, checks the canvas is actually painted, exercises
pause/resume, runs twice to populate the compare view, and fails on any console error. It
needs a one-time browser download:
npx playwright install chromiumPublishing
Release steps live in PUBLISHING.md.
Design notes
A few decisions worth stating outright, since they differ from the obvious approach.
Native worker_threads, not Piscina. The original brief named Piscina. It is not used:
Piscina 4.9 takes over the parent's parentPort at startup and exposes no inbound channel,
so the pause/resume/ramp control path that this engine is built around does not exist. Native
worker_threads gives a bidirectional MessagePort plus explicit SharedArrayBuffer
control, which is what the design actually needs.
Canvas charts, not Recharts. SVG charting libraries rebuild their element tree on every
update. Telemetry arrives 10× a second, and a tree that size visibly stutters. The charts
draw straight to a 2D context, on demand: a single coalesced repaint is scheduled when new
data arrives, the pointer moves, or the element resizes, reading the series from a ref. The
props reach that repaint through refs too, so a tick updates the pixels without tearing down
and re-creating the canvas. (Assigning canvas.width wipes the bitmap, so re-creating the
chart every tick makes it flicker to blank; the repaint also resizes only when the size
actually changes.) Because nothing re-schedules itself, an idle dashboard does no canvas work
at all. The bundle stays at ~44 kB of app code.
PDF via the browser, not a PDF library. Embedding a PDF library means shipping font files and eating a large dependency for a document that is fundamentally tabular. A print-optimised HTML page with inline SVG produces better output, is searchable, and costs nothing.
Errors are two numbers, not one. httpErrors counts 4xx/5xx responses; transportErrors
counts timeouts, refused connections, resets, DNS and TLS failures. Averaging them into a
single "errors" figure hides the difference between "the server said no" and "the server was
not there" — which are different incidents with different owners. The engine tracks them
separately all the way from the worker counters to the export, and attempts (not
requests) is the denominator, so a total-outage run reads as 100% errors rather than 0%.
A dead worker must not kill the server. EventEmitter throws on an unhandled 'error'
event, so routing worker failures through one would take down a healthy server because of one
bad worker. Failures surface as warning and drive the run to a clean ABORTED, so you
still get a report for whatever completed.
Author
N4OD — design, implementation, and maintenance of REDSTORM from the first commit.
License
MIT
