shutdown-check
v1.0.1
Published
Verify Node.js graceful shutdown under SIGTERM: drain in-flight HTTP requests, withdraw readiness, reject new work, and report CI failures.
Maintainers
Readme
shutdown-check — Node.js graceful shutdown tester
Test whether your Node.js HTTP service handles SIGTERM without dropping requests already in progress. shutdown-check launches your real command, waits for readiness, starts real HTTP work, sends the signal, and checks the responses and process exit. It reports a timestamped trace locally and JSON or JUnit results in CI.
This solves the question developers face during deploys: “Will my server actually finish current requests when it receives SIGTERM, or will users see dropped requests and 502s?” It is a black-box shutdown tester, not another shutdown-handler library. The standalone CLI and typed API have zero runtime dependencies.
Install · What it checks · Configuration · CI · FAQ · Report an issue
Requirements
Node.js 22 or later on macOS or Linux. Use a dedicated local port and isolated test data: the tool starts and terminates the configured process. The CLI targets local HTTP/1 services; it does not require Express, Fastify, NestJS, or any other framework.
Quick start
npm install --save-dev shutdown-check
npx shutdown-check initEdit shutdown-check.json to match your app, then run:
npx shutdown-check testExample result (timings and process ID vary):
PASS SC000: Graceful shutdown verified
Timeline:
+ 152 ms service ready — HTTP 200
+ 160 ms work confirmed active — 1 response(s) sent headers; bodies still in progress
+ 161 ms signal sent — SIGTERM
+ 498 ms work request finished — #1 HTTP 200
+ 501 ms process exited — code=0, signal=none
+ 503 ms shutdown verified — work completed and service exited before deadlineThe generated configuration looks like this:
{
"command": ["node", "server.js"],
"cwd": ".",
"baseUrl": "http://127.0.0.1:3000",
"readiness": { "path": "/health", "status": 200, "timeoutMs": 10000 },
"workload": {
"path": "/slow",
"method": "GET",
"status": 200,
"started": { "type": "response-headers", "timeoutMs": 5000 }
},
"shutdown": { "deadlineMs": 10000, "exitCode": 0 }
}The example assumes /slow sends response headers promptly, keeps the response open while doing work, and eventually returns HTTP 200. Replace it with an endpoint that exercises real in-flight work. A fast endpoint cannot prove graceful shutdown and will fail the start-barrier check.
If your app is built before launch, point command at its production start command, for example ["node", "dist/server.js"] or ["npm", "run", "start"]. Prefer a direct Node command when possible: wrapping a server in a shell or package manager may change how SIGTERM reaches it, which this tool can reveal.
For handlers that do not send headers until work finishes
Expose a test-only probe endpoint that reports whether the operation has started. The probe must initially return one status and switch to another while the slow request is still running:
"started": {
"type": "probe",
"path": "/test/work-active",
"inactiveStatus": 204,
"activeStatus": 200,
"timeoutMs": 5000,
"intervalMs": 50
}This is often the better choice for ordinary request handlers. The probe is checked before work starts and again while the request remains open, so the tool does not mistake a completed request for an in-flight one.
What it checks
- The configured local HTTP port is free before launch.
- Your command starts and the readiness route returns its expected status.
- Each workload request reaches an observable start barrier and remains in flight.
- The process receives
SIGTERM. - Optionally, readiness withdraws and a new request is rejected while old work is still draining.
- Optionally, a second
SIGTERMis delivered while work remains active. - Every complete workload response arrives with the expected status and optional
bodyIncludestext. - The process exits with the expected code before the deadline and the HTTP port closes. This catches launchers that exit but leave a child server running.
On failure, the CLI reports a diagnostic code, timeline, and the last 8 KiB of process output. Exit code 0 means pass, 1 means the behavioral check failed, and 2 means configuration or setup failed. Use --json for a machine-readable result, --junit result.xml for a CI report, --config FILE for a custom configuration path, or --version to check the installed release.
Diagnosing a failure
| Code | What happened | First thing to check |
| --- | --- | --- |
| SC001 | The local port was already occupied | Use a dedicated test port |
| SC100 / SC101 | Readiness timed out or the process exited early | Check the launch command, port, and service stderr |
| SC111 | Work was never observed in flight | Use a slow endpoint or a probe start barrier |
| SC200 / SC201 | In-flight work hung or was interrupted | Check signal handling and when you close HTTP connections |
| SC202 / SC203 | The final status or body differed | Check the configured expectation and response path |
| SC300 / SC301 | Exit timed out or returned the wrong code | Check cleanup promises, timers, and the exit code |
| SC302 | The launcher exited but an HTTP listener remained | Check whether a shell, npm script, or worker forwarded SIGTERM |
| SC310 / SC311 | Readiness stayed ready or new work was accepted | Withdraw readiness and reject fresh work during drain |
| SC312 | Work finished before new-request rejection could be checked | Use a longer-running test workload |
The timeline and captured stderr usually identify the failing stage. Do not share logs containing secrets in a public issue.
More demanding shutdown contract
Add these fields when you want to verify multiple requests, withdrawal from readiness, rejection of new traffic, and repeated signals:
{
"workload": { "concurrent": 3 },
"shutdown": {
"readinessWithdrawal": true,
"newRequests": {
"path": "/test/new-work",
"rejectStatuses": [503],
"allowConnectionRefused": true
},
"repeatSignalAfterMs": 100
}
}These are additional fields, not a complete configuration. concurrent may be 1–20 and needs the response-headers barrier above 1, so the tool can confirm every request started. newRequests requires readinessWithdrawal: true; it sends one GET after the readiness route stops returning its ready status and before the old requests finish. Use a safe, test-only GET route. Connection refusal is accepted by default; list application-level rejection statuses such as 503 when the server stays reachable during drain. A timeout is not counted as rejection.
Configuration notes
commandis an argument array, not a shell command. For example, use["node", "dist/server.js"], not"node dist/server.js".cwdis resolved relative to the configuration file. If omitted, it is the config file's directory.envcan add or override environment variables for the launched process.baseUrlmust be local HTTP onlocalhost,127.0.0.1, or[::1]. Its port should be dedicated to the test.readinesssupportspath,status(default 200),timeoutMs(default 10000), andintervalMs(default 100).workloadsupportspath,method(default GET), stringheaders, stringbody,status(default 200), optionalbodyIncludes,concurrent(default 1), and a requiredstartedbarrier.shutdownsupportsdeadlineMs(default 10000),exitCode(default 0),readinessWithdrawal(default false), optionalnewRequests, and optionalrepeatSignalAfterMs. Version 1.0 testsSIGTERMonly.
The package does not claim to verify database or queue draining, WebSocket/HTTP2 sessions, container orchestration, Windows signals, or downstream work after an HTTP response. Add application-specific probes before treating those as covered. A response-headers barrier proves a response is still open; a probe barrier can prove your actual operation has started.
Use from node:test or Vitest
import assert from "node:assert/strict";
import { checkShutdown, defineConfig } from "shutdown-check";
const config = defineConfig({
command: ["node", "dist/server.js"],
baseUrl: "http://127.0.0.1:3000",
readiness: { path: "/health" },
workload: {
path: "/slow",
started: { type: "response-headers" },
},
shutdown: { readinessWithdrawal: true },
});
const result = await checkShutdown(config);
assert.equal(result.pass, true, `${result.code}: ${result.message}`);checkShutdown starts and cleans up its own process. parseConfig, loadConfig, runCheck, and junitXml are also exported for custom runners. The public TypeScript declarations ship in the package.
CommonJS consumers can use const { checkShutdown } = require("shutdown-check"); ESM consumers use import { checkShutdown } from "shutdown-check".
CI
Run npx shutdown-check test --json --junit shutdown-result.xml after building your app, against a dedicated local port and isolated test data. Treat a nonzero exit as a failed check. The package's own integration fixtures are kept outside the publishable package project.
Version scope
The planned milestones are implemented together in this 1.0 codebase; intermediate versions were not published:
- 0.1 foundation: real process launch, readiness, in-flight barrier,
SIGTERM, response drain, exit deadline, timeline. - 0.2 traffic contract: concurrent requests, readiness withdrawal, and rejection of new traffic during drain.
- 0.3 resilience: repeated signal, orphaned child-server detection, and focused failure diagnostics.
- 1.0 integration: typed test-runner API, JSON/JUnit CI output, strict configuration, and separate black-box integration fixtures.
The 1.0.1 maintenance release adds CommonJS API support, --version, and clearer npm/GitHub documentation. See the changelog.
FAQ
Does this make my application shut down gracefully? No. It verifies your existing signal handling. You still need to stop accepting new work and finish work already in progress in your application.
Why does SC111 say the work was not active? Your request finished before SIGTERM. Use a test endpoint that keeps a response open, or use the probe start barrier to confirm a real operation began.
Does it test databases, queues, WebSockets, or Kubernetes? Not automatically. This release verifies local HTTP/1 request handling and process behavior. Add application-specific probes for downstream work; container behavior needs a separate test environment.
How do I report a problem? Open a GitHub issue with the diagnostic code, relevant timeline, Node version, OS, and a small reproducible service. Remove secrets from process output before sharing.
Package size and compatibility
The npm tarball includes the CLI, ESM and CommonJS APIs, TypeScript declarations, this README, and the MIT license—no test fixtures or runtime dependencies. Node.js 22+ on macOS and Linux is supported.
License
MIT © Sohail Khan
