@coderatechmw/health-kit
v0.2.0
Published
Drop-in /api/health endpoint and security event recording for Next.js apps. Zero runtime dependencies.
Maintainers
Readme
@coderatechmw/health-kit
A protected /api/health endpoint for your Next.js app, in about ten lines.
Zero runtime dependencies. This package installs into every app you own, so it
pulls in nothing — not Zod, not the Mongo driver, not an AWS SDK. It uses only web
standards (Request, Response, fetch, AbortController, Web Crypto), which
means it runs unchanged on the Node and Edge runtimes.
Install
pnpm add @coderatechmw/health-kitUse
// app/api/health/route.ts
import { createHealthHandler, mongoCheck, clerkCheck, r2Check, envCheck } from "@coderatechmw/health-kit";
import clientPromise from "@/lib/mongodb";
export const GET = createHealthHandler({
service: "my-app",
checks: [
mongoCheck(() => clientPromise),
clerkCheck(),
r2Check(),
envCheck(["MONGODB_URI", "CLERK_SECRET_KEY", "R2_ACCESS_KEY_ID"]),
],
});
// Health must never be served from cache.
export const dynamic = "force-dynamic";Then set HEALTH_TOKEN in the app's environment to a long random string, and
register the same token in the monitor.
node -e "console.log(require('node:crypto').randomBytes(32).toString('base64url'))"The two-tier response
The same URL answers differently depending on who is asking.
Anonymous — up or not, and nothing else. External uptime probes use this, which is why it isn't a 401:
{ "schemaVersion": 1, "status": "healthy", "timestamp": "2026-08-09T20:00:00.000Z" }With Authorization: Bearer $HEALTH_TOKEN — the full picture:
{
"schemaVersion": 1,
"status": "degraded",
"service": { "name": "my-app", "commitSha": "a1b2c3d", "environment": "production", "region": "iad1" },
"uptimeSec": 431.2,
"timestamp": "2026-08-09T20:00:00.000Z",
"durationMs": 84,
"checks": [
{ "name": "mongo", "status": "healthy", "durationMs": 12, "meta": { "latencyMs": 11 } },
{ "name": "r2", "status": "unhealthy", "durationMs": 3001, "message": "check timed out after 3000ms" }
]
}Without a token configured, the endpoint fails closed — nobody gets detail. A forgotten env var must never mean "publish internals to the world".
Status codes
| App status | HTTP | Why |
|---|---|---|
| healthy | 200 | — |
| degraded | 200 | Still serving. A 503 would have the platform pull it from rotation over something that doesn't warrant it. |
| unhealthy | 503 | Load balancers and external probes react correctly with no extra configuration. |
Built-in checks
| Check | What it does |
|---|---|
| mongoCheck(getClient) | Pings the DB and reports latency. Pass your existing client promise so it exercises the live pool — a pool that has silently died is invisible to an external probe. |
| envCheck([...names]) | Verifies vars are present and non-empty. Reports names only, never values. |
| httpCheck({ name, url }) | Probes an outbound dependency from inside the app, so it uses the same egress path and region your real requests do. |
| clerkCheck() · r2Check() · cloudinaryCheck() | Verify credentials are configured. Pass probeUrl to also make a real request. |
| customCheck(name, fn) | Your own logic. Return nothing for healthy, throw for unhealthy. |
Critical vs. non-critical
Checks are critical by default. Mark the ones your app survives without:
customCheck("analytics", pingAnalytics, { critical: false })A non-critical failure reports the app as degraded (HTTP 200, no page at 2am)
while its own row still shows unhealthy, so the dashboard tells you the truth
about what broke.
Detecting a leaked health token
The health token is a credential. If it escapes — a committed .env, a build
log, an error report — anyone holding it can read your dependency detail. This
counts who is using it.
The monitor sends an x-shm-probe header on every authenticated request it
makes. A request that carries the token but not that header did not come from
your monitor. That is the whole detection, and unlike a traffic baseline it has
no threshold to tune and no quiet Sunday to get wrong.
export const GET = createHealthHandler({
service: "my-app",
checks: [mongoCheck(() => clientPromise)],
tokenAudit: true,
});true counts in process memory. That is exact on a long-lived server, and a
floor rather than a total on serverless, where each instance only sees its
own requests — the counts are reported with shared: false so the monitor says
"at least N" instead of overstating what it knows.
For a real total on Vercel, Lambda or Workers, back it with anything shared:
import { createHealthHandler, TokenAuditor } from "@coderatechmw/health-kit";
import { Redis } from "@upstash/redis";
const redis = Redis.fromEnv();
export const GET = createHealthHandler({
service: "my-app",
tokenAudit: new TokenAuditor({
store: {
shared: true,
async increment(key, ttlSec) {
const value = await redis.incr(key);
if (value === 1) await redis.expire(key, ttlSec);
return value;
},
get: (key) => redis.get<number>(key),
set: (key, value, ttlSec) => redis.set(key, value, { ex: ttlSec }).then(() => undefined),
},
}),
});Three primitives, so a table, a KV namespace or a Map satisfies it just as well.
Notes worth having before you rely on it:
- Only authenticated requests are counted. The endpoint is public and your platform's own probes hit it constantly; counting those would bury the signal.
- The store never breaks the endpoint. If Redis is unreachable the count is skipped and the health response is served as normal. Failing to count is a far smaller problem than failing to answer.
- It is off by default, because it costs a write on every authenticated request and that is not a cost to impose on someone who did not ask for it.
- What it does not catch: someone replaying a captured request verbatim, header and all. That is a strictly harder position for an attacker to reach than holding a leaked string, which is the case this is for.
Security event recording
@coderatechmw/health-kit/security reports security events to the monitor's
detector, which turns patterns of them into incidents and alerts. Still zero
dependencies, and still opt-in: without configuration, every call is a no-op.
The quickest way in is the brief. On the app's page in the monitor, under Security event detection, copy the setup brief into the AI assistant working in this repository. It does everything below — including finding the places your code refuses access — and tells you which steps are yours. What follows is the same thing, by hand.
Set two variables in the app (the dashboard shows both when you create an ingest key):
SHM_SECURITY_ENDPOINT=https://your-detector.up.railway.app
SHM_SECURITY_KEY=shmk_…At the door: proxy.ts
// proxy.ts (Next.js 16) — or middleware.ts on older versions
import { withSecurity } from "@coderatechmw/health-kit/security";
export default withSecurity();
// Already have a proxy? Wrap it:
// export default withSecurity(clerkMiddleware());The proxy sees every request but not the response, so it records only what is
visible there: requests for scanner paths (/.env, /.git/config,
/wp-login.php…), path traversal, and absurdly long URLs. Ordinary pages
record nothing. Requests carrying the monitor's probe header are labelled as
the monitor's and never counted, so the exposed-paths check does not report
itself.
Where the app says no: recordSecurityEvent
import { recordSecurityEvent } from "@coderatechmw/health-kit/security";
if (order.ownerId !== userId) {
recordSecurityEvent({ type: "authz.fail", route: "/api/orders/[id]", userId, target: order.id, request });
return new Response("Not found", { status: 404 });
}Enumeration (one user walking through other people's objects) is only visible
from here. Other useful events: privilege.grant, privilege.api_key_created,
authn.login_fail (if you have your own sign-in), and custom.<name> for
expensive endpoints — exports, uploads, AI calls — whose volume you want watched.
What it sends, and what it never sends
Only an allowlist: event type, route template, method, status, client IP and
country, an opaque user id, a target id, a truncated user agent, and up to 200
characters of detail. It never reads request bodies, cookies, Authorization
headers or query strings — where passwords and tokens live. Concrete paths are
templated (/api/orders/8812 → /api/orders/[id]), so object ids do not leave
the app.
The client IP is only marked trustworthy when the platform sets it (Vercel's
x-vercel-forwarded-for). Elsewhere, pass trustProxy: { header: "cf-connecting-ip" }
to a SecurityRecorder, or addresses are labelled untrusted and the detector
will not count or block on them.
It never slows the app down
record is synchronous and touches only memory. Events are merged, batched and
sent in the background with a 1.5 s timeout; failures are swallowed and counted,
and the count is reported with the next batch so the monitor can say "partial
coverage". On serverless, the proxy wrapper hands the send to the platform's
waitUntil; in a route handler, pass Next's after:
import { after } from "next/server";
recordSecurityEvent({ type: "authz.fail", route, userId, request }, { waitUntil: after });Blocking
export default withSecurity({ enforceBlocklist: true });With this, the proxy refuses addresses the monitor has blocked (with a 403). The list is cached and refreshed in the background, never awaited, and fails open: an unreachable monitor never blocks a real user. Only platform-verified addresses are ever blocked. Apps on Vercel can instead have blocks applied at Vercel's firewall, configured from the dashboard.
Failure isolation
Every check runs under its own deadline (default 3s, checkTimeoutMs to change).
The check receives an AbortSignal — forward it to any fetch you make so the
request is genuinely cancelled — and the result is additionally raced against a
timer, so a check that ignores the signal, or a driver call that simply never
settles, still cannot hold the response open. A hung dependency becomes a timeout
row, never a hung endpoint.
Errors are reduced to error.message, truncated to 500 chars. Stacks are never
included: they routinely contain absolute filesystem paths, and driver errors can
carry a connection string.
Contributing
src/conformance.test.ts is the important file. This package deliberately defines
its own types rather than importing @shm/contracts at runtime, so that test is
what stops the two from drifting — it runs the real handler and validates the
output against the actual contract schema, boundary values included.
If you change the response shape, that test must change with it, and the contract
needs a new schemaVersion rather than an edit in place.
