@crawlytics/sensor-node
v1.0.2
Published
Fail-open Node sensor for Crawlytics, with Express and Next.js middleware.
Readme
@crawlytics/sensor-node
Fail-open Node sensors for Crawlytics: a core sensor, Express middleware, and Next.js middleware. Each reports requests to your Crawlytics server in the background.
Install
npm install @crawlytics/sensor-nodeIt is an ES module with no dependencies. It needs a global fetch, or one passed
as the fetch option. Tested on Node 22, where CommonJS apps can require() it
as well.
Core
import { createSensor } from "@crawlytics/sensor-node";
const sensor = createSensor({
key: process.env["CRAWLYTICS_KEY"] ?? "",
url: process.env["CRAWLYTICS_URL"] ?? "http://localhost:3000"
});
sensor.record({
bytes: 0,
ip: "203.0.113.10",
method: "GET",
path: "/robots.txt",
referer: "",
status: 200,
ts: new Date().toISOString(),
ua: "GPTBot/1.3"
});Express
import express from "express";
import { expressSensor } from "@crawlytics/sensor-node";
const app = express();
app.use(
expressSensor({
key: process.env["CRAWLYTICS_KEY"] ?? "",
url: process.env["CRAWLYTICS_URL"] ?? "http://localhost:3000"
})
);Next.js Middleware
import { nextSensor, type NextRequestLike } from "@crawlytics/sensor-node";
const recordCrawlyticsRequest = nextSensor({
key: process.env["CRAWLYTICS_KEY"] ?? "",
url: process.env["CRAWLYTICS_URL"] ?? "http://localhost:3000"
});
export function middleware(request: NextRequestLike) {
recordCrawlyticsRequest(request);
// In a real Next.js app, import NextResponse from "next/server" and return:
// return NextResponse.next();
}The Next.js adapter is request-side only because middleware runs before the
response is produced. It records method, path with its query string, user agent,
referer, and client IP, but response status and bytes are not available in
middleware and are sent as status: 0 and bytes: 0. Full response capture
requires route instrumentation, which is outside this middleware adapter.
Behaviour
The sensor buffers raw events and posts them to /api/ingest in the background.
It cannot throw into the host app: record() never throws and flush() never
rejects, whatever goes wrong inside them.
The query string is sent as part of path. The server uses it to recognise AI
assistant referrals whose referer was stripped (utm_source=chatgpt.com) and
drops it before storage.
When the server answers 429 or 5xx, or cannot be reached, the batch is kept and
nothing is sent until the server's Retry-After has passed (capped at one hour),
or, when it names no time, for 5 seconds, doubling to 60 while the refusals go on.
That holds every path: the timer, the flush record() starts at flushSize,
stop(), and an explicit flush(), which sends nothing during the wait. Events
still buffered when the process exits are lost. The buffer is bounded by
maxBuffer (10000 by default), and a batch still being sent counts toward it;
once it is full, new events are refused.
To hear what the server made of each batch, pass onDelivery:
createSensor({
key: process.env["CRAWLYTICS_KEY"] ?? "",
url: process.env["CRAWLYTICS_URL"] ?? "http://localhost:3000",
onDelivery: ({ accepted, rejected }) => {
if (rejected > 0) {
console.warn(`crawlytics: the server refused ${String(rejected)} events`);
}
}
});rejected counts the events the server could not read one at a time, or the
whole batch when the server will never take it (a wrong key answers 401). Those
are not sent again. A batch kept to be sent again is not reported. The same
option works on expressSensor and nextSensor.
License
AGPL-3.0-only. See LICENSE and NOTICE.
