@sherryy67/agent-radar
v0.1.0
Published
See which AI agents and crawlers actually read your site. Server-side collector for Express, Fastify, Node, Nginx and Apache.
Maintainers
Readme
agent-radar
See which AI agents and crawlers actually read your site.
npm install @sherryy67/agent-radarconst express = require("express");
const { aiBotMonitor } = require("@sherryy67/agent-radar");
const app = express();
app.use(aiBotMonitor({
siteId: "acme-one.com",
apiKey: process.env.AI_MONITOR_KEY,
endpoint: "https://radar.example.com/api/collect",
}));That is the whole integration. Two lines, no build step, no dependencies.
Why this can't be a <script> tag
GPTBot, ClaudeBot and PerplexityBot fetch your raw HTML and never
execute JavaScript. Every JS-tag analytics tool — GA, Plausible, PostHog — is
structurally blind to them. Collection has to happen server-side or at the edge.
That's the gap this fills.
What it costs you
On a human request: one .toLowerCase() and a substring scan, then next().
Nothing is allocated, no listener is attached, no I/O happens.
On a crawler request: the hit goes into an in-memory queue and the response continues. The queue drains on a timer, batched, with a 3s timeout. Nothing is ever awaited on the request path.
If the endpoint is slow, down, or misconfigured, your site does not notice. The queue is bounded (1000 events by default); past that, events are dropped and counted rather than growing until your server runs out of memory. Telemetry must not become an outage.
Mount it first
app.use(aiBotMonitor({ siteId, apiKey })); // ← before your routers
app.use("/api", apiRouter);Mounted last, it never sees requests that 404. "GPTBot is crawling URLs that don't exist" is one of the more useful things this can tell you.
It hooks res.on("finish"), so the status it reports is the one the crawler
actually got — a 403 to ClaudeBot is a finding, not a dropped row.
Other runtimes
Fastify:
await fastify.register(require("@sherryy67/agent-radar/fastify"), {
siteId: "acme-one.com",
apiKey: process.env.AI_MONITOR_KEY,
endpoint: "https://radar.example.com/api/collect",
});Bare http.createServer — and the escape hatch for Koa, Hapi, Hono on Node,
or anything else that ends up with a Node req/res pair:
const { withAgentRadar } = require("@sherryy67/agent-radar");
http.createServer(withAgentRadar(handler, { siteId, apiKey })).listen(3000);Nginx and Apache cannot make an outbound HTTP request from their config, so
they get a different shape: the server writes one extra, pre-filtered log file
and agent-radar-tail ships it. See collectors/nginx
and collectors/apache.
npm install -g @sherryy67/agent-radar
agent-radar-tail --site SITE_123 --key $AI_MONITOR_KEY \
--log /var/log/nginx/agent-radar.logIt tracks inode and byte offset, so it survives rotation and a restart costs you nothing. It's also entirely off the request path — if it dies, your web server doesn't know.
Cloudflare Workers, Pages, and Next.js edge middleware each need their own
copy-paste collector, because the edge runtime has no process.on, no timers
that outlive a request, and no memory shared between invocations. See
collectors/.
Options
| Option | Default | |
|---|---|---|
| siteId | required | From your dashboard. |
| apiKey | required | Ingest key. Keep it in an env var. |
| endpoint | required | Your deployment, e.g. https://radar.example.com/api/collect. |
| trustProxy | false | See below. Get this right. |
| flushIntervalMs | 5000 | How often the queue drains. |
| maxBatch | 100 | Events per POST; also triggers an early flush. |
| maxQueue | 1000 | Ceiling before events are dropped. |
| timeoutMs | 3000 | Per-POST timeout. |
| ignore | static assets | (path) => boolean. Return true to skip. |
| redactPath | — | (path) => string, applied before anything leaves. |
| onError | — | Without it, internal failures are silent. |
| sampleRate | 1 | Only useful at genuinely large crawler volume. |
| debug | false | Log activity to the console. |
trustProxy matters more here than in normal analytics
The client IP is CIDR-matched against OpenAI's and Anthropic's published egress
IP lists to decide whether a hit was really who it claimed to be. Getting the
IP wrong doesn't give you a missing row — it gives you a real ClaudeBot
reported as spoofed. A false accusation is worse than no data.
So the default is deliberately conservative:
false(default) — use the socket address. Correct when your app faces the internet directly.true— use the left-mostX-Forwarded-Forentry. A client can forge this by sending its ownX-Forwarded-For.3(a number) — take the 3rd entry from the right, i.e. skip 2 trusted hops. The only setting a client cannot forge. Use it when you know your topology.
Two things are handled for you: CF-Connecting-IP is trusted without
configuration (Cloudflare sets it and strips any client copy), and if your
Express app already calls app.set("trust proxy", …), that answer is used as-is.
Shutdown
The flush timer is unref'd, so this never keeps a process alive. Queued events
are flushed on SIGTERM, SIGINT and beforeExit. To be explicit:
const monitor = aiBotMonitor({ siteId, apiKey });
app.use(monitor);
process.on("SIGTERM", async () => {
await monitor.close();
server.close();
});monitor.stats() returns { queued, sent, dropped }.
Privacy
Full IPs are sent (verification needs them), then truncated to a /24 or /48 at
ingest and discarded. Query strings are stripped client-side and never leave
your server. Use redactPath if your path segments are themselves sensitive.
Requirements
Node 18+, for fetch and AbortController.
Tests
npm test