@clicsdev/crawler
v1.0.2
Published
Server-side verified AI crawler analytics for Clics
Maintainers
Readme
@clicsdev/crawler
Verified AI crawler analytics for Clics.
Server-side middleware for detecting answer fetchers, search crawlers, and training bots. The package submits only known candidates; Clics records an event only when the user-agent token and the request IP both match ranges published by the crawler operator.
Install
npm install @clicsdev/crawlerpnpm add @clicsdev/crawler
yarn add @clicsdev/crawler
bun add @clicsdev/crawlerConfigure
Create a project-scoped crawler token in Clics → Configure → Crawler tokens. Keep this token on the server; do not expose it in browser code or use a workspace API key.
CLICS_PROJECT_ID=your_project_id
CLICS_CRAWLER_TOKEN=your_crawler_tokenHow verification works
Your server sees the network request made by the crawler, including the source IP that a browser script cannot reliably access.
- The middleware checks the request user agent against the supported crawler tokens.
- It forwards a candidate event to Clics using your project-scoped ingest token.
- The Clics API checks the original client IP against the operator's published CIDR ranges.
- A crawler event is stored only when both values match.
The full URL, hostname, pathname, HTTP method, response status, provider, crawler, and category are stored. The IP address and full user agent are used for verification but are not written to analytics storage.
Automatic IP detection
No IP option is required. The SDK reads the original crawler IP automatically from the request information supplied by the supported platform:
- Next.js on Vercel uses Vercel's forwarded client IP.
- Cloudflare uses
CF-Connecting-IP. - Fastly, Fly.io, and trusted proxy headers are recognized when present.
- Express uses
request.ip.
If Express runs directly without a reverse proxy, request.ip already contains the remote IP. If Express runs behind Cloudflare, Nginx, or another reverse proxy, configure Express trust proxy for that known proxy so request.ip contains the crawler IP instead of the proxy IP. Do not enable broad proxy trust when clients can reach the Express server directly or can supply their own forwarding headers.
Next.js
Use this from proxy.ts in Next.js 16 or middleware.ts in earlier versions:
import { trackClicsCrawlerRequest } from "@clicsdev/crawler"
import {
NextResponse,
type NextFetchEvent,
type NextRequest,
} from "next/server"
export function proxy(request: NextRequest, event: NextFetchEvent) {
trackClicsCrawlerRequest(request, event, {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
})
return NextResponse.next()
}
export const config = {
matcher: ["/((?!api|_next/static|_next/image|favicon.ico).*)"],
}Passing event lets the tracker use waitUntil() without delaying the page response. Keep crawler-facing files such as robots.txt, llms.txt, and sitemaps inside the matcher.
TanStack Start
Cloudflare Workers
When the TanStack Start application runs on Cloudflare Workers, create a custom src/server.ts entrypoint:
import { createCloudflareHandler } from "@clicsdev/crawler"
import handler from "@tanstack/react-start/server-entry"
import { env } from "cloudflare:workers"
function isCrawlerFacingFile(pathname: string) {
return (
pathname === "/robots.txt" ||
pathname === "/llms.txt" ||
pathname === "/llms-full.txt" ||
(pathname.includes("sitemap") && pathname.endsWith(".xml"))
)
}
const fetch = createCloudflareHandler(
(request) => {
const pathname = new URL(request.url).pathname
if (isCrawlerFacingFile(pathname)) {
return env.ASSETS.fetch(request)
}
return handler.fetch(request)
},
{
projectId: env.CLICS_PROJECT_ID,
token: env.CLICS_CRAWLER_TOKEN,
}
)
export default { fetch }Open the wrangler.jsonc file at the root of the project. Add the ASSETS binding and these selective routes inside the existing assets object. Keep the existing asset directory and other settings unchanged:
{
// Keep your existing Wrangler configuration.
"assets": {
// Keep your existing directory and other asset options.
"binding": "ASSETS",
"run_worker_first": [
"/robots.txt",
"/llms.txt",
"/llms-full.txt",
"/sitemap.xml",
"/*sitemap*.xml",
],
},
}The final pattern also covers other sitemap XML files, including names such as /sitemap-1.xml, /post-sitemap.xml, and nested sitemap paths. Other static assets keep Cloudflare's normal asset-first behavior.
Cloudflare calls the exported fetch(request, env, context) handler with its real execution context. The Clics wrapper runs the selected static asset or TanStack handler, reads the final response status, and passes the tracking request to context.waitUntil() before returning the unchanged response.
You do not create waitUntil(), call the tracker elsewhere, or add await. Do not also register crawler tracking in src/start.ts, because that would submit the same request twice.
If you already have a custom src/server.ts, keep its existing behavior and wrap its current fetch handler with createCloudflareHandler.
Long-running Node.js server
For a persistent Node.js, Docker, or VPS deployment, wrap the TanStack server entry:
import { withCrawlerTracking } from "@clicsdev/crawler"
import handler from "@tanstack/react-start/server-entry"
const fetch = withCrawlerTracking((request) => handler.fetch(request), {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
})
export default { fetch }The Node.js process remains alive after returning the response, so no waitUntil() setup is required. The wrapper starts tracking after TanStack produces the response and returns that response unchanged.
For another serverless TanStack deployment, do not assume that TanStack's middleware context contains the platform execution context. Use the platform's documented background-task API when available. If the platform has no equivalent of waitUntil(), await response tracking when guaranteed delivery matters; this can add latency.
Cloudflare Workers and Pages
import { createCloudflareHandler } from "@clicsdev/crawler"
export default {
fetch: createCloudflareHandler(
(request, env, context) => app.fetch(request, env, context),
(env) => ({
projectId: env.CLICS_PROJECT_ID,
token: env.CLICS_CRAWLER_TOKEN,
})
),
}The adapter uses context.waitUntil() so reporting can finish after the response is returned.
Hono
Track after await next() so Clics receives the final response status:
import { Hono } from "hono"
import { trackClicsCrawlerResponse } from "@clicsdev/crawler"
const app = new Hono()
app.use("*", async (c, next) => {
await next()
trackClicsCrawlerResponse(c.req.raw, c.res, c.executionCtx, {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
})
})c.executionCtx lets the tracker finish in the background with waitUntil() without delaying the Hono response.
Express
import { createExpressMiddleware } from "@clicsdev/crawler"
app.use(
createExpressMiddleware({
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
})
)The tracker calls next() immediately. After Express finishes the response, it sends the event with the final HTTP status. When the application is behind a known reverse proxy, configure trust proxy according to that proxy's topology before registering the middleware.
Generic Fetch handlers
import { withCrawlerTracking } from "@clicsdev/crawler"
const handler = withCrawlerTracking((request) => app.fetch(request), {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
})The wrapper keeps every argument accepted by the original handler. If one of those per-request arguments provides waitUntil(), it is used automatically:
const handler = withCrawlerTracking(
(request, context) => app.fetch(request, context),
{
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
}
)Cloudflare, Fastly, Fly.io, Vercel, and standard forwarded client IP headers are recognized automatically. Express uses request.ip; configure trust proxy only for reverse proxies that you control.
Custom Docker, Cloud Run, or reverse proxies
Most runtimes provide the public request URL automatically. If the request URL contains an internal origin such as http://localhost:3000 or http://0.0.0.0, set the public origin explicitly:
trackClicsCrawlerRequest(request, context, {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
publicOrigin: "https://example.com",
})The SDK replaces only the protocol, hostname, and port. It preserves the complete pathname and query string exactly as received.
If your runtime exposes the client IP through a custom trusted source, provide it explicitly:
trackClicsCrawlerRequest(request, context, {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
getClientIp(request) {
return request.headers.get("your-trusted-client-ip-header")
},
})Supported crawlers
Every entry below requires an operator-published IP range match.
Clics uses a versioned snapshot of these ranges. Updates are reviewed and committed manually; the API does not fetch crawler ranges at request time.
| Category | Provider | User-agent token |
| -------------- | ------------ | ---------------------------- |
| AI answers | OpenAI | ChatGPT-User |
| AI answers | Anthropic | Claude-User |
| AI answers | Perplexity | Perplexity-User |
| AI answers | Google | Google-Agent |
| AI answers | Google | Google-GeminiNotebook |
| AI answers | Google | Google-NotebookLM (legacy) |
| AI answers | Google | Google-Read-Aloud |
| AI answers | Mistral | MistralAI-User |
| AI answers | Amazon | Amzn-User |
| AI answers | DuckDuckGo | DuckAssistBot |
| AI answers | Moonshot AI | Kimi-User |
| Search indexes | OpenAI | OAI-SearchBot |
| Search indexes | Anthropic | Claude-SearchBot |
| Search indexes | Perplexity | PerplexityBot |
| Search indexes | Google | Google-InspectionTool |
| Search indexes | Google | Googlebot |
| Search indexes | Microsoft | bingbot |
| Search indexes | Microsoft | msnbot |
| Search indexes | Mistral | MistralAI-Index |
| Search indexes | Amazon | Amzn-SearchBot |
| Search indexes | Moonshot AI | Kimi-SearchBot |
| Training | OpenAI | GPTBot |
| Training | Anthropic | ClaudeBot |
| Training | Google | GoogleOther |
| Training | Google | Google-CloudVertexBot |
| Training | Apple | Applebot |
| Training | Amazon | Amazonbot |
| Training | Moonshot AI | KimiBot |
| Training | Common Crawl | CCBot |
User-agent-only bots are intentionally excluded. Robots.txt control tokens such as Google-Extended and Applebot-Extended are also excluded because they do not appear as separate HTTP crawler user agents.
API
createCrawlerTracker(config)
Creates a tracker with:
reportRequest(request, responseStatus?)for Fetch API requestsreport(input)for manually constructed server-side observations
withCrawlerTracking(handler, config)
Wraps a Fetch API handler, reports after the handler returns, and automatically uses a per-request waitUntil() context when present.
Tracking primitives
trackClicsCrawlerRequest(request, config)ortrackClicsCrawlerRequest(request, context, config)for middleware that only sees the incoming requesttrackClicsCrawlerResponse(request, response, config)ortrackClicsCrawlerResponse(request, response, context, config)when the final response is available
Framework adapters
createTanStackStartMiddleware(config)createCloudflareHandler(handler, config)createExpressMiddleware(config)
Configuration
| Option | Required | Description |
| -------------- | -------- | ------------------------------------------------------------------- |
| projectId | Yes | Clics project receiving crawler events |
| token | Yes | Project-scoped crawler ingest token |
| publicOrigin | No | Public HTTP(S) origin when the runtime exposes an internal hostname |
| getClientIp | No | Override client IP detection for a trusted runtime or proxy |
| onError | No | Receives reporting errors without failing the page request |
Reporting is best effort and does not delay or change the application response. A transient server error is retried once with the same request ID so Clics can deduplicate the event.
Background delivery: no extra setup required
For the framework integrations above, copy the examples as written. You do not need to create waitUntil(), configure a background job, or add await around the Clics tracker.
The package chooses the delivery behavior from the framework arguments you already pass:
| Integration | Automatic behavior |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| Next.js | Uses waitUntil() from the supplied NextFetchEvent |
| Cloudflare Workers / Pages | Uses waitUntil() from the handler execution context |
| Hono on an edge runtime | Uses waitUntil() from c.executionCtx |
| TanStack Start on Cloudflare | createCloudflareHandler receives the Worker's execution context directly and passes tracking to it |
| TanStack Start on Node.js | The persistent Node.js process can finish the tracking request after returning the application response |
| Generic Fetch handler | withCrawlerTracking detects an available waitUntil() context in the handler arguments |
| Express / long-running Node.js | Tracks after the response finishes and lets the long-running Node.js process complete the tracking request |
Only manually await when you call a tracking primitive without a context, the serverless runtime can stop immediately after returning, it has no waitUntil(), and delivery must finish before the handler returns:
await trackClicsCrawlerResponse(request, response, {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
})This can add latency. Do not use this form with the Next.js, TanStack Start, Hono, Cloudflare, Express, or wrapped Fetch examples above.
Use onError only when reporting failures must be sent to your server logger:
trackClicsCrawlerRequest(request, event, {
projectId: process.env.CLICS_PROJECT_ID!,
token: process.env.CLICS_CRAWLER_TOKEN!,
onError(error) {
logger.warn("Clics crawler tracking failed", { error })
},
})Troubleshooting
- No events — confirm the middleware runs on the server and receives the original crawler IP. Use
getClientIpwhen your runtime exposes the IP through a custom trusted field. - Local tests do not appear — a crawler-looking user agent from your own IP is rejected by design.
- Events work behind one proxy but not another — verify that the proxy preserves the original client IP using its standard trusted header.
- 401 or 403 responses — create or rotate the crawler token for the same project configured in
projectId.
Links
License
MIT
