@convrs/ai-bot-sdk
v1.0.1
Published
Detect and track AI crawlers, answer agents, and search bots (GPTBot, ClaudeBot, PerplexityBot, and more) hitting your site or API — server-side SDK for Convrs.
Maintainers
Readme
@convrs/ai-bot-sdk
Server-side SDK to detect and track AI crawlers, answer agents, and search bots (GPTBot, ClaudeBot, PerplexityBot, Googlebot, and 20+ more vendors) hitting your site or API, and report those visits to Convrs.
Install
npm install @convrs/ai-bot-sdkQuick start (Next.js middleware)
// middleware.ts
import { createBotTrackingMiddleware } from "@convrs/ai-bot-sdk";
const trackBots = createBotTrackingMiddleware({
siteId: process.env.CONVRS_SITE_ID!,
});
export function middleware(request: Request, context: { waitUntil: (p: Promise<unknown>) => void }) {
trackBots(request, context);
}Express
import express from "express";
import { createExpressBotMiddleware } from "@convrs/ai-bot-sdk";
const app = express();
app.use(createExpressBotMiddleware({ siteId: process.env.CONVRS_SITE_ID! }));Wrapping a single route handler
import { withBotTracking } from "@convrs/ai-bot-sdk";
export const GET = withBotTracking(
async (request: Request) => new Response("ok"),
{ siteId: process.env.CONVRS_SITE_ID! }
);Just classify a User-Agent
import { classifyBotUserAgent } from "@convrs/ai-bot-sdk";
classifyBotUserAgent(req.headers.get("user-agent"));
// => { vendor: "OpenAI", agentName: "gptbot", category: "training_crawler" } | nullConfiguration
| Option | Default | Description |
| --- | --- | --- |
| siteId | — | Required. Your Convrs project id. |
| enabled | true | Master on/off switch. |
| endpoint | Convrs ingest URL | Override for self-hosting/testing. |
| timeoutMs | 1500 | Abort the tracking call after this long. |
| authToken | — | Bearer token for the ingest endpoint. |
| trackedMethods | ["GET","HEAD"] | HTTP methods eligible for tracking. |
| skipIndexCrawlers | false | Skip Googlebot/Bingbot-style crawlers. |
| skipAnswerAgents | false | Skip ChatGPT-User/Perplexity-User style real-time fetchers. |
| skipTrainingCrawlers | false | Skip GPTBot/CCBot-style dataset crawlers. |
| skipOtherBots | false | Skip anything not in the three categories above. |
| ignoredPathPrefixes / extraIgnoredPathPrefixes | see registry.ts | Paths never tracked (e.g. /api, /_next). |
| ignoredExtensions / extraIgnoredExtensions | see registry.ts | File extensions never tracked. |
| maxUrlLength | 8192 | Drop events for implausibly long URLs. |
| domain | request host | Force the reported domain. |
| publicOrigin | — | Rewrite request URL to a public origin (behind tunnels/proxies). |
| fetch | global fetch | Custom fetch implementation. |
| resolveIp | common proxy headers | Custom IP resolution. |
| filter | — | Final (url, bot) => boolean gate. |
| debug | false | Log internal warnings to console. |
Bot categories
answer_agent— real-time, on-demand fetches triggered by a user's chat prompt (e.g.ChatGPT-User,Perplexity-User).index_crawler— crawlers that build a search index (e.g.Googlebot,Bingbot).training_crawler— crawlers that harvest content for model training datasets (e.g.GPTBot,CCBot).other— everything else recognized as an AI-related bot but not cleanly in the above buckets.
License
MIT
