@maango/vercel
v1.2.0
Published
AI agent permissions for Vercel — block training crawlers, allow search and AI assistants, enforced at the edge.
Maintainers
Readme
@maango/vercel
GPTBot takes your content. Googlebot sends you traffic. Your robots.txt cannot tell them apart.
One line of middleware that can. It classifies every request against 594 known agent signatures, then blocks the crawlers that train on your work while letting the ones that send you readers through. Enforced at the edge, in real time, whether or not the crawler reads your robots.txt.
npm install @maango/vercelRuns on Vercel, Cloudflare Workers, Deno, Bun, Netlify, Node, and Express. Zero dependencies.
Just want the data? The bot registry ships with the package and is free to use: 579 agents classified by purpose, plus ready-made robots.txt files for 11 site types. See Free data below.
Quick start
Create middleware.ts in your project root:
// middleware.ts
export * from '@maango/vercel';Set one environment variable on your Vercel project:
MAANGO_VERTICAL=ecommerceThat's it. Redeploy. AI training crawlers (GPTBot, ClaudeBot, Bytespider, CCBot, Google-Extended, etc.) now get a structured 403; Googlebot and human visitors pass through unchanged.
What gets blocked, what doesn't
| Category | Examples | Default behavior | |---|---|---| | Search engines | Googlebot, Bingbot, DuckDuckBot, Yandex | ✅ Allowed. Your SEO is untouched. | | AI search bots | Claude-SearchBot, OAI-SearchBot, PerplexityBot, Applebot | ✅ Allowed. Referral traffic preserved. | | AI assistants | ChatGPT-User, Claude-User, Perplexity-User, Meta-ExternalFetcher | ✅ Browse-only. Can read, cannot submit forms. | | AI training crawlers | GPTBot, ClaudeBot, CCBot, Google-Extended, Bytespider | ❌ Blocked. Your content stays out of training corpora. | | Automation tools | curl, scrapers, headless browsers | 👀 Monitored. Logged, not blocked. | | Real humans | Chrome, Safari, Firefox, mobile browsers | ✅ Fast-pathed. Never touch the policy engine. |
Free data
The registry behind all of this is yours to use. It ships inside the package, under node_modules/@maango/vercel/assets/:
| File | What it is |
|---|---|
| ai-bots.json | 579 agents, 594 user-agent patterns, each tagged with its purpose, operator, and whether it honors robots.txt |
| ai-bots.csv | The same, for spreadsheets |
| robots/<vertical>.txt | Ready-to-paste robots.txt for 11 site types, denying training crawlers and welcoming the AI search bots that cite you |
| llms.txt | A starting point for the emerging llms.txt convention |
Free to use, including commercially, with attribution to maango.io. Generated from the same registry the middleware enforces with, so the files cannot drift from real behavior.
# Drop a ready-made robots.txt straight into your site
cp node_modules/@maango/vercel/assets/robots/publisher.txt public/robots.txtWhy a robots.txt is not enough
robots.txt is advisory. It is a sign on the door, not a lock. The registry tracks what each operator says about honoring it, and the numbers are not reassuring:
Of the 42 AI training crawlers in the registry, 6 claim to honor
robots.txt. 23 openly state that they do not. The rest do not say.
Every one of those 23 will read your carefully written robots.txt and crawl you anyway. That is not a bug in the protocol, it is the protocol. The only thing that stops them is a server that checks who is asking and answers with a 403, which is what the middleware above does.
Vertical presets
Pick the one closest to your site. Each preset comes with curated path-level rules.
| Vertical | What it's tuned for |
|---|---|
| ecommerce (default) | DTC stores, marketplaces, Shopify-style sites |
| saas | B2B web apps with logins, dashboards, billing |
| publisher | Blogs, news sites, paywalled content |
| marketing | Landing pages, lead-gen, agencies |
| regulated | Healthcare, legal, government (strictest) |
| personal | Portfolios, resumes, personal blogs |
| community | Forums, review sites, UGC platforms |
| healthcare | Hospitals, clinics, telehealth |
| legal | Law firms, legal directories |
| government | Federal/state/local government sites |
| education | Schools, universities, MOOCs |
Set it via env var:
MAANGO_VERTICAL=saasEnvironment variables
| Var | Purpose | Default |
|---|---|---|
| MAANGO_VERTICAL | Which preset to use | ecommerce |
| MAANGO_MODE | enforce (default) or monitor_only | enforce |
| MAANGO_FAIL_OPEN | If something goes wrong: allow (true) or block (false) | true |
| MAANGO_LOG_LEVEL | silent | errors | all | all |
| MAANGO_DISABLED | Kill switch: true bypasses middleware | unset |
| MAANGO_WEBHOOK_URL | If set, POST every event to your own analytics endpoint | unset |
| MAANGO_TELEMETRY | First-party telemetry to Maango: off to disable | on |
| MAANGO_INGEST_URL | Override the telemetry ingest endpoint | https://api.maango.io/v1/observations |
| MAANGO_SITE_KEY | Optional account key linking observations to your dashboard | unset |
| MAANGO_TELEMETRY_SAMPLE | Sample rate 0 to 1 for very high-volume sites | 1 |
| MAANGO_TELEMETRY_IP | Include a hashed + truncated IP: on to enable | off (geo only) |
What blocked agents see
HTTP/1.1 403 Forbidden
Content-Type: application/json
X-Maango-Decision: deny
X-Maango-Agent-Category: ai_training
X-Maango-Agent-Name: GPTBot
{
"error": "agent_action_denied",
"agent_category": "ai_training",
"agent_name": "GPTBot",
"documentation": "https://maango.io/agents/getting-allowed"
}Existing middleware? Compose it
// middleware.ts
import { NextResponse, type NextRequest, type NextFetchEvent } from 'next/server';
import { maango } from '@maango/vercel';
export const config = {
matcher: ['/((?!_next/static|_next/image|favicon.ico).*)'],
};
export async function middleware(req: NextRequest, event: NextFetchEvent) {
const decision = await maango(req, event);
if (decision.shouldEnforce) return decision.response;
// your own auth / redirects / a/b testing here
return NextResponse.next();
}Passing event lets telemetry flush after the response returns. Omit it and
telemetry still works, just best-effort.
Not on Vercel? Use the core
@maango/vercel/core is the same engine with no framework imports. It runs anywhere Web Standards do: Cloudflare Workers, Deno, Bun, Netlify Edge, Node 18+, Hono, Express. No Next.js required.
Cloudflare Worker. Pass the env your handler receives, since Workers have no process.env:
import { createMaango } from '@maango/vercel/core';
export default {
async fetch(request, env, ctx) {
const maango = createMaango({ env });
const decision = await maango(request, ctx);
if (decision.shouldEnforce) return decision.response;
return fetch(request);
},
};Express or any Node server. fromNodeRequest adapts Node's request shape:
import { createMaango, fromNodeRequest } from '@maango/vercel/core';
const maango = createMaango();
app.use(async (req, res, next) => {
const decision = await maango(fromNodeRequest(req));
if (!decision.shouldEnforce) return next();
res.status(decision.response.status).send(await decision.response.text());
});Deno, Bun, Hono, or anything with a Web-standard Request:
import { createMaango } from '@maango/vercel/core';
const maango = createMaango();
const decision = await maango(request);Passing the runtime's context (ctx on Workers, event on Next.js) lets telemetry flush after the response returns. Omit it and telemetry still works, just best-effort.
Onboard risk-free
Want to watch traffic for a week before enforcing?
MAANGO_MODE=monitor_onlyEvery decision is logged in your Vercel function logs, but no requests are blocked. Flip to enforce once you're comfortable.
Telemetry and privacy
By default the middleware sends Maango one small observation per non-human request, dispatched after the response via event.waitUntil, so it adds no latency. This powers your dashboard and the cross-web access graph.
What is reported (non-human requests only):
- The agent name and category we classified (for example
GPTBot,ai_training). - Our decision (for example
allow,deny,monitor) and the mode (enforceormonitor_only). - The host, the request method, and the path (normalized, never with a query string).
- A coarse country code and the package version.
What is never reported:
- Humans. Real browser visitors are fast-pathed and never reported.
- Query strings, request bodies, cookies, or any header beyond what is listed above.
- Raw IP addresses. IP is geo country only by default. With
MAANGO_TELEMETRY_IP=on, the IP is truncated to its network and salted-hashed at the edge before it is ever sent.
Turn it off entirely:
MAANGO_TELEMETRY=offMAANGO_DISABLED=true disables telemetry too. MAANGO_SITE_KEY is optional: set it to attribute observations to your Maango account and dashboard, or leave it unset to contribute anonymously. Telemetry is independent of MAANGO_MODE, so monitor-only and enforcing sites both contribute.
Compatibility
| Entry point | Runtimes |
|---|---|
| @maango/vercel | Next.js 13.4+ on the Vercel Edge Runtime, or self-hosted |
| @maango/vercel/core | Cloudflare Workers, Deno, Bun, Netlify Edge, Node 18+, Hono, Express |
next is an optional peer dependency, so nothing extra is installed when you only use the core.
License
This package is licensed under the Maango Plugin License (see LICENSE). For commercial inquiries: [email protected]
Built by Maango.
