npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@beeranked/ai-crawler-checker

v0.1.1

Published

See which AI crawlers a site's robots.txt allows or blocks, split into crawlers that affect AI answers and crawlers that affect model training.

Downloads

284

Readme

ai-crawler-checker

See which AI crawlers a website's robots.txt allows or blocks, and understand the difference that actually matters: crawlers that decide whether you can be cited in AI answers (ChatGPT search, Claude, Perplexity, Gemini grounding, Siri) versus crawlers that train models on your content. Blocking the wrong group quietly removes you from AI results while doing nothing you intended.

No browser, no API key, two small HTTP requests. Runs as a library, a CLI, or a self-hosted HTTP endpoint.

Why the two groups are separate

A lot of sites paste a big "block the AI bots" list into robots.txt to stay out of model training, and accidentally block the crawlers that put them in AI answers. Those are different bots from the same companies. This tool reads the robots.txt, checks every documented crawler against it (following RFC 9309), and tells you which group each verdict falls in, with a plain-language summary of what it costs you.

Every crawler in the catalog is documented by the operator that runs it, and each entry links to that operator's own page as its source.

Install

npm install @beeranked/ai-crawler-checker
# or run it once without installing:
npx @beeranked/ai-crawler-checker example.com

Requires Node 18 or newer (uses the built-in fetch).

CLI

ai-crawler-checker example.com
ai-crawler-checker https://example.com --json

Library

import { checkCrawlers } from '@beeranked/ai-crawler-checker';

const report = await checkCrawlers('example.com');
console.log(report.summary);   // { answerVisible, answerTotal, trainingAllowed, trainingTotal }
console.log(report.findings);  // [{ level: 'bad' | 'warn' | 'good' | 'info', text }]
console.log(report.bots);      // per-crawler verdicts with sources

On Node versions without a global fetch, pass one:

import { checkCrawlers } from '@beeranked/ai-crawler-checker';
import { fetch } from 'undici';
await checkCrawlers('example.com', { fetch });

The robots.txt primitives are exported too, if you only want the parser:

import { parseRobots, groupFor, verdict } from '@beeranked/ai-crawler-checker/robots';

Self-host as an HTTP endpoint

worker.js is a ready Cloudflare Worker:

npm install
cp wrangler.toml.example wrangler.toml
npx wrangler deploy

Current wrangler (v4) needs Node 22 or newer; on Node 18 or 20, deploy with npx wrangler@3 deploy instead.

Then POST / with { "url": "https://example.com" }, or GET /?url=..., and you get the full JSON report. It runs comfortably on the Cloudflare Workers free tier.

The crawler catalog

The catalog (src/catalog.js) is meant to stay current as operators publish or rename agents. A pull request that adds or updates a crawler should link to the operator's own documentation as the source; entries that cannot be sourced are not accepted, because the whole point is that the table is true.

Hosted version

Prefer to paste a URL and get a visual report? There is a free hosted version at beeranked.online/ai-crawler-checker, from the team that maintains this project.

License

MIT. See LICENSE.