npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

agent-404

v0.1.0

Published

Make 404 pages agent-friendly for AI crawlers & LLMs

Readme

agent-404

Self-healing 404 pages for AI agents and developers.

When you restructure documentation or deprecate an API route, AI coding assistants (Claude Code, Cursor, Copilot), RAG pipelines, and search bots continue to follow outdated URLs baked into their pre-training data. Standard 404 error pages return client-rendered HTML that bots never execute — causing models to hallucinate or give up.

agent-404 intercepts requests at the HTTP middleware layer, instantly returning ranked semantic destination routes in RFC 5988 Link alternate headers, schema.org/ItemList JSON-LD, and JSON payloads so agents self-correct in a single hop.


Quick Install (HTTP Layer — Recommended)

AI crawlers (GPTBot, ClaudeBot, PerplexityBot) do not execute client-side JavaScript. Intercepting at the HTTP layer ensures bots receive recovery metadata before the response body finishes.

Next.js (App Router & Pages)

npm install @agent404/next
// middleware.ts
import { agent404 } from "@agent404/next";

export const middleware = agent404({
  apiKey: process.env.AGENT404_PUBLIC_KEY!,
});

export const config = {
  matcher: ["/((?!api|_next/static|_next/image|favicon.ico).*)"],
};

Cloudflare Workers

npm install @agent404/cloudflare
// worker.ts
import { agent404Worker } from "@agent404/cloudflare";

export default agent404Worker({
  apiKey: "pk_your_public_key",
  origin: "https://docs.example.com",
});

Express / Node.js

npm install @agent404/express
// server.js
import { recoverExpress404 } from "@agent404/express";

app.use(async (req, res) => {
  const recovered = await recoverExpress404(req, "<h1>Not Found</h1>", {
    apiKey: process.env.AGENT404_PUBLIC_KEY,
  });
  res.status(404);
  recovered.headers.forEach((v, k) => res.setHeader(k, v));
  res.send(await recovered.text());
});

Also available: @agent404/netlify for Netlify Edge Functions and nginx (adapters/nginx.md).


Zero-Config Script Tag (Browsers Only)

For human visitors and headless browser agents (e.g. Browser-Use, Playwright, MultiOn), add a single script tag:

<script
  src="https://www.agent404.dev/agent-404.min.js"
  data-site-id="your-site-id"
  data-public-key="your-public-key"
  defer
></script>

Security Note: data-public-key is strictly read-only (/api/suggest). Never expose your secret API key in HTML. Page indexing is handled automatically via verified sitemap crawls.


How It Works

1. ClaudeBot / GPTBot requests moved endpoint:
   GET /docs/v1/authentication

2. agent-404 Edge Middleware intercepts HTTP 404:
   ├── Hybrid Matcher queries indexed sitemap (<25ms)
   └── Evaluates Path Jaccard + pgvector Cosine Similarity

3. Response delivered with structured recovery metadata:
   HTTP/1.1 404 Not Found
   Link: </docs/v2/auth>; rel="alternate"
   Content-Type: text/html

   <script type="application/ld+json">
   {
     "@context": "https://schema.org",
     "@type": "WebPage",
     "mainEntity": {
       "@type": "ItemList",
       "itemListElement": [{ "position": 1, "url": "https://yoursite.com/docs/v2/auth" }]
     }
   }
   </script>

4. AI Agent reads alternate relation and recovers in 1 hop.

4-Signal Hybrid Matching Engine

Every incoming dead URL is scored against your indexed sitemap using four weighted signals:

| Signal | Weight | Purpose & Catch Category | |---|---|---| | Path Segment Jaccard | 0.35 | Tokenized path overlap, version bumps (/v1/auth/v2/auth) | | pgvector Cosine Embeddings | 0.30 | 256d semantic vectors for zero-lexical overlap rewrites (/auth/security/tokens) | | Levenshtein Distance | 0.20 | Character-level typos, singular/plural differences (/payment/payments) | | Keyword & Heading Overlap | 0.15 | Matches tokens against page titles and H1/H2 metadata |

When embeddings are unconfigured or unavailable, the matcher falls back gracefully to a 3-signal heuristic (0.50 / 0.30 / 0.20).


API Reference

1. Register a Domain

curl -X POST https://www.agent404.dev/api/sites \
  -H "Content-Type: application/json" \
  -d '{"domain": "docs.yourcompany.com"}'

Returns id, apiKey (write secret), publicKey (read-only for middleware/HTML), and verificationToken.

2. Prove Domain Ownership & Start Crawl

Prove ownership via DNS TXT or .well-known before suggestions go live:

# DNS TXT: _agent404.docs.yourcompany.com = <verificationToken>
# OR https://docs.yourcompany.com/.well-known/agent-404.txt

curl -X POST https://www.agent404.dev/api/sites/<id>/verify

Once verified, agent-404 automatically fetches and indexes https://docs.yourcompany.com/sitemap.xml.

3. Query Suggestions

curl -X POST https://www.agent404.dev/api/suggest \
  -H "Content-Type: application/json" \
  -H "x-api-key: your-public-key" \
  -d '{"url": "https://docs.yourcompany.com/v1/old-endpoint"}'

Response:

{
  "deadUrl": "https://docs.yourcompany.com/v1/old-endpoint",
  "suggestions": [
    {
      "url": "https://docs.yourcompany.com/v2/new-endpoint",
      "title": "New Endpoint Reference",
      "score": 0.94,
      "matchType": "moved"
    }
  ],
  "jsonLd": {
    "@context": "https://schema.org",
    "@type": "WebPage",
    "name": "Page Not Found",
    "mainEntity": {
      "@type": "ItemList",
      "itemListElement": [
        {
          "@type": "ListItem",
          "position": 1,
          "url": "https://docs.yourcompany.com/v2/new-endpoint",
          "name": "New Endpoint Reference"
        }
      ]
    }
  }
}

Self-Hosting & Deployment

Deploy your own hosted instance with Neon Postgres and Auth0 passwordless authentication in under 2 minutes:

Environment Variables

| Variable | Required | Description | |---|---|---| | DATABASE_URL | Yes | Neon Postgres connection string with pgvector extension | | CRON_SECRET | Yes | Bearer secret for automated sitemap re-crawling (/api/cron) | | AUTH0_DOMAIN | Yes | Auth0 tenant domain (for passwordless owner dashboard) | | AUTH0_CLIENT_ID | Yes | Auth0 Regular Web App Client ID | | AUTH0_CLIENT_SECRET | Yes | Auth0 Regular Web App Client Secret | | AUTH0_SESSION_ENCRYPTION_KEY | Yes | 32+ character cookie encryption secret | | BASE_URL | Yes | Canonical app origin (e.g. https://www.agent404.dev) | | EMBEDDING_API_KEY | Optional | OpenRouter / OpenAI API key for 256d semantic vectors | | EMBEDDING_API_URL | Optional | Custom OpenAI-compatible embeddings endpoint | | EMBEDDING_MODEL | Optional | Custom embedding model (default: text-embedding-3-small) |

Local Development

# 1. Clone repository
git clone https://github.com/bharath31/agent-404.git
cd agent-404

# 2. Install dependencies
npm install

# 3. Configure local environment in .env.local
cp .env.example .env.local

# 4. Run database migrations
npm run db:migrate

# 5. Start dev server
npm run dev

# 6. Run test suite
npm test                 # Unit tests (193 passing)
npm run test:browser     # Playwright browser suite

Technology Stack

  • Framework: Hono with @hono/node-server (Vercel Node.js Serverless) & Cloudflare Workers
  • Database: Neon Postgres with pgvector
  • Embeddings: OpenAI text-embedding-3-small (256 dimensions)
  • Crawler: Streaming SAX sitemap parser with SSRF guard and DNS pinning
  • Client Overlay: Zero-dependency vanilla JS (<3KB gzipped)

Contributing & Workflow Rules

Any new change or feature must be started in an isolated Git worktree:

git fetch origin main
git worktree add ../agent-404-<feature> origin/main -b feat/<feature>

See AGENTS.md for full contributor guidelines.


License

MIT © Bharath Natarajan