npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@cyanheads/internet-archive-mcp-server

v0.3.0

Published

Search the Wayback Machine and IA library (40M+ items), fetch archived snapshots, retrieve item metadata and full text via MCP. STDIO or Streamable HTTP.

Readme

Version License Docker MCP SDK npm TypeScript Bun

Install in Claude Desktop Install in Cursor Install in VS Code

Framework


Overview

The Wayback Machine and Internet Archive library (40M+ items). Find and fetch archived snapshots of any URL, search the library by keyword and metadata, and retrieve item metadata, file manifests, and OCR text from any MCP client. Runs as a stdio process or a local Streamable HTTP server.

Tools

| Tool | Description | |:---|:---| | ia_find_snapshots | Find Wayback Machine snapshots of a URL, by closest timestamp or full capture history | | ia_get_snapshot | Fetch archived page content at a specific Wayback timestamp | | ia_search_items | Search the IA library (40M+ items) by keyword and metadata filters | | ia_get_item | Retrieve full metadata and file manifest for an Archive item | | ia_get_text | Retrieve readable OCR text from a text item, with paging |

Resources

| Resource | Description | |:---|:---| | ia://item/{identifier} | Metadata snapshot for an Archive item — title, creator, mediatype, description, subjects, collections, date, license, and file count |

All resource data is also reachable via ia_get_item.

Capability reference

ia_find_snapshots tool

  • closest mode: returns the nearest capture to a given timestamp via the Availability API, falling back to one CDX closest-capture query (preferring a 200 capture, 25 s deadline) when the Availability API has no answer
  • history mode: full capture list via the CDX API; filter by date range (from/to), HTTP status (status_filter), and MIME type
  • Default collapse of timestamp:8 (one capture per day); adjustable to timestamp:N, N=1–14
  • Up to 10,000 records per call (limit, default 100); resume_key pagination for large histories
  • Replay URLs are always https://web.archive.org/web/…
  • Typed errors: missing_timestamp (closest mode without a timestamp, answered before any lookup), no_snapshots (no matches), no_snapshot_available (closest mode, no capture near timestamp from either source), cdx_unavailable (CDX 5xx, 429, or unreadable response, or the closest-mode CDX check did not complete), availability_unavailable (closest mode, Availability API 5xx, 429, or unreadable response); a 429 is answered after one request with a hint to wait

ia_get_snapshot tool

  • Resolves to the nearest available capture when the exact timestamp has no snapshot; exact 14-digit timestamps skip resolution
  • Reports the capture Wayback served — replay_url, resolved_timestamp, and resolved_status follow Wayback's redirect when the requested timestamp is not itself a capture
  • Decodes the page with its declared charset (Content-Type, then <meta>, else UTF-8 when the bytes are valid UTF-8, else Wayback's guessed charset or windows-1252), then removes scripts, styles, comments, and tags and decodes character references, returning readable plain text alongside the replay URL
  • Reads at most the first 4 MiB of a page (a notice says when a page is longer); output capped at IA_MAX_SNAPSHOT_CHARS (default 50,000 characters)
  • Typed errors: no_snapshot_available, content_fetch_failed (Wayback unreachable during lookup or fetch)

ia_search_items tool

  • Solr query syntax plus structured filters: mediatype, collection, creator, language, and date range (date_from/date_to)
  • mediatype takes the ten Internet Archive media types — texts, movies, audio, software, image, data, web, collection, etree, account — case-insensitively, and resolves common near-misses (text, book, books → texts; movie, video, videos → movies; images → image; collections → collection)
  • Sort by relevance, date, or downloads (sort, Solr syntax; default downloads desc)
  • Up to 200 results per page (rows, default 50), 1-indexed page
  • Output carries total_found, page, rows for pagination; an empty page returns a notice rather than an error — naming the mediatype applied when nothing matched, or the last page when page is past the end
  • Typed error invalid_mediatype for any other mediatype, listing the accepted values, answered before any search request

ia_get_item tool

  • Returns title, creator, description, subject, collection, licenseurl, rights, and language when present in upstream metadata; creator, description, subject, collection, and language may be a string or a list
  • files[] is one page of the manifest in upstream order — format, size, md5, and a direct download_url per file; file_count is always the full manifest size
  • max_files (1–500, default 50) and file_offset (default 0) page through large items; when files remain, the response sets truncated and names the next file_offset
  • format keeps one file type (exact match, case-insensitive — e.g. DjVuTXT, Text PDF, VBR MP3) before paging and reports the match count as totalCount; a format with no matches returns an empty page and a notice listing the formats the item has
  • Typed error item_not_found for unknown identifiers

ia_get_text tool

  • max_chars (defaults to IA_MAX_SNAPSHOT_CHARS) and char_offset page through long documents; has_more signals additional text remains
  • Locates the best available text file — DjVuTXT preferred, falls back to plain text; source_file names the file fetched
  • Typed errors: item_not_found, no_text_file, download_forbidden (restricted collections)

ia://item/{identifier} resource

  • Returns application/json — title, creator, mediatype, description, subject, collection, date, licenseurl, rights, language, and file_count
  • identifier comes from ia_search_items results
  • Typed error item_not_found for unknown identifiers

Features

Built on @cyanheads/mcp-ts-core: stdio and Streamable HTTP transports, pluggable auth (none / jwt / oauth), swappable storage (in-memory, filesystem, Supabase, Cloudflare KV/R2/D1), structured logging with optional OpenTelemetry tracing.

Internet Archive-specific:

  • No credentials required — all four APIs are public
  • Three service layers: WaybackService (Availability + CDX), ArchiveSearchService (Solr), ArchiveMetadataService (Metadata + downloads)
  • CDX collapse-by-day default and configurable limit keep responses tractable for high-capture URLs
  • Identifies via a custom User-Agent on every request as required by IA's terms of use; configurable via IA_USER_AGENT

Agent-friendly output:

  • Pagination context on every list response — total_found, page, rows (search), resume_key (CDX history), and file_count plus the next file_offset (item files) so agents never have to guess whether results are complete
  • Typed error reasons (missing_timestamp, no_snapshots, no_snapshot_available, cdx_unavailable, availability_unavailable, content_fetch_failed, invalid_mediatype, item_not_found, no_text_file, download_forbidden) with recovery hints so callers can retry or explain to users without parsing text
  • Structured file manifests — ia_get_item returns file-level metadata (format, size, URL) and a format filter, so agents can pick the right file without paging through thumbnails

Getting started

No API key required — the Internet Archive's APIs are fully public.

Add the following to your MCP client configuration file:

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "bunx",
      "args": ["@cyanheads/internet-archive-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with npx (no Bun required):

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "npx",
      "args": ["-y", "@cyanheads/internet-archive-mcp-server@latest"],
      "env": {
        "MCP_TRANSPORT_TYPE": "stdio",
        "MCP_LOG_LEVEL": "info"
      }
    }
  }
}

Or with Docker:

{
  "mcpServers": {
    "internet-archive-mcp-server": {
      "type": "stdio",
      "command": "docker",
      "args": [
        "run", "-i", "--rm",
        "-e", "MCP_TRANSPORT_TYPE=stdio",
        "ghcr.io/cyanheads/internet-archive-mcp-server:latest"
      ]
    }
  }
}

For Streamable HTTP, set the transport and start the server:

MCP_TRANSPORT_TYPE=http MCP_HTTP_PORT=3010 bun run start:http
# Server listens at http://localhost:3010/mcp

Prerequisites

  • Bun v1.4.0 or higher (or Node.js v24+).
  • No external accounts or API keys required.

Installation

  1. Clone the repository:
git clone https://github.com/cyanheads/internet-archive-mcp-server.git
  1. Navigate into the directory:
cd internet-archive-mcp-server
  1. Install dependencies:
bun install
  1. Configure environment:
cp .env.example .env
# Optional: edit .env for custom User-Agent, timeouts, etc.

Configuration

All configuration is validated at startup via Zod schemas in src/config/server-config.ts.

| Variable | Description | Default | |:---------|:------------|:--------| | MCP_TRANSPORT_TYPE | Transport: stdio or http | stdio | | MCP_HTTP_PORT | HTTP server port | 3010 | | MCP_AUTH_MODE | Auth mode: none, jwt, or oauth | none | | MCP_LOG_LEVEL | Log level (debug, info, notice, warning, error) | info | | LOGS_DIR | Directory for log files (Node.js only) | <project-root>/logs | | STORAGE_PROVIDER_TYPE | Storage backend | in-memory | | OTEL_ENABLED | Enable OpenTelemetry instrumentation | false | | IA_USER_AGENT | Custom User-Agent for IA API requests | internet-archive-mcp-server/{version} (github.com/cyanheads/internet-archive-mcp-server) | | IA_REQUEST_TIMEOUT_MS | HTTP request timeout in milliseconds | 30000 | | IA_MAX_SNAPSHOT_CHARS | Default character cap for ia_get_text responses | 50000 |

See .env.example for the full list of optional overrides.

Running the server

Local development

  • Build and run:

    # One-time build
    bun run rebuild
    
    # Run the built server
    bun run start:stdio
    # or
    bun run start:http
  • Run checks and tests:

    bun run devcheck   # Lint, format, typecheck, security
    bun run test       # Vitest test suite
    bun run lint:mcp   # Validate MCP definitions against spec

Docker

docker build -t internet-archive-mcp-server .
docker run --rm -p 3010:3010 internet-archive-mcp-server

The Dockerfile defaults to HTTP transport, stateless session mode, and logs to /var/log/internet-archive-mcp-server. OpenTelemetry peer dependencies are installed by default — build with --build-arg OTEL_ENABLED=false to omit them.

Project structure

| Directory | Purpose | |:----------|:--------| | src/index.ts | createApp() entry point — registers tools, resource, and inits services. | | src/config | Server-specific environment variable parsing and validation with Zod. | | src/mcp-server/tools | Tool definitions (*.tool.ts). Five tools across Wayback and IA library. | | src/mcp-server/resources | Resource definitions. ia://item/{identifier} item metadata resource. | | src/services/wayback | WaybackService — Availability API + CDX API client, plus archived-page charset decoding and text extraction. | | src/services/archive-search | ArchiveSearchService — Solr Advanced Search client. | | src/services/archive-metadata | ArchiveMetadataService — Metadata API + file download client. | | tests/ | Unit and integration tests mirroring src/. |

Development guide

See CLAUDE.md for development guidelines and architectural rules. The short version:

  • Handlers throw, framework catches — no try/catch in tool logic
  • Use ctx.log for request-scoped logging, ctx.state for tenant-scoped storage
  • Register new tools and resources via the barrels in src/mcp-server/*/definitions/index.ts
  • Wrap external API calls: validate raw → normalize to domain type → return output schema; never fabricate missing fields

Contributing

Issues are welcome. Run checks and tests before submitting:

bun run devcheck
bun run test

License

Apache-2.0 — see LICENSE for details.