npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

ai-sdk-agents-universal-scraper-tool

v1.0.0

Published

AI SDK Agents Universal Scraper Tool

Downloads

15

Readme

AI SDK Robust Scraping + Crawling Tool

A TypeScript package providing robust web scraping and crawling tools for AI SDK agents. Supports multiple providers (Exa, Firecrawl, Cheerio) with automatic fallback on rate limits.

Installation

npm install ai-sdk-agents-universal-scraper-tool

Usage

Basic Scraping

import { scrapeTool } from "ai-sdk-agents-universal-scraper-tool";
import { generateText } from "ai";
import { openai } from "@ai-sdk/openai";

const result = await generateText({
  model: openai("gpt-4o-mini"),
  prompt: "Scrape and summarize the content from https://example.com",
  tools: {
    scrapeTool,
  },
});

Custom Scraping Tool

import { createScrapeTool } from "ai-sdk-agents-universal-scraper-tool";

// Create a tool with default provider and settings
const customScrapeTool = createScrapeTool({
  defaultProvider: "exa",
  defaultMaxChars: 5000,
  defaultMarkdown: true,
});

Crawling with Subpage Discovery

import {
  crawlTool,
  crawlBatchTool,
} from "ai-sdk-agents-universal-scraper-tool";

// Crawl a single page with subpage discovery
const crawlResult = await crawlTool.execute({
  url: "https://example.com",
  maxSubpages: 5,
  maxDepth: 2,
});

// Crawl multiple URLs in batch
const batchResult = await crawlBatchTool.execute({
  urls: ["https://example.com", "https://another.com"],
  maxSubpages: 3,
});

Features

  • Multiple Providers: Supports Exa, Firecrawl, and Cheerio
  • Automatic Fallback: Automatically falls back to alternative providers on rate limits
  • Flexible Output: Returns markdown, HTML, or plain text
  • Subpage Discovery: Crawl tools can discover and crawl linked subpages
  • Configurable: Customize defaults for provider, max characters, format, and more

Development

Setup

  1. Clone the repository
  2. Install dependencies:
pnpm install
  1. Create a .env.local file with your API keys:
# Exa API key (optional)
EXA_API_KEY=your_exa_api_key

# Firecrawl API key (optional)
FIRECRAWL_API_KEY=your_firecrawl_api_key

Note: The tools will automatically fall back to Cheerio (no API key required) if other providers are unavailable.

Testing

Test your tool locally:

pnpm test

Building

Build the package:

pnpm build

Publishing

Before publishing, update the package name in package.json to your desired package name.

The package automatically builds before publishing:

pnpm publish

Project structure

.
├── src/
│   ├── index.ts           # Tool exports
│   ├── scraper-tool.ts    # Scraping tool implementation
│   ├── crawler-tool.ts    # Crawling tool implementation
│   ├── scraper.ts         # Scraping logic
│   ├── crawler.ts         # Crawling logic
│   └── *.test.ts          # Test files
├── dist/                  # Build output (generated)
├── package.json
├── tsconfig.json
└── README.md

Exported Tools

  • scrapeTool - Default scraping tool instance
  • createScrapeTool() - Create a custom scraping tool with defaults
  • crawlTool - Default crawling tool instance (single URL)
  • createCrawlTool() - Create a custom crawling tool with defaults
  • crawlBatchTool - Default batch crawling tool instance (multiple URLs)
  • createCrawlBatchTool() - Create a custom batch crawling tool with defaults

Providers

Exa

  • Requires EXA_API_KEY environment variable
  • Supports live crawling with configurable preferences
  • High-quality content extraction

Firecrawl

  • Requires FIRECRAWL_API_KEY environment variable
  • Supports caching with configurable max age
  • Good for structured content extraction

Cheerio

  • No API key required (local processing)
  • Fast and reliable fallback option
  • Supports subpage crawling with configurable concurrency

License

ISC