npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

site-seo-pipeline

v0.3.0

Published

Sitemap-driven SEO pipeline: crawl pages, Google Suggest + SerpAPI SERP research, Gemini meta and below-the-fold suggestions

Readme

site-seo-pipeline

Turn your sitemap into SEO recommendations for each page: crawl the HTML, pull keyword and SERP context, then generate meta title, meta description, and below-the-fold copy (plus optional FAQs) with Google Gemini.

Runs on Node.js 18+ (ESM). Works on macOS, Linux, and Windows.

For contributors & AI coding tools

This repo includes instructions that Cursor, GitHub Copilot, Claude Code, Windsurf, and similar tools can use for context:

| File | Purpose | |------|---------| | AGENTS.md | Main agent guide — layout, conventions, where to edit, testing | | CONTRIBUTING.md | PR checklist and publish notes | | CLAUDE.md | Short pointer for Claude Code | | GEMINI.md | Short pointer for Gemini / AI Studio workflows | | .cursor/rules/site-seo-pipeline.mdc | Cursor project rules | | .github/copilot-instructions.md | GitHub Copilot workspace instructions | | .windsurf/rules/site-seo-pipeline.md | Windsurf rules | | .clinerules | Cline / compatible agents |

AGENTS.md, CONTRIBUTING.md, CLAUDE.md, GEMINI.md, and .clinerules are included in the npm package (package.json files) so they also appear under node_modules/site-seo-pipeline/ for tools that index dependencies.


What it does (in order)

| Step | What happens | Needs an API key? | |------|----------------|-------------------| | 1. Sitemap | Reads your sitemap.xml (including nested sitemap indexes) and lists URLs on your site | No | | 2. Crawl | Fetches each page and reads title, meta description, headings, etc. | No | | 3. Research | Google Suggest for related queries; optional SerpAPI for real Google results (titles, snippets, related searches, People Also Ask) | Suggest: no · SerpAPI: yes | | 4. Suggest | Sends page + research to Gemini and returns structured SEO fields | Yes (GEMINI_API_KEY) |

You can skip paid steps (--skip-serp, --skip-suggest) and still get crawl + free Suggest, or crawl-only output.


Installation

npm install site-seo-pipeline

Quick start (CLI)

1. Get keys (optional but recommended for full output):

2. Run:

export GEMINI_API_KEY="your-key"
export SERPAPI_API_KEY="your-key"

npx site-seo-pipeline \
  --sitemap https://www.example.com/sitemap.xml \
  --site-url https://www.example.com \
  --brand "Example Inc" \
  --limit 10 \
  --out ./seo-output.json

3. Open seo-output.json. Each entry has:

  • snapshot — what the crawler saw on the page
  • research — Suggest keywords + optional SERP snapshots
  • suggestion — Gemini’s metaTitle, metaDescription, belowTheFoldMarkdown, optional faqs
  • errors — per-page issues (e.g. crawl timeout, API error)

If you omit GEMINI_API_KEY, you’ll see a warning and no suggestion (everything else can still run). If you omit SERPAPI_API_KEY, SERP snapshots are skipped but Google Suggest still runs.


Starting workflows

  • CLI: follow Quick start (CLI)npx site-seo-pipeline with --sitemap, --site-url, and --brand (plus keys for full output).
  • Library: call runPipeline from site-seo-pipeline with seedSitemaps, site, limits, and optional API keys — see Use from your own code.

When to use Playwright (--render-js)

Default crawling uses plain HTTP + HTML parsing (fast, no browser). Use this if your pages are server-rendered or the important tags are in the first HTML response.

Use --render-js if the site is a client-side SPA and title/description only appear after JavaScript runs:

npm install playwright
npx playwright install chromium

npx site-seo-pipeline --sitemap https://example.com/sitemap.xml \
  --site-url https://example.com --brand "Example" --render-js

CLI reference

Required (or use env)

| Input | Flag | Env fallback | |--------|------|----------------| | Sitemap URL | --sitemap <url> | SEO_SITEMAP_URL or SITEMAP_URL |

Site / copy context

| Flag | Description | Env fallback | |------|-------------|--------------| | --site-url | Canonical origin (e.g. https://www.example.com) | SITE_URL, else derived from sitemap URL | | --brand | Brand name used in prompts | BRAND_NAME, else hostname | | --locale | Language hint (hl), e.g. en | LOCALE (default en) | | --region | Market (gl), e.g. us, in, gb | REGION (default us) | | --industry | Short description for the model | INDUSTRY | | --tone | Voice, e.g. professional | TONE | | --cta | Primary call-to-action phrase | PRIMARY_CTA |

Run control

| Flag | Default | Description | |------|---------|-------------| | --out <path> | ./seo-pipeline-output.json | Where to write JSON | | --limit <n> | 50 | Max URLs to process (after sitemap discovery) | | --crawl-concurrency <n> | 4 | Parallel page fetches | | --render-js | off | Use Playwright instead of HTTP fetch | | --skip-serp | off | Do not call SerpAPI | | --skip-suggest | off | Do not call Gemini | | --max-serp-calls <n> | 80 | Max SerpAPI requests for the whole run (each URL may use up to 2 when budget remains) | | --gemini-model <id> | gemini-2.5-flash | Override if your key doesn’t support the default (e.g. gemini-1.5-flash) |

Environment variables (summary)

| Variable | Role | |----------|------| | GEMINI_API_KEY | Gemini suggestions | | SERPAPI_API_KEY | Google SERP via SerpAPI | | SITE_URL, BRAND_NAME, LOCALE, REGION, INDUSTRY, TONE, PRIMARY_CTA | Same as flags above | | SEO_SITEMAP_URL / SITEMAP_URL | Sitemap if --sitemap omitted | | MAX_SERP_CALLS | Default for --max-serp-calls |

The CLI loads a .env file from the current working directory when you run the binary.


API reference

  • CLI: all flags and env fallbacks are listed in CLI reference below.
  • Programmatic: import from site-seo-pipeline — see Use from your own code for runPipeline and composed steps. Source exports live in src/index.ts (published build: dist/).

Configuration

Use CLI flags and environment variables together (flags win when both are set). The CLI reads a .env file in the current working directory.

| Area | Where it’s documented | |------|------------------------| | Site / brand / locale | Site / copy context | | Concurrency, limits, skips | Run control | | Env var names | Environment variables (summary) |


Use from your own code

Your project must use ESM ("type": "module" or .mjs) or a bundler that resolves ESM.

Option A — one call: runPipeline

import { runPipeline } from "site-seo-pipeline";

const results = await runPipeline({
  site: {
    brandName: "Example Inc",
    siteUrl: "https://www.example.com",
    locale: "en",
    region: "us",
    industry: "Event marketplace",
  },
  seedSitemaps: ["https://www.example.com/sitemap.xml"],
  limit: 25,
  crawlConcurrency: 4,
  geminiApiKey: process.env.GEMINI_API_KEY,
  serpApiKey: process.env.SERPAPI_API_KEY,
  // Optional:
  // skipSerp: true,
  // skipSuggest: true,
  // renderJs: true,
  // maxSerpCalls: 40,
  // includeSitemapUrl: (url) => !url.includes("/admin"),
});

for (const row of results) {
  console.log(row.url, row.suggestion?.metaTitle, row.errors);
}

Pass geminiApiKey / serpApiKey the same way as env vars in the CLI; if omitted, those steps are skipped (same behavior as missing env).

Option B — compose steps yourself

Useful if you already have URLs, or you only want crawl/research without Gemini.

import {
  discoverPageUrlsFromSitemaps,
  originOnly,
  fetchPageSnapshot,
  gatherPageResearch,
  suggestSeoWithGemini,
} from "site-seo-pipeline";

const siteUrl = "https://www.example.com";
const base = originOnly(siteUrl);

const urls = await discoverPageUrlsFromSitemaps(
  [`${base}/sitemap.xml`],
  base,
  { maxFetches: 500 }
);

const snapshot = await fetchPageSnapshot(urls[0]!);

const research = await gatherPageResearch(snapshot, {
  site: {
    brandName: "Example",
    siteUrl,
    locale: "en",
    region: "us",
  },
  serpApiKey: process.env.SERPAPI_API_KEY,
  skipSerp: !process.env.SERPAPI_API_KEY,
  suggestDelayMs: 400,
});

const suggestion = await suggestSeoWithGemini(
  { brandName: "Example", siteUrl },
  snapshot,
  research,
  { apiKey: process.env.GEMINI_API_KEY!, model: "gemini-2.5-flash" }
);

Other exports include playwrightPageSnapshot, serpApiGoogleSearch, googleSuggest, parseModelJson, and mapPool — see src/index.ts.


Output shape (JSON)

Each array element is roughly:

{
  "url": "https://www.example.com/page",
  "snapshot": {
    "url": "...",
    "mode": "fetch",
    "title": "...",
    "metaDescription": "...",
    "h1Tags": ["..."],
    "h2Tags": ["..."],
    "wordCount": 0,
    "error": "only if crawl failed"
  },
  "research": {
    "queriesUsed": ["..."],
    "googleSuggest": ["..."],
    "serpByQuery": [
      {
        "query": "...",
        "organic": [{ "position": 1, "title": "...", "link": "...", "snippet": "..." }],
        "relatedSearches": ["..."],
        "peopleAlsoAsk": [{ "question": "...", "snippet": "..." }]
      }
    ]
  },
  "suggestion": {
    "metaTitle": "...",
    "metaDescription": "...",
    "belowTheFoldMarkdown": "## Section\\n...",
    "faqs": [{ "question": "...", "answer": "..." }],
    "serpAngle": "..."
  },
  "errors": []
}

suggestion may be missing if Gemini was skipped or failed; check errors for that page.


Costs and limits (practical notes)

  • Google Suggest is free but rate-sensitive; the pipeline spaces requests (suggestDelayMs, default ~350ms in runPipeline).
  • SerpAPI is billed per search; use --max-serp-calls / maxSerpCalls to cap spend.
  • Gemini is billed by your Google AI plan; use --limit while testing.

Troubleshooting

| Issue | What to try | |--------|-------------| | Empty title / meta on many pages | Use --render-js and install Playwright + Chromium | | No suggestions | Set GEMINI_API_KEY or pass geminiApiKey; check errors for API messages | | Model not found | --gemini-model gemini-1.5-flash (or another model enabled for your key) | | SerpAPI errors | Check quota and key; use --skip-serp to test without SERP |


Testing

npm test
npm run test:coverage

Development

git clone <your-repo>
cd site-seo-pipeline
npm install
npm run build
npm test

Contributing

See CONTRIBUTING.md for the PR checklist and publish notes.


License

MIT