google-crawl-simulator
v1.0.1
Published
Simulates how Googlebot crawls and indexes a page: robots.txt rules, raw HTML crawl (wave 1), JS-rendered crawl (wave 2, like Google's Web Rendering Service), and SEO signal extraction (title/meta/canonical/hreflang/structured data/alt text/links/status c
Readme
google-crawl-simulator
Simulates how Googlebot actually crawls and indexes a page, replicating Google's real two-wave process:
- robots.txt check — fetches
/robots.txt, parsesUser-agent/Allow/Disallowgroups the way Google's spec resolves them (longest match wins, 5xx = full disallow, 404 = full allow), and decides whether Googlebot may fetch the URL. - Wave 1 — raw HTML crawl — fetches the URL with a Googlebot user agent, follows redirects manually (reporting the full chain), and parses the raw (non-JS-executed) HTML.
- Wave 2 — JS rendering — loads the page in headless Chromium via Playwright (Google's Web Rendering Service is also an evergreen Chromium), with a configurable render-timeout to mimic Google's finite render/crawl budget, and captures failed requests + console errors that could break rendering.
- SEO/indexability signal extraction — for both raw and rendered HTML: title, meta description, canonical, meta robots (noindex/nofollow), lang, viewport, headings, image alt text, internal/external/nofollow links, JSON-LD structured data, Open Graph/Twitter Card tags, and visible word count.
- Raw vs. rendered diff — flags when content, title, meta description, or canonical only appear after JS execution, which is exactly the situation that delays or risks indexing in the real Google pipeline.
Setup
npm install(postinstall downloads Playwright's Chromium automatically.)
Usage
npx google-crawl-simulator http://localhost:3000 [options]Options:
-m, --mobile— use the Googlebot Smartphone UA + mobile viewport (Google is mobile-first)--no-render— skip JS rendering, only do the raw HTML crawl--timeout <ms>— render timeout (default 15000), simulating Google's render budget--json <file>— dump the full structured result to a JSON file
Example:
node src/index.js --mobile https://example.com --json report.jsonYou can also install it as a global command:
npm link
gcrawl https://example.comWhat it does NOT simulate
- PageRank/link-graph-based crawl prioritization or actual crawl scheduling/frequency.
- Google's exact render queue delay (real Wave 2 can lag Wave 1 by hours to weeks — this tool renders immediately).
- Signals that require Search Console data (actual indexing status, coverage errors, manual actions).
