n8n-nodes-henrks-webscraping
v1.0.4
Published
n8n community node for web scraping. Extract clean content from pages, crawl websites following internal links.
Maintainers
Readme
n8n-nodes-henrks-webscraping
This is an n8n community node for web scraping. It extracts clean content from web pages using Mozilla Readability and can crawl websites following internal links.
n8n is a fair-code licensed workflow automation platform.
Features
- Extract Clean Content: Automatically removes headers, footers, navigation, ads, and other boilerplate using Mozilla Readability
- Crawl Websites: Follow internal links to scrape multiple pages automatically
- Custom CSS Selectors: Extract specific data using CSS selectors
- Configurable Limits: Set max depth, max pages, and delay between requests
- Domain Restriction: Stay within the same domain or include subdomains
Installation
Follow the installation guide in the n8n community nodes documentation.
npm install n8n-nodes-henrks-webscrapingOperations
Scrape Page
Extract content from a single web page.
| Operation | Description | |-----------|-------------| | Extract Content | Uses Mozilla Readability to extract the main content, removing boilerplate | | Custom Selectors | Extract specific data using CSS selectors |
Crawl Site
Crawl multiple pages starting from a URL, following internal links.
| Operation | Description | |-----------|-------------| | Crawl Pages | Follows internal links up to a specified depth and page limit |
Configuration
Scrape Page Options
| Option | Description | Default | |--------|-------------|---------| | URL | The page URL to scrape | Required | | Include Raw HTML | Include original HTML in output | false | | Include Text Content | Include plain text (no HTML) | true | | Timeout | Request timeout in milliseconds | 30000 | | User Agent | Custom User-Agent header | Browser default |
Crawl Site Options
| Option | Description | Default | |--------|-------------|---------| | Start URL | The starting URL to begin crawling | Required | | Max Depth | Maximum depth of links to follow | 2 | | Max Pages | Maximum number of pages to crawl | 50 | | Same Domain Only | Only follow links within the same domain | true | | Include Subdomains | Include subdomains in domain filtering | false | | Delay Between Requests | Milliseconds to wait between requests | 1000 | | Include/Exclude Patterns | Regex patterns to filter URLs | - |
Output
Each scraped page returns:
{
"url": "https://example.com/page",
"title": "Page Title",
"content": "<p>Clean HTML content...</p>",
"textContent": "Plain text content without HTML",
"author": "Author name (if found)",
"date": "Publication date (if found)",
"siteName": "Site name",
"excerpt": "Page description/summary",
"length": 1234,
"lang": "en",
"crawledAt": "2024-01-13T10:30:00.000Z"
}Compatibility
- n8n version: 1.0.0+
- Node.js version: 18.10+
