@kreisler/js-scraper
v2.0.0
Published
A lightweight TypeScript/JavaScript web scraping library with Cloudflare bypass support (no jsdom)
Maintainers
Readme
@kreisler/js-scraper
A powerful TypeScript/JavaScript web scraping library with Cloudflare bypass support.
Features
- 🚀 Easy to use - Simple API for web scraping
- 🛡️ Cloudflare bypass - Built-in support for Cloudflare protection
- 🔄 Automatic retries - Configurable retry logic with exponential backoff
- 📊 Metadata extraction - Get response metadata (status, headers, content type, size, response time)
- 🎯 TypeScript support - Full TypeScript type definitions
- ⚡ Zero dependencies - Minimal footprint with essential libraries only
Installation
npm install @kreisler/js-scraper
# or
yarn add @kreisler/js-scraper
# or
pnpm add @kreisler/js-scraperQuick Start
Basic Scraping With Metadata
import { scrapeWithMetadata } from '@kreisler/js-scraper'
const response = await scrapeWithMetadata('https://example.com')
console.log(`Status: ${response.statusCode}`)
console.log(`Content-Type: ${response.contentType}`)
console.log(`Response Time: ${response.responseTime}ms`)
console.log(`Content Size: ${response.size} bytes`)
console.log(`Content: ${response.content}`)Advanced Usage
import { RequestService } from '@kreisler/js-scraper'
// Custom options
const html = await RequestService.fetchData({
url: 'https://example.com',
method: 'GET',
timeout: 30000,
retries: 3,
headers: {
'Custom-Header': 'value'
}
})
console.log(html)API Reference
scrapeWithMetadata(url, options?)
Fetch content with metadata.
Parameters:
url(string) - The URL to fetchoptions(object, optional)timeout(number) - Request timeout in msheaders(object) - Custom headers
Returns: ParsedResponse object with:
content(string) - HTML contentstatusCode(number) - HTTP status codeheaders(object) - Response headerscontentType(string) - Content typesize(number) - Content size in bytesresponseTime(number) - Response time in ms
RequestService.fetchData(options)
Low-level fetch with retry logic.
Parameters: FetchOptions
url(string)method(string) - 'GET', 'POST', 'HEAD'headers(object, optional)timeout(number, optional)retries(number, optional)
Returns: Promise - HTML content
RequestService.fetchWithMetadata(options)
Fetch with full metadata.
Returns: Promise
Error Handling
import { scrapeWithMetadata, FetchError } from '@kreisler/js-scraper'
try {
const result = await scrapeWithMetadata('https://example.com')
} catch (error) {
if (error instanceof FetchError) {
console.error(`Failed to fetch ${error.url}: ${error.message}`)
console.error(`Status: ${error.statusCode}`)
}
}Configuration
Disable SSL Certificate Verification (Development Only)
process.env.NODE_TLS_REJECT_UNAUTHORIZED = '0'Custom User-Agent
import { scrapeWithMetadata } from '@kreisler/js-scraper'
const result = await scrapeWithMetadata('https://example.com', {
headers: {
'User-Agent': 'Custom User Agent'
}
})License
MIT
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
