@scraping-proxy/client
v0.23.0
Published
Typed HTTP client for the scraping-proxy API. Requires Node.js `>=20`.
Readme
@scraping-proxy/client
Typed HTTP client for the scraping-proxy API. Requires Node.js >=20.
What It Provides
ScrapingProxyClient— main client classScrapingProxyClientOptions— constructor options typegetAuthState(filePath)— read a browser auth state file from diskSubscription— re-exported from@adonisjs/transmit-client- Types:
ScrapeRequest,ScrapeRequestRaw,ScrapeRequestWithSelectors,JobResult,JobResultRaw,JobResultHtmlBrowser,JobStatus,ScrapeMode,ScrapeOptions,SelectorSpec,SelectorInput,GroupedSelectorSpec,Selectors,ScrapeMeta,RateLimitGuidance,BrowserAction,BrowserStorageState
Installation
pnpm add @scraping-proxy/clientBasic Usage
import { ScrapingProxyClient } from '@scraping-proxy/client';
const client = new ScrapingProxyClient({
baseUrl: 'https://proxy.example.com',
apiToken: 'your-api-token', // Settings → API Tokens in the webapp
});
const job = await client.scrapeAndWait({
url: 'https://example.com',
scrapeMode: 'html',
selectors: {
title: 'h1',
link: { selector: 'a.cta', attribute: 'href' },
},
});
console.log(job.result); // { title: string[], link: string[] }
console.log(job.status); // 'done'
console.log(job.meta); // { statusCode: 200, proxyConfigName: '...', durationMs: ..., attempts: 1, rateLimit: {...} }Configuration
ScrapingProxyClient requires:
baseUrl— base URL of the scraping-proxy server.apiToken— bearer token created under Settings → API Tokens.
const client = new ScrapingProxyClient({
baseUrl: 'https://proxy.example.com',
apiToken: 'your-api-token',
});Scrape Modes
| scrapeMode | Description |
| ------------ | ------------------------------------------------------------- |
| "raw" | Returns raw HTTP response body. selectors not allowed. |
| "html" | Static HTML scraping with CSS selector extraction. |
| "browser" | Headless browser for JS-rendered pages. Supports authState. |
Retries
requestOptions.retries (0-10, default 0) retries a failed attempt (fetch error, or non-2xx/non-ok result) before the job fails — same proxy reused across attempts, a fixed delay between them. job.meta.attempts reflects how many were actually made. This is separate from the queue's own automatic retry on no_proxy/rate_limited — that one isn't configurable and doesn't count toward attempts.
const job = await client.scrapeAndWait({
url: 'https://flaky.example.com',
scrapeMode: 'html',
requestOptions: { retries: 3 },
});
console.log(job.meta?.attempts); // up to 4Browser Actions
requestOptions.actions ("browser" mode only) runs interaction steps in order, after navigation, before selector extraction — click a cookie-consent button, scroll to trigger lazy-loaded content, wait for something to appear. Runs once through, no looping — not a fit for open-ended infinite scroll.
const job = await client.scrapeAndWait({
url: 'https://example.com/listing',
scrapeMode: 'browser',
requestOptions: {
actions: [
{ type: 'click', selector: '.cookie-consent-accept' },
{ type: 'scroll', to: 'bottom' },
{ type: 'waitForSelector', selector: '.lazy-loaded-content' },
],
},
selectors: { title: 'h1' },
});Action types: { type: 'click', selector, timeout? }, { type: 'waitForSelector', selector, timeout? }, { type: 'waitForTimeout', ms }, { type: 'scroll', to: 'top' | 'bottom' | <CSS selector>, timeout? }.
API Methods
scrapeAndWait(request)
Enqueue a job and wait for completion. Selector keys are narrowed at compile time when a literal object is passed.
const job = await client.scrapeAndWait({
url: 'https://example.com',
scrapeMode: 'html',
selectors: { title: 'h1' },
});
// job.result.title is string[]
const raw = await client.scrapeAndWait({
url: 'https://example.com',
scrapeMode: 'raw',
});
// raw.result is string | nullGrouped selectors (repeated containers)
A flat selector ({ selector: 'h2' }, 'h2') returns string[] — every match on the whole page, in DOM order. That's fine for independent fields, but it desyncs when the same key needs to correlate with a sibling field across a repeated element (a listing card, a key/value table row) and one field is missing on a single instance.
Use a container + fields spec instead: it scopes each field to one matched container instance, so a missing field becomes null in that instance's slot rather than shifting every later array entry by one.
const job = await client.scrapeAndWait({
url: 'https://example.com/listing',
scrapeMode: 'html',
selectors: {
items: {
container: '.result-card',
fields: {
title: 'h2',
link: { selector: 'a', attribute: 'href' },
image: { selector: 'img', attribute: 'src' },
},
},
},
});
// job.result.items: Array<{ title: string | null; link: string | null; image: string | null }>Each field resolves to a single string | null (the first match inside that container instance) — not an array. container accepts a fallback array (string[]) same as a flat selector. Works in both "html" and "browser" mode.
scrapeAndWaitMany(requests)
Batch-enqueue multiple jobs and wait for all to complete. Jobs run concurrently on the server.
const jobs = await client.scrapeAndWaitMany([
{ url: 'https://a.com', scrapeMode: 'html', selectors: { title: 'h1' } },
{ url: 'https://b.com', scrapeMode: 'html', selectors: { title: 'h1' } },
]);scrape(request)
Enqueue a job and return immediately with the job ID. Use waitForJob to await completion later.
const { data } = await client.scrape({
url: 'https://example.com',
scrapeMode: 'raw',
});
console.log(data.jobId); // UUIDscrapeMany(requests)
Batch-enqueue multiple jobs and return immediately with an array of job IDs.
const jobs = await client.scrapeMany([...]);waitForJob(jobId)
Wait for a job to reach a terminal state (done, failed, or cancelled) via SSE.
const result = await client.waitForJob(jobId);getJob(jobId)
Fetch the current state of a job without waiting.
const { data } = await client.getJob(jobId);
console.log(data.status); // 'pending' | 'running' | 'done' | 'failed' | ...cancelJob(jobId)
Cancel a single non-running job. Returns true if cancelled.
const cancelled = await client.cancelJob(jobId);cancelJobs(jobIds)
Cancel multiple non-running jobs. Returns the count actually cancelled.
const { cancelled } = await client.cancelJobs([id1, id2]);startProxies()
Warms every proxy in the shared, admin-managed pool by unpausing idle ones. Pool membership itself (how many proxies exist) is provisioned by the server admin, not by this call — there's no count argument, it always warms the whole pool.
const { warm, total } = await client.startProxies();session(callback)
Runs callback inside a scrape session: opens the session (warming the whole pool), keeps it alive with an automatic heartbeat, and always closes it afterwards — so the shared pool pauses once the last concurrent session's work ends. Prefer this over openSession/heartbeatSession/closeSession directly.
const jobs = await client.session(() =>
client.scrapeAndWaitMany([
{ url: 'https://a.com', scrapeMode: 'raw' },
{ url: 'https://b.com', scrapeMode: 'raw' },
])
);openSession() / heartbeatSession(sessionId) / closeSession(sessionId)
Lower-level session primitives if you need to manage the lifecycle yourself instead of session(). openSession() takes no arguments — it warms the whole pool, same as startProxies().
const { sessionId } = await client.openSession();
// ... periodically: await client.heartbeatSession(sessionId);
await client.closeSession(sessionId);subscribeToJob(jobId)
Subscribe to the SSE channel of a specific job. Resolves the channel name automatically.
const sub = await client.subscribeToJob(jobId);
await sub.create();
sub.onMessage<{ jobStatus: string }>((msg) => console.log(msg.jobStatus));subscribe(channel)
Subscribe to a raw Transmit channel. Prefer subscribeToJob for job events.
const sub = client.subscribe(`jobs-list/${userId}`);closeEventStream()
Close the underlying SSE connection. It reconnects on the next subscribeToJob / subscribe call.
client.closeEventStream();Auth State (Browser Mode)
Use the scraping-proxy-auth CLI to capture a logged-in browser state, then pass it to scrape calls that require authentication.
import { getAuthState } from '@scraping-proxy/client';
const authState = await getAuthState('state.json');
const job = await client.scrapeAndWait({
url: 'https://protected.example.com/dashboard',
scrapeMode: 'browser',
authState,
});Proxy Lifecycle Example
const jobs = await client.session(() =>
client.scrapeAndWaitMany([
{ url: 'https://a.com', scrapeMode: 'raw' },
{ url: 'https://b.com', scrapeMode: 'raw' },
])
);
console.log(jobs.map((j) => j.result));Error Handling
All methods throw when the API responds with a non-2xx status or when a job completes with status: "failed".
try {
const job = await client.scrapeAndWait({
url: 'https://example.com',
scrapeMode: 'raw',
});
if (job.status === 'failed') {
console.error('Job failed:', job.error);
}
} catch (error) {
console.error('HTTP error:', error);
}