pendra
v0.17.2
Published
Official TypeScript SDK for the Pendra LLM inference API
Maintainers
Readme
Pendra TypeScript SDK
Official TypeScript/JavaScript SDK for the Pendra LLM inference API. Mirrors the OpenAI SDK interface for easy migration.
Requires Node.js 18+ (uses native fetch). Zero runtime dependencies.
Installation
npm install pendraQuick Start
import Pendra from 'pendra';
const client = new Pendra({
apiKey: 'pdr_sk_...', // or set PENDRA_API_KEY env var
});
const response = await client.chat.completions.create({
model: 'qwen3.5:0.8b',
messages: [{ role: 'user', content: 'Hello!' }],
});
console.log(response.choices[0].message.content);Your first request — full sequence
Pendra serves inference from workers you (or your org) run, so a brand-new account needs three things in place before that chat.completions.create() call returns a 200:
- A worker connected. Install Pendra on any host with a GPU or CPU and run
pendra setup. The wizard walks through pasting a worker key from console.pendra.ai/workers and connecting to the API. - A model on disk. The worker only serves models it has locally. From the worker host run, e.g.,
pendra models install qwen3.5:0.8b. Browse console.pendra.ai/models for the full catalogue. - An API key. Create one at console.pendra.ai/api-keys and pass it as
apiKeyabove.
If your call returns 404 Model 'X' is in the catalogue but no connected worker has it installed yet, skip back to step 2 — that's the API telling you the model is known but hasn't been pulled onto a worker yet.
Streaming
const stream = await client.chat.completions.create({
model: 'qwen3.5:0.8b',
messages: [{ role: 'user', content: 'Write a poem' }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
}Reasoning models
A reasoning model's chain-of-thought comes back separately from its answer, on
reasoning_content (mirrored on reasoning, and streamed as
delta.reasoning_content). On a hybrid model you can skip the thinking
entirely and get straight to the answer:
const response = await client.chat.completions.create({
model: 'qwen3.6:27b',
messages: [{ role: 'user', content: "Classify: 'ship it'. positive or negative?" }],
max_tokens: 5,
enable_thinking: false, // or reasoning_effort: 'none'
});
console.log(response.choices[0].message.reasoning_content); // the working, when thinking is on
console.log(response.choices[0].message.content); // the answerNotices, timings and other response extras
Replies carry a Pendra-specific pendra field alongside choices. Read
response.pendra?.notice when an answer looks wrong for no obvious reason —
truncated_during_reasoning is the model saying it spent the whole
max_tokens budget thinking and never reached an answer, which otherwise just
looks like an empty message.content:
if (response.pendra?.notice?.code === 'truncated_during_reasoning') {
console.log(response.pendra.notice.message);
// Where the budget went:
console.log(response.usage?.completion_tokens_details?.reasoning_tokens);
// Retry with a bigger max_tokens, or with enable_thinking: false.
}The other codes are truncated_during_structured_output and
strict_schema_not_enforced. The same field carries pendra.web_tool_steps
when the serving worker has web tools enabled.
When streaming, the last few chunks report on the request rather than
continuing the answer — one carries chunk.usage, one carries
chunk.pendra?.timing (ttft_ms, tokens_per_second, queue_wait_ms,
worker_queue_wait_ms), and a notice arrives on its own chunk too:
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
if (chunk.pendra?.notice) console.warn(chunk.pendra.notice.message);
if (chunk.pendra?.timing) console.log(`${chunk.pendra.timing.tokens_per_second} tok/s`);
}Those chunks have no text, which is why the loop above reads content with
chunk.choices[0]?.delta?.content rather than indexing choices[0]
unconditionally.
Which worker served a request
Every reply names the worker that served it, as pendra.worker (id, the
name you gave it in the console, and the worker version that served the
request; either of the last two can be null), alongside a request_id worth
quoting to support:
const response = await client.chat.completions.create({ ... });
console.log(response.pendra?.worker?.id); // 'wrk-1a2b3c'
console.log(response.pendra?.worker?.name); // 'gpu-box-1'
console.log(response.pendra?.worker?.version); // '3.115.1'
console.log(response.pendra?.request_id);A stream reports it on its last chunk, so read it from the stream once you've
finished iterating (it is null until then):
const stream = await client.chat.completions.create({ ..., stream: true });
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '');
}
console.log(stream.worker?.id, stream.requestId);Embeddings, reranking, image generation, transcriptions and speech carry the same field:
const embeddings = await client.embeddings.create({ model: 'nomic-embed-text:latest', input: 'hi' });
console.log(embeddings.pendra?.worker?.id);It works the same with private inference turned on.
Parameters
chat.completions.create() passes everything you give it to the API —
including fields these types don't declare yet, so a Chat Completions
parameter newer than your SDK version still reaches the model. TypeScript
rejects an undeclared key written straight into the call (that's what catches
typos), so pass a newer one through a variable or a cast when you mean it. The
only field that isn't sent as JSON is worker_id, which travels as the
X-Pendra-Worker-Id routing header.
List Models
const models = await client.models.list();
models.forEach((m) => console.log(m.id));Image Generation
Generate images from a text prompt. Returns base64-encoded PNGs by default.
import { writeFileSync } from 'node:fs';
const response = await client.images.generations.create({
model: 'x/z-image-turbo',
prompt: 'A red London double-decker bus at sunset',
size: '1024x1024',
});
const b64 = response.data[0].b64_json;
if (b64) {
writeFileSync('bus.png', Buffer.from(b64, 'base64'));
}Image generation is non-streaming — the response is returned as a single JSON payload once the worker finishes.
Text to speech
Turn text into spoken audio. You get back a complete WAV file:
import { writeFileSync } from 'node:fs';
const speech = await client.audio.speech.create({
model: 'qwen3-tts',
input: 'Hello from Pendra.',
// voice: 'Vivian', // optional: leave it out for the model's default voice
});
writeFileSync('speech.wav', speech.audio);
console.log(speech.duration_ms); // length of the audio, in ms
console.log(speech.pendra?.worker?.id); // the worker that generated itCode written for the OpenAI SDK keeps working:
Buffer.from(await speech.arrayBuffer()) returns the same bytes.
Speech isn't available with private inference yet. On a client with
privateInference turned on, speech.create() throws a
PrivateInferenceError and sends nothing.
Configuration
| Option | Env var | Default |
| --------- | ---------------- | -------------------------- |
| apiKey | PENDRA_API_KEY | — |
| baseURL | — | https://api.pendra.ai |
| timeout | — | 120000 (ms) |
For a streaming request, timeout covers waiting for the response to start and
then any gap between chunks, not the length of the whole stream, so long
generations are not cut off. Breaking out of a stream's for await loop early
cancels the request.
Migrating from OpenAI
- import OpenAI from 'openai';
- const client = new OpenAI({ apiKey: '...' });
+ import Pendra from 'pendra';
+ const client = new Pendra({ apiKey: 'pdr_sk_...' });
// Everything else stays the same
const res = await client.chat.completions.create({
model: 'qwen3.5:0.8b',
messages: [{ role: 'user', content: 'Hi' }],
});Error Handling
import { AuthenticationError, RateLimitError, APIStatusError } from 'pendra';
try {
await client.chat.completions.create({ ... });
} catch (err) {
if (err instanceof AuthenticationError) {
console.error('Bad API key');
} else if (err instanceof RateLimitError) {
console.error('Slow down');
} else if (err instanceof APIStatusError) {
console.error(`API error ${err.status}: ${err.message}`);
}
}Self-Hosted Workers
Run inference on your own GPUs with a single command. Your prompts and completions never leave your infrastructure.
curl -fsSL https://get.pendra.ai/worker | bashSee the Workers documentation for full setup instructions.
Licence
Apache-2.0
