xai-sdk-js
v0.4.0
Published
TypeScript/JavaScript SDK for the xAI API (gRPC)
Maintainers
Readme
xai-sdk-js
TypeScript/JavaScript SDK for the xAI API, ported from the official xai-sdk Python package.
Talks to api.x.ai over gRPC (Connect-ES). Async-first. Node.js 18+, plus a fetch-only build for Cloudflare Workers and other edge runtimes.
Install
npm install xai-sdk-js
# or
bun add xai-sdk-js
# or
pnpm add xai-sdk-jsAPI key resolution
The client resolves secrets in this order (first hit wins):
- Constructor options:
new Client({ apiKey: "..." }) - Process environment:
export XAI_API_KEY=...(andXAI_MANAGEMENT_KEY) - Dotenv files in the working directory (or
envDir):
| File | Typical use |
| --- | --- |
| .env | defaults |
| .env.<mode> | .env.production, .env.staging, .env.development (mode = XAI_ENV or NODE_ENV) |
| other .env.* | custom env names |
| .env.local | local overrides (usually gitignored) |
| .env.<mode>.local | mode-specific local overrides |
Existing shell/process.env values are never overwritten by files. No dotenv dependency is required.
# shell
export XAI_API_KEY=xai-...
export XAI_MANAGEMENT_KEY=... # optional, Collections management
# or a file
echo 'XAI_API_KEY=xai-...' >> .env.localQuick start
import { Client, user, system } from "xai-sdk-js";
const client = new Client(); // reads XAI_API_KEY
const chat = client.chat.create({
model: "grok-4",
messages: [system("You are helpful.")],
});
chat.append(user("Explain black holes in one sentence."));
const response = await chat.sample();
console.log(response.content);
console.log("cost USD:", response.costUsd);Multi-turn + prompt cache (lower cost)
xAI sticky prompt-cache routing needs a stable conversation id on every chat RPC as header x-grok-conv-id. Pass it as conversationId on chat.create — the SDK sends it automatically on sample/stream/defer/parse/compact.
import { Client, system, user } from "xai-sdk-js";
const client = new Client();
const conversationId = "thread_abc123"; // stable per app conversation
let previousResponseId: string | undefined;
// Turn 1
{
const chat = client.chat.create({
model: "grok-4",
conversationId,
storeMessages: true,
messages: [system("You are helpful."), user("What is prompt caching?")],
});
const res = await chat.sample();
previousResponseId = res.id;
console.log(res.content);
console.log("cached tokens:", res.usage?.cachedPromptTextTokens);
}
// Turn 2 — only the new user message; server holds prior turns via previousResponseId
{
const chat = client.chat.create({
model: "grok-4",
conversationId,
storeMessages: true,
previousResponseId,
messages: [user("Show a short example.")],
});
const res = await chat.sample();
previousResponseId = res.id;
console.log("cached tokens:", res.usage?.cachedPromptTextTokens);
}Best practices:
- Always pass a stable
conversationIdfor multi-turn (do not mint a new one each request). - Set
storeMessages: truewhen chaining withpreviousResponseId. - Follow-ups: send only the new user message when using
previousResponseId(don’t resend full history). - If you resend full history instead of chaining, don’t edit/reorder earlier turns.
- Watch
usage.cachedPromptTextTokens— stuck at0usually means the id/header/routing is wrong. - One-shot helpers (titles, classifiers): omit
conversationIdor use a unique id; keepstoreMessages: false(the default). - Front-load static content (system prompt, few-shot examples, reference docs) so the stable prefix is as long as possible.
- On reasoning models, replay
reasoningContentfrom prior turns (chat.append(response)does this), or setuseEncryptedContent: true. Dropping it is the top cause of cache misses. - Cache hits are best-effort. Entries get evicted, so the app must still work at full price.
You can also set client-wide metadata x-grok-conv-id, but per-chat conversationId is correct for concurrent threads — a per-chat value wins over client metadata on that header.
Context compaction (long agent loops)
Once a conversation gets long, every turn re-pays input tokens for the whole history. chat.compact() folds the messages into one opaque blob and replaces the chat's messages with it in place, so later sample() calls run on top of the compacted context.
const chat = client.chat.create({
model: "grok-4.6",
conversationId,
useEncryptedContent: true, // keeps prior reasoning through the compaction
});
chat.append(system("You are helpful. Keep answers brief."));
for (let turn = 1; turn <= 100; turn++) {
chat.append(user(nextUserMessage()));
const res = await chat.sample();
chat.append(res);
if (turn % 5 === 0) {
const c = await chat.compact();
console.log(`compacted, dropped ${c.droppedMessageCount} messages`);
}
}Rules:
- Treat
encryptedContentas opaque. Never parse, edit, or merge blobs. - The compaction item becomes the new head. Only append new turns after it.
- Compaction shrinks a conversation; it cannot rescue one already over the context limit.
- Re-compacting later is fine.
- The compaction call itself costs tokens (
compact.usage), so compact every N turns, not every turn. - Set
useEncryptedContent: trueon reasoning models so prior reasoning survives.
Streaming
const chat = client.chat.create({ model: "grok-4" });
chat.append(user("Write a short poem about space."));
for await (const [response, chunk] of chat.stream()) {
const delta = chunk.content;
if (delta) process.stdout.write(delta);
}Tools & search
import { Client, user, webSearch, SearchParameters, webSource } from "xai-sdk-js";
const client = new Client();
const chat = client.chat.create({
model: "grok-4",
tools: [webSearch()],
searchParameters: new SearchParameters({
mode: "auto",
sources: [webSource()],
returnCitations: true,
}),
});
chat.append(user("What are the latest developments from xAI?"));
const res = await chat.sample();
console.log(res.content);Images
const img = await client.image.sample("A watercolor fox under starlight", "grok-imagine-image", {
aspectRatio: "16:9",
resolution: "2k",
});
console.log(img.url);Video (deferred + poll)
const video = await client.video.generate("A drone shot over misty mountains", "grok-imagine-video", {
aspectRatio: "16:9",
duration: 5,
});
console.log(video.url);
// Pin first, last and mid-clip frames (grok-imagine-video-1.5 only)
const pinned = await client.video.generate("The sketch becomes clay, then bronze", "grok-imagine-video-1.5", {
imageUrl: "https://example.com/sketch.jpg",
lastFrameUrl: "https://example.com/bronze.jpg",
keyframes: [{ imageUrl: "https://example.com/clay.jpg", timestamp: 4 }],
duration: 8,
});Files
const file = await client.files.upload("./notes.pdf");
console.log(file.id, file.filename);
const bytes = await client.files.content(file.id);Batch
import { user } from "xai-sdk-js";
const batch = await client.batch.create("capitals");
const chats = ["UK", "USA", "Egypt"].map((country) => {
const c = client.chat.create({
model: "grok-4",
batchRequestId: `capital_${country}`,
});
c.append(user(`Capital of ${country}?`));
return c;
});
await client.batch.add(batch.batchId, chats);
const page = await client.batch.listBatchResults(batch.batchId);
for (const r of page.succeeded) {
console.log(r.batchRequestId, r.response.content);
}Auth / models / tokenize
const info = await client.auth.getApiKeyInfo();
const models = await client.models.listLanguageModels();
const tokens = await client.tokenize.tokenizeText("hello world", "grok-4");Collections (needs management key)
const client = new Client({ managementApiKey: process.env.XAI_MANAGEMENT_KEY });
const col = await client.collections.create("docs", { modelName: "grok-embedding" });
await client.collections.uploadDocument(col.collectionId, "readme.md", "# Hello", {
waitForIndexing: true,
});
const hits = await client.collections.search("hello", [col.collectionId], { limit: 5 });Cloudflare Workers / edge runtimes
The default entry point uses gRPC over HTTP/2 (node:http2), which Cloudflare Workers ship only as a non-functional stub. Import the web entry point instead — it speaks gRPC-Web over fetch and pulls in no Node built-ins.
import { Client, user } from "xai-sdk-js/web";
export default {
async fetch(request: Request, env: { XAI_API_KEY: string }) {
const client = new Client({ apiKey: env.XAI_API_KEY });
const chat = client.chat.create({ model: "grok-4" });
chat.append(user("Hello from the edge."));
const res = await chat.sample();
return new Response(res.content);
},
};The API surface is identical to the Node entry point, with two differences:
- No dotenv loading. Pass
apiKeyexplicitly, or bindXAI_API_KEYas a Worker secret.loadEnvFiles/parseEnvFileare not exported. client.files.upload("./path.pdf")needs a filesystem. Pass aUint8Array,Blob, orFileinstead.
Bundlers that honor the worker or browser export conditions pick this build automatically from a plain xai-sdk-js import. The explicit xai-sdk-js/web subpath always works.
Client options
new Client({
apiKey: "...", // else XAI_API_KEY env / .env*
managementApiKey: "...", // else XAI_MANAGEMENT_KEY env / .env*
envDir: process.cwd(), // directory scanned for .env* files
apiHost: "api.x.ai",
managementApiHost: "management-api.x.ai",
timeoutMs: 27 * 60 * 1000,
metadata: { "x-custom": "value" },
useInsecureChannel: false, // local/testing only
});API surface
| Property | Description |
| --- | --- |
| client.chat | Conversations: create, sample / stream, deferred, parse, compact, stored completions |
| client.image | Image generation / editing |
| client.video | Video generate / extend (deferred) |
| client.files | Upload, list, get, delete, content, public URLs |
| client.batch | Batch create / add / list / results |
| client.collections | RAG collections (management API) + document search |
| client.models | List/get language, embedding, image models |
| client.tokenize | Tokenize text |
| client.auth | API key info |
Message helpers: user, system, assistant, developer, toolResult, text, image, file, tool, requiredTool.
Tool helpers: webSearch, xSearch, codeExecution, collectionsSearch, mcp, functionTool.
Cost helper: costUsdFromUsage(usage) / response.costUsd.
Proto types
Generated protobuf types live under the package build output and can be imported if you need raw messages:
import type { GetChatCompletionResponse } from "xai-sdk-js";(Most apps only need the high-level clients.)
Development
bun install
bun run gen # buf generate from proto/
bun run typecheck
bun test
bun run build # tsup → dist/Releasing (npm + GitHub)
The npm package is linked to this repo via package.json repository / homepage / bugs.
Publishing is automated by .github/workflows/release.yml:
- Add repo secret
NPM_TOKEN(npm access token with publish access)
GitHub → Settings → Secrets and variables → Actions → New repository secret - Bump and ship a tag:
# bump package.json + src/version.ts, commit, tag vX.Y.Z, push
./scripts/release.sh 0.1.2
# or tag the version already in package.json
./scripts/release.shPushing tag v* runs CI build/tests, npm publish, and creates a GitHub Release with notes.
Creating a GitHub Release from an existing v* tag also triggers publish (skips npm if that version already exists).
Let CI own the publish. Running npm publish locally for a version CI is already building makes the CI job fail with E403 cannot publish over the previously published versions — npm accepted the first upload and rejects the second.
License
Apache-2.0 — same as the Python SDK and xAI protos where applicable.
