@tokz/openai
v0.4.1
Published
Intent-aware Tokz compression for OpenAI Chat Completions, Responses, and Agents SDK lifecycles.
Downloads
3,438
Maintainers
Readme
@tokz/openai
Drop-in tokz compression for the official openai SDK. It covers Chat Completions and Responses function outputs, derives bounded task intent from the current request and tool call, and exports OpenAI Agents lifecycle hooks.
Install
npm install @tokz/openai @tokz/sdk openaiUsage
import OpenAI from "openai";
import { withTokz } from "@tokz/openai";
const openai = withTokz(new OpenAI({ apiKey: process.env.OPENAI_API_KEY! }), {
apiKey: process.env.TOKZ_API_KEY!, // required
});
const bigKubectlJson = JSON.stringify({
items: Array.from({ length: 200 }, (_, i) => ({
metadata: { name: `web-${i}` },
status: { phase: i % 37 === 0 ? "CrashLoopBackOff" : "Running" },
})),
});
// Everything works identically — streaming, tools, tool_choice, response_format…
const completion = await openai.chat.completions.create({
model: "gpt-4o",
messages: [
{ role: "user", content: "How many pods are running?" },
{ role: "tool", tool_call_id: "call_1", content: bigKubectlJson }, // ← compressed
],
});Before / after
A 4,000-token kubectl get pods -o json tool result, at the default targetRatio: 0.45:
| | tokens | |---|---| | before | ~4,000 | | after | ~1,800 |
Structural responses are byte offsets into your original text. Prose responses contain extractive source runs with a provenance map; neither path paraphrases content.
Options
import OpenAI from "openai";
import { withTokz } from "@tokz/openai";
const openai = withTokz(new OpenAI(), {
apiKey: process.env.TOKZ_API_KEY!, // required
baseUrl: "https://api.tokz.dev", // default
targetRatio: 0.45, // fraction of bytes kept
autoCompress: ["tool"], // which message roles to compress
firewall: true, // protect system/control selection
onSavings: (originalTokens, compressedTokens, originalCost, compressedCost) => {
console.log(`${originalTokens} -> ${compressedTokens} tokens`);
},
});Firewall mode sends the full text context in one segmented Tokz request: system and developer text is protected, user text is control, and tool outputs are data. Only data is compressible. The wrapper falls back to the original OpenAI payload if the firewall request fails.
onSavings token counts are estimates (~4 bytes/token) and cost is derived from a built-in input-price table keyed on the request model. Directional, not billing-grade. Compression results are also cached per client by default — see cache/cacheBytes — so a resent transcript doesn't pay to re-compress tool results it already compressed.
Docs
Full documentation: https://tokz.dev/docs
