@journal.one/llm-economy-proxy
v0.3.0
Published
LLM providers charge less for non-interactive work, but the discounted paths differ by provider. OpenAI offers a synchronous **Flex** processing tier (~50% off standard) that you select with one request parameter. Anthropic's **Message Batches** API is al
Readme
LLM Economy Proxy
LLM providers charge less for non-interactive work, but the discounted paths differ by provider. OpenAI offers a synchronous Flex processing tier (~50% off standard) that you select with one request parameter. Anthropic's Message Batches API is also ~50% off but harder to use: you upload jobs, poll for completion, download results, and match responses back to requests.
This library lets you keep using familiar non-streaming OpenAI and Anthropic API shapes while it routes eligible requests through each provider's cheapest non-interactive path behind the scenes — OpenAI Flex and Anthropic Message Batches.
Use it for non-interactive work where lower cost is more important than peak latency: summarization jobs, classification, evaluations, data enrichment, code review sweeps, and other background LLM workflows.
Do not use this for interactive chat, streaming, or user flows that need the fastest possible responses.
What It Provides
- OpenAI-shaped calls:
openai.responses.create(...)andopenai.chat.completions.create(...). - Anthropic-shaped calls:
anthropic.messages.create(...). - Discounted execution for latency-tolerant requests: OpenAI Flex (synchronous) and Anthropic Message Batches (asynchronous).
- Optional promotion to the standard tier on timeout.
- TypeScript and JavaScript support.
- OpenTelemetry hooks for customer-controlled telemetry.
Latency Reality
OpenAI requests run through the synchronous Flex tier (service_tier: "flex"), retried on capacity pressure, so they return in seconds. Anthropic requests run through Message Batches, which release results together and trade minutes of latency for the discount. Local live tests, OpenAI on June 25, 2026 and Anthropic on June 24, 2026, using Journal models:
| Test | OpenAI gpt-5.5 (Flex) | Anthropic claude-opus-4-7 (Batches) |
|---|---:|---:|
| 100 requests submitted together | 100/100 in ~4.5s (mean 2.2s, p95 2.9s) | 100/100 in ~2m |
| One request at a time | 100/100, median 1.2s, worst 24.4s (~3.3m total) | 27 succeeded in 30m cap, median 60s, worst 90s |
Use promoteOnTimeout to fall back to the standard interactive tier instead of failing if the discounted path cannot complete within timeout: for OpenAI this drops Flex for standard service; for Anthropic it sends a direct call.
Full benchmark notes: docs/benchmarks.md.
Install
pnpm add @journal.one/llm-economy-proxyHello World
import { createOpenAIProxy } from "@journal.one/llm-economy-proxy";
const openai = createOpenAIProxy({
apiKey: process.env.OPENAI_API_KEY!,
mode: "economy",
// Retry the Flex tier for up to 30 minutes if capacity is unavailable.
// If the timeout is reached, fall back to the standard tier for this request.
timeout: 30 * 60_000,
promoteOnTimeout: true,
});
const response = await openai.responses.create({
model: "gpt-5.5",
input: "Say hello in one short sentence.",
max_output_tokens: 16,
});
console.log(response);Anthropic example:
import { createAnthropicProxy } from "@journal.one/llm-economy-proxy";
const anthropic = createAnthropicProxy({
apiKey: process.env.ANTHROPIC_API_KEY!,
mode: "economy",
});
const message = await anthropic.messages.create({
model: "claude-opus-4-7",
max_tokens: 16,
messages: [{ role: "user", content: "Say hello in one short sentence." }],
});
console.log(message);Migrating One Call
Keep your existing OpenAI or Anthropic client for interactive paths. Move only latency-tolerant calls to the economy proxy.
Before:
export async function classifyDocument(text: string) {
return openai.responses.create({
model: "gpt-5.5",
input: `Classify this document:\n\n${text}`,
max_output_tokens: 64,
});
}After:
export async function classifyDocument(text: string) {
return economyOpenai.responses.create(
{
model: "gpt-5.5",
input: `Classify this document:\n\n${text}`,
max_output_tokens: 64,
},
{
timeout: 30 * 60_000,
promoteOnTimeout: true,
}
);
}With promoteOnTimeout: true, the library keeps trying the discounted path until timeout — for OpenAI, retrying the Flex tier while capacity is unavailable. If it still has not completed, it sends the same request through the standard interactive tier and returns that response.
Try It Locally
pnpm install
pnpm buildCreate a local env file:
cp .env.example .env
# Add OPENAI_API_KEY and/or ANTHROPIC_API_KEY.Run checks:
pnpm typecheck
pnpm testRun live benchmarks:
pnpm live:benchmark
pnpm live:benchmark:serialBenchmark results are written to ignored JSON files under live-results/.
More Docs
- Requirements: docs/requirements.md
- Design and advanced integration details: docs/design.md
- Benchmarks: docs/benchmarks.md
- Publishing: packaging/npm/README.md
Package versions are tracked in package.json and published releases should be tagged as vX.Y.Z. Publish with:
./packaging/npm/publish.sh