vyaya
v0.1.0
Published
Opt-in server-side retry budgets and cancellation for the Vyaya proxy.
Readme
vyaya
Opt-in server-side request controls for applications using the Vyaya proxy. Includes bounded HTTP retries, caller cancellation, and an operation deadline that stays active while a response stream is consumed. Prompts, models, token limits, and successful response streams are unchanged.
Implemented locally, not published to a registry. Build with
pnpm --filter vyaya build; use a workspace dependency or make a local
package archive with pnpm --filter vyaya pack.
import OpenAI from "openai";
import { createVyayaFetch } from "vyaya";
const baseURL = "http://localhost:8787/v1";
const client = new OpenAI({
baseURL,
apiKey: "proxy-managed",
maxRetries: 0, // Required: Vyaya owns the retry budget.
fetch: createVyayaFetch({
baseURL,
apiKey: workspaceKey, // From your application's server-side configuration.
featureTag: "support", // Must be allowed by your workspace.
policy: {
id: "support-retry-v1",
mode: "enforce",
maxAttempts: 2, // First attempt plus at most one retry.
timeoutMs: 15_000, // Includes backoff and streaming consumption.
},
}),
});
const controller = new AbortController();
// Connect controller.abort() to your application's disconnect event.
const stream = await client.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: "Hello" }],
stream: true,
stream_options: { include_usage: true },
}, { signal: controller.signal });
for await (const chunk of stream) {
// Deliver chunk; abort if this response is no longer needed.
}Contract
observeforwards one attempt with attribution and caller cancellation. No added retry or deadline; it does not simulate another retry engine.enforcepermits 1–5 attempts and a 1–300,000 ms deadline per fetch invocation. Disable provider SDK retries and avoid additional application retry loops.- Only HTTP 429, 502, 503, and 504 may be retried. These failures can still have provider costs; opt in only when retries are appropriate for the workload.
- Network errors, permanent HTTP errors, and stream failures are not retried. Transport exceptions and final provider responses are preserved.
- Replaying requires a string body in RequestInit. Request objects or streaming bodies are sent once without cloning, buffering, or reading their contents.
- Backoff doubles from 500 ms. Retry-After can extend the delay; if it exceeds the remaining deadline, the response is returned instead of retrying early.
- Attempts receive distinct IDs correlated to the first attempt. Policy labels and feature tags are captured by the proxy. No separate SDK telemetry queue.
- Credentials go only to the configured proxy's supported endpoints. Redirects are rejected. HTTPS is required outside loopback development. Server use only.
- Cancellation requires standards-compliant fetch; it does not guarantee that the provider stopped generation or billing. Never promise saved tokens from aborts.
Measure and roll back
Give every changed policy a new ID. Review matching models, endpoints, features,
and price versions in /measurements or GET /api/measurements?days=30. Compare
success/error/disconnect counts, latency, and coverage alongside cost estimates.
Use controlled cohorts or a reviewed baseline before attributing an improvement
to a policy. Fewer attempts alone do not prove equivalent quality or dollar savings.
Switch to observe or remove the adapter to roll back. Keep the provider retry
setting explicit. Central policy rollout, application outcome ingestion, quality
evaluation, and causal savings calculations are not included yet.
