@autonome-research/fleet-router
v0.2.0
Published
Client-side request routing across heterogeneous local LLM inference endpoints: context-window eligibility, capacity-normalized least-loaded selection, health probing. Zero dependencies.
Maintainers
Readme
fleet-router
Client-side request routing across heterogeneous local LLM inference endpoints — the case where your instances differ in context window, output budget, multimodal overhead, and concurrency capacity, and your requests differ enough in size that "send it anywhere" produces 400s and queues.
Extracted from a production batch pipeline (305-session litigation transcription + vision-description corpus) that ran a Qwen3-VL fleet across two 98 GB cards (262k ctx, 16 seats each) and one 32 GB card (49k ctx, 2 seats).
Why not an existing tool?
A mid-2026 survey of the space (LiteLLM Router, vLLM production-stack, SGLang router, Ray Serve, Envoy/Kong/Portkey gateways, llm-d/KServe, GPUStack, Paddler, npm ecosystem) found no tool that combines:
- Token-estimated context eligibility — per request, only endpoints whose window fits (prompt + output reserve + per-endpoint overhead like video-frame tokens) are candidates. Gateways rate-limit on tokens; none place on tokens.
- Capacity-normalized selection — lowest
inflight/capacity, so a 16-seat instance absorbs 8× the traffic of a 2-seat one. LiteLLM's least-busy uses raw counts; everything else is static weights. - Embedded TypeScript, zero infra — a library in your pipeline process, not a Python sidecar (+Redis) or a k8s gateway data path.
- Local-only — no cloud calls, suitable for confidential workloads.
The industry's routing investment is going the other way (KV-cache-aware gateways for homogeneous replica pools), so this niche is likely to stay open.
Usage
import { FleetRouter, waitHealthy } from 'fleet-router';
const router = new FleetRouter([
{ url: 'http://localhost:30010/v1', maxContext: 262_144, capacity: 16, outputReserve: 16_384, overheadTokens: 40_000 },
{ url: 'http://localhost:30011/v1', maxContext: 262_144, capacity: 16, outputReserve: 16_384, overheadTokens: 40_000 },
{ url: 'http://localhost:30009/v1', maxContext: 49_152, capacity: 3, outputReserve: 8_192, overheadTokens: 19_000 },
]);
const result = await router.withEndpoint({ promptChars: transcript.length }, async (lease) => {
const prompt = transcript.length > lease.promptBudgetChars
? transcript.slice(0, lease.promptBudgetChars) + '\n[...truncated...]'
: transcript;
return callOpenAIChat(lease.endpoint.url, { prompt, max_tokens: lease.maxTokens });
});On a connection error (engine mid-restart under k8s), await
waitHealthy(lease.endpoint.url) then retry.
Status
v0.2.0 (unpublished; name and home TBD). The full three-axis surface from
DESIGN.md is implemented and tested: outcome-fed circuit breaker
(failureThreshold/cooldownMs), failover in withEndpoint (exclusion-set
retries), blocking acquire with per-endpoint hardCap and AbortSignal,
capability tags/require, rendezvous-hash affinityKey with spill-over,
and pluggable estimateTokens/score. 18 unit tests.
Deferred by design: latency-EWMA selection (needs collected outcome timing), cost weighting, cross-process state. See DESIGN.md for rationale.
