ag-ai
v0.3.4
Published
Provider-neutral AI client and cluster runtime for the rpi-mega cluster.
Readme
ag-ai
Provider-neutral AI client and cluster runtime for the rpi-mega cluster.
Target layout (subpath exports)
| Subpath | Contents |
| --------------- | --------------------------------------------------------------------- |
| ag-ai | Client SDK: generate, generateText, promptDirect, promptImage, models |
| ag-ai/client | Client SDK entry point |
| ag-ai/k8s | Cluster plugin descriptor (ClusterPlugin) |
| ag-ai/runtime | Cluster runtime: generation protocol, store, server, bridge (ESM) |
| ag-ai/image | Neutral image-output adapter (provider per implementation) |
The client compiles to CommonJS. The runtime is hand-authored ESM and shipped as-is, because it runs inside the pod with no build step.
The client-only dependencies (@google/genai, node-cache) are declared as
optionalDependencies; the cluster runtime install omits them because the
runtime imports only Node builtins.
Boundaries
- No prompts. Prompt templates belong to consumers (gec.dev, bolt.garden, …).
- No secrets, secret names, SSM paths, or cluster addresses.
- Provider keys are declared as logical names and resolved by
rpi-mega. - Codex is not used. The executor is OpenCode; the image adapter is separate because OpenCode has no image-generation output.
Gemini Fallback
Gemini is always the primary provider when no endpoint is selected. A client
can opt into one explicit cluster fallback after every eligible Gemini
model/key combination is quota- or capacity-exhausted, or is still in its
cooldown window:
const client = createAIClient({
fallback: {
endpoint: { protocol: 'ag-ai', url: clusterURL, token: clusterToken },
model: 'gpt-5.6-luna',
reasoningEffort: 'xhigh',
},
});The fallback endpoint is never inferred. It is used only after aggregate
Gemini exhaustion, not for validation, authentication, transport, or other
nonquota failures. The Gemini model id and reasoning settings are not copied
to the cluster request; inputs, outputSchema, and other request fields are
preserved while the configured fallback model and reasoning effort are used.
onFallback receives an explicit exhaustion event, and onGenerated reports
the actual selected model and provider (google or cluster). Cluster
network operations and retry waits for this fallback are individually capped
at 3000 ms; model generation time is not included in that cap.
Quota (429) and capacity (5xx) failures only put a model/key on a short
cooldown (60 s doubling to 15 min, or the provider's retry delay). Models that
return 404 are hidden for an hour. Without a fallback, a request whose whole
pool is cooling down waits for the soonest model to recover and retries, at
most MAX_COOLDOWN_WAITS times and never past the request timeout; with a
fallback it falls back immediately instead.
Tooling
Formatting and linting use e7-oxfmt / e7-oxlint from
eslint-config-e7npm. Dependency updates run through Renovate via
github>andreigec/renovate.
