@ai-agent-forge/plugin-image-generation
v0.88.1
Published
Image generation capability plugin for Agent Forge (image_generate / image_edit tools)
Maintainers
Readme
@agent-forge/plugin-image-generation
First-party image generation capability plugin for Agent Forge. Registers the
image_generate and image_edit tools (full-control hosts) against the
host-configured provider's OpenAI-compatible images API:
- Credentials (
apiKey+baseUrl) are resolved at runtime through the publicCredentialAPI(api.credentials.resolve({ providerId })); the plugin keeps no credentials of its own. provider/modelare resolved through the capability-slot chain (设计供应商能力槽位与主备切换设计.md§3.2, 优先级从高到低):- tool parameters — explicit
provider/modelon the call; - settings slot —
capabilities.imageGenerationinsettings.json(host-injected;modelmay be omitted: the provider's registered image model is used — a single hit applies directly, multiple hits pick the deterministic first by name and the result notes the source); - env —
AGENT_FORGE_IMAGEGEN_PROVIDER/AGENT_FORGE_IMAGEGEN_MODEL(legacy compatibility, below the settings slot); - primary-provider inference — when the session's current provider has
an image model registered in
models.json("output": ["image"]), it is used automatically (follows mid-session model switches); - single-candidate inference — when exactly one provider registers
image models in
models.json, it is used with zero settings; - fail-closed —
config_missingwith a configuration example; multiple candidates are listed instead of guessed.
- tool parameters — explicit
- Generated and edited image files are written under the workspace
generated-images/directory, named and typed by their actual format (magic-byte sniff: png/jpg/webp/gif; unknown falls back to png); the tool result returns their workspace-relative paths only (see "Images never ride the conversation" below) — the model reads a returned file to view it. - Cancellation cascades through
api.sessionAbortSignal()and the invocation signal; requests time out after 180s and error messages never echo the API key or request headers.
Configuring an image provider
settings.json (host capability slot):
{
"capabilities": {
"imageGeneration": { "provider": "myimg", "model": "gpt-image-1" }
}
}models.json registers the provider and its image model (the output field
marks image-capable models; cost requires all four rate keys):
{
"providers": {
"myimg": {
"baseUrl": "https://img.example.com/v1",
"api": "openai-completions",
"models": [
{
"id": "gpt-image-1",
"reasoning": false,
"input": ["text"],
"output": ["image"],
"cost": { "input": 0, "output": 0, "cacheRead": 0, "cacheWrite": 0 },
"contextWindow": 8192,
"maxTokens": 4096
}
]
}
}
}Save the provider key with /login (it is stored in auth.json); models.json
rejects an apiKey field at load. The plugin resolves the key and baseUrl
through the host's credential API: the stored key lives in auth.json, and
models.json is not a key source — --api-key/runtime-injected and
provider-registered keys are resolved too.
With only the models.json registration (no capabilities settings), the
plugin auto-uses the single registered image provider (inference layers 4/5).
Registering an image model in models.json implies that provider really has
an /images/generations endpoint; builtin chat catalog models never take part
in the inference. The slot travels to the plugin through the host's
assembly-time configOverride injection (api.config.imageGeneration +
api.config.imageModelCatalog snapshot) — the plugin never reads host
settings files itself.
Proxy support
Every outbound request (API call and provider-CDN image download) honors the
standard proxy environment variables — HTTPS_PROXY, HTTP_PROXY, NO_PROXY
(both cases work, same convention as curl/git/npm) — via undici's
EnvHttpProxyAgent, injected per request. Node's fetch ignores these
variables by default, so if your machine reaches the internet only through a
VPN/proxy client, export its local HTTP port before running the host, e.g.:
export HTTPS_PROXY=http://127.0.0.1:7897
export NO_PROXY=api.shenwenai.com # optional: keep the gateway directWindows system-proxy (registry) settings are not read; export the port manually. The global dispatcher is never replaced — only this plugin's requests are affected.
Notes:
HTTP(S)_PROXYmust point at an http proxy (most VPN clients expose a mixed http/socks port; use that port with anhttp://URL).socks5://URLs are rejected with an explicit error — undici speaks http proxies only.ALL_PROXYis not read (undici limitation); setHTTP_PROXY/HTTPS_PROXY.- The host itself already honors these variables for model traffic through
its own global dispatcher (http-dispatcher's
EnvHttpProxyAgent); this plugin's wrapper exists so image traffic also works in hosts that do not install that global dispatcher (embedded hosts), and to reject socks URLs with an actionable message. - Enterprise proxies that re-sign TLS require
NODE_EXTRA_CA_CERTSwith the proxy's root certificate.
Images never ride the conversation (read on demand)
Image payloads are never inlined into the conversation — they would be
re-sent to the model on every later turn. The tool result carries conclusions
only: {files, model, usage} plus a note telling the model to read a file.
When the model actually needs to look at what it produced (e.g. to revise it),
it reads the file — the host read tool serves images as attachments and
auto-resizes them, which is both cheaper and more precise than carrying raw
payloads in history. To let your conversation model see images at all, declare
"input": ["text", "image"] for it in the provider's models.json entry
(the host drops image parts for text-only models with a readable note).
