hero-run-ai
v0.2.4
Published
Use Hero Run's multi-gateway AI router (375+ models across 10 gateways, paid in $HERO) from any app: an OpenAI-compatible client, a Vercel AI SDK provider, and a drop-in React chat widget.
Maintainers
Readme
hero-run-ai
Use Hero Run's multi-gateway AI router from any app. One key, 375+ models across 10 gateways, paid per call in $HERO. Use model "auto" and Hero Run reads your prompt and routes it to a right-sized model (cheap prompt → fast cheap model, hard prompt → frontier reasoner) at one flat price.
Three ways to use it:
- Core client: zero dependencies, works anywhere
fetchexists. - Vercel AI SDK provider: plugs into
generateText/streamText. - React widget: a drop-in
<HeroRunChat/>chat box.
All three call the same OpenAI-compatible endpoint, so you can also just point the plain openai SDK at https://herorunai.com/v1.
Install
npm install hero-run-aiMint a key at herorunai.com/keys and set HERO_RUN_KEY.
1. Core client
import { createHeroRun } from "hero-run-ai";
const hero = createHeroRun({ apiKey: process.env.HERO_RUN_KEY! });
// one-shot
const { text, hero: meta } = await hero.chat([{ role: "user", content: "Explain CRDTs in one line" }]);
console.log(text, "· routed", meta?.routed_tier, "→", meta?.resolved_model);
// streaming (yields text deltas; the return value is the final result)
const gen = hero.stream([{ role: "user", content: "Count to 5" }]);
for (;;) { const { value, done } = await gen.next(); if (done) break; process.stdout.write(value); }
await hero.models(); // string[] of model ids (includes "auto")
await hero.balance(); // { balance, deposited, spent } in $HERO2. Vercel AI SDK
import { generateText, streamText } from "ai";
import { createHeroRunProvider } from "hero-run-ai/ai-sdk";
const hero = createHeroRunProvider({ apiKey: process.env.HERO_RUN_KEY });
const { text } = await generateText({ model: hero("auto"), prompt: "Write a haiku about routing" });Requires the peer dep: npm install @ai-sdk/openai-compatible ai.
3. React widget
import { HeroRunChat } from "hero-run-ai/react";
export default function Page() {
return (
<div style={{ height: 480 }}>
<HeroRunChat apiKey={process.env.NEXT_PUBLIC_HERO_KEY!} model="auto" system="You are concise." />
</div>
);
}The widget streams the answer and shows which model the router picked and the $HERO cost. Requires react >= 18.
Client-side keys are visible to the browser. For public apps, proxy through your backend (or issue scoped keys) instead of shipping a full key.
Images, video and audio
The same client generates media, billed against the same key:
const { image, charged } = await hero.generate({
kind: "image", // "image" | "video" | "audio"
prompt: "a single orange hexagon on white",
model: "flux-2-klein-4b", // or "auto" to let the router pick
});The result carries a URL in image, video or audio depending on kind, plus charged in
$HERO. Media costs far more per call than text does, so check balance() before generating in a
loop rather than after.
Any OpenAI SDK
No wrapper needed. Hero Run is OpenAI-compatible:
from openai import OpenAI
client = OpenAI(base_url="https://herorunai.com/v1", api_key="hr_live_...")
client.chat.completions.create(model="auto", messages=[{"role": "user", "content": "hi"}])Works with LangChain, LlamaIndex, CrewAI, and anything that accepts an OpenAI base URL.
Tool calling (agents)
chat() accepts OpenAI-style tools, so you can build an agent on any model in the catalog.
When the model wants a tool, text may be empty and toolCalls is populated: run them, push the
results back as tool messages, and call again until finishReason is no longer "tool_calls".
const tools = [{
type: "function",
function: {
name: "read_file",
description: "Read a UTF-8 text file.",
parameters: { type: "object", properties: { path: { type: "string" } }, required: ["path"] },
},
}];
const messages = [{ role: "user", content: "What does package.json declare as the entry point?" }];
for (;;) {
const r = await hero.chat(messages, { model: "auto", tools });
messages.push(r.message!); // append the assistant turn verbatim
if (!r.toolCalls?.length) { console.log(r.text); break; }
for (const c of r.toolCalls) {
const args = JSON.parse(c.function.arguments);
messages.push({ role: "tool", tool_call_id: c.id, content: await runYourTool(c.function.name, args) });
}
}finishReason is worth checking: "length" means the answer was cut off at the token budget
rather than finished, so an agent loop can tell a truncated reply from a complete one instead of
treating half an answer as the whole thing.
License
MIT
