elektric-ai
v0.7.3
Published
Official TypeScript SDK for the Elektric inference API
Maintainers
Readme
Elektric TypeScript SDK
Canonical documentation: Quickstart · AI coding agents · llms.txt
One provider-neutral TypeScript interface for chat, media, tools, Web, jobs, and realtime.
Install
npm install elektric-aiThis is the official elektric-ai package on npm. Do not substitute a similarly named package.
Configure and make a request
import { Elektric } from "elektric-ai";
const elektric = new Elektric({ apiKey: process.env.ELEKTRIC_API_KEY! });
const result = await elektric.chat({
userId: "customer-456",
conversationId: "support-123",
message: "What is our refund policy?",
});
console.log(result.message, result.requestId);The production URL, https://elektric.ai, is built in. baseUrl is available for development and testing. elektric-auto is implicit in this simple method. It does not send provider, retrieval, or Context controls, and the SDK never automatically retries an AI execution.
Auto and Direct models
Omit model or use elektric-auto for automatic routing. The SDK preserves its existing explicit Auto payload when the field is omitted.
const auto = await elektric.chat({
messages: [{ role: "user", content: "Explain this design." }],
});
const direct = await elektric.chat({
model: "<public-model-id>",
messages: [{ role: "user", content: "Explain this design." }],
});
console.log(direct.model); // The public identity returned by Elektric.<public-model-id> is a placeholder, not an available model. Named Direct models are live. Direct requests use your selected model with a 1% Elektric fee; Auto selects the model with a 5% fee. The Playground remains Auto-only. Discover available IDs with await elektric.models.list() (GET /v1/models) instead of maintaining a static list.
Named models execute the selected model. Auto and Direct use the same Elektric API key and base URL. The SDK forwards explicit model strings unchanged, including legacy electric-auto; the server validates availability and compatibility. Unknown names surface the server's existing error code and message.
Both the short message form and chat.stream() accept model. Existing Context IDs and tools serialize the same way with Direct. Response model comes from the server, never from the request. It remains non-enumerable on normalized TypeScript results for compatibility. Legacy responses without model metadata expose undefined; stream events expose an optional model when their server chunk contains it.
Persistent conversations
const result = await elektric.chat({
conversationId: "support-123",
userId: "customer-456",
message: "What about enterprise?",
});Both IDs are required by the native beta chat method. Conversation, Memory, and Knowledge are independently configurable in Project Tools and off by default. With Conversation enabled, the same userId and conversationId continues a non-streaming thread. With Memory enabled, selected user information can carry across threads. Ready Knowledge is retrieved when enabled and relevant.
const conversation = await elektric.conversations.get({ conversationId: "support-123" });
const recent = await elektric.conversations.list({ userId: "customer-456", limit: 20 });
await elektric.conversations.delete({ conversationId: "support-123" });Conversation lists are newest-updated-first and cursor-paginated. GET returns up to 100 chronological messages; pass messageCursor to continue. Deleting a Conversation does not delete Memory, Knowledge, or other Conversations. A later chat may deterministically create a new thread with the same external ID.
Memory is durable user-specific information across Conversations. When enabled, it can update automatically after successful non-streaming chat. Inspect or correct it with elektric.memory.list/get/update/delete; these methods are scoped to the API-key project and require userId. Deleting a card marks it inactive; its summary and provenance remain stored. Manual create is not part of the current public card contract.
const memories = await elektric.memory.list({ userId: "customer-456" });
const card = await elektric.memory.get({ userId: "customer-456", memoryId: memories.data[0].id });
await elektric.memory.update({ userId: "customer-456", memoryId: card.id, summary: "User prefers concise prose." });
await elektric.memory.delete({ userId: "customer-456", memoryId: card.id });Deleting Memory deactivates that card immediately but does not erase its source Conversation or prevent History from recalling legitimate past events. Conversation is the current thread, History finds details from older threads, and Knowledge contains project/company reference material.
OpenAI-compatible requests
The existing wrapper also accepts model: "<public-model-id>". It forwards the exact string and returns the server completion unchanged.
const raw = await elektric.chat.completions.create({
model: "elektric-auto",
conversation_id: "support-123",
messages: [{ role: "user", content: "What about enterprise?" }],
});The convenience result exposes normalized text, id, usage, and requestId, plus the original completion at result.raw. Direct HTTP and OpenAI-compatible clients use the same /v1/chat/completions backend.
Streaming
for await (const event of elektric.chat.stream({ userId: "customer-456", conversationId: "support-123", message: "Explain our policy." })) {
if (event.type === "text_delta") process.stdout.write(event.text);
}Pass signal to cancel locally. Streaming accepts conversation and user IDs, but the current backend does not load or persist same-thread Conversation context or update Memory for streaming requests. Enabled Memory reads, historical retrieval and Knowledge may still run. Cancellation does not imply that upstream work or billing was stopped. The deprecated chatStream() alias remains for compatibility.
Files and multimodal input
const file = await elektric.files.upload({ file: pdfBytes, mediaType: "application/pdf" });
const answer = await elektric.chat({ messages: [{ role: "user", content: [
{ type: "text", text: "Summarize this." },
{ type: "file", fileId: file.id },
] }] });Use files for uploads and assets for normalized generated media metadata/downloads. Assets expose opaque Elektric IDs, never storage or upstream IDs.
Video generation
SDK 0.5.0 adds public asynchronous Video helpers:
const video = await elektric.videos.create({
model: "grok-imagine-video-1.5",
prompt: "An orange paper airplane crosses a pale blue studio",
seconds: 5,
resolution: "720p",
aspectRatio: "16:9",
});
const completed = await elektric.videos.waitForCompletion(video.id);
const bytes = await elektric.videos.downloadContent(completed.id);Use grok-imagine-video-1.5 for 480p or 720p and grok-imagine-video for 480p. Both are text-to-video only and accept 16:9 or 9:16. Video is Direct-only at provider cost +1%, with exact-model execution and no fallback. The service polls jobs independently of your client. Provider media is temporary and Elektric does not retain generated video by default.
Transcription
SDK 0.6.0 adds public exact-product transcription:
const transcript = await elektric.audio.transcriptions.create({
model: "gpt-transcribe",
file: audioBytes,
mediaType: "audio/wav",
filename: "meeting.wav",
});
console.log(transcript.text);Available products are gpt-transcribe, gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize, whisper-1, gemini-3.5-transcribe, and xai-speech-to-text. Transcription is Direct-only at provider cost +1%, requires an exact product ID, and never falls back. Timestamp and diarization support differs by product; inspect authenticated GET /v1/models before requesting either feature.
Tools and Web
const weather = elektric.tool<{ city: string }>({
name: "weather", description: "Read the current weather.",
inputSchema: { type: "object", properties: { city: { type: "string" } }, required: ["city"] },
});
const response = await elektric.chat({ messages, tools: [weather], web: true });
console.log(response.text, response.sources, response.toolCalls);Your application authorizes and executes tool calls, then returns an ElektricToolResult in the next request. Elektric does not execute customer tools.
Images, audio, video, and jobs
// Images in SDK 0.4.0; discover available models with models.list().
const result = await elektric.images.generate({ model: "gpt-image-2", prompt: "A lighthouse at sunrise" });
const image = result.data[0]; // b64_json, or a temporary URL where supported
const edited = await elektric.images.edit({ model: "gpt-image-2", prompt: "Make it blue", image: pngBlob });Images cost provider cost +1%, with no Auto or fallback, and require an exact model. Edits accept one PNG/JPEG Blob or data URI; outputs are not automatically uploaded to Elektric storage. n and size are model-specific; unsupported settings fail before execution.
Use audio.transcribe, audio.speech, images.generate, images.edit, video.analyze, and video.generate. Async operations return an ElektricJob; job.wait() is sugar for elektric.jobs.wait(job.id).
Embeddings and realtime
Embedding requests require an exact model: elektric.embeddings.create({ model: "text-embedding-3-small", input: "Elektric" }). Profiles are discoverable at elektric.embeddingProfiles and remain available at elektric.embeddings.profiles for compatibility. elektric.realtime.connect() returns a normalized session with sendText, sendAudio, commitAudio, interrupt, on, and close.
Errors and runtimes
import { ElektricError } from "elektric-ai";
try { await elektric.chat({ messages }); }
catch (error) { if (error instanceof ElektricError) console.error(error.code, error.requestId); }Errors expose safe Elektric code, status, requestId, and optional details; upstream errors are never returned. Supported environments are ESM, fetch-compatible Node.js 18+, Cloudflare Workers, and other server/edge runtimes with standard Web APIs. Realtime additionally requires WebSocket support. Use ELEKTRIC_API_KEY in server-side code. Never ship an Elektric secret key in a public browser bundle; a browser app should call your backend, which calls Elektric.
The native SDK is recommended for the complete Elektric platform. OpenAI compatibility is available for rapidly migrating existing OpenAI-compatible code. See Elektric documentation.
Observe exporter (local preview)
The unpublished elektric-ai/observe subpath contains the Node.js 18+ Observe exporter foundation. It accepts only an explicitly supplied dedicated Observe token; it never reads environment variables or reuses the inference API key.
import { observe } from "elektric-ai/observe";
observe({ token: process.env.ELEKTRIC_OBSERVE_TOKEN });One initialization automatically hooks certified OpenAI and Anthropic public methods. Provider clients created before or after initialization keep calling their providers directly. The production endpoint is built in. A Benchmark-issued token safely selects benchmark content capture; legacy and metadata-scoped tokens remain metadata-only.
observe({
token: process.env.ELEKTRIC_OBSERVE_TOKEN,
endpoint: "http://localhost:8787", // optional local-development override
});
Automatic instrumentation is version-gated to OpenAI >=7.4.0 <8, Anthropic
>=0.126.0 <0.127.0, and Google GenAI exactly 2.23.0. Google clients created before
or after observe({ token }) are captured without wrappers or import changes. Google support
uses two version-certified internal generation methods on Models.prototype; these are private
SDK implementation details, not a supported Google extension contract. Unknown versions, missing
metadata (including some bundled layouts), or missing methods disable automatic Google hooks and
report a diagnostic while leaving provider calls untouched.
Google observations are per internal generation request. Automatic function-calling loops can produce multiple observations, one per model request; validation failures before a generation request are not observed. HTTP retries inside a request do not create additional observations. Explicit wrappers observe the whole public call and suppress automatic hooks across async work and stream iteration. Non-generation APIs are outside this coverage.
The OpenAI resource hook cannot safely access its owning client's public baseURL, so automatic
calls use the conservative openai_compatible provider identity. Explicit wrappers remain available
for advanced use. No global fetch, HTTP, socket, or raw network interception is used. Google version
verification reads only its resolved package metadata; no directory or source scanning is performed.
Troubleshooting missing traffic
Keep the normal setup to observe({ token }). If expected requests are missing, retain its returned
observer and inspect observer.instrumentationDiagnostics(). Each entry separates hook status,
sdkDetected, versionSupport, and traffic (detected or not_detected). Traffic means this
observer has queued a validated observation, not that the server received it. A missing observation
does not imply the provider call failed. These diagnostics stay local and contain no prompts,
responses, credentials, client configuration, or request objects. Existing exporter counters describe
delivery separately. OpenAI-compatible traffic is grouped under OpenAI SDK coverage, not vendor identity.
If Google TypeScript traffic is missing, use this advanced fallback with the observer returned by your existing initialization:
const google = observer.instrument(existingGoogleClient);Use google.models.generateContent(...) and google.models.generateContentStream(...) for those
requests. Wrap each client, whether created before or after observe(). Repeated wrapping with the
same observer/configuration returns the same wrapper. The underlying client, credentials, routing,
request parameters, and returned provider values are preserved. No additional Elektric package is needed.
The explicit Google adapter's certified range remains >=2.23.0 <2.24.0; the internal automatic hook
is restricted to exactly 2.23.0 until other versions are tested. instrumentationDiagnostics()
reports installed, missing, unsupported_version, or error separately from local traffic.
Observe v1 rejects browser and non-Node initialization. Export happens asynchronously on a one-second cadence, at 20 events or near 240 KiB, whichever comes first. The in-memory queue holds at most 1,000 events and drops the newest event when full so observation cannot block production inference. Delivery uses three bounded attempts with request timeouts. Runtime export and serialization failures are contained and reflected only in content-free diagnostics.
No global shutdown handlers are installed. Call flush() from your application’s existing graceful-shutdown path where the runtime provides one. Serverless runtimes may terminate before best-effort buffered delivery completes. Custom endpoints require HTTPS, except loopback HTTP for local development.
The local OpenAI adapter supports OpenAI SDK >=7.4.0 <8 on the runtime supported by that SDK (currently Node.js 22+). It is explicit and side-band: requests continue directly to OpenAI, and the adapter never reads the OpenAI API key.
import OpenAI from "openai";
import { observe, observeOpenAI } from "elektric-ai/observe";
const observer = observe({ token: process.env.ELEKTRIC_OBSERVE_TOKEN });
const openai = observeOpenAI(new OpenAI({ apiKey: process.env.OPENAI_API_KEY }), observer);
const response = await openai.responses.create({ model: "gpt-5.6", input: "Hello" });Prompt content is off for legacy and metadata-scoped tokens. A Benchmark-issued token selects
benchmark_content without repeating that server-authorized setting in application code; only the
last unambiguous user text is then eligible for content.user_prompt. An explicit captureProfile
remains available for advanced use, while the server independently enforces the connection profile.
The certified surface is responses.create() and chat.completions.create(), each in non-streaming
and stream: true async-iteration forms. OpenAI APIPromise helpers, responses.stream(), Chat
helper runners, and all other OpenAI resources are outside this adapter surface.
OpenRouter is locally certified through the same OpenAI SDK adapter for Chat Completions and Responses, including non-streaming and stream: true calls. Use the explicit adapter to retain exact OpenRouter identity; Observe never reads the API key, default headers, or request headers.
import OpenAI from "openai";
import { observe, observeOpenAI } from "elektric-ai/observe";
const observer = observe({ token: process.env.ELEKTRIC_OBSERVE_TOKEN });
const openrouter = observeOpenAI(
new OpenAI({
baseURL: "https://openrouter.ai/api/v1",
apiKey: process.env.OPENROUTER_API_KEY,
defaultHeaders: {
"HTTP-Referer": "https://your-app.example",
"X-OpenRouter-Title": "Your app",
},
}),
observer,
{ provider: "openrouter", captureProfile: "benchmark_content" },
);
const completion = await openrouter.chat.completions.create({
model: "anthropic/claude-sonnet-4.5",
messages: [{ role: "user", content: "Hello" }],
});Other endpoints using the certified OpenAI SDK request, response, and stream shapes may opt into { provider: "openai_compatible" }. This is experimental shape compatibility, not certification of a provider. Custom endpoints must set this option so their traffic is not labeled as OpenAI. Raw fetch, provider-specific SDKs, altered streaming protocols, OpenRouter's own SDK, and arbitrary APIs that merely resemble OpenAI are unsupported. OpenRouter routing metadata, cost, request IDs, provider attribution, and custom headers are deliberately ignored.
The local Anthropic adapter supports @anthropic-ai/sdk >=0.126.0 <0.127.0 on Node.js 20 LTS or later. It wraps only messages.create() in non-streaming and stream: true async-iteration forms. Requests remain direct to Anthropic, and the adapter never reads the Anthropic API key.
import Anthropic from "@anthropic-ai/sdk";
import { observe, observeAnthropic } from "elektric-ai/observe";
const observer = observe({ token: process.env.ELEKTRIC_OBSERVE_TOKEN });
const anthropic = observeAnthropic(new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY }), observer);
const message = await anthropic.messages.create({
model: "claude-sonnet-5",
max_tokens: 256,
messages: [{ role: "user", content: "Hello" }],
});The same metadata default and explicit benchmark-content policy applies. The wrapper excludes system text, prior turns, assistant text, tool definitions and payloads, image/document data, provider responses, and streamed output. messages.stream(), parse(), countTokens(), batches, beta resources, and other Anthropic APIs are outside the certified surface.
The local Google adapter supports @google/genai >=2.23.0 <2.24.0 on Node.js 20 or later. The legacy @google/generative-ai package is not supported. The adapter wraps only models.generateContent() and models.generateContentStream().
import { GoogleGenAI } from "@google/genai";
import { observe, observeGoogle } from "elektric-ai/observe";
const observer = observe({ token: process.env.ELEKTRIC_OBSERVE_TOKEN });
const google = observeGoogle(new GoogleGenAI({ apiKey: process.env.GOOGLE_API_KEY }), observer);
const response = await google.models.generateContent({
model: "gemini-3.5-flash",
contents: "Hello",
});Requests remain direct to Google, and the adapter never reads the Google API key. Metadata capture is the default. Explicit benchmark-content capture can include only the final direct user text; system instructions, prior turns, model output, tools, function payloads, media/file data, raw errors, and streamed output stay excluded. Images, video, embeddings, files, caches, batches, live APIs, chats, and every other Google surface are unsupported.
Knowledge management
Knowledge is project-level reference material used automatically when relevant. Use knowledge.add, list, get, and delete; poll until status === "ready". Upload accepts Web bytes and Node Buffer with a filename/MIME type. Retry is deferred: delete and re-upload failed sources.
Streaming Web sources
Pass web: true to elektric.chat.stream() or chatStream(). Streams emit normalized source events in addition to existing text/content and finish events. Sources may arrive before, during, or after text and are de-duplicated by URL; the underlying SSE ends after the finish chunk with [DONE].
Context storage and data controls
Conversation, Memory, and Knowledge are independently configurable in Project Tools and off by default for new Projects. Identifiers alone do not enable them. Disabling a tool preserves existing data. Memory deletion deactivates a card rather than erasing its summary and provenance. Streaming skips same-thread Conversation loading/persistence and automatic Memory updates; enabled Memory reads, historical retrieval and Knowledge can still run.
See Security and Data, Context and the Privacy Policy before sending personal or confidential information.
Auto selects a model through Elektric and applies a 5% fee. Direct executes the exact selected model with a 1% fee and no cross-model fallback. Available capabilities vary by model.
Embeddings (0.3.0)
Use text-embedding-3-small (1536 dimensions) or text-embedding-3-large (3072 dimensions). Discover currently available products with await elektric.models.list().
const vectors = await elektric.embeddings.create({
model: "text-embedding-3-small",
input: ["first", "second", "third"],
dimensions: 256,
encoding_format: "float",
});Each input returns one vector, in order. Embeddings cost provider usage +1%. An exact model is required; there is no embedding Auto or fallback. Keep one model and dimension configuration per vector index; changing models requires re-embedding your corpus.
