npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@sinch/functions-runtime

v0.7.0

Published

Node.js SDK and local development runtime for Sinch Functions: voice, messaging and HTTP serverless functions

Readme

@sinch/functions-runtime

The Node.js SDK for Sinch Functions: serverless functions that answer phone calls, reply to messages and serve HTTP endpoints on Sinch infrastructure.

A function is one file, function.ts, that exports handlers. The runtime hosts it, gives every handler a context with configuration, cache, storage, a SQLite database and pre-configured Sinch API clients, and turns voice handlers' return values into the wire body the platform expects.

Install

Functions are created and run with the Sinch CLI, which installs this package for you:

npm install -g @sinch/cli
sinch auth login
sinch functions init      # pick a template
sinch functions dev       # local server with hot reload and a tunnel for callbacks
sinch functions deploy

Requires Node.js 24 or later. See the quickstart for the full walk-through.

A voice function

Voice API v2 is what context.voice is and what a new function should be written against.

import { onCall } from '@sinch/functions-runtime/voice';

export const voiceWebhook = onCall({
  incoming: (call, builder) =>
    builder.answer().say('Thanks for calling.').dialPhone('+15550001111'),
  completed: (call) => {
    console.log('call ended', call.callId);
  },
});

onCall takes handlers keyed by call lifecycle stage — incoming, answered, manage, completed, a webhooks map for named webhook commands, and fallback — and returns the voiceWebhook export. Each handler is called as (call, builder, context, request): the call, a command builder already primed with it, the function context, and the raw request. call is flat — call.from, call.callId, with call.event beside them, and call.menu on a menu — and typed per slot: incoming gets an IncomingCall, manage a MenuCall whose menu is guaranteed once call.event === 'call.menu', completed a CompletedCall. Return the builder (.build() is optional) and the runtime emits the wire body; return nothing and it answers 204. call.call — how the call was read before this flatten — is deprecated, not gone: it still compiles and still returns call itself (call.call.callId === call.callId), so code written against the old nested shape keeps working. It never appears on the wire (not present on JSON.stringify(call) or in a body built from call). Read the field directly off call instead — new code should not use call.call.

The builder's dialPhone, dialSip, dialStream and dialRelay put another party on the call: they bridge the caller to the new leg and end each leg when the other goes. Because the builder knows the call, they fill in what used to be written by hand — the leg names, the bridge, the teardown, and from (the Sinch number for a phone, the caller's number for SIP). They never answer the call: answer() stays explicit. onNoAnswer covers busy, rejected, timed-out and failed at once, and takes a whole flow, so a fallback is another dial:

import { dialPhone, say } from '@sinch/functions-runtime/voice';

incoming: (call, builder) =>
  builder.answer().dialPhone(agent, {
    onNoAnswer: dialPhone(fallback, { onNoAnswer: say('Nobody is free.').hangup() }),
  }),

say, dialPhone and the rest also exist as standalone functions that start a new flow, for the places that take one — a menu option('1', say('Connecting.').dialPhone(sales)), onNoAnswer, onFail. In a menu, option takes keypresses literally (* and # included); match takes a regular expression. A menu of keypresses can also be written as one object — menu('main', { prompt, options: { '1': dialPhone(sales), '2': dialPhone(support) }, onFail }) — whose options keys go through option; a match needs the callback form. bridgeTo(destination, options) is the mechanism under the dial verbs, for a hand-built destination or a named bridge or leg. commands() still starts a free flow for batch bodies and tests; a dial verb there needs { liveLeg }, since no call is known.

Routing calls to the function

An inbound call reaches the function through a Voice v2 service. sinch functions init picks one and writes its id to .env as VOICE_SERVICE_ID; sinch functions deploy then points that service's webhook at the deployed function. Every call.* event for the service arrives at that one URL, so voiceWebhook is the only export the platform needs.

Those events are signed. Each one carries Authorization: service <serviceId>:<signature> and an x-timestamp, signed with the service secret over the raw body, the Content-Type, the timestamp and the path. The runtime verifies that signature under the same WEBHOOK_PROTECTION modes as the v1 callbacks (never, deploy, always) as soon as VOICE_SERVICE_SECRET holds the Base64 secret, and rejects a request that fails with 401. Set VOICE_SERVICE_ID as well and the id in the header has to match it.

Sinch does not hand out a service's secret anywhere yet. Until it does, a function with protection on but no VOICE_SERVICE_SECRET logs one warning per process and serves the webhook — verification switches itself on the day the secret is set, with no code change. List voiceWebhook in auth if you want Basic auth on the endpoint in the meantime, or if you intend to drive it yourself.

Placing calls

context.voice is the v2 client. It dials, bridges legs, attaches a WebSocket media stream to a live call, and patches a call that is already up.

await context.voice.call('+15559876543', {
  from: '+15551234567',
  onAnswer: (c) => c.say('Your appointment is confirmed.').hangup(),
});

// Dial a human and bridge a second leg to your own audio socket
await context.voice.callWithStream('+15559876543', '/media', { from: '+15551234567' });

transferToPhone, transferToSip, transferToStream, transferToRelay and transferToAgent move a live leg: out of the bridge it is on and into a new one with the destination, whoever it was talking to staying behind. They ring for 30 seconds; if nobody answers, or the new party hangs up later, the moved leg is dropped. The moved leg is taken to be caller (origin for an outbound call, when call is passed); name it with liveLeg otherwise:

await context.voice.transferToPhone(callId, '+15550001111');
await context.voice.transferToAgent(callId, 'grok', grokNumber);

transfer is the other operation: it dials a third party into the bridge the call is already on.

v2 authenticates with the project Access Key pair (PROJECT_ID_API_KEY / PROJECT_ID_API_SECRET), not the v1 application key.

Everything a voice function needs imports from @sinch/functions-runtime/voice: onCall, the builder and its standalone factories, voiceRelay, realtimeAgent, connectAgent, the client, and their types. The package root keeps context, config and Voice v1. ./voice/v2 is the same entry under its older name, and onCall, commands, createClient, Client and the core v2 types are also on the root — every path resolves the same declarations, so they mix freely.

Connecting a call to an AI agent

connectAgent bridges an inbound call to a SIP-based AI voice agent — ElevenLabs or xAI Grok — in one call, instead of hand-writing the answer/bridge/dial/teardown commands it takes:

import { AgentProvider } from '@sinch/functions-runtime';
import { onCall, connectAgent } from '@sinch/functions-runtime/voice';

export const voiceWebhook = onCall({
  incoming: (call) => connectAgent(AgentProvider.ElevenLabs, '+15551234567', { call }),
});

The second argument is the number in the SIP URI's user part — for Grok the Direct SIP phone number registered in the xAI console, for ElevenLabs the number attached to the agent's phone integration. Neither provider routes by agent id, and it must be an E.164 number you configured, never a value taken from a caller-controlled webhook field. Passing call (the incoming event) defaults from to the caller's number and adds the X-Sinch-Call-Id / X-Caller-Number / X-Dialed-Number SIP headers the agent sees as per-call context. Stream-based providers — AgentProvider.Sinch, or a future OpenAI/Deepgram provider — are a different call shape and throw; use @sinch/agents for those.

Return the result directly from incoming, as above — callName and the agent-leg teardown on hangup are only honoured on a response to call.incoming, so forwarding just its .commands elsewhere silently drops the teardown and leaves the SIP leg billing until maxDuration.

The builder's dialAgent does the same without that restriction, because it knows which webhook it answers: it names the caller leg and adds the teardown only on call.incoming, and it never answers the call itself. So an agent can be one option of a menu, and a realtimeAgent handle dials the same way. Every agent leg dialled through dialAgent — SIP or a realtimeAgent handle's own stream — is capped by the MAX_CALL_DURATION function variable (seconds), defaulting to 1800 when it is unset; an explicit maxDuration on the dial itself still wins:

import { dialAgent, dialPhone } from '@sinch/functions-runtime/voice';

incoming: (call, builder) =>
  builder.answer().menu('main', (m) =>
    m.prompt('Press 1 for sales, 2 for our assistant.')
      .option('1', dialPhone(sales))
      .option('2', dialAgent('grok', grokNumber))),

Talking to a caller in text

voiceRelay puts a Voice Relay leg on the call: Sinch runs speech-to-text and text-to-speech, and the function only exchanges text. It owns the /relay socket, the handshake, the greeting and the one-time token that proves a connecting socket is a leg this function dialled; the author writes the conversation:

import { onCall, voiceRelay } from '@sinch/functions-runtime/voice';

const relay = voiceRelay({
  greeting: 'Hi, how can I help?',
  onConnected(call) {
    call.onSpeech(async (text) => call.send(await answer(text)));
    call.onDigits((digits) => call.send(`You pressed ${digits}.`));
    call.onEnded(() => console.log('relay closed', call.id));
  },
});

export default {
  setup: relay.setup,
  voiceWebhook: onCall({ incoming: (call, builder) => builder.answer().dialRelay(relay) }),
};

voice, language, interruptions and path (default /relay) are the other options. voice, language and greeting also take (call, context) => value, resolved once per call as realtimeAgent does; greeting may be async, while voice and language go into the dial itself and must return a string. RELAY_URL overrides the socket address for local development. call.hangup() ends the call, and call.transferToPhone(...) and the rest of the transferTo* family move the caller on.

Hooking a call into a realtime voice agent

connectAgent hands the call away — the vendor's own agent runtime takes it from there. realtimeAgent is the opposite: this function stays on the line for the whole call, relaying audio to a vendor's realtime voice API and dispatching its tool calls in-process.

All four adapters are raw ws — no vendor SDK, and no optional peer dependency to install. ElevenLabs is not among them — its wire protocol has no way to take a tool schema (or, without an opt-in dashboard setting, a prompt) at connect time, so it cannot honor realtimeAgent's "switch provider with no code change" promise; it stays a connectAgent-only vendor (see above).

| provider | Transport | Verified | Notes | | --------------- | ------------------------------------------------------------------------------- | -------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | 'deepgram' | Raw ws | Against the shipped realtime-deepgram template | No resampling — Deepgram accepts the platform's own negotiated rate. Every declared tool defaults to defer_until_eot: true, so a handler waits for the caller to actually finish speaking rather than running speculatively. UserStartedSpeaking is gated against an estimated playout window (generation still active, or already-sent audio not yet finished playing) rather than treated as barge-in unconditionally — Deepgram's own docs say it fires for every new utterance, not only a genuine interruption; see adapters/deepgram.ts's file header. | | 'openai' | Raw ws (not @openai/agents) | Against GPT-Live's current docs and a live probe session — see below | GPT-Live (gpt-live-1), Responses delegation, 8kHz audio/pcmu both ways. No barge-in signal. | | 'grok' | Raw ws, shares the (previous) OpenAI Realtime-shaped adapter's implementation | Against xAI's docs describing the endpoint as OpenAI-compatible | 8kHz audio/pcmu both ways. greeting is paraphrased, not verbatim — it is sent as response.create instructions, not a dedicated field. | | 'gemini' | Raw ws (not @google/genai) | Against a live Gemini Live session | Fixed 16kHz in / 24kHz out, resampled through a stateful anti-aliasing resampler. Reconnects transparently (capped retries) on a session-window close, using Live's own resumption handle. Known vendor-side issue: Gemini can go silent for the rest of a call after a barge-in — see below. | | 'voice-relay' | Voice Relay (Sinch STT/TTS), not a vendor realtime session | Against the adapter's own test suite, not yet a live call | Runs any model behind an OpenAI Responses-compatible API — OpenAI, Azure OpenAI, or a gateway such as OpenRouter or LiteLLM — as a text loop on top of voiceRelay. See "Talking to an LLM over Voice Relay" below. |

'openai' targets OpenAI's GPT-Live API (wss://api.openai.com/v1/live/sessions, model gpt-live-1), not the Realtime API 'grok' still shares an implementation with (openai-compatible.ts) — the two vendors no longer share adapter code. GPT-Live delegates tool selection to a separate backend model (delegation: { type: 'responses', responses: { model, tools } } — default gpt-5.6-luna, override via providerOptions.openai.backendModel); this function's own RealtimeTool.handler still runs the call in-process exactly like every other adapter, Responses delegation just gets GPT-Live to ask for it in a shape this adapter can dispatch. GPT-Live is genuinely full-duplex and sends no barge-in event of any kind — this adapter never calls onBargeIn, and the cross-vendor contract test has an explicit, documented allowance for it being the one provider whose normalised hook sequence has no bargeIn step. Turn boundaries are synthesised (GPT-Live's transcript-delta events are explicitly documented as not marking turn boundaries): the agent side flushes on the response.completed Responses event, the caller side flushes after a short silence gap — see adapters/openai.ts's file header for both heuristics.

'openai' and 'gemini' have both been confirmed on live calls; 'deepgram' and 'grok' have not — "verified" above otherwise means checked against the vendor's current documentation or a previously-working template, not a real session. 'openai' was run twice against a real GPT-Live session — once probing the raw protocol, once driving the actual adapter through RealtimeAdapter's own interface (see its file header for the observed sequence, timings, and what those runs left unverified). 'gemini' was probed live against a real Gemini Live session, driving the actual adapter through the same RealtimeAdapter interface; a greeting, a tool call, a barge-in mid-response, and a forced reconnect all worked. Each adapter's own file header (packages/runtime-shared/src/ai/realtime/adapters/*.ts) has the specific list of what to check live. A cross-vendor contract test (packages/runtime-shared/src/ai/realtime/contract.test.ts) drives one agent through all four adapters' own scripted fake vendor sockets and asserts the normalised author-hook sequence (barge-in aside, for 'openai'), the tool-result wire format, and the 20ms/8kHz audio frame size are identical across providers — the one place this repo checks vendor-to-vendor consistency directly, rather than only each adapter's own behavior in isolation.

/stream requires a one-time token connect() mints per call — nothing an author does, but worth knowing when reading logs: a connection that never answers and logs "missing, unknown, expired, or already-consumed session token" reached /stream without going through this agent's own call.incoming plan (a stale worker, a health-check probe, or a caller-controlled URL). The token is checked from the WebSocket upgrade headers immediately when present — the normal case for a STREAM leg this agent dialled itself — rather than waiting on the first frame to arrive; a socket with no token on its upgrade gets 2 seconds to supply one in its connect frame before it is closed, and no more than 32 such token-less sockets may be waiting at once, so the public endpoint cannot be used to hold an unbounded number of unauthenticated sockets open. The platform is also told answer only once the vendor session actually acks ready (Deepgram SettingsApplied, GPT-Live session.started, Grok session.updated, Gemini setupComplete) — a vendor outage or a bad API key rejects the call outright instead of answering into silence.

import { realtimeAgent, tool } from '@sinch/functions-runtime';

const checkAvailability = tool<{ date: string; partySize: number }>({
  name: 'check_availability',
  description: 'Check table availability for a given date and party size.',
  parameters: {
    type: 'object',
    properties: { date: { type: 'string' }, partySize: { type: 'number' } },
    required: ['date', 'partySize'],
  },
  handler: async ({ date, partySize }) => {
    // ... look up availability ...
    return { available: true, times: ['18:00', '19:30'] };
  },
});

const bookTable = tool<{ date: string; time: string; partySize: number; name: string }>({
  name: 'book_table',
  description: 'Book a table once the caller has confirmed a date, time and party size.',
  parameters: {
    type: 'object',
    properties: {
      date: { type: 'string' },
      time: { type: 'string' },
      partySize: { type: 'number' },
      name: { type: 'string' },
    },
    required: ['date', 'time', 'partySize', 'name'],
  },
  handler: async ({ date, time, partySize, name }, ctx) => {
    const confirmationId = `BK-${Date.now().toString(36).toUpperCase()}`;
    await ctx.storage.write(
      `bookings/${confirmationId}.json`,
      JSON.stringify({ date, time, partySize, name })
    );
    return { confirmationId, message: `Booked for ${name}, ${partySize} guests, ${date} ${time}.` };
  },
});

export default realtimeAgent({
  provider: 'deepgram',
  // voice: 'aura-2-thalia-en', // vendor specific, omit for the vendor default
  prompt: 'You are a reservations agent for Acme Cafe. Be brief and friendly.',
  // A longer prompt usually belongs in a file, not inlined:
  // prompt: (call, ctx) => ctx.assets('prompt.md'),
  greeting: async (call, ctx) => {
    const name = await lookupCallerName(call.from, ctx);
    return name
      ? `Hello, ${name}! Thanks for calling Acme Cafe.`
      : 'Hello! Thanks for calling Acme Cafe.';
  },
  tools: [checkAvailability, bookTable],
  // Transfer to the host stand for anything this agent cannot resolve — injects a
  // `transfer_to_human` tool the model can call; `hangup: true` (the default) injects
  // `hangup_call` the same way. Both are ordinary PerCallValue options, so this can be a
  // (call, context) => number lookup instead of a fixed string.
  transferTo: '+15550001234',
  providerOptions: {
    deepgram: { sttModel: 'nova-3', llmProvider: 'open_ai', llmModel: 'gpt-5.4-nano' },
  },
  onTranscript: (entry, ctx) => {
    console.log(`[${ctx.callId}] ${entry.role}: ${entry.text}`);
  },
  onEnd: async (ctx) => {
    await ctx.storage.write(
      `calls/${ctx.callId}.json`,
      JSON.stringify({ endedAt: new Date().toISOString(), transcript: ctx.transcript })
    );
  },
});

provider is 'openai' | 'gemini' | 'deepgram' | 'grok' | 'voice-relay', so a typo is caught at compile time rather than only on the first call. providerOptions is keyed by provider and only reaches the matching adapter — providerOptions.deepgram above does nothing if provider is ever changed to 'openai'. The STREAM leg realtimeAgent dials itself is capped by the same MAX_CALL_DURATION function variable dialAgent reads (default 1800 seconds).

Per-call overrides. prompt, greeting, voice, transferTo and language each take either a plain value or a (call, context) => value | Promise<value> function, resolved once per call — before the vendor session opens, so the greeting above can say "Hello, Maria" instead of a generic line. language is a BCP-47 tag (e.g. 'en-US', 'es-MX') that pins the provider's speech recognition instead of letting it auto-detect; defaults to 'en-US'. timezone (plain string, not per-call — a property of the deployment, not the caller) defaults to 'America/Chicago' and drives the current date/time line prepended to every provider's resolved prompt. model (in providerOptions) is not per-call; only prompt/greeting/voice/ transferTo/language are. prompt: (call, ctx) => ctx.assets(...) is how a longer prompt comes from a file in assets/ instead of being inlined — ctx.assets reads it as plain text, so any {{placeholder}} substitution the prompt needs is ordinary string work in that same function, not a templating feature this helper provides.

Voice. Unset means no voice is sent to the vendor at all — every adapter omits the field and lets the vendor apply its own default, rather than this repo picking one. Setting voice (a portable id beyond a small mapped set — it otherwise falls through to the adapter unchanged) is the only way a specific voice is used.

Built-in tools. hangup (default true) injects a hangup_call tool; transferTo — a number or a (call, context) => number | undefined — injects transfer_to_human (args: reason, summary). A per-call resolver that returns undefined (say, only some callers should be able to reach a human) omits the tool for that call, same as leaving transferTo unset entirely. Neither built-in is declared in tools above; both call the same ctx.hangup()/ctx.transfer(number) primitives a custom tool would, and an author tool of the same name wins over the injected one. Set hangup: false to drop the standard one — a template with its own goodbye-and-hang-up tool does not want both.

Whenever a hangup_call tool is offered (built-in, or an author override of the same name), the prompt gets an explicit CALL TERMINATION instruction appended — tool_choice stays 'auto' everywhere, so a directive tool description alone does not reliably make a model call it; a live probe found GPT-Live saying goodbye and then never ending the call without this. GPT-Live's own frontend voice model never calls a tool itself, only delegates to the backend that does, so it gets a differently-worded variant telling it to delegate instead — its backend (delegation.responses.instructions) gets its own short rule, since that's the model that actually calls hangup_call once delegation happens. Nothing to configure; this only ever adds to the prompt, never replaces it.

GPT-Live backend tuning, via providerOptions.openai — backendServiceTier (service_tier, default 'priority'), backendReasoningEffort (reasoning.effort, default 'none'), backendVerbosity (text.verbosity, default 'low'), backendParallelToolCalls (parallel_tool_calls, default true), backendMaxOutputTokens (max_output_tokens, default 300) — alongside the existing backendModel/backendInstructions. Defaults favor a spoken conversation; no temperature field is ever sent to the backend, so it never conflicts with reasoning per OpenAI's own delegation guide.

What's normalised: the tool contract, ctx.transcript and onTranscript (every turn from either party, and every completed tool call, in order — see onEnd above), onBargeIn/ onError/onEnd, call control from inside a tool handler (ctx.hangup(), ctx.transfer(number) — the same move-the-leg mechanism as the voice builder's transferTo*/voiceRelay's call.transferToPhone: busy, unanswered or failed drops the caller rather than returning them here), audio format and barge-in clearing, and prompt/greeting delivery.

Usage and turn timing. ctx.usage (inputTokens/outputTokens/audioSeconds) is the call's running total so far, updated as the vendor reports it — tokens/seconds only, no dollar prices. Each TranscriptEntry also carries, where the provider reports it: latencyMs (time from the caller's turn finishing to this reply's first spoken sentence, or a tool's own run time), totalMs (agent turns only: time to the whole reply completing, including any tool rounds — unset if the turn was abandoned by a barge-in), speechStartAt/speechEndAt (user turns; filled by an energy-VAD fallback on raw caller audio when a provider doesn't report its own boundary), and playbackStartAt/playbackEndAt/interrupted (agent turns; left unset when a provider reports no playback signal). Fields a provider doesn't support stay unset rather than guessed at.

What isn't normalised: voice and model ids beyond a small portable set, a vendor's own split STT/LLM/TTS choices (Deepgram's sttModel/llmProvider/llmModel above, for instance), and turn-detection/latency feel — tune those per vendor through providerOptions. Audio defaults to 8kHz telephony, matching the Voice v2 platform's own default for calls that traverse the PSTN; pass audio: 'wideband' for a media path that never does.

Secrets — DEEPGRAM_API_KEY for the example above — are read through context.config, the same way every other function reads a secret, never inlined in providerOptions.

realtimeAgent returns a VoiceFunction, the same export shape a plain voiceWebhook/setup function has: it answers the caller, bridges them, and dials a STREAM leg at this function's own /stream — no dial.stream or raw ws.WebSocket to write by hand.

The one-liner export default realtimeAgent({...}) above owns incoming itself. When a function needs its own incoming logic first — a business-hours check, a lookup that decides whether to answer at all — the same object also exposes connect, the plan-building half voiceWebhook uses internally, so an author can fall through to it instead of reimplementing the answer → bridge → dial STREAM-leg plan:

import { onCall, realtimeAgent } from '@sinch/functions-runtime/voice';

const agent = realtimeAgent({ provider: 'deepgram', prompt: '...' /* tools, hooks, etc. */ });

export const voiceWebhook = onCall({
  incoming: async (call, builder, context) => {
    if (isAfterHours()) {
      return builder.say('Sorry, we are closed. Please call back during business hours.').hangup();
    }
    return agent.connect(call, builder, context);
  },
});

Talking to an LLM over Voice Relay

provider: 'voice-relay' runs realtimeAgent on top of voiceRelay instead of a vendor's own realtime audio session: Sinch does the speech (STT/TTS), and this function runs the model as a text loop over that same connection. From the author's side it behaves like every other provider — the same options, ctx, onTranscript, injected hangup_call/transfer_to_human tools, tool calling and barge-in:

export default realtimeAgent({
  provider: 'voice-relay',
  model: 'gpt-4o-mini',
  prompt: 'You are a reservations agent for Acme Cafe. Be brief and friendly.',
  tools: [checkAvailability, bookTable],
  transferTo: '+15550001234',
});

model works with any model behind an OpenAI Responses-compatible API — OpenAI, Azure OpenAI, or a gateway such as OpenRouter or LiteLLM — sent straight through with no prefix or registry lookup. baseUrl (or the OPENAI_BASE_URL function variable, when baseUrl is unset) points it at a server other than https://api.openai.com/v1. The credential is always the caller's own OPENAI_API_KEY function secret, read fresh per call rather than cached globally.

Low-latency options, for a model and account that support them:

export default realtimeAgent({
  provider: 'voice-relay',
  model: 'gpt-5.4-nano',
  serviceTier: 'priority',
  reasoningEffort: 'none',
  prompt: '...',
});

serviceTier ('auto' / 'default' / 'flex' / 'priority' / 'fast') maps to service_tier, reasoningEffort ('none' / 'minimal' / 'low' / 'medium' / 'high') to reasoning.effort, verbosity ('low' / 'medium' / 'high') to text.verbosity — all omitted from the request when unset. LLM_SERVICE_TIER/LLM_REASONING_EFFORT/LLM_VERBOSITY function variables set them without a code change; an explicit option here wins over the variable, which wins over the built-in default.

What the default depends on, and why: with no baseUrl (or api.openai.com explicitly), serviceTier defaults to 'fast' and every turn uses store: true with previous_response_id chaining — only the new turn is sent, not the whole conversation. Azure OpenAI (*.openai.azure.com) gets the same chaining but no serviceTier default — that setting is calibrated against OpenAI's own account tiers. Any other baseUrl (a gateway — OpenRouter, LiteLLM, a customer's own proxy) is not assumed to support store/chaining at all: every turn resends the full conversation as input, store: false, no previous_response_id, no serviceTier unless set explicitly. reasoningEffort/verbosity default to 'none'/'low' only for a gpt-5*/o-series model (gpt-4o-mini confirmed to reject them outright with a 400) — that default applies the same way regardless of baseUrl. An o-series model is the one exception: it rejects reasoning.effort: 'none' outright, so the defaulted value is withheld for one (reasoning is omitted entirely) — an explicit option or variable is still sent, on any host and for any model.

A persistent WebSocket (plain ws, not the openai package) is opened for the call on every host — OpenAI, Azure OpenAI and a gateway alike — instead of a new HTTP/SSE connection per turn. What a live WS failure does next depends on the host: OpenAI and Azure OpenAI never fall back to HTTP — a failure opens a fresh WebSocket and retries the same turn once, and chaining state (previous_response_id) survives the reconnect untouched. A gateway is not assumed to speak this protocol at all: its first failure (the upgrade itself, or the first request being rejected) latches, and every turn for the rest of that call goes over HTTP/SSE instead. A barge-in is never treated as a failure on either path.

Voice v1 (legacy)

v1 still works, and a function already written against it needs no changes beyond the client: the SDK namespace that used to be context.voice is now context.voice.v1, so context.voice.callouts.tts(...) becomes context.voice.v1.callouts.tts(...). The ice/ace/pie/dice/notify callbacks below are untouched.

import type { FunctionContext, IceCallback, PieCallback } from '@sinch/functions-runtime';
import { IceSvamlBuilder, PieSvamlBuilder, createMenu } from '@sinch/functions-runtime';

export default {
  async ice(context: FunctionContext, event: IceCallback) {
    const menu = createMenu()
      .prompt('Press 1 for sales, 2 for support.')
      .option('1', 'return(sales)')
      .option('2', 'return(support)')
      .timeout(5000)
      .repeats(2)
      .build();

    return new IceSvamlBuilder().runMenu(menu).build();
  },

  async pie(context: FunctionContext, event: PieCallback) {
    if (event.menuResult?.value === 'sales') {
      return new PieSvamlBuilder()
        .say('Connecting you to sales.')
        .connectPstn('+15551234567')
        .build();
    }
    return new PieSvamlBuilder().say('Goodbye.').hangup().build();
  },
};

Callbacks

| Handler | Fires when | Returns | | -------- | --------------------------------------------------- | ------- | | ice | An incoming call reaches your number | SVAML | | ace | An outbound call is answered | SVAML | | pie | The caller pressed keys or spoke after a menu | SVAML | | dice | The call ended | nothing | | notify | A notification event arrives (recording, and so on) | nothing |

Only ice is required. The platform may still post the others; the runtime answers 200 for any handler you did not implement.

Handlers receive camelCased payloads typed as IceCallback, AceCallback, PieCallback, DiceCallback and NotifyCallback. The call identifier arrives as callId.

SVAML builders

Each builder chains instructions (say, play, playFiles, setCookie, sendDtmf, startRecording, stopRecording), ends with one action and finishes with build(). The actions available depend on the callback:

| Builder | Actions | | ----------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ | | IceSvamlBuilder | hangup, continue, connectPstn, connectSip, connectMxp, connectConf, runMenu, park, connectAgent; also the answer instruction | | AceSvamlBuilder | hangup, continue | | PieSvamlBuilder | hangup, continue, connectPstn, runMenu, park, connectAgent |

new IceSvamlBuilder()
  .say('Please hold.')
  .connectPstn('+15551234567', {
    cli: '+15559876543',
    timeout: 30,
    enableAce: true,
    enableDice: true,
  })
  .build();

Menus

createMenu() builds a DTMF menu with prompts, options, timeouts and repeats. MenuTemplates has ready-made ones: business(name), yesNo(question), language(), afterHours(name) and numericInput(prompt, digits).

HTTP handlers

Any other exported async function is an HTTP endpoint named after the export. export async function notes(...) answers requests to /notes.

import type { FunctionContext, FunctionRequest, FunctionResponse } from '@sinch/functions-runtime';

export async function notes(
  context: FunctionContext,
  request: FunctionRequest
): Promise<FunctionResponse> {
  if (request.method === 'POST') {
    const { key, content } = request.body as { key: string; content: string };
    await context.storage.write(`notes/${key}.txt`, content);
    return { statusCode: 201, body: { key } };
  }
  return { statusCode: 200, body: await context.storage.list('notes/') };
}

FunctionRequest carries method, path, query, headers, body and params. Return { statusCode, headers?, body? }, or use createJsonResponse, createResponse and createErrorResponse.

Webhooks from other Sinch products are plain HTTP handlers too. An inbound Conversation API message, for example:

export async function conversationWebhook(
  context: FunctionContext,
  request: FunctionRequest
): Promise<FunctionResponse> {
  const body = request.body as Record<string, unknown>;
  const contactId = body.contact_id as string | undefined;
  const text = (body.message as any)?.contact_message?.text_message?.text as string | undefined;
  if (!contactId || !text) return { statusCode: 200, body: { ok: true } };

  await context.conversation!.messages.send({
    sendMessageRequestBody: {
      app_id: context.env!.CONVERSATION_APP_ID!,
      recipient: { contact_id: contactId },
      message: { text_message: { text: 'Thanks for your message.' } },
    },
  });
  return { statusCode: 200, body: { ok: true } };
}

Helpers such as getText, getChannel, getContactId, isTextMessage and ConversationMessageBuilder are exported for working with Conversation API payloads.

Protecting endpoints

Export auth to gate handlers behind your project API key/secret, and optionally ADMIN_USER/ADMIN_PASSWORD (a second variable/secret pair you set) as an alternative credential. '*' protects every endpoint; an array names the ones to protect, including 'default' for the root handler (or the static public/index.html, if that's what's served there) and a static file's path, e.g. 'public/admin.html', for anything else under public/. Unlisted static assets (CSS, JS, images) stay public.

export const auth = ['notes', 'config', 'default'];

A rejected request gets one of two things: a browser navigating there (an Accept: text/html GET) is redirected to a branded /login page; anything else (curl, a script, a fetch() call) gets a 401 JSON body with no WWW-Authenticate header, so a browser never pops its own Basic Auth dialog. curl -u user:pass still works unchanged — Basic credentials are sent pre-emptively regardless of the challenge header.

Signing in via /login sets an HttpOnly session cookie (HMAC of the password, no session store) as an alternative to sending Basic on every request; it's cleared by POST /logout. Both routes, and the login page itself, are built in — export your own login and/or logout handler to replace either one.

Startup and WebSockets

Export setup to run code once before the server accepts requests, or to accept WebSocket connections (for example a connectStream media stream).

import type { SinchRuntime } from '@sinch/functions-runtime';

export function setup(runtime: SinchRuntime) {
  runtime.onStartup(async (context) => {
    await context.cache.set('ready', true);
  });
  runtime.onWebSocket('/stream', (ws, req) => {
    ws.on('message', (frame) => {
      /* audio */
    });
  });
}

The context

Every handler receives a FunctionContext.

| Member | What it is | | ---------------------- | --------------------------------------------------------------------------------------------------------------------------------------- | | context.config | projectId, functionName, environment, variables, and getVariable/requireVariable/getSecret/requireSecret/hasVariable… | | context.env | Environment variables, including secrets | | context.cache | Key-value store with TTLs: get, set, has, delete, extend, keys, getMany. In memory locally, shared in production | | context.storage | Durable files: write, read, list, exists, delete. Local filesystem in dev, object storage in production | | context.database | Path to a per-function SQLite file, replicated in production. Open it with node:sqlite or any SQLite driver | | context.assets(name) | Reads a file from the function's assets/ directory | | context.voice | Voice API v2 client: call, callWithStream, bridge, transfer, patch, onCall. Always present | | context.voice.v1 | Voice API v1 client, when VOICE_APPLICATION_KEY and VOICE_APPLICATION_SECRET are set | | context.conversation | Conversation API client, when CONVERSATION_APP_ID is set | | context.sms | SMS API client, when SMS_SERVICE_PLAN_ID is set | | context.numbers | Numbers API client, when ENABLE_NUMBERS_API=true |

The API clients also need PROJECT_ID, PROJECT_ID_API_KEY and PROJECT_ID_API_SECRET, which sinch auth login provides. A client whose variables are missing is undefined — except context.voice, which is always there and reports missing credentials when a request is actually sent. Full list: SDK environment variables.

import { DatabaseSync } from 'node:sqlite';

const db = new DatabaseSync(context.database);
db.exec('CREATE TABLE IF NOT EXISTS visits (day TEXT PRIMARY KEY, n INTEGER NOT NULL DEFAULT 0)');

context.config carries getVariable, requireVariable, getSecret, requireSecret, hasVariable, hasSecret, isDevelopment, isProduction and getApplicationCredentials directly: context.config.getVariable('COMPANY_NAME', 'Acme') returns a string, and without a default it returns null when unset. createConfig(context) and createUniversalConfig(context) still return the same methods on a wrapper.

Local development and production

This package is the local development runtime and the SDK you import from. sinch functions dev runs your function on http://localhost:3000 (override with PORT) with hot reload, an in-memory cache, filesystem storage and a tunnel that forwards live Sinch callbacks to your machine.

While the tunnel is up it repoints the callbacks at it: the v1 voice application callback when VOICE_APPLICATION_KEY is set, and the Voice v2 service webhook when VOICE_SERVICE_ID is set.

When you deploy, the platform installs @sinch/functions-runtime-prod in the container. It exposes the same API, so nothing in function.ts changes between environments. Do not install the prod package yourself.

Limits

| Limit | Value | | ------------------------------------- | ----------- | | Request body | 10 MB | | Request or response body kept in logs | First 64 KB |

The 10 MB cap applies locally and deployed, so a request that works in sinch functions dev works in production. A deployed function rejects a larger body with 413 and a JSON body naming the limit; the platform router in front of it enforces the same cap, so a 413 without that JSON body came from the router. Deployed functions keep the first 64 KB of text bodies (JSON, XML, plain text, form data) in their logs and append …[truncated — N bytes total] past that. Binary bodies are never logged. See limits.

TypeScript

Everything is typed. Import types alongside the builders:

import type {
  FunctionContext,
  IceCallback,
  SvamletResponse,
  VoiceFunction,
} from '@sinch/functions-runtime';

Documentation

License

MIT