@juspay/neurolink
v12.41.10
Published
The pipe layer of an AI nervous system: one interface connecting provider neurons to your application, across three inference types — generate, stream and decide. `decide` returns typed, calibrated judgements from a non-generative model (~400ms, ~$0.00002
Keywords
Readme
NeuroLink
The pipe layer for the AI nervous system.
AI intelligence flows as streams — tokens, tool calls, memory, voice, documents. NeuroLink is the vascular layer that carries these streams from where they are generated (LLM providers: the neurons) to where they are needed (connectors: the organs).
import { NeuroLink } from "@juspay/neurolink";
const pipe = new NeuroLink();
// Everything is a stream
const result = await pipe.stream({ input: { text: "Hello" } });
for await (const chunk of result.stream) {
if ("content" in chunk) {
process.stdout.write(chunk.content);
}
}
// Or skip text entirely: a calibrated decision, not a token stream
const decision = await pipe.tryDecide({
// null if no decision provider is set
state: { ticket: "Refund request, $42, first occurrence" },
questions: {
autoApprove: {
type: "boolean",
instructions: "Approve without human review.",
},
},
});
// decision?.answers.autoApprove.probability -> 0.91→ Docs · → Quick Start · → npm · → Blog
🧠 What is NeuroLink?
NeuroLink is the pipe layer of an AI nervous system. Providers — OpenAI, Anthropic, Google, AWS, Azure, Mistral, local runtimes like Ollama, and dozens more — are the neurons: each generates a different kind of intelligence, at a different cost and latency. NeuroLink is the vascular layer that carries that intelligence, as a stream, to the applications — the organs — that consume it, across three inference types: generate and stream produce text, decide produces a calibrated boolean/choice/score judgment instead. A curated model registry (64 models, 132 aliases) backs metadata, routing, and context-window checks out of the box, and hundreds more models are reachable through aggregator providers — 100+ via LiteLLM, 300+ via OpenRouter.
Extracted from production systems at Juspay, NeuroLink provides a practical, TypeScript-first way to plug any application into that nervous system. Switch which neuron answers a request with a single parameter change — OpenAI, Anthropic, Google, AWS Bedrock, Azure, a local runtime, or any provider you add. decide is the third inference type — a typed, calibrated judgment instead of text — for the model-routing and gating decisions generate/stream were never meant to make, powered by a purpose-built decision model (TypeSafe Jev, or the open-weights Laya or XOR) rather than a general-purpose LLM: with Jev, routing decisions land in ~400ms for about $0.00002, instead of a full generation call.
Why NeuroLink? Three genuine inference types, not one dressed up three ways — generate and stream produce text; decide produces a calibrated boolean/choice/score judgment, and which types a provider serves is declared per-provider via inferenceKinds rather than inferred from behavior. Every neuron plugs into the same pipe, including 3 fully local runtimes (Ollama, LM Studio, llama.cpp) with per-request credential overrides, and MCP support covers all 4 transports (stdio, HTTP, SSE, WebSocket). Every AI-driven optimization the pipe performs — model routing, context compaction, tool selection — fails open: no key configured behaves exactly like NeuroLink without it, and routing uses asymmetric confidence thresholds (upgrade at 0.3, downgrade at 0.6) rather than a single cutoff, because a wrong downgrade costs more than a wrong upgrade. Switch providers with a single parameter change, leverage built-in tools plus any MCP-compliant tool server, deploy with confidence using enterprise features like Redis memory and multi-provider failover, and optimize costs automatically with intelligent routing. Use it via our professional CLI or TypeScript SDK—whichever fits your workflow.
Where we're headed: We're building for the future of AI—edge-first execution and continuous streaming architectures that make AI practically free and universally available. Read our vision →
What's New
| Feature | Version | Description | Guide |
| ------------------------------------------------------- | ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------- |
| decide Inference Type + TypeSafe Jev + Laya + XOR | next | A third inference type alongside generate/stream: typed, calibrated judgments (boolean, choice, score) via neurolink.decide() / tryDecide(), one parallel pass (~400ms and ~$0.00002/decision on Jev). First provider is TypeSafe Jev (TYPESAFE_API_KEY, also reachable via the Vercel AI Gateway); Laya (LAYA_API_KEY + LAYA_BASE_URL), an open-weights model you run yourself, and XOR (XOR_API_KEY + XOR_BASE_URL), Juspay's open-weights model, follow it — the first one configured of TypeSafe, Laya and XOR runs. Used internally for model routing, context budgeting, relevance compaction and tool routing — fail-open and a no-op without a key. Per-query RAG planning is opt-in via RAGPipeline. | Decide Guide |
| 7 More Catalog Providers | v12.11.0–v12.16.0 | Baseten, GMI Cloud, Inception Labs, io.net Intelligence, Mancer, Upstage and API Route onboarded as Tier-2 catalog entries — one JSON file each, roster live-verified against the provider's own /v1/models. | Tier 2 Onboarding |
| Claude-on-Vertex Proxy Fallback | v12.18.0 | The Anthropic proxy pool can fall back to Claude served on Google Vertex, so an agentic turn survives losing its primary backend mid-conversation instead of failing the turn. | Claude Proxy |
| Native-Loop V3 Conversation Reclaim | v12.17.0 | Reclaims V3 conversations without splitting tool-call/tool-result pairs — the pairing a provider rejects the whole request over. | Claude Proxy Architecture |
| Multi-Modal Embeddings | v12.15.0 | embed() / embedMany() accept images alongside text on providers whose embedding models are multi-modal, for cross-modal retrieval in RAG and custom vector search. | Embeddings Guide |
| Grok Build Auto-Configuration | v12.14.0 | The proxy configures Grok Build automatically, deriving context windows and backends from the model catalog rather than hardcoded values. | Proxy CLI Onboarding |
| Anthropic Execution-Control Contract | v12.13.0 | Truthful stream termination plus an opt-in execution-control contract, so a stream that stopped early reports why instead of looking like a clean finish. | Claude Proxy |
| Catalog Tool Declarations Honoured at Runtime | v12.12.0 | A Tier-2 catalog entry declaring tools: false (e.g. Mancer) no longer has tools offered to it at runtime — the JSON declaration is enforced, not just documented. | Tier 2 Onboarding |
| Artifact Stores: Redis, Custom, Range Reads, Search | v12.10.0 | Artifacts can be backed by Redis or a custom store, read by byte range, and searched — instead of being held only in process memory. | Claude Proxy |
| Local CLI Spend Reading | v12.6.0–v12.9.0 | Reads token usage directly from other coding CLIs' own local stores — Cursor, Grok Build, Hermes Agent and three more — and names them in proxy traffic, so spend is attributed per client. | Proxy CLI Onboarding |
| Native OpenAI Audio Streaming | v12.7.0 | OpenAI TTS audio streams natively rather than being buffered to completion first. | TTS Guide |
| HITL Pending-Confirmation State | v12.5.0 | Exposes whether a human-in-the-loop confirmation is still outstanding, so a caller can distinguish 'waiting on a human' from 'finished'. | Task Manager |
| OpenCode + Gemini CLI Proxy Clients | v12.4.0 | OpenCode's generated config is actually loadable, and Gemini CLI is onboarded as a proxy client. | OpenCode Proxy | Proxy CLI Onboarding |
| SambaNova Provider | v12.3.0 | RDU-accelerated open-weight flagships: Llama 3.3 70B (default), GPT-OSS 120B, DeepSeek V3.x, MiniMax, Gemma 4 (vision) — OpenAI-compatible Tier 2 catalog entry. Note: new SambaNova accounts require purchased credits. | SambaNova Guide |
| Cerebras Provider | v12.1.0 | Wafer-scale inference at ~3000 tok/s: GPT-OSS 120B (default) + Gemma 4 31B, OpenAI-compatible Tier 2 catalog entry, live-verified end to end (generate, stream, tools, structured output). | Cerebras Guide |
| Avatar / Music Modalities + 12 Providers | v9.65.0 | New output: { mode: "avatar" \| "music" } dispatch with handlers for D-ID, HeyGen, Replicate-MuseTalk (avatar) and Beatoven, ElevenLabs Music, Lyria, Replicate-MusicGen (music). Plus Fish Audio TTS, Kling/Runway/Replicate video, xAI/Groq/Cohere/Together/Fireworks/Perplexity/Cloudflare LLMs, Voyage/Jina embeddings, Stability/Ideogram/Recraft/Replicate image-gen. | Provider Integration |
| Multi-Provider Voice (TTS/STT) | v9.62.0 | 6 TTS providers (OpenAI TTS, ElevenLabs, Google TTS, Azure TTS, Fish Audio, Cartesia) + 4 STT providers (Whisper, Deepgram, Azure STT, Google STT) + 2 realtime APIs (OpenAI Realtime, Gemini Live). | TTS Guide | STT Guide | Realtime Guide |
| 4 New Providers | v9.60.0 | DeepSeek (V3/R1), NVIDIA NIM (400+ catalog), LM Studio (local), llama.cpp (GGUF local). | Provider Setup |
| ModelAccessDeniedError | v9.59.0 | Typed ModelAccessDeniedError + sdk.checkCredentials() API for proactive credential validation before first call. | Error Reference |
| Provider Fallback Policy | v9.58.0 | providerFallback callback + modelChain config for centralized multi-provider fallback logic. | Advanced Guide |
| Per-Request Credentials | v9.52.0 | Pass credentials per-call or per-instance for all providers. Per-call overrides instance; instance overrides env vars. | Credentials Guide |
| AutoResearch | v9.53.0 | Autonomous AI experiment engine: proposes code changes, runs experiments, evaluates metrics — unattended for hours. | AutoResearch Guide |
| Gemini 3 Multi-turn Tool Fix | v9.49.0 | Fixed multi-step agentic tool calling on Vertex AI Gemini 3. Correct thoughtSignature replay, stepIndex grouping, executionId session isolation, 5-min timeout. | Vertex AI Guide |
| MCP Enhancements | v9.16.0 | Tool routing (6 strategies), result caching (LRU/FIFO/LFU), request batching, annotations, elicitation protocol, multi-server management. | MCP Enhancements Guide |
| Memory | v9.12.0 | Per-user condensed memory across conversations. LLM-powered condensation with S3, Redis, or SQLite. | Memory Guide |
| Context Window Management | v9.2.0 | 5-stage compaction pipeline with budget gate at 80% usage, per-provider token estimation. | Context Compaction Guide |
| Tool Execution Control | v9.3.0 | prepareStep and toolChoice for per-step tool enforcement in multi-step agentic loops. | API Reference |
| File Processor System | v9.1.0 | 17+ file type processors with ProcessorRegistry, security sanitization, SVG text injection. | File Processors Guide |
| RAG with generate()/stream() | v9.2.0 | Pass rag: { files } for automatic document chunking, embedding, and AI-powered search. 10 chunking strategies, hybrid search, reranking, and a choice of 4 vector stores (in-memory, Chroma, PgVector, Pinecone). | RAG Guide |
// decide() — a third inference type: calibrated judgments, not text (next)
// Enable with TYPESAFE_API_KEY (or AI_GATEWAY_API_KEY via Vercel AI Gateway).
import { NeuroLink, readDecisionChoice } from "@juspay/neurolink";
const neurolink = new NeuroLink();
const result = await neurolink.tryDecide({
// null if no decision provider is configured
state: ticketText,
questions: {
team: {
type: "choice",
instructions: "Which team should handle this?",
criteria: { billing: "Payments", technical: "Bugs", sales: "Pricing" },
},
urgent: { type: "boolean", instructions: "Is this urgent?" },
},
});
const team = result && readDecisionChoice(result.answers, "team");
if (team && team.confidence > 0.7) {
route(team.choice); // "billing" | "technical" | "sales", plus a full ranking
}
// Multi-Provider Voice (v9.62.0) — TTS + STT
// Voice is configured via the `tts` / `stt` options on generate() / stream(),
// not via dedicated synthesizeSpeech / transcribeAudio methods.
// Text in, audio out (TTS)
const result = await neurolink.generate({
input: { text: "Hello from NeuroLink" },
provider: "vertex",
tts: {
enabled: true,
voice: "en-US-Neural2-C",
format: "mp3",
output: "./output.mp3", // optional: save to disk
provider: "elevenlabs", // optional override: openai-tts | elevenlabs | google-ai | vertex | azure-tts | fish-audio | cartesia
},
});
// result.audio: { buffer: Buffer, format: "mp3", ... }
// Audio in (STT), text out
const transcript = await neurolink.generate({
input: { text: "Transcribe and summarize" },
provider: "openai",
stt: {
enabled: true,
audio: audioBuffer, // Buffer of the audio file
provider: "whisper", // whisper | deepgram | google-stt | azure-stt
language: "en-US",
},
});
// Real-time bidirectional voice (OpenAI Realtime / Gemini Live)
import { RealtimeProcessor } from "@juspay/neurolink";
await RealtimeProcessor.connect(
"openai-realtime",
{ provider: "openai-realtime", model: "gpt-4o-realtime-preview" },
{ onAudio, onTranscript, onError, onFunctionCall },
);
// AutoResearch — autonomous experiment loop (v9.53.0)
import { resolveConfig, ResearchWorker } from "@juspay/neurolink/autoresearch";
const config = resolveConfig({
repoPath: "/path/to/repo",
mutablePaths: ["train.py"],
runCommand: "python3 train.py",
metric: {
name: "val_bpb",
direction: "lower",
pattern: "^val_bpb:\\s+([\\d.]+)",
},
});
const worker = new ResearchWorker(config);
await worker.initialize("experiment-1");
const result = await worker.runExperimentCycle("Try lower learning rate");
// Provider Fallback Policy (v9.58.0) — fires only on ModelAccessDeniedError
import { NeuroLink, ModelAccessDeniedError } from "@juspay/neurolink";
const neurolink = new NeuroLink({
// Async callback. Single error arg. Return null to give up,
// or { provider?, model? } to retry with a substitute.
providerFallback: async (error) => {
if (
error instanceof ModelAccessDeniedError &&
error.allowedModels?.length
) {
return { model: error.allowedModels[0] };
}
return null;
},
// Sugar over providerFallback: if no callback is set, NeuroLink walks this list
// on each access denial. modelChain is `string[]` only (model names; same provider).
modelChain: ["claude-opus-4-7", "claude-sonnet-4-6", "gpt-4o"],
});- Sharp image compression (v9.50.0) – Automatic image compression for AI providers via the sharp library; reduces upload bandwidth and bypasses provider size limits.
- Redis URL/TLS (v9.49.0) – Redis URL-based connections with TLS support for secure conversation memory in production.
- TaskManager (v9.41.0) – Scheduled and self-running AI tasks; cron-style execution with state checkpointing.
- Multi-user memory retrieval (v9.40.0) – Per-user memory storage and retrieval with customizable prompts.
- Evaluation Scoring (14 scorers) (v9.37.0) – Modular evaluation system with 14 scorers, pipelines, and CLI for offline quality assessment.
- Browser-compatible bundle (v9.34.0) – Client-side SDK bundle for browser use; no Node.js dependency for the core API.
- Per-call memory control (v9.33.0) – Read/write memory control per
generate()andstream()call. - Server Adapters (v8.43.0) – HTTP server with Hono, Express, Fastify, Koa. Foreground/background modes, route management, OpenAPI generation. → Guide
- External TracerProvider (v8.43.0) – Integrate NeuroLink with existing OpenTelemetry setups. → Guide
- Title Generation Events (v8.38.0) –
conversation:titleGeneratedevent +NEUROLINK_TITLE_PROMPTcustom titles. → Guide - Video Generation with Veo (v8.32.0) – Video generation via Google Veo 3.1 on Vertex AI. 720p/1080p, portrait/landscape. → Guide
- Image Generation (v8.31.0) – Native image generation with Gemini and Imagen models. → Guide
- HTTP/Streamable HTTP Transport (v8.29.0) – Remote MCP servers via HTTP with auth headers, retry, rate limiting. → Guide
- PPT Generation – 35 slide types, 5 themes, optional AI-generated images. Works across supported AI providers. → Guide
- Structured Output with Zod – Type-safe JSON via
schema+output.format: "json". → Guide - CSV & PDF File Support – Attach CSV/PDF with auto-detection. PDF: native visual analysis on Vertex, Anthropic, Bedrock, AI Studio. → CSV | PDF
- LiteLLM, SageMaker & OpenRouter – 100+ models via LiteLLM, custom endpoints on SageMaker, 300+ via OpenRouter. → LiteLLM | SageMaker
- HITL & Guardrails – Human-in-the-loop approval workflows and content filtering. → HITL | Guardrails
- Redis Conversation Export – Export full session history as JSON for analytics and audit. → Guide
Decide: Calibrated Judgments, Not Text
NeuroLink now supports decision models — a third inference type, with two providers today: a hosted one and an open-weights one you can run yourself.
decide sits alongside generate and stream. Instead of tokens, a decision
model takes one state plus a map of named typed questions and returns one
typed, calibrated answer per question, all in a single parallel pass — no
text output anywhere, so nothing has to be parsed back out of prose.
Public API: neurolink.decide() and the fail-open neurolink.tryDecide()
(returns null instead of throwing). This is a TypeScript SDK surface —
there is no neurolink decide CLI command today.
| Primitive | Answer shape | Use it for |
| --------- | -------------------------------------------------------------------------------------- | ------------------------------------------------------------------- |
| boolean | A probability, 0–1 (no confidence of its own — gate on distance from 0.5) | Yes/no gates: approve, drop, include, flag |
| choice | An option + the full probability distribution + a confidence | Routing to one of N options — the distribution also ranks all N |
| score | A probability-weighted index into an ordered rubric + a confidence, and a legend | Position on a scale: severity, priority, quality tier |
A score is probability-weighted, so it can land between rubric levels —
useful for sorting a queue, not just bucketing it.
What you can build with it
The model is fast, cheap and calibrated, but ~68% accurate (see the trade-off below). That combination fits work that is batched, gated and reversible — where a wrong answer is caught by a threshold or a human, not shipped to a user.
| Use case | Primitive | Why it fits |
| ---------------------------------- | -------------------------------- | ---------------------------------------------------------------------------- |
| Ticket / helpdesk triage | choice team + score priority | ranked gives a fallback team order; a human still sees the ticket |
| Content moderation, first pass | one boolean per item | Only a confident "yes" auto-hides; everything else escalates |
| Lead or severity queues | score over an ordered rubric | The between-levels score sorts a queue rather than bucketing it |
| Shortlisting & reranking | one choice over N candidates | One request ranks the whole catalogue — SKUs, canned replies, search results |
| Your own model-tier gate | boolean or choice | Decide cheap-vs-capable per message before you call a text model |
| Spam / fraud pre-screen | boolean with asymmetric bars | A high bar to auto-reject, a lower one to flag for review |
Do not use it for a final answer a user reads, an irreversible action with no
confirmation step, or anything needing a rationale — a decision carries a
probability, never an explanation. Those belong to generate.
Enabling it
Set TYPESAFE_API_KEY — that one env var is the whole switch.
TypeSafe Jev is the first decision provider
(AIProviderName.TYPESAFE, aliases jev / typesafe-ai). It is also reachable
through the Vercel AI Gateway via AI_GATEWAY_API_KEY; force one transport
with TYPESAFE_TRANSPORT=direct|gateway. Per-request credentials work as they do
for every other provider.
Laya is the second decision
provider — Convai Innovations' Apache-2.0, open-weights "System One" model, for
when you'd rather run the decision model on your own infrastructure than call a
hosted one. There is no built-in endpoint: set LAYA_API_KEY and
LAYA_BASE_URL to point at a Laya server you run yourself, or a LiteLLM proxy
with a pass-through route to one (or pass credentials.laya to
new NeuroLink({ credentials }), or per call). Its encoders read a much
shorter state than TypeSafe's — about 768 tokens on the default
typed-decisions checkpoint (320 on english/auto), against TypeSafe's
~33,000 — so it fits short, structured decisions rather than long context.
Every built-in consumer below asks for the first configured decision provider,
and TypeSafe is listed first: with both configured, TypeSafe runs; with only
Laya's key and base URL set, Laya runs.
XOR is Juspay's Apache-2.0,
open-weights decision model (xor-1.1); setup is on
Hugging Face. It has no built-in endpoint
either: set XOR_API_KEY and XOR_BASE_URL to the origin of a deployment or of
a LiteLLM proxy route (or pass credentials.xor). Built-in consumers use the
first configured provider in the order TypeSafe, Laya, XOR; XOR counts only with
both its key and base URL set, and a call can always name it with
provider: "xor". It is also the decision provider that accepts images and
video, which TypeSafe and Laya refuse.
import {
NeuroLink,
readDecisionChoice,
gateDecisionBoolean,
} from "@juspay/neurolink";
const neurolink = new NeuroLink();
const result = await neurolink.tryDecide({
state: { ticket: "Customer reports a failed $42 payment, first occurrence." },
questions: {
team: {
type: "choice",
instructions: "Which team should handle this?",
criteria: { billing: "Payments", technical: "Bugs", sales: "Pricing" },
},
autoRefund: {
type: "boolean",
instructions: "Approve the refund without human review.",
},
},
});
const team = result && readDecisionChoice(result.answers, "team");
if (team && team.confidence > 0.7) {
route(team.choice); // team.ranked is the full ordering, not just the winner
}
// A boolean has no confidence of its own, so gate on BOTH the probability and
// its distance from a coin flip. `undefined` means "not sure" — not "no".
const refund = result && gateDecisionBoolean(result.answers, "autoRefund");
if (refund === true) autoRefund();
else queueForHuman();Where NeuroLink uses it itself
Five decide() calls across the codebase, each fail-open and a no-op
without a key — so nothing changes in its absence:
| # | Call site | What it asks | Guide |
| --- | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 1 | routing/classifierStrategies.ts | Difficulty, required capabilities, risk, how much context is needed, and which model to pick — all in one request | routing · catalogue · context budget |
| 2 | context/contextDecision.ts | One yes/no per earlier message: is this still needed for the current request? | relevance compaction |
| 3 | context/contextDecision.ts | Does this generated summary preserve every decision and open question? | relevance compaction |
| 4 | core/toolRoutingDecision.ts | One yes/no per MCP server: does the request need it? Drops only on a confident "no" | tool routing |
| 5 | rag/retrieval/searchDecision.ts | Per query: topK breadth, and whether to use hybrid / graph / rerank | RAG planning |
Worth being precise about two of these, because the grouping is easy to misread:
- Call 1 is a single request that does the work of three features. Model
routing, the registry-derived catalogue and the per-request
compactionThresholdall read different answers out of the same call — the context budget is not a second round trip, andmodelCatalog.tsnever callsdecide()at all; it renders the candidate lines that call 1's model question chooses between. That is the batch-never-fan-out rule applied to NeuroLink's own code. - Call 5 is opt-in wiring, not automatic. Per-query planning lives in
RAGPipeline, which therag: { files }shortcut ongenerate()/stream()does not construct. Build aRAGPipelineyourself and pass a decide function to get it; the shortcut path is unchanged.
Every one goes through the same decide() / tryDecide() core, so each gets
telemetry even on failure — a dedicated model.decision span, never folded
into generation metrics. That matters because a decision path that has silently
stopped working (rate-limited, timed out, provider down) would otherwise look
identical to one that was never configured; the span is what tells the two
apart.
What it costs, and its limits
Measured against the live API — don't extrapolate past these:
- Latency is flat in question count: 1 question ~393ms, 400 questions ~465ms. Concurrent requests queue, so batch every question into one call — never fan out.
- ~$0.042 per million input tokens, output reported but billed at zero — about $0.00002 per decision.
- Two input ceilings:
state+ the longest single question ≈ 33,000 tokens;state+ all questions ≈ 64,000 tokens. - Accuracy is the trade-off: ~68% on TypeSafe's own 711-case benchmark vs. ~73% for a frontier model (TypeSafe's published figures, not our measurement). Use it for decisions that are gated and reversible — routing, dropping, budgeting — never for a final answer a user will see.
Decide Guide · TypeSafe Provider Guide · Laya Provider Guide · XOR Provider Guide
Enterprise Security: Human-in-the-Loop (HITL)
NeuroLink includes a HITL (Human-in-the-Loop) system for regulated industries and high-stakes AI operations:
| Capability | Description | Use Case | | --------------------------- | ----------------------------------------------------------------------- | ------------------------------------------ | | Tool Approval Workflows | Require human approval before AI executes sensitive tools | Financial transactions, data modifications | | Output Validation | Route AI outputs through human review pipelines | Medical diagnosis, legal documents | | Confidence Thresholds | Automatically trigger human review below confidence level | Critical business decisions | | Complete Audit Trail | Audit logging to support your compliance program (HIPAA / SOC 2 / GDPR) | Regulated industries |
import { NeuroLink } from "@juspay/neurolink";
const neurolink = new NeuroLink({
hitl: {
enabled: true,
requireApproval: ["writeFile", "executeCode", "sendEmail"],
confidenceThreshold: 0.85,
reviewCallback: async (action, context) => {
// Custom review logic - integrate with your approval system
return await yourApprovalSystem.requestReview(action);
},
},
});
// AI pauses for human approval before executing sensitive tools
const result = await neurolink.generate({
input: { text: "Send quarterly report to stakeholders" },
});Enterprise HITL Guide | Quick Start
📚 Quick Start Guide
This guide will have you generating AI responses in under 5 minutes using either the SDK or CLI.
Installation
Choose your preferred package manager:
# npm
npm install @juspay/neurolink
# pnpm (recommended)
pnpm add @juspay/neurolink
# yarn
yarn add @juspay/neurolink
# CLI only (no installation needed)
npx @juspay/neurolink --helpConfiguration
NeuroLink works with a broad set of AI providers — and local runtimes that need no API key at all. You'll need at least one to get started:
Option 1: Interactive Setup (Recommended)
# Run the setup wizard to configure providers
pnpm dlx @juspay/neurolink setupThe wizard will guide you through:
- Selecting your preferred AI providers
- Validating API keys
- Setting up configuration files
Option 2: Manual Configuration
Create a .env file in your project root:
# Choose one or more providers
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
GOOGLE_AI_API_KEY=...Free Tier Options:
- Google AI Studio: Get a free API key at aistudio.google.com
- Mistral AI: Free tier available at console.mistral.ai
- Ollama: 100% free local models (requires Ollama installation)
Your First API Call (SDK)
Basic Text Generation:
import { NeuroLink } from "@juspay/neurolink";
// Initialize (auto-selects best available provider from your .env)
const neurolink = new NeuroLink();
// Generate a response
const result = await neurolink.generate({
input: { text: "Explain quantum computing in simple terms" },
});
console.log(result.content);Streaming Responses:
// Stream tokens in real-time
const stream = await neurolink.stream({
input: { text: "Write a haiku about code" },
});
for await (const chunk of stream.stream) {
if ("content" in chunk) process.stdout.write(chunk.content);
}Multimodal Input (Images + Text):
const result = await neurolink.generate({
input: {
text: "What's in this image?",
images: ["./photo.jpg"],
},
});Using Tools:
// Built-in tools are automatically available
const result = await neurolink.generate({
input: {
text: "What time is it and what files are in the current directory?",
},
// AI can call getCurrentTime and listDirectory tools
});Your First API Call (CLI)
Basic Generation:
# Simple text generation
npx @juspay/neurolink generate "Explain TypeScript generics"
# Specify provider and model
npx @juspay/neurolink generate "Hello!" --provider openai --model gpt-4o
# Stream responses
npx @juspay/neurolink stream "Write a story about AI" --provider anthropicMultimodal Input:
# Analyze images
npx @juspay/neurolink generate "Describe this image" --image photo.jpg
# Process PDFs
npx @juspay/neurolink generate "Summarize this document" --pdf report.pdf
# Combine multiple file types
npx @juspay/neurolink generate "Analyze this data" --file data.xlsx --file config.jsonInteractive Loop Mode:
# Start an interactive session with persistent context
npx @juspay/neurolink loop
# Inside loop mode:
> set provider anthropic
> set model claude-opus-4
> generate "Hello, Claude!"
> history # View conversation history
> exitCommon Use Cases
RAG (Retrieval-Augmented Generation):
// Automatically chunk, embed, and search documents
const result = await neurolink.generate({
input: { text: "What are the key features mentioned in the documentation?" },
rag: {
files: ["./docs/guide.md", "./docs/api.md"],
chunkSize: 512,
topK: 5,
},
});Structured Output with Zod:
import { z } from "zod";
const schema = z.object({
name: z.string(),
age: z.number(),
email: z.string().email(),
});
const result = await neurolink.generate({
input: {
text: "Extract user info: John Doe, 30 years old, [email protected]",
},
schema,
output: { format: "json" },
});
// Parse the structured JSON from result.content
const parsed = schema.parse(JSON.parse(result.content));
console.log(parsed); // { name: "John Doe", age: 30, email: "[email protected]" }External MCP Servers (GitHub, Slack, etc.):
// Connect to GitHub MCP server
await neurolink.addExternalMCPServer("github", {
command: "npx",
args: ["-y", "@modelcontextprotocol/server-github"],
transport: "stdio",
env: { GITHUB_TOKEN: process.env.GITHUB_TOKEN },
});
// AI can now interact with GitHub
const result = await neurolink.generate({
input: { text: 'Create an issue titled "Bug: login fails"' },
});Next Steps
- Complete Documentation - Comprehensive guides and API reference
- Provider Setup Guide - Configure a provider
- SDK API Reference - Full TypeScript API documentation
- CLI Command Reference - Complete CLI documentation
- Example Projects - Real-world integration examples
- Advanced Features - Middleware, observability, workflows
Troubleshooting
Issue: "Provider not configured"
- Run
npx @juspay/neurolink setupor add provider API key to.env
Issue: Rate limit errors
- Configure multiple providers for redundancy — NeuroLink auto-selects the best available
- Use
provider: "litellm"with LiteLLM to proxy across many providers
Issue: Large context overflows
- Enable conversation memory with compaction:
new NeuroLink({ conversationMemory: { enabled: true } }) - Use
ragoption to search documents instead of sending full content
Need help? Check our Troubleshooting Guide or open an issue.
🌟 Complete Feature Set
NeuroLink is a comprehensive AI development platform. Every feature below is shipped and documented.
🤖 AI Provider Integration
Provider neurons behind one API - Switch providers with a single parameter change. Nearly all serve generate/stream; TypeSafe Jev, Laya and XOR serve decide instead. Tool support: 31 native tool-calling, 3 model-dependent, the rest serve no tools at all (embedding-, media- and decision-only). 3 are fully local runtimes (Ollama, LM Studio, llama.cpp) and 4 need zero configuration to start (those three plus LiteLLM) — no cloud account, no API key. 9 providers (OpenAI, Google AI Studio, Google Vertex, Amazon Bedrock, Cohere, Ollama, LiteLLM, Voyage, Jina) expose embed()/embedMany() natively for RAG and custom vector search.
| Provider | Models | Free Tier | Tool Support | Status | Documentation | | --------------------- | -------------------------------------------------------------------------- | --------------- | ------------ | ------------- | ----------------------------------------------------------------------------------------------------------------------------- | | OpenAI | GPT-4o, GPT-4o-mini, o1 | ❌ | ✅ Full | ✅ Production | Setup Guide | | Anthropic | Claude 4.6 Opus/Sonnet, Claude 4.5 Opus/Sonnet/Haiku, Claude 4 Opus/Sonnet | ❌ | ✅ Full | ✅ Production | Setup Guide | Subscription Guide | | Google AI Studio | Gemini 3 Flash/Pro, Gemini 2.5 Flash/Pro | ✅ Free Tier | ✅ Full | ✅ Production | Setup Guide | | AWS Bedrock | Claude, Titan, Llama, Nova | ❌ | ✅ Full | ✅ Production | Setup Guide | | Google Vertex | Gemini 3/2.5 (gemini-3-*-preview) | ❌ | ✅ Full | ✅ Production | Setup Guide | | Azure OpenAI | GPT-4, GPT-4o, o1 | ❌ | ✅ Full | ✅ Production | Setup Guide | | LiteLLM | 100+ models unified | Varies | ✅ Full | ✅ Production | Setup Guide | | AWS SageMaker | Custom deployed models | ❌ | ✅ Full | ✅ Production | Setup Guide | | Mistral AI | Mistral Large, Small | ✅ Free Tier | ✅ Full | ✅ Production | Setup Guide
