nexus-ai-pro
v1.25.0
Published
Typed AI framework for TypeScript: provider routing and failover, durable agent graphs and workflows, retrieval over any vector store, MCP, evals, guardrails, and cost control
Downloads
3,074
Maintainers
Keywords
Readme
nexus-ai-pro
The typed AI framework for TypeScript. One API for provider routing and failover, agent graphs and workflows that run branches in parallel and survive a restart, retrieval over any vector store, MCP, durable background operations, guardrails, cost control, images, voice, and evals — with a studio a team can share.
Import the whole runtime, or one piece: every capability has its own entry point with a size budget CI enforces, and the packaging guide publishes what each one costs. A graph-only application loads 49 KB and installs no third-party package at all.
- NPM: https://www.npmjs.com/package/nexus-ai-pro
- GitHub: https://github.com/mkhitar-abrahamyan/nexus-ai
- Contributing, testing, and release procedure: CONTRIBUTING.md
- Delivered work and what is planned next: ROADMAP.md
Install
npm install nexus-ai-proRequires Node.js 22 or newer. The package is authored as ESM and ships both an ESM and a CommonJS build,
so import and require() both work — including from NestJS and other apps compiled with
"module": "commonjs". Types are shared between both builds.
The main Nexus pipeline is Node-oriented; the isolated realtime subpaths use portable structural
interfaces and can be bundled for modern browsers.
Install only the provider SDKs your app uses. The OpenAI SDK also powers OpenAI-compatible adapters such as OpenRouter, Groq, Mistral, DeepSeek, Azure OpenAI, LM Studio, and llama.cpp.
npm install openai @anthropic-ai/sdk ollamaThe Google Gemini provider uses Node.js fetch, so it does not need an extra SDK.
Realtime transports use injected or platform WebRTC, WebSocket, and fetch interfaces and do not
require the OpenAI SDK.
Quick Start
import { createNexus } from 'nexus-ai-pro';
const ai = createNexus({
provider: 'openai',
apiKey: process.env.OPENAI_API_KEY!,
model: 'gpt-5.4-mini',
security: 'standard',
});
const response = await ai.complete({
model: 'auto',
messages: [{ role: 'user', content: 'Explain RAG in 3 bullet points.' }],
});
console.log(response.content);
console.log(response.meta.providerUsed);
console.log(response.meta.modelUsed);For a fuller typed config, use the builder:
import { createNexusConfig } from 'nexus-ai-pro';
const ai = createNexusConfig()
.openai(process.env.OPENAI_API_KEY!)
.anthropic(process.env.ANTHROPIC_API_KEY!)
.deepseek(process.env.DEEPSEEK_API_KEY!)
.lmstudio({ baseUrl: 'http://localhost:1234/v1' })
.auto('quality')
.security('standard')
.retry({ enabled: true, maxRetries: 2 })
.create();You can still pass a plain NexusAIConfig to new NexusAI(...) when you want full object-literal control.
Why Use It
nexus-ai-pro is useful when an app needs more than a direct SDK call:
- route across multiple providers and model aliases
- fail over on provider errors, timeouts, or unhealthy providers
- stream through one normalized interface
- protect inputs and outputs with guardrails
- compact long conversations before they hit model limits
- estimate tokens and cost before provider calls
- cache exact or semantically similar prompts
- add tools, agents, RAG context, evals, batch jobs, and queues
- embed text through the same routing, caching, batching, budget, retry, and metrics as a completion
- run long operations that survive a restart, with leases, retries, dead-lettering, and signed webhooks
- build stateful graphs with cycles, fan-out, subgraphs, human approval, and resumable checkpoints
- make ordinary async functions durable step by step, and recover a crashed run at the step it died on
- load files, web pages, sitemaps, and Git repositories into retrieval, over memory, Postgres, SQLite, Redis, Qdrant, Pinecone, Weaviate, or Chroma
- retrieve with hybrid keyword and vector search, reranking, and diversity, and measure it with evals
- reach tools from many MCP servers under one configuration, with allowlists and credentials from the environment
- version prompts, instructions, tools, and skills together, and promote them through evaluation gates
- find failing and slow runs, detect regressions, and review evaluated fix proposals in a shared studio
- serve assistants as revisions, with canaries that roll back on a regression, workers that scale on the queue under Kubernetes, and per-tenant quotas, rate limits, and budgets
- trip routing away from a failing provider, and share one rate-limit budget across workers
- reach the providers' half-price asynchronous batch tier behind one operation handle
- persist generated media to disk or S3 with tenant isolation, retention, and checksums
- build persistent realtime voice agents with interruption, live tools, and normalized conversation state
- keep TypeScript types around every request and response
Guides
Each feature has its own guide, covering every export of its entry points with a reference generated from the doc comments. The guides live in the repository, so these links go to GitHub.
| Guide | Covers |
| --- | --- |
| The client | Requests, context windows, token optimization, cost checks, reasoning, prompt caching, streaming |
| Providers and routing | The twelve completion providers, routing and failover, the model registry |
| Guardrails | Injection, PII, secrets, schema validation, output redaction |
| Graphs | Parallel branches, interrupts, checkpoints, time travel, subgraphs, caching, diagrams |
| Agents and tools | Graph-based agents with approvals and middleware, and the simple tool loop |
| Long-term memory | Namespaced memory with semantic search, in memory or Redis |
| MCP | Borrowing tools from MCP servers, a registry of many, and serving your own |
| Traces | Run trees, queries, feedback, comparison, and alerts |
| Evaluation | Datasets, evaluators, experiments, comparisons with a verdict, review queues, LLM judges |
| Context hub | Prompts, instructions, tools, and skills versioned together, promoted through gates, and moved between projects |
| Insights | Failing and slow runs clustered, regressions between time windows, and evaluated fix proposals |
| Prompts | Typed templates, content versions, gated promotion, rollback, A/B splits, serving through outages |
| Postgres | One adapter family for operations, memory, traces, evaluation, circuits, and prompts |
| SQLite | Durable operations, checkpoints, and memory on one machine, over any SQLite driver |
| Durable operations and jobs | Operations that survive a restart, webhooks, and in-process job helpers |
| Resilience and observability | Circuit breaking, distributed rate limits, metrics, logs, and audit |
| Batch tiers | The providers' discounted batch APIs, resumable from any process |
| Caching | Exact and semantic response caches, with memory, Redis, and SQLite adapters |
| Grounding | Retrieval, citations, verification, self-consistency, and knowledge graphs |
| Loaders | Files, Markdown, HTML, CSV, JSON, PDF, web pages, sitemaps, and Git repositories into retrieval |
| Retrieval | Hybrid search, reranking, MMR, parent documents, and Redis, Pinecone, Weaviate, and Chroma stores |
| Embeddings and retrieval | Embeddings as a routed operation family, and RAG helpers |
| Images (experimental) | Generation and masked edits over three backends, input safety, moderation, asset stores |
| Voice | Transcription, speech, voice turns, and voice sessions |
| Realtime voice | Browser and server realtime sessions with barge-in, tools, and exports |
| Telephony | Calls, webhooks, phone numbers, and phone agents on realtime sessions |
| Testing | Recording and replaying provider traffic, and conformance suites |
| Studio | A shared UI for traces, threads, approvals, experiments, prompts, deployments, costs, and health, as a separate package |
| Agent server | Self-hosted HTTP server for assistants: threads, durable runs, resumable streams, cron, a worker queue, and metrics |
| Deployments | Revisions and canaries with automatic rollback, autoscaling on Kubernetes and Helm, and per-tenant limits |
| Workflows | Ready-made chains and domain workflows |
| Command line | The nexus command |
| Packaging and install weight | Every entry point and what it costs |
Examples
npm run example:minimal
npm run example:create-nexus
npm run example:config-builder
npm run example:custom-provider
npm run example:feature-flags
npm run example:basic
npm run example:security
npm run example:optimizer
npm run example:agent
npm run build && npx tsx examples/realtime-scheduling.ts
npx tsc -p tsconfig.examples.browser.jsonRepository example files cover:
- OpenTelemetry with Node HTTP and Express
- BullMQ workers
- OCR/PDF ingestion
- persisted vector stores
- classifier calibration
- domain workflows
- custom providers
- deterministic realtime scheduling with tools, confirmation, metrics, and conversation exports
- import examples under
examples/exports
Test Commands
npm install
npm run check
npm run check:releasecheck runs formatting and lint gates, source/test/example type checks, the build, unit and mock
conformance tests, coverage thresholds, package import checks, and an external type-consumer test.
check:release additionally verifies the dry-run tarball and a clean packed-package install.
npm run docs:check fails when any public export, or any public member of an exported class,
interface, or enum, has no doc comment, and npm run docs:guides fails when a guide does not name
every export of the entry points it covers, or its generated reference is stale. Both hold at 100%;
--list names what is missing, and npm run docs:update regenerates the references.
Real provider conformance, and the vector store contract against real servers, are opt-in:
npm run test:conformance:real
npm run test:vectors:liveRoadmap
See ROADMAP.md for what has shipped and what is planned. It is a design proposal, not a compatibility promise; the guarantees live in API_STABILITY.md.
Shipped so far:
- completion controls and capability negotiation;
- durable operations, provider batch tiers, and distributed resilience, now including shared circuit state and a Postgres adapter family;
- embeddings, graphs, long-term memory, agents on graphs, and MCP;
- queryable traces, an evaluation platform, and a CI gate on top of it;
- record and replay of provider traffic;
- prompt versioning: typed templates, content versions, gated promotion, and serving through outages;
- a self-hosted agent server, and a local studio as a separate package;
- per-entry-point size budgets;
- vector stores behind one contract, and deprecation warnings at run time;
- durable functional workflows, step-level recovery in the server, SQLite persistence, and diagrams as images;
- document loaders, five more vector stores, hybrid retrieval with reranking, and an MCP registry;
- a shared studio with accounts, roles, an audit log, and comments; a context hub; insights with evaluated fix proposals; and evaluation caching;
- self-managed deployment at scale: revisions and canaries with automatic rollback, a worker queue that autoscales on Kubernetes, and per-tenant limits;
- the bridge to 2.0: every capability on a subpath of its own, everything 2.0 removes deprecated, and a codemod that moves your imports.
Next: 2.0.0 consolidates. It slims the root import to the core client; npx nexus migrate src --write
moves your imports today, and MIGRATING.md
covers the rest. The image family leaves experimental once recorded live
conformance passes on all three backends.
Before Production
Each guide ends with its limitations. The ones worth reading first: the client guide on latency, traces, and shared adapters; the security guide on what guardrails do not replace; and the grounding guide on embeddings and verification. Voice, telephony, images, realtime, local models, and custom providers are optional layers; enable only what you use.
Donations
- USDT Tron:
TNbS2ub2Wys6j8yrv57bWg3Ke21ZNwt115
