npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

ai-sdk-openai-guardrails

v0.1.0

Published

Run OpenAI Guardrails inside the Vercel AI SDK, as language model middleware. Works with any AI SDK provider.

Readme

ai-sdk-openai-guardrails

Run OpenAI Guardrails inside the Vercel AI SDK, as language model middleware. Works with any AI SDK provider, not just OpenAI.

const model = wrapLanguageModel({
  model: anthropic('claude-sonnet-4-5'),
  middleware: guardrailsMiddleware({ config, guardrailLlm: new OpenAI() }),
});

Moderation, jailbreak detection, PII redaction, secret-key leak detection, and prompt-injection detection then run on every generateText, streamText, and agent step that uses that model.

Why this package exists

@openai/guardrails ships two layers. The first is GuardrailsOpenAI, a drop-in replacement for the openai client. It cannot work with the AI SDK: AI SDK providers talk to model APIs through their own fetch and never take an openai client. GuardrailAgent is bound to @openai/agents for the same reason.

The layer underneath is a model-agnostic runtime that operates on plain text. This package plugs that runtime into the AI SDK middleware lifecycle, and does the work the drop-in client does for OpenAI's own request shapes, which nobody does for the AI SDK:

  • Map the bundle's three stages onto transformParams, wrapGenerate, and wrapStream.
  • Apply pre-flight PII redactions back into the prompt, so the provider receives masked text.
  • Convert an AI SDK prompt (tool calls and tool results included) into the conversation history that the history-aware checks read. Without this, Prompt Injection Detection has no tool traffic to inspect.
  • Run the output stage against a stream without reordering its parts.

Install

npm install ai-sdk-openai-guardrails @openai/guardrails

ai and @openai/guardrails are peer dependencies. The package itself has no runtime dependencies.

| Peer | Range | |---|---| | ai | ^6.0.0 \|\| ^7.0.0 | | @openai/guardrails | >=0.2.0 <1 |

Quick start

Create a pipeline bundle with the wizard at guardrails.openai.com, or write it by hand:

import { openai } from '@ai-sdk/openai';
import { generateText, wrapLanguageModel } from 'ai';
import OpenAI from 'openai';
import { GuardrailTripwireError, guardrailsMiddleware } from 'ai-sdk-openai-guardrails';

const model = wrapLanguageModel({
  model: openai('gpt-5'),
  middleware: guardrailsMiddleware({
    config: {
      version: 1,
      pre_flight: {
        version: 1,
        guardrails: [
          {
            name: 'Contains PII',
            config: {
              entities: ['US_SSN', 'EMAIL_ADDRESS'],
              block: false,
              detect_encoded_pii: true,
            },
          },
        ],
      },
      input: {
        version: 1,
        guardrails: [
          { name: 'Moderation', config: { categories: ['hate', 'violence'] } },
          { name: 'Jailbreak', config: { model: 'gpt-5-mini', confidence_threshold: 0.7 } },
        ],
      },
      output: {
        version: 1,
        guardrails: [{ name: 'Secret Keys', config: { threshold: 'balanced' } }],
      },
    },
    guardrailLlm: new OpenAI(),
  }),
});

try {
  const { text } = await generateText({ model, prompt: 'My SSN is 123-45-6789. Help me file.' });
  console.log(text);
} catch (error) {
  if (GuardrailTripwireError.isInstance(error)) {
    console.error(`Blocked at ${error.stage} by ${error.guardrailNames.join(', ')}`);
  } else {
    throw error;
  }
}

The model sees My SSN is <US_SSN>. Help me file. A tripwire at any stage throws GuardrailTripwireError.

config also accepts the bundle as a JSON string, so a file read or an environment variable can be passed straight in.

How the stages map

| Bundle stage | Middleware hook | Runs on | Effect | |---|---|---|---| | pre_flight | transformParams | Prompt text | Redactions are applied to the prompt before the provider call. A tripwire fails the call. | | input | transformParams | Prompt text, after redaction | A tripwire fails the call, before the provider is billed. | | output | wrapGenerate / wrapStream | Generated text | A tripwire fails the call or errors the stream. |

Unused stages are skipped. They are not instantiated and they cost nothing.

What counts as prompt text

By default the pre-flight and input stages see the last user message, matching the OpenAI drop-in client. Set inputScope: 'all-messages' to cover every user and assistant turn plus all tool output, excluding the system prompt. Tool output is a common injection vector. Use all-messages with Prompt Injection Detection.

Streaming

Output guardrails need a block of text before they can run. stream.mode chooses how that delay is handled.

| Mode | Behavior | Use when | |---|---|---| | buffer (default) | Holds every part, checks the completed text once, then releases. | No unchecked text may reach the reader. | | chunk | Checks the text accumulated so far every chunkChars characters (default 200) and releases the parts behind each passing check. | You want tokens to arrive as they are generated, and can accept that a late tripwire arrives after some text has already gone out. | | off | Passes the stream through unchecked. Input stages still run. | Output checks belong somewhere else in your stack. |

guardrailsMiddleware({ config, guardrailLlm, stream: { mode: 'chunk', chunkChars: 200 } });

In chunk mode each check runs on the whole block accumulated so far, not just the new characters, so an LLM-backed output check is called repeatedly on growing text. Prefer the local checks (Secret Keys, Contains PII, Competitors, Keyword Filter) there.

Parts are always released in arrival order, and only once every character of text ahead of them has passed, so a tool call a provider interleaves with text can never overtake that text. A tripwire errors the stream, which cancels the upstream request.

Cost of each guardrail

Names and configuration come from @openai/guardrails, not from this package. The difference that matters for cost is which checks call OpenAI.

| Guardrail | Engine | Calls OpenAI | |---|---|---| | Keyword Filter | regex | no | | Competitors | regex | no | | Secret Keys | regex | no | | URL Filter | regex plus URL parsing | no | | Contains PII | regex | no | | Moderation | Moderation API | yes | | Jailbreak | LLM | yes | | NSFW Text | LLM | yes | | Off Topic Prompts | LLM | yes | | Custom Prompt Check | LLM | yes | | Prompt Injection Detection | LLM | yes | | Hallucination Detection | Responses API file search | yes, plus a vector store |

The local checks add no network call and no latency worth measuring. The rest cost one OpenAI call per stage per model call, including when the answering model is Anthropic, Google, or a local one.

Contains PII in the JavaScript port is regex-based and needs no Presidio service, unlike the Python original.

Options

| Option | Default | What it does | |---|---|---| | config | required | Pipeline bundle, as an object or a JSON string. | | guardrailLlm | required | An openai client for the checks that call OpenAI. Typed structurally, so any openai major works. | | inputScope | 'latest-user' | 'latest-user' or 'all-messages'. See above. | | maskInput | true | Apply pre-flight redactions to the prompt. | | stream.mode | 'buffer' | 'buffer', 'chunk' or 'off'. | | stream.chunkChars | 200 | Characters of new text per check in chunk mode. | | onCheckError | 'throw' | 'throw' or 'ignore'. See below. | | onResults | none | Called once per stage that ran, tripwire or not. Awaited. | | context | none | Extra fields merged into the guardrail context, for custom checks. |

Failing closed

If a check cannot run (the Moderation API is down, or the guardrail model rejects the request), the default is to block the call and throw GuardrailCheckFailedError. A pipeline that cannot decide whether content is safe does not pass it through.

@openai/guardrails defaults to the opposite: runGuardrails swallows the failure and returns executionFailed: true with the tripwire clear. Set onCheckError: 'ignore' for that behavior, and read event.results[].executionFailed in onResults to see what happened.

Observability

guardrailsMiddleware({
  config,
  guardrailLlm,
  onResults: ({ stage, text, results, triggered }) => {
    metrics.increment('guardrails.stage', { stage, triggered: triggered.length });
    for (const result of results) {
      if (result.executionFailed) {
        logger.warn({ guardrail: result.info?.guardrail_name }, 'guardrail could not run');
      }
    }
  },
});

onResults is awaited, so an error thrown there fails the model call.

Agent loops and tool injection

The middleware runs per model call. In a multi-step agent loop that is once per step.

Tool output is checked on later steps: by step two, the first step's tool results are in the prompt, so an input-stage check sees them. Prompt Injection Detection depends on this. The middleware converts the AI SDK prompt into function_call and function_call_output conversation entries, which is the shape the check matches on.

guardrailsMiddleware({
  config: {
    version: 1,
    input: {
      version: 1,
      guardrails: [
        {
          name: 'Prompt Injection Detection',
          config: { model: 'gpt-5-mini', confidence_threshold: 0.7 },
        },
      ],
    },
  },
  guardrailLlm: new OpenAI(),
  inputScope: 'all-messages',
});

The input stage also repeats. A four-step agent run means four input-stage passes. With LLM-backed checks that is four extra OpenAI calls. You can accept that cost, run the expensive checks once at your application boundary with runGuardrails, or keep the middleware bundle to the local checks.

Errors

Both error classes carry a Symbol.for marker, so isInstance still works when duplicate copies of this package share a dependency tree. instanceof does not.

import { GuardrailCheckFailedError, GuardrailTripwireError } from 'ai-sdk-openai-guardrails';

if (GuardrailTripwireError.isInstance(error)) {
  error.stage; // 'pre_flight' | 'input' | 'output'
  error.guardrailNames; // ['Jailbreak']
  error.results; // the GuardrailResult objects that fired, with their info payloads
}

if (GuardrailCheckFailedError.isInstance(error)) {
  error.stage;
  error.guardrailName;
  error.cause; // the underlying failure
}

Helpers

Exported for anyone running checks outside the middleware, against runGuardrails directly:

import {
  extractContentText,
  extractPromptText,
  toGuardrailConversation,
} from 'ai-sdk-openai-guardrails';

extractPromptText(prompt, 'latest-user'); // string
extractContentText(result.content); // string
toGuardrailConversation(prompt); // conversation entries for the history-aware checks

AI SDK version support

One build serves ai@6 (provider spec v3) and ai@7 (spec v4). Prompt text, tool parts, the stream envelope, and the middleware hooks are the same in both specs, so the types are structural instead of imported from a pinned @ai-sdk/provider. src/compat.test.ts checks assignability both ways and runs the middleware through ai@6 wrapLanguageModel and generateText.

Limitations

  • Output checks look at text. Tool call arguments are not checked.
  • Once chunk mode has released text it cannot recall it. Only buffer mode guarantees that nothing unchecked reaches the reader.
  • Pre-flight redaction rewrites text parts and the system prompt, never tool results: those are JSON the model has to parse, and masking them would corrupt the payload.
  • Hallucination Detection needs a vector store you have already populated. This package forwards that config. It does not create a vector store.

License

MIT