npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@yadsh/dsh-model-safety-gate

v0.2.6

Published

Independent two-layer safety gate around the DeepSeek Harness agent loop: deterministic and model-classifier verdicts for prompts, streamed output, tools, and tool results

Readme

@yadsh/dsh-model-safety-gate

Independent defense-in-depth safety gate for DeepSeek Harness: a deterministic scanner plus an isolated small-model classifier check user prompts, streamed model output (text and reasoning), tool calls, and tool results before they reach the model, the user, or the execution layer.

The gate is an additional decision layer. It does not replace the DSH sandbox, the permission system, or approval gates, and it never patches DSH core.

Features

  • Input guard (agent/pre-step): prompts are normalized and scanned before the main-model request; matched prompts can be warned about or blocked outright, with a sanitized reason published to the session.
  • Output stream guard (llm/stream): text-delta and reasoning-delta channels are checked with rolling windows. In the default buffered mode chunks are quarantined until their window passes, so blocked content never reaches the UI or session; blocking aborts the upstream provider request.
  • Tool gate (tools/pre-execute): fully assembled tool calls are checked before execution — allow, defer to the native DSH approval flow, or deny.
  • Indirect-injection guard (tools/post-execute): tool results from untrusted sources are scanned; hits raise the turn risk state, which tightens decisions for subsequent sensitive tool calls.
  • Two classification layers: L0 is a fast local scanner (injection patterns, jailbreak markers, secrets/credential formats, Unicode obfuscation, zero-width characters, suspicious encoding, repeated-payload floods, user patterns). L1 is an isolated safety classifier on a small model — a DSH provider/model pair, an OpenAI-compatible endpoint, or off.
  • Monotonic safety merge: deterministic red lines cannot be weakened by the classifier; verdicts only escalate.
  • Failure modes: classifier timeout or malformed output follows the configured failure mode — closed, open, rules-only (default), ask.
  • Safety ≠ usefulness: quality verdicts (unclear, spam, low-information) warn by default and never block unless explicitly opted in.
  • Sanitized audit: plugin logs and counters record decisions with content hashes, never raw blocked content or secret values (raw logging is opt-in). Audit records include the session id but never enter the Harness session journal.
  • Classifier isolation: classifier calls run under a process-local bypass marker, so moderating a generation never recursively moderates the moderator; the classifier has no tools.

Install

Install the published npm package by name:

dsh plugin --profile web add @yadsh/dsh-model-safety-gate

Configuration

All options are optional; defaults are shown.

enabled: true # master switch: false silences every surface at once
mode:
  warn # off | audit | warn | enforce — default decision profile
  # off scans nothing; audit records findings but never enforces, including
  # turn-risk escalation

classifier:
  backend: none # none | dsh | openai-compatible
  provider: "" # backend: dsh — DSH provider id for the classifier model
  model: "" # backend: dsh — model id
  baseURL: "" # backend: openai-compatible — endpoint base URL
  apiKey: "" # backend: openai-compatible — API key (kept out of logs)
  timeoutMs: 3000 # classifier request timeout
  maxTokens: 128 # bounded structured response
  temperature: 0
  failureMode: rules-only # closed | open | rules-only | ask — on timeout/error/malformed
  requireLocal: false # true forbids remote (openai-compatible) endpoints

input:
  enabled: true # gate user prompts on agent/pre-step
  safetyAction: block # allow | warn | block for safety verdicts
  qualityAction: warn # allow | warn | block for quality-only verdicts (block = opt-in)

output:
  enabled: true # gate main-model streaming output
  mode: buffered # observe | interrupt | buffered
  text: true # check the visible-answer channel
  reasoning: true # check the reasoning channel when the provider streams it
  checkEveryChars: 512 # new quarantined chars between classifier snapshots
  windowChars: 1536 # snapshot window size sent to the classifier
  lookbehindChars: 768 # preceding context included with each window
  minCheckIntervalMs: 250
  maxBufferedChars: 8192 # overflow fails closed in buffered mode

tools:
  enabled: true # gate tool calls on tools/pre-execute
  semanticClassifier: true
  unanswerableAsk: deny # deny | ask — an escalation the session's approval policy refuses before asking anyone

toolResults:
  enabled: true # scan tool results on tools/post-execute
  classifyUntrustedSources: true

audit:
  enabled: true
  includeRawContent: false # opt-in raw content logging (default: hashes only)

ui:
  enabled: true
  showWarnings: true

allowSessionOverride: true # false forbids per-session downgrade of the global mode

Switching the gate off

enabled: false and mode: off are the same decision spelled twice, and both mean the gate does nothing at all: no scan, no classifier call, no audit record, no blocked decision — on the input, streaming-output, tool-call and tool-result surfaces alike. Switching either one at runtime reaches the very next check; there is nothing to restart.

Use mode: off when the profile is what changes between environments and enabled: false when the plugin itself should be inert; audit is the middle setting for a deployment that wants the findings without the enforcement.

Escalations the deployment cannot answer

A tool call the gate escalates is resolved by DSH, not by the gate: the call becomes an ask decision, the tool runtime passes it to the approval service, and the outcome carries no reason — the runtime turns it into its own sentence. So the refusal a model sees is the user rejected tool "X", even when nobody was asked, and the rule that actually fired is lost.

That is exactly what a session whose effective approval policy is never produces: the service returns the refusal before dispatching to any answerer, deterministically. A QA lockdown pins that policy, so every escalation there reached the model as a human "no".

tools.unanswerableAsk: deny (default) refuses such an escalation with the gate's own verdict instead, categories included:

Blocked by dsh-model-safety-gate (unsafe_tool_intent): this call needs
confirmation, but the session's approval policy is "never", so the request
could only ever be refused without asking anyone

The gate reads the policy the same way the approval service does — the session's logged override first, else the deployment default — and only refuses when that read completes: a host composing no approval service, or a policy the gate cannot read, keeps the native ask, which the runtime still resolves through the real seam. Set ask for a deployment whose own gate answers asks ahead of that policy (for example a QA surface with interaction.approvals: interactive, where the operator decides).

Deployment presets

| Profile | Input | Output | Tools | Failure mode | | -------- | ----- | --------------------------- | ----- | --------------------------------- | | Personal | warn | interrupt | ask | rules-only | | Balanced | block | buffered | ask | rules-only | | Strict | block | buffered (text + reasoning) | block | closed, session override disabled |

Privacy

If the classifier backend is openai-compatible, prompts, streamed output, and reasoning content are sent to that endpoint. The settings card states this in place, next to the endpoint it would use; set classifier.requireLocal: true to forbid remote endpoints entirely.

classifier.apiKey is declared a secret slot: configuration surfaces receive only whether a key is configured, and the literal is never returned to a browser. The safetyGate Remote projection redacts it as well.

Settings card

Settings → Plugins → Plugin configuration → Model Safety Gate edits this plugin's configuration through the model-safety-gate settings namespace. The namespace is installed as the gate's configuration source, so every change re-resolves the running gate on the spot: turning the gate off stops the next check, and switching the mode to enforce blocks the next matching prompt. Values the schema cannot express (a dsh backend without a provider, an uncompilable customBlockPatterns entry) are refused at write time instead of being stored and ignored.

The card also reports what the running gate is doing through the safetyGate Typert Remote — effective mode, whether the classifier is actually wired, the process-lifetime counters, and the last 50 sanitized verdicts. That view never carries the classifier key, and it carries content only when raw logging is switched on.

Chat moderation banners and a per-session shield control are not built yet (design SPEC Phase 6); ui.* and allowSessionOverride are accepted keys with no effect today.

Harness API

ctx.safetyGate.inspect(); // effective config, metrics, recent verdicts

The plugin registers one Cordis service (safetyGate) and publishes four stable plugin-log categories (safety-gate/check|block|warn|classifier-error).

Compatibility

  • DeepSeek Harness >=0.1.5-rc.2 <0.2.0 (channel next), extension points: agent/pre-step, llm/stream, tools/pre-execute, tools/post-execute.
  • Node.js ^22.19.0 || >=24.0.0.
  • See compatibility.json for the machine-readable manifest.

Development

pnpm nx run dsh-model-safety-gate:lint
pnpm nx run dsh-model-safety-gate:typecheck
pnpm nx run dsh-model-safety-gate:test
pnpm nx run dsh-model-safety-gate:build
pnpm nx run dsh-model-safety-gate:verify

License

MIT — see LICENSE. Architecture credits are listed in NOTICE.md.