npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

hitl-agent

v0.2.1

Published

LLM agent + LLM-as-Judge + Human-in-the-Loop review + learning loop, as an embeddable library

Readme

hitl-agent

Wrap an LLM agent in a confidence-gated quality pipeline. Every answer is validated by an LLM judge; low-confidence answers are escalated to a human reviewer; human corrections feed a learning loop that versions and improves the agent's policy over time. Ships with an operator dashboard you can mount into your own app.

infer → judge → route → (human review?) → respond → record → learn

Judge-before-respond is absolute: no unvalidated content ever reaches your caller. When the judge is uncertain, the request suspends and waits for a human instead of guessing.

Install

npm install hitl-agent

Provider SDKs are optional peer dependencies — install the one(s) you use:

npm install @langchain/anthropic      # Claude (default provider)
npm install @langchain/google-genai   # Gemini

The core runs with zero external infrastructure: in-memory / JSON-file stores and an in-process checkpointer. Requires Node.js >= 20. Ships ESM + CommonJS with TypeScript types.

Quick start

import { createAgent } from "hitl-agent";

const agent = createAgent({
  policy: {
    systemPrompt: "You are a helpful support agent.",
    responseFormat: { kind: "text" }, // or { kind: "options", options: [...] }
  },
  judge: { rubric: "Is the answer correct, complete, and safe?" },
  thresholds: { certainty: 0.8 },
  hitl: { timeoutMs: 30 * 60_000, onTimeout: "fallbackToRefusal" },
  validationSet: [
    { query: "What are your hours?", gradingHint: "Must mention 9-5." },
  ],
});

// Resolves after the full pipeline — including human review, which may take a while.
const res = await agent.handle("How do I reset my password?");
console.log(res.answer, res.answeredBy); // "agent" | "human" | "fallback"

model is optional. Omitted, the agent runs on Anthropic's default tiers and the judge on the agent's provider at judge tier. Everything is overridable, and agent/judge providers mix freely:

createAgent({/* ... */}); // Anthropic defaults everywhere
createAgent({ model: { provider: "google" } /* ... */ }); // Gemini agent + Gemini judge
createAgent({
  model: { provider: "google", model: "gemini-2.5-pro" }, // Gemini agent...
  judge: { rubric, model: { provider: "anthropic" } }, // ...Claude judge
});
createAgent({ model: myLangChainChatModel /* ... */ }); // any LangChain BaseChatModel

Handling responses

// Synchronous: await the whole pipeline (may block on human review).
const res = await agent.handle("...");

// Async: fire-and-subscribe, for webhook-style hosts.
const { requestId } = await agent.submit("...");
agent.onResponse((r) => {
  /* ... */
});
const res2 = await agent.waitFor(requestId);

Human-in-the-loop

When the judge's certainty falls below the threshold, the request lands in the review queue. Resolve it from the built-in dashboard, or wire your own channel (Slack, email, ...):

agent.on("review:requested", ({ data }) => notifySlack(data.item));

await agent.resolveReview(requestId, {
  decision: "override", // "approve" | "override" | "reject"
  answer: "Here's how to reset your password: ...",
  reasoning: "Required for approve/override — teaches the agent *why*.",
});

The reasoning on a human decision is threaded into the taught few-shot examples, so the agent learns not just the corrected verdict but the rationale behind it.

Learning loop

Human corrections feed an optimizer (few-shot injection by default) that proposes a new policy version. A regression gate re-evaluates the candidate against your validation set before it can be promoted — a correction that would regress known-good answers never ships.

await agent.optimizer.runNow(); // trigger a proposal now
await agent.policy.history(); // list policy versions
await agent.policy.rollback("v3"); // revert to an earlier version
await agent.validationSet.promote(exampleId, { gradingHint: "..." });

Dashboard

An operator dashboard (review queue, requests, policy history) ships inside the package. Serve it standalone or mount it into an existing app:

await agent.ui.listen(3141); // standalone server
app.use("/agent", agent.ui.middleware()); // embed (Express/Connect-style)

Everything the dashboard does goes through this same public API — there are no privileged backdoors. Enable auth by configuring credentials (ui: { enabled: true, auth: "builtin" } plus AGENTIC_HITL_USER / AGENTIC_HITL_PASS); on localhost you can run auth: "none" for a frictionless local tour.

Example — a Chuck Norris joke referee

A compact end-to-end configuration: the agent classifies whether a submission is a genuine Chuck Norris joke, a weighted-checklist judge grades classification correctness, uncertain calls escalate to a human, and each human decision teaches the agent.

import { createAgent, JsonFileStore, type JudgeCheck } from "hitl-agent";

const SYSTEM_PROMPT = [
  "You are a strict Chuck Norris joke referee. The user submits a candidate joke; you decide",
  "whether it is a genuine Chuck Norris joke: it must explicitly feature Chuck Norris and",
  "portray him with exaggerated, larger-than-life toughness, power, or invincibility.",
  "",
  "Respond with the verdict as the very first word:",
  "YES - <one sentence reason>   or   NO - <one sentence reason>",
].join("\n");

// A weighted checklist scores several independent dimensions instead of one blunt rubric.
const JUDGE_CHECKS: JudgeCheck[] = [
  {
    id: "features_chuck_norris",
    weight: 2,
    description:
      "Does it actually feature Chuck Norris with exaggerated power?",
  },
  {
    id: "verdict_matches",
    weight: 2,
    description: "Does the referee's YES/NO verdict match that assessment?",
  },
  {
    id: "reason_accurate",
    weight: 1,
    description: "Is the stated reason accurate and relevant?",
  },
];

const agent = createAgent({
  policy: {
    systemPrompt: SYSTEM_PROMPT,
    responseFormat: { kind: "options", options: ["YES", "NO"] },
  },
  model: { provider: "google" }, // or "anthropic", or a LangChain model
  judge: {
    rubric:
      "Grade whether the YES/NO classification is CORRECT, not how funny the joke is.",
    checks: JUDGE_CHECKS,
  },
  thresholds: { certainty: 0.8, hardFail: 0.4 },
  validationSet: [
    {
      query: "Chuck Norris counted to infinity. Twice.",
      expectedAnswer:
        "YES - Exaggerates Chuck Norris's power in the classic style.",
      gradingHint:
        "Only the YES/NO verdict must match; reason wording may differ.",
      mustPass: true,
    },
    {
      query: "Why did the chicken cross the road? To get to the other side.",
      expectedAnswer:
        "NO - Classic chicken joke, does not feature Chuck Norris.",
      gradingHint:
        "Only the YES/NO verdict must match; reason wording may differ.",
      mustPass: true,
    },
  ],
  learning: {
    mode: "auto",
    trigger: { everyNPromotedExamples: 1 },
    maxFewShots: 8,
  },
  store: new JsonFileStore(".hitl-agent-data/store.json"),
  ui: { enabled: true, auth: "none" },
});

const res = await agent.handle("Chuck Norris can divide by zero.");
console.log(res.answer); // "YES - ..."  (or suspends for human review if the judge is unsure)

Public API at a glance

  • createAgent(config) → Agent
  • agent.handle / agent.submit / agent.waitFor / agent.onResponse
  • agent.on(event, cb) — events include review:requested, feedback:recorded, ...
  • agent.resolveReview(requestId, decision)
  • agent.optimizer, agent.policy, agent.validationSet, agent.teachingPool
  • agent.ui.listen(port) / agent.ui.middleware()
  • Stores: MemoryStore, JsonFileStore (implement FeedbackStore for your own)
  • Audit logs: MemoryLearningAuditLog, JsonFileLearningAuditLog
  • Testing / offline: MockChatModel, seededRandom

Using this with a coding assistant

This package ships AGENTS.md — a concise, LLM-oriented reference covering the mental model, the invariants, and accurate API shapes. It installs with the package, so your assistant can read it straight out of node_modules.

Assistants don't go looking there on their own, so paste this into your project's CLAUDE.md (or AGENTS.md / .cursorrules):

## hitl-agent

When writing code against `hitl-agent`, first read
`./node_modules/hitl-agent/AGENTS.md` — it carries the pipeline's invariants
(judge-before-respond, escalate-on-failure) and the accurate config and API
shapes. `./node_modules/hitl-agent/dist/index.d.ts` is the authoritative
contract.

License

Apache-2.0