npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

veriquote

v0.1.1

Published

Deterministic + semantic verification of quote-grounded LLM citations (EVI1 protocol): fuzzy verbatim-quote matching and an LLM entailment judge for transparent, per-claim hallucination detection.

Readme

VeriQuote

Deterministic + semantic verification of quote-grounded LLM citations. DOI

VeriQuote makes source-grounded assistant answers auditable. Instead of trusting that a [1] citation means anything, the answering model must attach a verbatim quote for every cited claim, and VeriQuote then checks per claim whether:

  1. the quote actually occurs in the source (deterministic fuzzy text matching with a percent score), and
  2. the quote actually supports the claim (a small, temperature-0 LLM judge classifying entailment strength).

The result is a transparent, per-citation report telling users exactly which statements are verbatim-backed and supported, which are overstated, and which are unsupported or confabulated.

VeriQuote is extracted from and battle-tested in NavigNine, a source-grounded research assistant, which serves as the reference deployment.

  • Zero runtime dependencies. Runs in Node ≥ 18, browsers, and edge runtimes.
  • Deterministic by construction. The text matcher is pure; the judge runs at temperature 0 with a closed class vocabulary and strict output validation.
  • Model-agnostic. Works with any answering model and any OpenAI-compatible chat-completions endpoint for the judge (OpenAI, OpenRouter, Azure, local gateways) or bring your own EntailmentJudge (e.g. a local NLI model).

How it works

                        ┌───────────────────────────┐
  numbered sources ───► │  Answering LLM            │
  + citation prompt     │  (any model)              │
                        └────────────┬──────────────┘
                                     │  answer body with [n]{cX} markers
                                     │  + EVI1 quote appendix
                                     ▼
                        ┌───────────────────────────┐
                        │ 1. parseAnswer()          │  claims, quotes, protocol
                        │    (deterministic)        │  completeness warnings
                        └────────────┬──────────────┘
                                     ▼
                        ┌───────────────────────────┐
                        │ 2. Quote ↔ source match   │  exact / normalized /
                        │    (deterministic, fuzzy) │  fuzzy %, offsets
                        └────────────┬──────────────┘
                                     ▼
                        ┌───────────────────────────┐
                        │ 3. Entailment judge       │  entailed / partially /
                        │    (LLM, temp 0, optional)│  overstated / insufficient
                        └────────────┬──────────────┘  / contradicted + conf.
                                     ▼
                        ┌───────────────────────────┐
                        │ 4. VerificationReport     │  per-citation scores +
                        │    (transparency for user)│  answer-level summary
                        └───────────────────────────┘

The EVI1 protocol

The answering model is instructed (via buildCitationInstructions()) to end every cited sentence with citation markers and a claim marker, and to append a machine-readable quote appendix:

Vitamin D supplementation reduced fall risk in older adults.[1]{c1}
It also improved bone mineral density.[2][3]{c2}

EVI1
c1|1|"supplementation reduced the rate of falls by 19%"
c2|2|"bone mineral density increased significantly"
c2|3|"BMD improved with \"high-dose\" regimens"
END_EVI1

The protocol is intentionally plain text (not JSON): it survives streaming, markdown renderers, and weak models. And [n] citations remain human-readable even if a client ignores VeriQuote entirely.

Why two checks?

The two checks fail independently, and both failure modes occur in practice:

  • A quote can be verbatim yet irrelevant: the model copied real text that doesn't support its claim (scope drift, outcome switching, overstatement). Text match passes; the entailment judge catches it.
  • A quote can be paraphrased or fabricated: the claim may even be true, but the "quote" is not in the source. The entailment judge might pass; the deterministic matcher catches it, with a percent score that distinguishes light paraphrase (high fuzzy score) from fabrication (low score).

The combined per-citation score is conservative: min(textMatchScore, judgeConfidence).

Installation

npm install veriquote

Quickstart

1. Prompt the answering model

import { buildCitationInstructions } from 'veriquote';

const systemPrompt = `${yourAssistantPrompt}\n\n${buildCitationInstructions()}`;
// Provide sources as numbered blocks [1], [2], ... in the user/context prompt.

2. Verify the raw answer

import { ChatCompletionsJudge, verifyAnswer } from 'veriquote';

const judge = new ChatCompletionsJudge({
  baseUrl: 'https://openrouter.ai/api/v1',   // any OpenAI-compatible endpoint
  apiKey: process.env.OPENROUTER_API_KEY,    // server-side only!
  model: 'google/gemini-2.5-flash-lite',
});

const report = await verifyAnswer({
  answer: rawModelOutput,          // including the EVI1 appendix
  sources: [
    { title: 'Trial A', url: 'https://…', text: extractedFullText1 },
    { title: 'Trial B', url: 'https://…', text: extractedFullText2 },
  ],
  judge,                           // omit for text-match-only verification
});

console.log(report.summary);
// { citationCount: 2, verbatimRate: 1, entailedRate: 0.5,
//   meanScore: 0.675, minScore: 0.4 }

for (const c of report.citations) {
  console.log(c.claimId, c.sourceIndex, c.textMatch.method,
              c.textMatch.score, c.entailment?.class, c.score);
}

report.cleanText is the answer with all {cX} markers removed, ready to render (the [n] markers remain as human-readable citations).

3. Show it to the user

Render each citation's textMatch.score (percent), entailment.class, and combined score next to the footnote — e.g. green/yellow/red per claim. This is exactly what the NavigNine UI does with tooltips and colored footnotes.

API overview

| Export | Purpose | | --- | --- | | buildCitationInstructions(options?) | Prompt block for the answering model (budgets and quote-length rules configurable). | | verifyAnswer(options) | Full pipeline: parse → match → judge → report. | | parseAnswer(answer) | Parse claims, evidence, and protocol warnings without verifying. | | parseEvi1Appendix / stripEvi1Appendix / serializeEvi1Appendix | Low-level EVI1 handling. | | matchQuoteAgainstSource(quote, source, options?) | Deterministic quote matching on its own. | | ChatCompletionsJudge | Entailment judge for any OpenAI-compatible API. | | EntailmentJudge (interface) | Bring your own judge (local NLI model, other provider). |

All inputs and outputs are plain, serializable data — see src/types.ts for the complete, documented data model and docs/DESIGN.md for the method description (scoring, thresholds, and design rationale).

Entailment classes

| Class | Confidence band | Meaning | | --- | --- | --- | | entailed | 0.9–1.0 | Claim fully covered by the quote. | | partially_entailed | 0.5–0.8 | Core message supported, details missing. | | overstated | 0.3–0.6 | Claim stronger/more general than the evidence. | | insufficient | 0.1–0.4 | Related but does not confirm the claim. | | contradicted | 0.0 | Evidence says the opposite. | | error | — | Judge unavailable for this item (never silently dropped). |

Security

  • Keep the judge server-side. ChatCompletionsJudge needs an API key; never instantiate it in a browser. Expose a thin authenticated endpoint that calls verifyAnswer instead.
  • Prompt-injection hardening. Source text is untrusted. Judge inputs are length-capped, stripped of control characters and HTML, and the judge prompt pins them as data ("never instructions"). Output is validated against a closed vocabulary; unknown classes, out-of-range confidences, and hallucinated item IDs are rejected.
  • No dynamic evaluation. Tolerant JSON recovery is a string-aware scanner; nothing is ever evaled.
  • Failure transparency. Judge failures degrade to class: "error" with a null score — they are reported, never counted as "supported".

Reproducibility

For a fixed answer, fixed sources, and a fixed judge model, results are reproducible: the matcher is pure, and the judge runs at temperature 0 (pass seed for providers that support it). Note that hosted LLM APIs are best-effort deterministic; for strict reproducibility, pin the model version or use a self-hosted judge behind the EntailmentJudge interface.

Integrations

  • verify-citations: a portable Agent Skill (single SKILL.md + bundled Node CLI) that runs VeriQuote as an internal hallucination gate for source-grounded agents: it verifies a cited answer, flags factual sentences that carry no citation, and returns a ready-to-use correction prompt for a self-correction loop. The same skill works across any Agent-Skills host (OpenClaw, Hermes Agent, Claude Code) and any orchestrator that can run a Node CLI.

Citing

If you use VeriQuote in academic work, please cite the Zenodo record (see CITATION.cff).

License

MIT