proveml
v0.9.0
Published
Deterministic claim-level verification for AI-generated text with Markdown-native inline markup.
Maintainers
Readme
ProveML
Prototype. ProveML is in its prototyping phase, version 0.x. The markup, the verifier's API and the review page all still move between releases, and the canonicalisation contract has not been audited. Use it to try the idea and to tell us where it breaks; do not build on it yet as if it were stable.
ProveML sits at one boundary: between an AI and the person who signs what it wrote. It makes the prose checkable, claim by claim, and makes that person's sign-off checkable by everyone else. It is not a control or a gate for agents. In front of a gate, a structured certificate checked by a predicate over the same data refuses the same actions; what ProveML adds there is the readable record, nothing more.
An AI writes: "Revenue was $416 billion and the margin is healthy."
Both halves can be wrong, and nothing in the sentence tells you which. ProveML lets the model mark its claims so a machine can check them:
@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD}.
?[healthy: IS_PROFITABLE]{The margin is healthy}.Every marked claim is checked against your data by lookup and comparison: no model in the verification loop. What the data cannot support is visible as such, instead of reading like everything else.
See it in 10 seconds
npx proveml demo Apple Inc. reported revenue of 416161000000 USD
────────── ────────────────
Alphabet reported revenue of 350018000000 USD.
╌╌╌╌╌╌╌╌ ╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌╌
? company:goog not in store
5/7 claims verifiedThen check your own:
npx proveml verify --input report.md --facts facts.jsonThe three constructs
| Construct | Means | Checked by |
|---|---|---|
| @[type:id]{Name} | this text refers to that record | name equality against type:id.name |
| %[field]{value} | this value is that field | string equality against type:id.field |
| %[type:id.field]{value} | same, naming the record itself | for sentences where the subject is not the nearest entity |
| ?[label: THRESHOLD]{text} | this judgment holds | arithmetic against a registered threshold |
Anything outside a construct is ordinary prose and is never claimed to be checked. The boundary is the point.
Why a threshold registry
%[...] covers numbers, but reports also say "critically low" and
"healthy". Those are claims too, and a model will happily invent them.
So qualitative wording must name a threshold that a domain expert defined
outside the generation:
IS_LOW_PASS: { field: 'passRate', op: 'lt', value: 25, label: 'critically low' }The model may use ?[low: IS_LOW_PASS]{critically low} and nothing else. It
cannot invent the magnitude, the direction, or the cutoff.
The prompt
What a model has to be told is short, and it is measured: under a prompt that only named the constructs, Claude Opus 5 verified 86% of its claims on the first pass; with the binding rule and the cutoff rule added, 100%. The package generates that prompt from your store and registry, so the field names the model sees are the ones the verifier resolves:
import { promptFor } from 'proveml/prompt';
const system = promptFor({ store, thresholds, role: 'You write monthly investor letters.' });npx proveml prompt --facts facts.json --thresholds registry.jsonReview: the judgements no machine can make
Vera is the collaborator built on this: a skill that co-writes a report against archived sources, and a review page that puts each reading beside its proof, marks the ones that need a person, takes a yes or a no, and folds the judgements into one signed review root. The page ships in this package (src/review-page.js), the skill as vera; docs/deck.html walks through the whole chain in fourteen slides.
A hand-back can leave as a signed credential rather than a JSON file: review --await --key reviewer.jwk --issuer did:web:you.example --credential review.jwt signs the review root with the reviewer's Ed25519 key (proveml reviewer-key makes one) and writes a compact JWS a stranger can check with verifyReviewCredential and the public key alone. The root is recomputed from the payload before the signature is consulted, so a payload that does not fold to its own root fails first. No dependencies; node:crypto does the signing.
A review that sits beside an action rides in the Proof-of-Control token and inherits its anchor. A standalone review, one with no gateway to ride in, needs its own way to prove when it existed, and it gets one path, not a menu: Hedera Consensus Service. review --await --anchor anchor.json submits the review root, the output root and the credential's hash as one message on a topic (through the Hiero SDK, an LF Decentralized Trust project and an optional install; the operator account lives in ~/.config/proveml/hedera-operator.json, or --operator), then reads it back from the public mirror node and keeps only what the mirror returned: consensus timestamp, sequence number, running hash. proveml anchor-verify anchor.json rechecks that record against the mirror node with nothing installed, and proveml/anchor exposes the same as anchorReview, verifyAnchor and anchorSigner. Other ledgers stay as examples in proveml-demos.
Some links in a chain of evidence are not lookups: whether a stored value is a
fair reading of a quote, whether a report may go out. proveml/review gives
those human judgements the verifier's discipline: a judgement is saved under a
hash of exactly the content it approved (reviewId(...parts)), so when the
content changes the judgement dies with it, and summarize reports the
orphans: every checkmark that would have silently lied on a hand-kept list.
A shared browser widget (REVIEW_CSS, REVIEW_JS) turns any page into a
checklist with progress, an unjudged filter, and JSON export for committing;
a committed review can be baked back in as window.PROVEML_REVIEW_COMMITTED.
Coverage: what is not a claim
A verification rate counts only what is inside markup, so a report that marks up one number and leaves nine in plain prose would score 100%. The verifier therefore also reports coverage: every standalone number in the prose that no construct covers.
Ylan scores 5% and missed 62 days.
──── ─ ┄┄
· not a claim
2/2 claims verified, 1 number outside any claimverifyProveml always returns coverage and unmarked; with { strict: true }
(CLI: --strict) each unmarked number is a finding. Years, list markers and
numbers inside code are not counted.
Two readings. The default is for prose: a bare year, a list marker, a fiscal
form and a digit glued to letters (3BS, FY2025) are not claims. Pass
coverage: 'certificate' and the verifier reads the text as an agent's warrant
for an action: every numeral in any script counts, nothing is exempt, and the
words inside a judgement's braces are scanned too. A certificate has no
business carrying a number the verifier did not see.
Every judgement in details also reports the entity it bound to and the
condition as written, so a caller that grades facts by provenance can grade
the right record instead of guessing from position.
What it does not do
- It verifies consistency with your data, not truth. Wrong data, wrong verified claims.
- It checks what is inside markup. Prose outside is unchecked, by design.
- Values must match exactly:
18.5, not18.50and not "about 18". The reader can still see "$391.0 billion": declarerevenue._display: 'currency:USD:1'in the store and the renderer formats the verified value, keeping the canonical one on the element. Rules:grouped[:digits],compact[:digits],currency:CODE[:digits],percent[:digits], optionally prefixedlocale=nl-BE;. - Derived values (differences, counts) need to exist in the store to be claimable.
- An entity verifies when the name the reader sees equals the name at the id the model chose; nothing checks that the id is the right record. So the verifier reports whether that name is unique in the store (
subjectUnique),doctorwarns on duplicate names, and--strictmakes an ambiguous subject a finding. - If the store carries a unit (
revenue._unit), the claim must carry it too:%[revenue]{416161000000 USD}. THRESHOLD(path)may point at another entity, never at another field:IS_STRONG(student:100.absent)is unverifiable, because the registry decides which field a judgment is about.- Constructs inside fenced code blocks, code spans, or preceded by a backslash are not claims. The verifier and the markdown-it plugin agree on this; both judge through the same core.
- Threshold names are uppercase letters, digits and underscores starting with a letter (
IS_ABOVE_30). A registry key outside that shape throws at load instead of sitting there unreachable.
Fact-store format
If your data is already structured, the ProveML target format is intentionally small:
{
'company:aapl.name': 'Apple Inc.',
'company:aapl.revenue': 416161000000,
'company:aapl.revenue._unit': 'USD',
'company:aapl.netIncome': 112010000000,
'company:aapl.netIncome._unit': 'USD',
'company:aapl.eps': 7.49,
'company:aapl.eps._unit': 'USD/shares',
}Use one small mapping step to flatten your source data into stable paths like entityType:entityId.field. See the full guide in docs/fact-store.md.
Install
Most adopters only need the package itself:
npm install provemlIf you specifically want the markdown-it plugin, add markdown-it too:
npm install proveml markdown-itFor a zero-install trial path:
npx proveml examplenpx proveml doctor --facts facts.jsonnpx proveml verify --input report.md --facts facts.jsonDevelopment setup
This repo is currently set up primarily as a reference implementation repo.
git clone https://github.com/abovebeyond-ai/proveml.git
cd proveml
npm install
npm testmarkdown-it is a peer dependency for consumers and a dev dependency here so the test suite runs out of the box.
Usage
markdown-it plugin
import markdownIt from 'markdown-it';
import provemlPlugin from 'proveml';
const factStore = {
'company:aapl.name': 'Apple Inc.',
'company:aapl.revenue': 416161000000,
'company:aapl.revenue._unit': 'USD',
'company:aapl.netIncome': 112010000000,
'company:aapl.netIncome._unit': 'USD',
};
const md = markdownIt();
md.use(provemlPlugin, { factStore });
const env = {};
const html = md.render(
'@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD} with net income of %[netIncome]{112010000000 USD}.',
env
);
console.log(html);
console.log(env.proveml);standalone verifier
import { stripProveml, verifyProveml } from 'proveml/verify';
const result = verifyProveml(
'@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD} with net income of %[netIncome]{112010000000 USD}.',
{
'company:aapl.name': 'Apple Inc.',
'company:aapl.revenue': 416161000000,
'company:aapl.revenue._unit': 'USD',
'company:aapl.netIncome': 112010000000,
'company:aapl.netIncome._unit': 'USD',
},
{ snapshot: 'sec-edgar-fy2024' }
);
console.log(result.total); // 2
console.log(result.verified); // 2
console.log(result.snapshot); // sec-edgar-fy2024
console.log(stripProveml('@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD}.'));
// Apple Inc. reported revenue of 416161000000 USD.The second argument can be either:
- a plain fact-store object
- an adapter with
resolve(path) -> { found, value, unit?, trust? }
When trust metadata is present, verification details gain additive fields such as trustStatus, trustBackend, trustIssuer, and trustProofRef.
optional trust adapters
import { verifyProveml } from 'proveml/verify';
const adapter = {
resolve(path) {
if (path === 'company:aapl.name') {
return {
found: true,
value: 'Apple Inc.',
trust: { status: 'verified', backend: 'sd-jwt', issuer: 'did:example:issuer-7' }
};
}
if (path === 'company:aapl.revenue') {
return {
found: true,
value: 416161000000,
unit: 'USD',
trust: {
status: 'verified',
backend: 'sd-jwt',
issuer: 'did:example:issuer-7',
proofRef: 'jwt:sha256:abc123'
}
};
}
return { found: false };
}
};
const result = verifyProveml(
'@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD}.',
adapter
);
console.log(result.details[1].status); // verified
console.log(result.details[1].trustStatus); // verified
console.log(result.details[1].trustBackend); // sd-jwtUse this when you want ProveML to stay responsible for claim-to-fact matching while a separate trust layer authenticates where the facts came from.
tiny CLI
npx proveml strip --input report.md > plain.md
npx proveml doctor --facts facts.json
npx proveml verify --input report.md --facts facts.json
npx proveml render --input report.md --facts facts.json --css > report.html
npx proveml example --jsonThe CLI is intentionally small:
stripremoves only the ProveML syntax and keeps the visible textdoctorchecks fact-store shape before you start debugging markupverifychecks ProveML markup against a fact storerenderreturns embeddable HTMLexampleprints a copyable built-in example for quick trials
Use strip when you want plain persisted text after verification. Deciding what content should or should not be persisted at all remains application policy rather than ProveML policy.
embeddable HTML renderer
import { renderProveml, attachHover, PROVEML_CLASSNAMES } from 'proveml/render';
import 'proveml/style.css';
const factStore = {
'company:aapl.name': 'Apple Inc.',
'company:aapl.revenue': 416161000000,
'company:aapl.revenue._unit': 'USD',
'company:aapl.netIncome': 112010000000,
'company:aapl.netIncome._unit': 'USD',
};
const { html, verification } = renderProveml(
'@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD} with net income of %[netIncome]{112010000000 USD}.',
factStore,
{ showProofPaths: true }
);
document.getElementById('output').innerHTML = html;
attachHover(document.getElementById('output'));
console.log(verification.verified, verification.total);
console.log(PROVEML_CLASSNAMES.fact); // "proveml-fact"Use the plugin when you want full markdown-it integration. Use proveml/render when you want one small renderer that can be embedded directly in ordinary HTML. proveml/render is the canonical name; proveml/render-html and proveml/renderer resolve to the same renderer and will go in 1.0.
If trust metadata is present and showProofPaths is enabled, the audit proof output includes both the fact path and the trust status/backend.
Styling is optional:
proveml/renderemits stable semantic class names such asproveml-entity,proveml-fact, andproveml-inferenceproveml/renderalso emitsdata-trust-*attributes and trust classes such asproveml-trust-verifiedwhen adapters provide source-authentication metadataproveml/style.cssis the reference theme, not a required UI layer- you can override the look with ordinary CSS variables such as
--proveml-entity-color,--proveml-danger-color, and--proveml-paragraph-gap - if you need zero-build embedding,
proveml/renderalso exportsPROVEML_CSS
Package + skill
The core runtime is the npm package. For agent tooling, the smoothest setup is usually to keep proveml as the single JavaScript implementation and put a thin skill wrapper on top of it. That gives agents an ergonomic entry point without creating a second implementation surface.
Use this pattern when:
- a JavaScript agent is drafting text from structured records
- important claims should be deterministically checked before release
- qualitative language should come from registered thresholds rather than free prose
Before applying ProveML, ask the user first. For weaker models, prefer the exact wording below.
Recommended agent check-in:
I can use ProveML here to turn the important claims into deterministically checked markup against your structured data. If you want, I can do that and return a verifiable answer.
Default agent flow:
- Notice that the task involves structured records or audit-ready claims.
- Send the check-in sentence above.
- Wait for the user's approval.
- Then use
provemlas the runtime and optionally a thin skill wrapper for ergonomics.
Recommended JS-first setup:
npm install provemlimport { verifyProveml } from 'proveml/verify';
import { renderProveml } from 'proveml/render';
import { plainAdapter } from 'proveml/trust-adapter';That combination is the intended adoption path:
- npm package for the real runtime
npx proveml ...for quick trialsllms.txtfor discovery- a thin skill for agent ergonomics
- optional reference CSS when you want default rendering styles quickly
Minimal executable example:
@[company:aapl]{Apple Inc.} reported revenue of %[revenue]{416161000000 USD} with net income of %[netIncome]{112010000000 USD}.
@[company:msft]{Microsoft Corporation} reported total assets of %[assets]{619003000000 USD}.Tests
The main test command runs four suites:
- plugin verification behavior
- grammar conformance
- paper-example regression
- detection-rate regression
Run them with:
npm testOr individually:
npm run test:plugin
npm run test:grammar
npm run test:examples
npm run test:detectionCompanion research repo
The paper, reference-audit workflow, benchmarks, datasets, and experiment outputs have been split into a companion proveml-research repository so this package repo can stay small and focused.
Documentation
docs/index.html: human-friendly docsdocs/agent-reference.md: reference for LLM agentsdocs/fact-store.md: fact-store guidedocs/deck.html: the deck, fourteen slides with speaker notesllms.txt: agent discovery file- Paper, benchmarks and experiments: proveml-research
