glove-classifier
v0.1.0
Published
Structured-decision (classifier) models for the Glove agent framework — TypeSafe's Jev, open System One models (Kev, Laya, Von, Rizzo, Decider), zero-shot label scorers, an LLM fallback, confidence-gated cascades, and agent/REPL/browser/Foundry integratio
Maintainers
Readme
glove-classifier
Structured-decision models for Glove.
A classifier model doesn't write text. It takes a state (the thing being judged) and a map of typed questions, and returns one typed answer per question, with the probability distribution behind it. TypeSafe's Jev is the flagship example: a "System One" model that answers in 70–500 ms, charges only for input tokens, and can't return a malformed answer.
Jev can't sit behind a Glove ModelAdapter, because it doesn't chat, stream
or call tools. This package gives classifier models their own adapter
contract, ClassifierAdapter, and ships:
| Export | What it is |
| --- | --- |
| jev() / typesafe() | TypeSafe System One over fetch. No SDK dependency. |
| kev() · laya() · von() · rizzo() · decider() | Open typed-decision models you host, all serving the same /v1/systemone API. |
| systemOne({ baseURL, model }) | Any other /v1/systemone-compatible server. |
| huggingfaceZeroShot() · gliclass() · labelScorer() | Zero-shot "text + labels → scores" classifiers (NLI models, GLiClass, your own). |
| llmClassifier() | Any Glove ModelAdapter answering the same typed questions. |
| cascade() | Asks a fast classifier first and re-asks only its low-confidence questions of a stronger one. |
| classifierTool() | A glove_classify tool: the agent writes its own questions. |
| defineClassifierTool() | A fixed-question tool: you write the questions, the agent supplies the input. |
| noul / choice / score | Question builders with typed answers. |
| gate() / answerConfidence() | Confidence-gated routing: act, review or escalate. |
| mountClassifier() | The full agent toolset: classify, batch, sources (data the agent classifies without reading), presets and a catalog. |
| classifierFns() | Functions for the REPLs (glove-js / -python / -lisp via the scratchpad catalog). |
| classifierEnv() | env:classifier for a working environment (glove-classifier/env). |
| withClassifier() | Adds a judge operation to a glove-execution browser adapter. |
| classifierPredicate() / classifyInbound() | Foundry inbound-transmission triage (glove-classifier/foundry). |
| classifyMany() / answersMatch() | Classify many items in parallel; filter results with where. |
pnpm add glove-classifierQuestion types
| Type | Asks | Answer |
| --- | --- | --- |
| noul | Is this statement true? | noul: the probability of yes (0–1) |
| choice | Which of these labels? | choice, probabilities per label, confidence |
| score | Where on this rubric? | score (can land between levels), probabilities per level, confidence |
The names and wire shapes match TypeSafe's API, so a request you build here can
be sent to POST /v1/systemone as it is.
Jev
import { jev, noul, choice, score } from "glove-classifier";
const model = jev(); // reads TYPESAFE_API_KEY; model "jev-latest"
const { answers, model: served, usage } = await model.classify({
state: "Help! My payouts have been failing for 3 days.",
questions: {
is_urgent: noul("Does this convey urgency?"),
department: choice("Which team should handle this?", {
billing: "Payments, invoicing, refunds",
technical: "Bugs, outages, integrations",
sales: "Pricing, upgrades, new accounts",
}),
frustration: score("How frustrated is the customer?", ["Calm", "Frustrated", "Very angry"]),
},
});
answers.department.choice; // "billing" | "technical" | "sales" (typed from the labels)
answers.department.confidence; // 0.81
answers.frustration.score; // 1.05
answers.is_urgent.noul; // 0.95
served; // "jev-1.13.0", the versioned model that answeredOptions: apiKey (defaults to TYPESAFE_API_KEY), baseURL (TYPESAFE_BASE_URL),
model (TYPESAFE_DEFAULT_MODEL, then jev-latest), timeout (per attempt,
default 10 s), maxRetries (default 2), backoffMs, headers and fetch.
Responses with status 408, 429 and 5xx (including 529 Overloaded) are retried
with exponential backoff, and Retry-After is honoured. An aborted
signal raises glove-core's AbortError. Every other failure raises a
ClassifierError whose code is one of invalid_request, auth,
rate_limited, provider, bad_response or connection.
model.listModels() returns the ids and aliases your account can use. If you
tuned confidence thresholds against a specific version, pin that versioned id
(for example jev-1.13.0), because aliases move when a new version ships.
State can be a string, or an object or array that names its parts. Keep each question atomic: split a broad judgement into several questions and combine the answers in code. All the questions are evaluated in parallel in one call, so adding more costs almost nothing.
Other classifier models
Jev defined the POST /v1/systemone contract, and open models now serve it on your own hardware. The presets below are the same client (SystemOneClassifier) with each project's documented defaults. Start the server as its README describes, then:
import { kev, laya, von, rizzo, decider, systemOne } from "glove-classifier";
const local = laya(); // http://127.0.0.1:8000
const gpu = kev({ baseURL: "http://gpu-box:8009" }); // any address
const other = systemOne({ baseURL: "https://decisions.internal", model: "my-model", apiKey });| Preset | Model | Default address | Notes |
| --- | --- | --- | --- |
| jev() | TypeSafe Jev (hosted) | https://api.typesafe.ai | Calibrated. Needs TYPESAFE_API_KEY. |
| kev() | Kev: Qwen3.5 + LoRA, 0.8B–27B | http://127.0.0.1:8009 | KEV_API_KEY if the server sets one. |
| laya() | Laya: ModernBERT / mmBERT encoders, ~400M | http://127.0.0.1:8000 | Runs on CPU. Score levels need descriptions. model: english, multilingual, typed-decisions, or auto. |
| von() | Von: ModernBERT-large, 395M | http://localhost:8000 | |
| rizzo() | Rizzo Flow: Spark-X2.5, 1.7B/4B | http://127.0.0.1:8017 | At most 26 choice labels. Uncalibrated by default. |
| decider() | Decider: Qwen-based, 0.8B–35B | http://127.0.0.1:8000 | English only, 32k context. |
Each preset reads <NAME>_BASE_URL and <NAME>_API_KEY (for example LAYA_BASE_URL), and every option can be overridden. The client tolerates the ways these servers differ:
- A missing
confidenceis computed from the probabilities. - A missing
legend,usageormodelis filled in. - Extra fields are ignored.
- Each server's option limits are checked locally, so an oversized question fails with a clear message instead of a server error.
The defaults follow each project's README as of September 2026. Pin the address and model in production.
Zero-shot label scorers. Older classifier models take a text and candidate labels and return a score per label. labelScorer() maps the three question types onto that:
- Yes/no: the question becomes the hypothesis.
- Choice: each label is scored as
label: description. - Score: each level's description is scored.
These models don't read instructions, so put the meaning in the labels.
import { huggingfaceZeroShot, gliclass, labelScorer } from "glove-classifier";
huggingfaceZeroShot({ model: "facebook/bart-large-mnli" }); // HF Inference API, HF_TOKEN
gliclass({ baseURL: "http://localhost:8000" }); // python -m gliclass.serve
labelScorer({ name: "my-setfit", score: async ({ text, labels }) => myModel(text, labels) });Every one of these is a ClassifierAdapter, so tools, REPL functions, cascades and Foundry predicates work with them unchanged. A common setup is a self-hosted model as the primary of a cascade() with Jev or an LLM as the fallback.
Measured. On the labelled inbox in examples/classifier-inbox (80 messages × 3 questions):
| Classifier | Refund | Urgent | Kind | Time per message | | --- | --: | --: | --: | --: | | Jev (hosted) | 100% | 97.5% | 100% | 18 ms | | Laya (self-hosted, 4 vCPU, no GPU) | 75% | 88.8% | 81.3% | 2.85 s |
Laya's English checkpoint reads 512 tokens, so these long emails are truncated. Its authors report about 40 ms per question on a T4 GPU.
Acting on confidence
import { gate } from "glove-classifier";
switch (gate(answers.department, { act: 0.8, review: 0.5 })) {
case "act": return route(answers.department.choice);
case "review": return confirmWithUser(answers.department.choice);
case "escalate": return handToHuman();
}answerConfidence(answer) works for all three answer types. A noul has no
confidence field, so its certainty is its distance from a coin flip,
|2p − 1|.
Any LLM as a classifier
import { createAdapter } from "glove-core/models/providers";
import { llmClassifier } from "glove-classifier";
const model = llmClassifier({
model: createAdapter({ provider: "openai", model: "gpt-4.1-mini", stream: false }),
});The LLM is asked for a probability distribution per question, returned as
JSON, and the reply is turned into the same typed answers. Replies that can't
be parsed are retried (maxAttempts, default 2). Calls are serialized, because
a ModelAdapter's system prompt is shared state. The probabilities are
self-reported, not calibrated.
Cascade
import { cascade, jev, llmClassifier } from "glove-classifier";
const model = cascade({
primary: jev(),
fallback: llmClassifier({ model: reasoningModel }),
threshold: 0.6, // re-ask answers with confidence below this; or pass shouldEscalate()
});
const result = await model.classify({ state, questions });
result.escalated; // ids the fallback answeredIn an agent: mountClassifier
import { mountClassifier, jev, llmClassifier, noul, choice } from "glove-classifier";
const classifiers = mountClassifier(glove, {
classifier: jev(),
classifiers: { careful: llmClassifier({ model: reasoningModel }) }, // the agent can pick by name
presets: {
triage: {
description: "Support triage",
questions: {
team: choice("Which team should handle this?", ["billing", "technical", "sales"]),
urgent: noul("Does the sender need help today?"),
},
},
},
sources: {
inbox: {
description: "Unread support email",
load: async () => (await mail.unread()).map((m) => ({ id: m.id, label: m.subject, state: m.body })),
},
},
});
// Presets and sources can change at any time. The agent finds them through the catalog tool.
classifiers.addSource("crm", { description: "Open CRM notes", load: loadNotes });| Tool | What it does |
| --- | --- |
| glove_classify | Judges one state, with its own questions and/or a preset. |
| glove_classify_batch | Judges many items the agent already holds. where keeps only the matches. |
| glove_classify_source | Judges every item in a host source and returns only ids, labels and answers. The content never enters the agent's context. |
| glove_classify_catalog | Lists presets, sources and named classifiers. |
A where condition is { question, choice?, min?, max? }, and a list of conditions must all hold:
- noul: matches when the yes-probability is at least
min(default 0.5). - choice: matches when
choiceis the chosen label. Withmin/max, it tests that label's probability instead. - score: matches when the score is within
minandmax.
Results are capped by limit (resultLimit, default 100); a truncated result carries a note telling the agent how to narrow or widen it. usage() returns the running totals.
For a single fixed judgement, fold one tool yourself. classifierTool() is the open tool alone, and defineClassifierTool() asks your questions over the agent's input:
glove.fold(defineClassifierTool({
name: "triage_ticket",
description: "Route a support ticket to a team.",
classifier: jev(),
questions: { team: choice("Which team?", ["billing", "technical", "sales"]) },
format: (a) => ({ team: a.team.choice }),
}));In code: REPLs, working environments and browsers
An agent that writes code can move data around without reading it. Only the program's return value enters its context. A classifier supplies the judgement that step needs, such as "which of these 400 emails ask for a refund?", and the program returns three ids.
REPLs (glove-js, glove-python, glove-lisp, over the scratchpad catalog):
import { JsSession, mountJs } from "glove-js";
import { classifierFns, jev } from "glove-classifier";
const session = JsSession.create();
session.registerAll(classifierFns(jev())); // classifier.classify / many / is / pick / rate
mountJs(glove, { session });// what the agent writes
const hits = classifier.many({
items: emails.map(e => ({ id: e.id, label: e.subject, state: e.body })),
questions: { refund: { type: "noul", instructions: "Does the sender ask for a refund?" } },
where: { question: "refund", min: 0.7 },
});
hits.map(h => h.id)REPL programs call host functions one at a time, so many runs a whole batch in parallel inside a single call. Programs get plain values: answers.refund is the yes-probability, answers.team is the chosen label, and answers.urgency is the level, so hits.filter(h => h.answers.refund > 0.5) works as written. The full typed answers are under details. Questions can be plain strings, which become yes/no questions. The five functions are:
| Function | Returns |
| --- | --- |
| classify({ state, questions }) | { answers, confidence, details }. |
| many({ items, questions, where? }) | { id, label, answers, confidence, details } per item. Items that fail carry an error. |
| is({ state, question }) | The yes-probability. |
| pick({ state, question, labels }) | { choice, confidence, probabilities } |
| rate({ state, question, levels }) | { score, confidence } |
Working environment:
import { createWorkingEnvironment } from "glove-working-environment";
import { email } from "glove-env-email";
import { classifierEnv } from "glove-classifier/env";
createWorkingEnvironment({ stdlib: [email(), classifierEnv(jev())] });
// scripts: import { many, is, pick } from 'env:classifier'The module ships a README and a classifier-triage skill under /skills.
Browser (glove-execution):
import { mountBrowser } from "glove-execution";
import { stationBrowser } from "glove-execution/station";
import { withClassifier } from "glove-classifier";
mountBrowser(glove, { adapter: withClassifier(stationBrowser({ client }), { classifier: jev() }) });
// in a workflow: const { answers } = await browser.judge({ sessionId, questions: { done: { type: "noul", instructions: "Did the order go through?" } } })judge observes the page, classifies what it sees, and returns only the answers. The DOM never comes back. Use state to trim the observation first, and maxStateChars (default 100 000) to cap it.
Foundry: triaging inbound transmissions
Each inbound event passes through its transmission's classify step and then each playbook's predicates, all before any agent starts. Putting a classifier there means agents start only for events that need them.
import { defineTransmissionPredicate } from "glove-foundry";
import { classifierPredicate, classifyInbound } from "glove-classifier/foundry";
// predicates/urgent.predicate.ts: the playbook wakes only for urgent tickets
export default defineTransmissionPredicate(classifierPredicate({
classifier: jev(),
questions: { urgent: noul("Does the sender need help today?") },
where: { question: "urgent", min: 0.7 }, // a playbook may override: predicate parameters { min: 0.9 }
state: (event: Ticket) => ({ subject: event.subject, body: event.body }),
}));
// in the transmission: resolve which event an inbound message is
inbound: {
// ...config, event, adapter
classify: classifyInbound({
classifier: jev(),
question: choice("What is this message?", ["refund", "bug", "other"]),
events: { refund: refundRequested, bug: bugReported },
fallback: generalInquiry,
minConfidence: 0.6,
state: (event: Ticket) => event.body,
}),
},Both helpers return Effects that fail with ClassifierError. Neither needs anything from glove-foundry at runtime.
Measured
examples/classifier-inbox benchmarks all of this on a labelled 80-message inbox. With gpt-4.1-mini as the agent, classifying the inbox as a source left 3,705 tokens in the agent's context, against 24,455 for reading it. Refund F1 went from 0.93 to 1.00, and the planted customer data reached the agent in 0 of 3 runs instead of 3 of 3. Jev judged 80 messages × 3 questions in 1.35 s for $0.0024, against 16.3 s and $0.028 for an LLM.
Bring your own classifier
A server that speaks /v1/systemone needs no code: use systemOne({ baseURL, model }). A model that scores labels needs one function: labelScorer({ name, score }). Anything else implements ClassifierAdapter, which has a name and a classify({ state, questions }, { signal }) method returning one answer per question id, with the same type as the question. answerFromDistribution(question, { label: p, … }) builds a well-formed answer from any probability distribution. Tools, cascades and gating then work with it unchanged.
License
MIT
