llama-system1
v1.1.0
Published
Low latency Llama.cpp '/v1/systemone' shim
Readme
llama-system1
Llama.cpp prompt template hack to perform fast one-shot single output token classification.
Usage
As a fully qualified /v1/systemone frontend:
$ npx llama-system1 8080 https://127.0.0.1:9931
Listening on http://127.0.0.1:8080/v1/systemone -> http://127.0.0.1:9931Programmatic api
import { infer } from 'llama-system1'
const input = {
state: 'There are 5 bananas on the table.',
questions: {
easy: {
instructions: 'Were fruits mentioned?',
type: 'noul'
}
}
}
const output = await infer(null, input)
console.log(output)Used on qwen3.5:9b produces output:
{
answers: { easy: { type: 'noul', noul: 0.9999997270783773 } },
usage: { input_tokens: 96, output_tokens: 2 }
}Performance
Measured locally against the 231 public JevBench decisions
| Metric | Result | | --- | ---: | | Accuracy | 69.70% (161 / 231) | | Brier score | 0.5656 | | Expected calibration error | 0.2599 | | Schema validity | 100% (231 / 231) | | Median latency | 48.1 ms | | p95 latency | 297.3 ms | | Ordinal MAE | 0.3113 |
License
ISC
All wrongs reversed - 2026 - Tony Ivanov
