npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@farhadarjmand/forecast-calibration

v0.1.1

Published

TypeScript binary forecast metrics: Brier score, log loss, ROC AUC, reliability diagrams, ECE and seeded day-block bootstrap intervals.

Downloads

317

Readme

Forecast Calibration

Small, deterministic diagnostics for binary probability forecasts: Brier score, log loss, tie-aware ROC AUC, reliability bins, ECE, and a seeded day-block bootstrap.

Experimental v0.1.1 · TypeScript · ESM · Node.js 22+ · MIT · zero runtime dependencies

CI npm License: MIT

API reference · Changelog · Documentation map · Report an issue

When to use this

Evaluate binary classifier probabilities, prepare reliability diagrams, and compare frozen forecasts against an independently specified constant baseline. Applicable to machine-learning evaluation and probabilistic forecasting, not only finance.

Diagnostics do not train or recalibrate a model. They cannot detect data leakage, establish profitability, or certify calibration.

Get started

Install the ESM package:

npm install @farhadarjmand/forecast-calibration

To build and test from source:

git clone https://github.com/farhad-arjmand/forecast-calibration.git
cd forecast-calibration
npm ci --ignore-scripts
npm run check
import { evaluateCalibration } from '@farhadarjmand/forecast-calibration';

const report = evaluateCalibration([
  { p: 0.8, y: 1, day: '2024-01-01' },
  { p: 0.2, y: 0, day: '2024-01-01' },
  { p: 0.6, y: 0, day: '2024-01-02' },
  { p: 0.7, y: 1, day: '2024-01-02' },
], { baselineProbability: 0.5, bootstrapResamples: 1000, seed: 7 });

console.log(report.brier, report.auc, report.brierSkillInterval95);

Package: @farhadarjmand/forecast-calibration.

Input and exclusions

Each sample has a finite numeric p in [0,1], binary numeric y, and a real UTC calendar date day formatted YYYY-MM-DD. Impossible dates such as 2024-02-30 are rejected. Dates should group samples by the forecast decision day, not by when the eventual outcome became known.

evaluateCalibration screens unknown input, counts each rejected row under its first failure (label, probability, then day), and excludes it. It never turns an unknown label into zero. Individual metric functions require valid samples and throw on malformed input.

Forecasts must have been frozen before their outcomes. This library cannot verify your training split, label availability, leakage, or baseline provenance.

Metrics

| Output | Meaning | | --- | --- | | brier | Mean (p − y)²; lower is better. | | brierSkill | 1 − Brier / constant-baseline Brier; null if baseline error is zero or the ratio is not representable as a finite number. | | logLoss | Natural-log loss, clipping only this metric to epsilon (default 1e-12). | | auc | Rank discrimination, O(n log n); ties receive half credit. Null for one class, 0.5 for constant scores when both classes exist. | | bins / ece | Approximately equal-count reliability bins without splitting tied scores; ECE depends on binning. | | decomposition | Reliability − resolution + uncertainty equals binned Brier, not necessarily raw Brier. rawMinusBinned explicitly records the residual. | | brierSkillInterval95 | Percentile interval from resampling entire decision-day blocks. |

The baseline probability is supplied explicitly; do not choose it from the evaluation labels. Endpoint probabilities 0 and 1 are allowed. Null means undefined or unavailable, not zero.

API

  • screenForecasts(raw)
  • brierScore(samples)
  • logLoss(samples, epsilon?)
  • rocAuc(samples)
  • reliabilityBins(samples, count = 10)
  • evaluateCalibration(raw, options)

Options: required baselineProbability; optional bins (1–1000), bootstrapResamples (0–10000, default 1000), unsigned 32-bit seed (default 1), and logLossEpsilon.

A bootstrap resamples the same number of day blocks with replacement, retaining every row of each sampled day. It estimates a sample-weighted statistic, not an equal-day-weighted one. Day keys are sorted before resampling. At least two blocks are required to emit an interval; two is a mathematical minimum, not a statistical adequacy claim. Invalid zero-denominator or non-finite draws are counted via requested/valid and PARTIAL status. Setting resamples to zero disables intervals.

Blocks must be representative and sufficiently independent for the intended inference. Serial dependence across days, overlapping labels, selection bias, tuning, and multiple tests require an appropriate study design outside this package. Small samples can produce misleadingly narrow or degenerate intervals.

What this does not do

It does not train a calibrator, certify probabilities, infer a trading edge, select a strategy, or promote a model. A positive skill interval does not prove calibration; discrimination and calibration are distinct. No universal sample-size threshold or automatic “calibrated” label is emitted.

Verification

Tests compare metrics against hand values and an independent pairwise AUC oracle, cover tie handling, strict date screening, raw/binned decomposition, deterministic block resampling, empty samples, and undefined baselines. Fixtures are synthetic.

See Provenance, Contributing, and MIT license.

Integration and reproducibility

ESM named imports only; tested on Node.js 22 and 24. Types are bundled. Browser and CommonJS support are not claimed. The package includes docs/API.md and llms.txt so humans and coding assistants can inspect the installed version's contract offline. A documentation map does not guarantee search ranking or AI indexing.

For contributors, npm run test:package installs a freshly packed tarball in a temporary consumer, executes the README example, checks a functional assertion and type-checks imports by the public package name. Pin the package version and retain your input identity and options when comparing results.