@elastic/statistics
v0.0.2
Published
Pure statistical functions for LLM evaluations
Readme
@elastic/statistics
Pure statistical functions for TypeScript and JavaScript.
Install
npm install @elastic/statisticsUsage
Paired t-test
import { pairedT } from '@elastic/statistics';
const { statistic, pValue } = pairedT([0.9, 0.85, 0.7, 0.95, 0.8], [0.8, 0.75, 0.72, 0.88, 0.7]);
// statistic and pValue are null when fewer than two pairs are givenWilcoxon signed-rank test
Mirrors scipy.stats.wilcoxon;
p-values match scipy to floating-point precision.
import { wilcoxonSignedRank } from '@elastic/statistics';
const target = [0.9, 0.85, 0.7, 0.95, 0.8, 0.6, 0.75];
const baseline = [0.8, 0.75, 0.72, 0.88, 0.7, 0.6, 0.7];
const { statistic, pValue, zStatistic } = wilcoxonSignedRank(target, baseline);
// { statistic: 1, pValue: 0.0625, zStatistic: null }
// Is target stochastically greater than baseline? Use the normal approximation with continuity correction.
wilcoxonSignedRank(target, baseline, {
alternative: 'greater',
method: 'asymptotic',
correction: true,
});| Option | Values | Default | Notes |
| ------------- | ------------------------------------ | ------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| alternative | 'two-sided', 'less', 'greater' | 'two-sided' | Hypothesis about the differences target - baseline. |
| zeroMethod | 'wilcox', 'pratt', 'zsplit' | 'wilcox' | Discard zero differences, rank them but drop their ranks, or split their ranks between both sums. |
| method | 'auto', 'exact', 'asymptotic' | 'auto' | 'auto' uses the exact distribution up to 50 pairs without ties or zeros, an exhaustive permutation test for up to 13 pairs otherwise, and the normal approximation above that. |
| correction | boolean | false | Continuity correction for the normal approximation. |
statistic is the smaller of the positive and negative rank sums for 'two-sided', and the positive
rank sum otherwise. zStatistic is set whenever the normal approximation is used. All three are null
when the test cannot be computed (no pairs, NaN in a sample, or no variance left for the normal
approximation, e.g. every difference is zero); the function never throws for such inputs.
Differences are computed in floating point, so pairs that should tie can receive distinct ranks
because of roundoff (0.825 - 0.775 !== 0.5 - 0.45). Round scores first when exact ties matter.
Scripts
npm ci
npm run build
npm test
npm run lintSemver policy
This package follows semantic versioning:
- Patch — bug fixes that do not change the public API
- Minor — new functions or non-breaking additions
- Major — breaking changes to exported types or function signatures
Until 1.0.0, the 0.x series may introduce breaking changes in minor versions.
License
Elastic License 2.0. This may be revisited when the package is made public.
