@wlearn/core
v0.3.2
Published
Runtime core for wlearn: matrix helpers, bundle format, registry, pipeline
Maintainers
Readme
@wlearn/core
Runtime core for wlearn: matrix helpers, bundle format, model registry, pipeline, scalers, metrics, and cross-validation. No WASM. No heavy dependencies.
Part of wlearn (GitHub, all packages).
Install
npm install @wlearn/core @wlearn/liblinear@wlearn/core contains no model backend. @wlearn/liblinear is included above
because the runnable example uses it.
Quick start
const { readFileSync, writeFileSync } = require('fs')
const { Pipeline, load, accuracy } = require('@wlearn/core')
const { LinearModel } = require('@wlearn/liblinear')
async function main() {
const X = [[-2, -2], [-1, -1], [1, 1], [2, 2]]
const y = new Int32Array([0, 0, 1, 1])
const XTest = [[-1.5, -1.5], [1.5, 1.5]]
const yTest = new Int32Array([0, 1])
const model = await LinearModel.create({ task: 'classification' })
const pipe = new Pipeline([['clf', model]])
pipe.fit(X, y)
const preds = pipe.predict(XTest)
console.log('accuracy:', accuracy(yTest, preds)) // 1
writeFileSync('pipeline.wlrn', pipe.save())
// Importing @wlearn/liblinear above registered its bundle loaders.
const restored = await load(readFileSync('pipeline.wlrn'))
console.log(Array.from(restored.predict(XTest))) // [0, 1]
}
main().catch(error => {
console.error(error)
process.exitCode = 1
})API
Matrix utilities
Convert user input to typed arrays for WASM consumption.
normalizeX(X)--number[][] | DenseMatrixto contiguousDenseMatrixnormalizeY(y)-- preservesInt32Array,Float32Array, andFloat64Array; convertsnumber[]toFloat64ArraymakeDense(data, rows, cols)-- createDenseMatrixfrom typed arrayvalidateMatrix(m)-- validate matrix structure and dimensions
Bundle format
Portable binary format for model artifacts. Language-agnostic, deterministic.
encodeBundle(manifest, artifacts)-- encode toUint8ArraydecodeBundle(bytes)-- decode to{ manifest, toc, blobs }validateBundle(bytes)-- decode + verify SHA-256 hashes
Artifacts are { id, mediaType?, data: Uint8Array } objects. The manifest includes typeId, params, and metadata. See @wlearn/types for full shapes.
Canonical v1 writers always emit bundleVersion, requires, params, and an
artifacts declaration matching the TOC. Each TOC record has exactly id,
offset, length, sha256, and mediaType; each artifact declaration has the
same fields except offset. Nested WLRN artifacts must themselves be canonical.
Use validateBundle(bytes, { allowLegacyManifest: false }) to prove canonical
conformance.
The default reader is a compatibility-ingestion path for historical v1 bundles.
It additionally permits missing requires, params, artifacts, and TOC
mediaType, non-canonical TOC ordering, and extensions on fixed TOC/declaration
records. It still enforces the v1 header and typeId, bounded UTF-8/portable JSON,
exact blob coverage, hashes, recursion limits, and every field that is present.
Because historical bundles may omit requires, they cannot guarantee complete
nested-loader preflight. Writers never emit this legacy shape. Legacy reads remain
supported through the current major release; any removal requires a major release,
a migration tool, and advance deprecation notice.
Registry
Global loader dispatcher. Model packages register themselves on import. A fresh
process must import the matching package before calling generic load(); core
does not eagerly import optional backends.
register(typeId, loaderFn, { acceptsContext?, sync? })-- register a deserializer; both flags default tofalseload(bytes)-- async: decode bundle, dispatch to registered loaderloadSync(bytes)-- sync variant, only for loaders explicitly registered withsync: truegetRegistry()-- inspect registered loadersassertRequiredLoaders(manifest)-- preflight all declared nested loaders
load(bytes, { loaderOptions }) snapshots plain nested objects and arrays before
the first asynchronous loader runs. Every recursive child receives the same
read-only snapshot, so caller mutation during an await cannot change load
policy. Runtime identity objects such as typed cancellation flags are preserved.
Pipeline
Sequential composition of transformers and estimators.
new Pipeline(steps)--stepsis[name, estimator][]pipe.fit(X, y)-- fit all steps in orderpipe.predict(X)-- transform + predictpipe.classes-- fitted class order forwarded by the final classifierpipe.capabilities-- final-estimator capability descriptorpipe.score(X, y)-- transform + scorepipe.save()/Pipeline.load(bytes)-- serialize/deserialize WLRN bytespipe.dispose()-- deterministic cleanup for long-running loops
pipe.setParams({ stepName: params }) invalidates the fitted pipeline before
mutating the first selected child. Call fit() again before inference, including
when a child parameter setter throws partway through the update.
Unknown step names are rejected, and every selected child is checked for a
callable parameter setter before any fitted state or child configuration changes.
Base estimators still fit synchronously. If a Pipeline contains an orchestration
composite with asynchronous fit (for example a voting or stacking ensemble),
pipe.fit() Promise-lifts that operation and becomes fitted only after it resolves;
use await pipe.fit(X, y) for code that accepts either kind of child.
Concurrent fit(), setParams(), and dispose() calls are rejected while an
asynchronous Pipeline fit is pending, preventing late completion from resurrecting
or leaking a disposed composite.
The package publishes index.d.ts; npm run test:types checks the public
TypeScript surface against representative registry, bundle, scaler, and pipeline
usage.
The current Pipeline executes a sequential list of steps. TensorRef and graph interfaces describe future routing; they are not the current Pipeline executor. Matrix normalization and CV require dense inputs; sparse-capable model APIs must handle CSR through their declared capability.
Pipeline routes predictQuantiles, predictInterval, predictSet,
predictRegion and predictDistribution through its fitted transforms to a
capable final estimator. Arguments and asynchronous results are preserved.
Measures can require structured prediction fields; metadata.predictionArgs
supplies positional arguments after X when scoreEstimator invokes that method.
The core provides validation and routing; numerical uncertainty methods live in
an optional package.
Preprocessing
StandardScaler-- zero mean and population variance (ddof=0)MinMaxScaler-- scale the fitted range to [0, 1]
Core 0.3 removes the legacy state-only Preprocessor. Import it from
@wlearn/preprocess and replace new Preprocessor(config) with
await Preprocessor.create(config). The adapter wraps Tranfi prepared transforms
and implements WLRN save()/load registration. Refit legacy getState() data
from its original training inputs; that state was not a portable WLRN artifact.
New scaler artifacts use the corrected standard_scaler@2 and
minmax_scaler@2 contracts. Both runtimes retain @1 loaders so existing
constant-column artifacts preserve their historical inference behavior.
Metrics
Classification: accuracy, confusionMatrix, precisionScore, recallScore, f1Score, logLoss, rocAuc
Regression: r2Score, meanSquaredError, meanAbsoluteError
Metrics accept sampleWeight / sample_weight where defined. Classification metrics support binary, micro, macro, and weighted averaging. rocAuc supports binary scores plus multiclass multiClass: 'ovr' | 'ovo'; undefined metrics can throw, warn, or return NaN.
const { accuracy, f1Score, r2Score } = require('@wlearn/core')
accuracy(yTrue, yPred) // number
f1Score(yTrue, yPred, { average: 'macro' }) // number
r2Score(yTrue, yPred) // numberCross-validation
kFold(n, k?, opts?)-- k-fold split indicesstratifiedKFold(y, k?, opts?)-- stratified k-foldtrainTestSplit(n, opts?)-- single train/test splitcrossValScore(ModelClass, X, y, opts?)-- evaluate with CVgetScorer(name)-- get scoring function by name ('accuracy','r2','neg_mse')
cv accepts a fold count, explicit { train, test } row-index arrays, or a
ResamplingPlan. resolveCv checks indices and train/test separation. AutoML,
OOF, stacking, and bagging use these same folds; OOF consumers require each row
in test folds exactly once. Bagging repeats may reuse an explicit complete plan.
Temporal/index/period split generators are experimental, outside the stable
CV contract. Their partial coverage is supported for evaluation, not materialized
OOF training. This does not add time-series modeling or uncertainty estimation.
getScorer resolves the Measure registry, including log_loss, roc_auc, mse,
and mae. Measure definitions declare their required response and whether to
minimize or maximize. Plain two-array callable scorers keep response predictions
and maximization. scoreEstimator requests the correct prediction method and
validates a finite score. Probability rows must contain finite values in [0, 1]
and sum to one within 1e-6; declared classes must be unique.
Ecosystem primitives
These are the stable objects for apps, AutoML, benchmarks, and agents. They are optional. If you only want to fit, predict, score, and save a model, use the estimator and pipeline APIs above.
createTask()/validateTask()-- dataset, feature schema, labels, groups, row rolescreatePrediction()/validatePrediction()-- responses, probabilities, class order, truthlistMeasures()/evaluateMetricSet()-- discover and score metrics, including weighted and multiclass measurescreateResamplingPlan()-- deterministic holdout, k-fold, stratified, group, time-series, and sliding row/index/period splitsArchive-- candidate records, scores, timings, errors, and archive-level leaderboards
const { createPrediction, createResamplingPlan, evaluateMetricSet, Archive } = require('@wlearn/core')
const pred = createPrediction({ truth: yTest, response, proba, classes })
const scores = evaluateMetricSet(['accuracy', 'log_loss', 'roc_auc_ovr'], pred)
const plan = createResamplingPlan({ strategy: 'sliding_period', index: dates, period: 'day', lookback: 14 })
const archive = new Archive({ measures: ['accuracy'], primaryMeasure: 'accuracy' })
archive.add({ trialId: 'xgb-0', candidateId: 'xgb-depth6', scores, status: 'ok' })
archive.leaderboard()Structured predictions and multiple targets
createPrediction validates flat row-major arrays with explicit metadata:
| Field | Layout | Required metadata |
| --- | --- | --- |
| quantiles | row, target, level | quantileLevels |
| interval | row, target, coverage, lower/upper | coverageLevels |
| sets (classification) | row, coverage, class | coverageLevels, classes |
| sets (multilabel) | row, coverage, label, state | coverageLevels, targetCount |
| samples | row, draw, target | sampleCount, sampleKind: 'outcome' or 'mean' |
Samples may declare sampleDependence: 'joint' when columns in each draw belong
together, or 'marginal' when no joint interpretation is supplied. Array shape
alone does not establish dependence. Optional sampleWeights are nonnegative
relative weights for draws, shared across rows; omission means uniform weights.
Numerical consumers normalize them. Posterior-mean samples do not represent
future outcomes without an observation-noise model.
Use rows, targetCount (default 1), and optional unique targetNames to
identify axes. Levels are strictly increasing; quantiles cannot cross. Infinite
interval endpoints express conservative support, and [Infinity, -Infinity]
represents the empty set. Ellipsoid regions carry centers, a shared precision
matrix and radii; the numerical producer must validate positive definiteness.
Matrix targets require the explicit task kind multioutput or multilabel.
normalizeTargets, subsetTargets, Task and crossValScore preserve those axes.
Weights have one entry per row. Multilabel probabilities are independent columns;
multiclass probability rows sum to one. Multioutput R2, MSE and MAE average the
per-target scores. Multilabel measures include subset_accuracy, hamming_loss
and multilabel_log_loss.
createModelClass
Factory for building unified model classes from separate classifier/regressor implementations. Handles automatic task detection, async WASM pre-loading, and lifecycle management.
const { createModelClass } = require('@wlearn/core')
// Task-agnostic model (same class handles both tasks)
const XGBModel = createModelClass(XGBModelImpl, XGBModelImpl, {
name: 'XGBModel',
load: loadXGB // async WASM loader, called in create()
})
// Split model (separate classifier/regressor classes)
const MLPModel = createModelClass(MLPClassifier, MLPRegressor, {
name: 'MLPModel'
})The returned class supports:
Model.create(params)-- async factory. Passtask: 'classification'ortask: 'regression'to select explicitly, or omit to auto-detect fromyatfit()time.model.fit(X, y)-- trains the model. Auto-detects task from labels if not set.model.predict(X),model.predictProba(X),model.score(X, y)-- proxied to inner model.model.save()/Model.load(bytes)-- serialize/deserialize WLRN bytes.model.dispose()-- deterministic cleanup for long-running loops.model.task-- the detected or specified task.- Extra methods and getters from the inner classes are discovered and proxied automatically.
Auto-detection rules: if y is Int32Array, task is classification. Otherwise, if any value is non-integer, task is regression. If all values are integers and there are 2 to 20 unique values, task is classification; otherwise regression.
Model packages with additional fitting entry points can declare
fitMethods: { fitSpecial: 'regression' } in createModelClass options. Those
methods use the same fit-state, failure, disposal and Promise handling as fit.
Other extra methods remain fitted-only queries.
Browser bundles
Independent browser bundles share one public core API and loader registry per
JavaScript realm. They must embed exactly the same core version; mixing core
versions throws an actionable RegistryError. Rebuild every browser bundle when
updating core. This also preserves error-constructor identity and custom Measure
registrations across composed packages. Workers and frames have their own realms.
Errors
WlearnError, BundleError, RegistryError, ValidationError, NotFittedError, DisposedError
Utilities
sha256Sync(data)-- pure JS SHA-256makeLCG(seed?)-- deterministic LCG random number generatorshuffle(arr, rng)-- in-place shuffleisPromiseLike(x)/lift(x, fn)-- MaybePromise utilities
License
Apache-2.0
