@johnhenry/math-plus-data
v0.0.4
Published
Async dataset pipelines for math-plus (issue #22) — the v2 `data` namespace: a curated Dataset facade over @johnhenry/iteration (chunk/batch/shuffle/epochs/mapConcurrent/prefetch/fold, AbortSignal cancellation) producing Tensor batches for tensor-autograd
Readme
@johnhenry/math-plus-data
Async dataset pipelines for math-plus (issue #22): a curated Dataset facade
over @johnhenry/iteration —
chunk/batch/shuffle/epochs/mapConcurrent/prefetch/fold with AbortSignal
cancellation — producing Tensor batches shaped exactly for
@johnhenry/math-plus-tensor-autograd's trainer.fit.
Install
npm install @johnhenry/math-plus-dataQuick start
import { collate, fromAsync } from "@johnhenry/math-plus-data";
const ds = fromAsync([1, 2, 3, 4, 5, 6, 7, 8])
.map((x) => x * 10)
.filter((x) => x % 20 === 0)
.drop(1)
.take(2);
await ds.toArray(); // [40, 60] — and re-iterable, so again: [40, 60]
// Batching straight into the trainer's Batch shape
const samples = [{ x: [0, 1], y: [0] }, { x: [1, 2], y: [2] } /* ... */];
const pipeline = fromAsync(samples)
.epochs(60, { reshuffle: { seed: 42 } })
.batch(16, { collate: collate.xy({ dtype: "f64" }) });
const { lossHistory } = await trainer.fit(pipeline); // tensor-autogradAPI surface
fromAsync(source)— source is anIterable,AsyncIterable, or a factory returning one.Datasetmethods:map,filter,mapConcurrent(fn, { concurrency, ordered?, signal? }),prefetch(n),chunk(n),batch(n, { collate? }),shuffle({ seed?, bufferSize? }),take,drop,abortable(signal),epochs(n, { reshuffle? }),fold,toArray.collate.vectors({ dtype? })→[batch, dim]Tensor;collate.scalars({ dtype? })→[batch];collate.xy({ dtype? })→{ x, y }— tensor-autograd'sBatchexactly.
Traps
- One-shot vs re-iterable. Arrays/Sets are re-iterable; a bare
AsyncIterableis assumed one-shot — a second pass throws with a hint, and.epochs()refuses one-shot sources up front. Pass a factory (fromAsync(() => stream())) for multi-pass pipelines. collatedefaults tof32, matchingnn.*parameters' default dtype (f32 since@johnhenry/math-plus-tensor-autograd#123). tensor-core has no implicit dtype promotion by design, so if you build a model with{ dtype: "f64" }parameters, passcollate.xy({ dtype: "f64" })too.- Shuffle: omitting
seedis non-reproducible.bufferSizedefaults toInfinity(full materialize + Fisher-Yates); a finite buffer is the tf.data streaming shuffle with mixing quality bounded by the buffer. shuffle()on top ofepochs()shuffles the concatenated stream. For per-epoch reshuffling useepochs(n, { reshuffle: { seed } })(derivesseed + epochIndex).- Ragged batches are loud:
collate.vectorsthrowsRangeErroron differing sample lengths — no silent padding. - Curated facade, enforced by a test: no
count*, no rawgroup/reduce*re-exports (groupcollides with dataframegroupByvocabulary;foldis the terminal reduce). Power users can import@johnhenry/iterationdirectly — the facade is the supported surface, not a wall.
Concurrency and cancellation
mapConcurrent delegates to iteration's mapConcurrentAsync: order
preserved by default, at most concurrency in flight, source pulled only
with spare capacity, source closed via return() on early exit or error.
prefetch(n) is a bounded read-ahead buffer. Cancellation is plain
AbortSignal end to end, rejecting with the signal's reason.
Provenance
Part of the math-plus monorepo; family docs at https://opensource.johnhenry.me/math/.
