@tangent.to/nn
v0.1.1
Published
Neural networks on the tangent tape: a functional layer API (branches, dropout, concrete dropout) differentiated by @tangent.to/grad, trained with @tangent.to/opt, exportable as JSON.
Maintainers
Readme
@tangent.to/nn
Neural networks on the tangent tape. A functional layer API in the spirit of Keras, with branches, dropout and concrete dropout, differentiated by grad, trained with opt's update rules, seeded by proba, whose densities are the likelihood losses; exportable as one JSON file for an app.
import nn from "@tangent.to/nn";
const x = nn.input(37); // 24 soil + 12 tissue + cycle
const soil = nn.dense(32, { activation: "relu" })(nn.slice(x, 0, 24));
const leaf = nn.dense(16, { activation: "relu" })(nn.slice(x, 24, 12));
const h = nn.concreteDropout(nn.dense(32, { activation: "tanh" }))(
nn.concat([soil, leaf, nn.slice(x, 36, 1)]));
const out = nn.dense(2)(h); // [mean, log sd]
const model = nn.model(x, out, { loss: "gaussianNLL", normalizeY: true, seed: 42 });
await model.fit(X, y, {
epochs: 300, batchSize: 32, validationSplit: 0.1,
callbacks: [nn.callbacks.earlyStopping({ patience: 20 })],
onEpochEnd: (epoch, logs) => console.log(epoch, logs.valLoss),
});
model.predict(Xnew); // point prediction, dropout off
model.predict(Xnew, { samples: 200 }); // { mean, std, epistemic, aleatoric }
model.predict(Xnew, { returnStd: true }); // the noise column alone
model.predictGradient(xnew); // d mean / d x, as the GP gives it
JSON.stringify(model); // architecture + weights, for an appWhat is in it
Layers. input(d), dense(units, { activation, useBias, init }) with
linear, relu, tanh, sigmoid, softplus, elu, gelu or a function
of the Var; concat, slice, add (a residual connection); dropout(rate);
concreteDropout(denseLayer, { perFeature, temperature, priorLengthScale }),
which learns its rate, one per input feature with perFeature, a relevance
per variable; lambda(fn), the escape hatch, a function of the input Var in
grad ops. A layer applied twice shares its weights.
Losses. mse, mae, huber(delta), and the likelihood losses over
proba's densities: gaussianNLL on a [mean, log sd] output, studentTNLL({
nu }) for robust regression, poissonNLL on a log-rate for counts,
binaryCrossEntropy on a logit. Every loss takes (yPred, yTrue, w), a
weight per row: sampleWeight in fit is how a missing target is left out.
fit. Asynchronous, yielding every epoch (fitSync for a synchronous caller); adam, sgd, momentum,
rmsprop by mini-batch, or lbfgs on the full batch with the tape's exact
gradient, which on a few hundred rows is the right tool. normalizeY,
validationSplit or validationData, metrics, clipNorm, an
AbortSignal, earlyStopping that restores the best weights, a learning-rate
schedule, and a non-finite loss that says what usually causes it.
Around the model. clone() for a cross-validation fold or a bootstrap
refit; ensemble(model, X, y, { members }); checkGradients(model, X, y)
for whoever writes a loss or a lambda layer; summary(); toJSON and
fromJSON.
The design note, docs/DESIGN-0.1.md, records why
this is a package of its own, what it borrows from the rest of the suite, the
two grad releases it depends on, and the order of work.
Where it stops
Rank two: matrices and vectors, so dense networks. Recurrent and
convolutional layers, normalization layers and a multiclass softmax wait for
grad 0.4 (reductions by axis, a column broadcast, gather); see section 7 of
the note. No GPU. Monte Carlo dropout's intervals are not calibrated on their
own; the model reports the epistemic and aleatoric parts so a conformal
factor from held-out residuals can be applied, and ensemble is the better
default when five fits are affordable.
Part of the tangent suite. GPL-3.0, like mc and ds.
