npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@unlocalhosted/browsergrad-grad

v0.5.2

Published

A small, readable tensor + autograd library that runs inside Pyodide. PyTorch-flavored API, NumPy-backed, designed to be educational source code. Not PyTorch — `import browsergrad_grad as grad`.

Readme

@unlocalhosted/browsergrad-grad

npm License: MIT

A small, readable tensor + autograd library that runs inside Pyodide.

import browsergrad_grad as grad
import browsergrad_grad.functional as F

x = grad.Tensor([1.0, 2.0, 3.0], requires_grad=True)
y = (x * x).sum()
y.backward()
print(x.grad.tolist())   # [2.0, 4.0, 6.0]

Status: v0.5.2 — stable. Comprehensive layer set for CNNs and Transformers: ConvTranspose2d, Conv3d/2d/1d, BatchNorm3d/2d/1d, GroupNorm/InstanceNorm2d, LayerNorm, MaxPool/AvgPool, AdaptiveAvgPool2d, Dropout/Dropout2d, Embedding, MultiHeadAttention, RNN/LSTM/GRU, Flatten + all common activations. Optimizers: SGD/Adam/AdamW plus LR schedulers. Module.train()/eval(), hooks, state_dict/load_state_dict, torch-alias compatibility shims, and end-to-end training checks for MLP, CNN, sequence-CNN, and transformer-block.

The lazy-IR successor is browsergrad-jit — same PyTorch surface, but ops build a UOp graph that fusion / symbolic backward / AMP / gradient checkpointing / functional transforms / ONNX export / WebGPU realizer-bridge all hook into. Use grad for stable curriculum content; use jit when you want fusion + GPU acceleration + the broader toolkit. Both coexist in the same Pyodide session.

What this is

PyTorch-flavored API, NumPy-backed, deliberately not PyTorch. The Python module is named browsergrad_grad, not torch; call install_torch_alias() when you want the supported torch shim. Unsupported PyTorch APIs fail loudly instead of pretending to work.

The library is meant to be legible source code. Tensor/autograd, functional ops, optimizers, and nn chunks live as editable Python files under src/python/; codegen embeds them into TypeScript for installation into Pyodide.

What this is not

  • ❌ PyTorch. We don't try to match its full API.
  • ❌ A full polyfill. install_torch_alias() supports common tutorial code, but CUDA, distributed, compile/fx/jit, ONNX, quantization, and multi-process loaders remain explicit browser-scope refusals.
  • ❌ Production-fast. NumPy-on-CPU by default. A forward-only device= escape hatch can dispatch matmul / softmax / layernorm / unmasked 2D attention through @unlocalhosted/browsergrad-kernels, but throughput-oriented model execution still belongs in browsergrad-jit via its WebGPU realizer-bridge.
  • ❌ A general framework. It's a teaching artifact sized to fit in your head.

Install

npm install @unlocalhosted/browsergrad-grad

The npm package includes @unlocalhosted/browsergrad-kernels as a runtime dependency so the explicit device= bridge is available without a second install. Kernels owns its compatible @unlocalhosted/browsergrad-semantic-core dependency; applications do not need to install semantic-core directly.

Pyodide is a standard optional peer. Applications using the direct Node adapter install a compatible pyodide@^0.26.4; installGrad itself continues to accept runtime-managed and other duck-typed Pyodide targets.

Usage

import { createSession } from "@unlocalhosted/browsergrad-runtime";
import { installGrad } from "@unlocalhosted/browsergrad-grad";

const session = await createSession({
  pyodideIndexURL: "/pyodide/v0.26.4/",
  packages: ["numpy"],
});
await installGrad(session);

await session.exec({
  code: `
    import browsergrad_grad as grad
    import browsergrad_grad.functional as F

    # Tiny regression: y = 3x + 1, learn it.
    X = grad.randn(32, 1, seed=0)
    y_true = X * 3.0 + 1.0

    model = grad.nn.Linear(1, 1)
    opt = grad.optim.SGD(model.parameters(), lr=0.1)

    for step in range(200):
        opt.zero_grad()
        y_hat = model(X)
        loss = F.mse_loss(y_hat, y_true)
        loss.backward()
        opt.step()

    print(f"learned: y ≈ {model.weight.item():.2f} x + {model.bias.item():.2f}")
  `,
  onStdout: (s) => console.log(s),
});

Works with any Pyodide target — not just our runtime. Anything with an async exec({code}) method works:

await installGrad({
  exec: async ({ code }) => pyodide.runPythonAsync(code),
});

For Node scripts and CI where you loadPyodide() directly, use the shipped adapter at the ./node-adapter subpath — it wraps Pyodide's FS.writeFile + FS.mkdirTree to go through installViaFs (faster than installViaExec):

import { loadPyodide } from "pyodide";
import { installGrad } from "@unlocalhosted/browsergrad-grad";
import { createNodePyodideTarget } from "@unlocalhosted/browsergrad-grad/node-adapter";

const py = await loadPyodide();
await py.loadPackage(["numpy"]);
await installGrad(createNodePyodideTarget(py));

pyodide is an optional peer — direct-adapter consumers bring their own compatible version. The adapter has no other dependencies.

Framework platform support

Platform code can inspect Grad's verified eager contracts without installing Grad into Pyodide:

import {
  frameworkPlatformSupportSource,
} from "@unlocalhosted/browsergrad-grad";

const gradSupport = frameworkPlatformSupportSource();
console.log(gradSupport.frameworkId);       // browsergrad.grad
console.log(gradSupport.operations.length); // 22

The source is generated from the frozen executable compatibility inventory, not inferred from method presence. Its ten decision fields distinguish CPU execution or refusal, eager autograd, unavailable lazy transforms/export, tensor planning, WebGPU, residency, and materialization. It reports contracts, not current device availability or terminal execution evidence.

Optional WebGPU forward dispatch:

import { createDevice } from "@unlocalhosted/browsergrad-kernels";
import { createGradKernelDeviceBridge } from "@unlocalhosted/browsergrad-grad/kernel-device";

const device = await createDevice();
const gradDevice = createGradKernelDeviceBridge(device);
py.globals.set("grad_device", gradDevice);

await py.runPythonAsync(`
import browsergrad_grad as grad
import browsergrad_grad.functional as F

a = grad.Tensor([[1., 2.], [3., 4.]])
b = grad.Tensor([[5., 6.], [7., 8.]])
y = grad.matmul(a, b, device=grad_device)
p = F.softmax(y, dim=-1, device=grad_device)
`);

device= is intentionally small: 2D grad.matmul / grad.mm, last-dim F.softmax, last-dim nn.LayerNorm(..., device=...), and unmasked default F.scaled_dot_product_attention(..., device=...) for 2D Q/K/V. Backward still uses BrowserGrad's CPU formulas after the GPU forward result is materialized.

Python API surface (v0.5)

import browsergrad_grad as grad

# Construction
t = grad.Tensor([1, 2, 3], requires_grad=False)
z = grad.zeros(3, 4)
o = grad.ones(2, 2)
r = grad.randn(5, 5, seed=42)

# Properties
t.shape, t.ndim, t.size, t.data    # numpy view
t.numpy(), t.tolist(), t.item()    # exports
t.detach()                         # storage-sharing leaf, no autograd
t.to("float64")                    # owning differentiable floating cast
t.to("cpu"), t.cpu()               # CPU identity
t.to("cuda"), t.cuda()             # explicit NotImplementedError

# Arithmetic — broadcasts in v0.2
a + b, a - b, a * b, a / b, -a
a @ b                              # any rank ≥ 2, batch dims broadcast
a ** 2.0                           # scalar power only
a.exp(), a.log()                   # elementwise

# Shape
a.reshape(*shape), a.view(*shape), a.transpose(d0, d1), a.T   # 2D only
a.contiguous()                     # identity if C-order, otherwise owning copy

# Reductions (axis-aware)
t.sum(), t.sum(axis=1, keepdims=True)
t.mean(axis=-1)

# Autograd
loss.backward()                    # accumulates into .grad of every leaf

# Functional
import browsergrad_grad.functional as F
F.relu(x), F.leaky_relu(x, 0.01), F.sigmoid(x), F.tanh(x), F.gelu(x)
F.softmax(x, dim=-1), F.log_softmax(x, dim=-1)
F.scaled_dot_product_attention(q, k, v, attn_mask=None, is_causal=False)
F.mse_loss(y_hat, y)               # regression
F.cross_entropy_loss(logits, targets)   # classification (fused, stable)
F.nll_loss(log_probs, targets)

# Neural net building blocks
import browsergrad_grad.nn as nn
nn.Module                          # base — auto-tracks Tensor params
nn.Linear(in_features, out_features, bias=True)
nn.Conv2d(in_channels, out_channels, kernel_size, stride=1, padding=0, dilation=1, groups=1, bias=True)
nn.ConvTranspose2d(in_channels, out_channels, kernel_size, stride=1, padding=0, output_padding=0, groups=1, bias=True, dilation=1)
nn.Conv1d(in_channels, out_channels, kernel_size, stride=1, padding=0, bias=True)
nn.Conv3d(in_channels, out_channels, kernel_size, stride=1, padding=0, dilation=1, groups=1, bias=True)
nn.MaxPool2d(kernel_size, stride=None, padding=0)
nn.AvgPool2d(kernel_size, stride=None, padding=0)
nn.AdaptiveAvgPool2d(output_size)
nn.BatchNorm2d(num_features, eps=1e-5, momentum=0.1, affine=True)
nn.BatchNorm1d(num_features, eps=1e-5, momentum=0.1, affine=True)   # (N,C) or (N,C,L)
nn.BatchNorm3d(num_features, eps=1e-5, momentum=0.1, affine=True)
nn.GroupNorm(num_groups, num_channels, eps=1e-5, affine=True)
nn.InstanceNorm2d(num_features, eps=1e-5, affine=False)
nn.LayerNorm(normalized_shape, eps=1e-5, device=None)
nn.Embedding(num_embeddings, embedding_dim)
nn.MultiHeadAttention(embed_dim, num_heads, bias=True)              # (N, S, D)
nn.RNN(input_size, hidden_size, num_layers=1, bias=True, batch_first=False, dropout=0.0, bidirectional=False)
nn.LSTM(input_size, hidden_size, num_layers=1, bias=True, batch_first=False, dropout=0.0, bidirectional=False)
nn.GRU(input_size, hidden_size, num_layers=1, bias=True, batch_first=False, dropout=0.0, bidirectional=False)
nn.Dropout(p=0.5)
nn.Dropout2d(p=0.5)                                                 # channel-wise
nn.Flatten(start_dim=1, end_dim=-1)
nn.Sequential(m1, m2, m3)
nn.ReLU(), nn.LeakyReLU(0.01), nn.Sigmoid(), nn.Tanh(), nn.GELU()
# Mode control (cross-cutting):
model.train()    # train-mode behavior (BN uses batch stats; Dropout drops)
model.eval()     # eval-mode behavior  (BN uses running stats; Dropout identity)

# Optimization
import browsergrad_grad.optim as optim
optim.SGD(params, lr=0.01, momentum=0.0, weight_decay=0.0)
optim.Adam(params, lr=1e-3, betas=(0.9, 0.999), eps=1e-8, weight_decay=0.0)
optim.AdamW(params, lr=1e-3, betas=(0.9, 0.999), eps=1e-8, weight_decay=1e-2)

Current capability status

Done since the original v0.3 cut:

  • Conv2d now uses im2col + batched matmul, supports tuple kernel_size / stride / padding, dilation, and groups.

  • ConvTranspose2d, Conv3d/2d/1d, BatchNorm3d/2d/1d, GroupNorm, InstanceNorm2d, Dropout/Dropout2d, RNN/LSTM/GRU (multi-layer + bidirectional), Embedding, and MultiHeadAttention are in.

  • Norm backward correctness now covers full statistics-aware input gradients for BatchNorm3d, GroupNorm, and InstanceNorm2d.

  • Module ergonomics now cover train() / eval(), hooks, buffers, state_dict, load_state_dict, and torch compatibility shims.

  • WebGPU forward dispatch is in as an explicit device= path over @unlocalhosted/browsergrad-kernels for matmul, softmax, layernorm, and attention.

Remaining explicit limits:

  • Direct eager GPU scope. device= is forward-only and intentionally explicit. It does not make all eager ops GPU-backed, and it materializes results before CPU autograd.

Design notes

  • No _ctx-mutability shenanigans. Each op captures the data it needs at forward time and binds it in a closure. Backward functions are pure.
  • Contiguous means contiguous. Non-C-order storage is copied into owning C-order storage with a dtype-preserving identity-gradient edge.
  • Expand means a view. Broadcast singleton axes use zero strides over the source storage. Dtype/layout and bidirectional mutation are preserved, while backward reduces expanded axes to the original shape.
  • Global gradient context exists. grad.no_grad() disables graph construction for inference sections; .detach() returns a distinct storage-sharing leaf with no autograd history.
  • Floating casts stay differentiable. .to() casts among float16/float32/float64 retain an autograd edge and restore the source dtype in backward. Bool/integer casts are detached.
  • Device requests are truthful. Eager tensor storage is CPU/Pyodide-backed. CPU requests preserve identity; unavailable CUDA/MPS/XPU/Meta requests fail before execution and never masquerade as a transfer. torch.tensor(device=...) follows the same CPU-only boundary. nn.Module.to("cpu") preserves identity, while module device or dtype conversion requests reject until recursive parameter conversion exists.
  • NumPy ownership is directional. from_numpy wraps supported writable arrays with exact zero-copy dtype/stride preservation. .numpy() and np.asarray(tensor) return owning snapshots so exported arrays cannot mutate tensor storage behind autograd.
  • The eager dtype registry is closed. String requests use the documented BrowserGrad/PyTorch aliases. NumPy dtype objects and scalar types are accepted only for bool, float16/32/64, int8/16/32/64, and uint8/16/32/64 storage; unsupported storage rejects before allocation.
  • Constructor surfaces are distinct. Direct Tensor(...) is the float32-default educational constructor. torch.tensor(...) is an owning leaf-copy adapter with bounded PyTorch-shaped inference; it preserves admitted NumPy/Tensor dtypes and never aliases caller storage.
  • Reverse-mode only. No forward-mode, no functional transforms (vmap, etc.).
  • Tensor.__slots__. Slot-based attribute layout to keep memory predictable for tensors in long training loops.

API reference

See src/python/*.ts — every Python module is embedded as a *_PY template literal in its own TS file. That's where the source code lives; reading those files is the documentation.

License

MIT