npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

datareader-spss

v0.2.0

Published

Correct, zero-dependency SPSS .sav/.zsav reader for the browser and Node — validated against R haven.

Readme

datareader-spss

Correct, zero-dependency SPSS .sav/.zsav reader for the browser and Node — validated against R haven.

datareader-spss parses IBM SPSS system files (.sav) and their compressed variant (.zsav) into plain JavaScript values: numbers, strings, Dates, variable metadata, value labels, and declared missing values. It is written in pure TypeScript against Web platform APIs (ArrayBuffer, DataView, TextDecoder, DecompressionStream) — no runtime dependencies, and the same build runs in the browser and in Node.

Why

SPSS .sav files are everywhere in the social sciences, market research, and public data, but the JavaScript ecosystem lacked a reader that decodes real files correctly — most tools stumble on RLE compression, ZSAV (zlib) blocks, very-long strings, encodings, or the exact bit-level meaning of system- vs. user-missing values. Getting any of those wrong silently corrupts data.

datareader-spss was built to be correct first: every construct is validated value-for-value against R's haven/ReadStat, the de-facto reference implementation. It uses only Web APIs, so it ships as a single ESM module with zero dependencies and no native addons — drop it into a browser upload flow or a Node script alike.

Install

npm i datareader-spss
bun add datareader-spss
pnpm add datareader-spss
yarn add datareader-spss

Usage

Browser — a picked/dropped File

import { readSav, applyUserMissing } from "datareader-spss";

const input = document.querySelector<HTMLInputElement>("#file");
input.addEventListener("change", async () => {
  const file = input.files?.[0];
  if (!file) return;

  const parsed = await readSav(await file.arrayBuffer());
  const sheet = parsed.sheets[0];

  console.log(sheet.variables.map((v) => v.name)); // column names
  console.log(sheet.rows.length); // row count

  // Replace every declared user-missing cell with null before analysis:
  const clean = applyUserMissing(sheet);
  console.log(clean.rows[0]);
});

React — a file input

import { useState } from "react";
import { readSav, applyUserMissing, type Sheet } from "datareader-spss";

export function SavImporter() {
  const [sheet, setSheet] = useState<Sheet | null>(null);

  async function onFile(e: React.ChangeEvent<HTMLInputElement>) {
    const file = e.target.files?.[0];
    if (!file) return;
    const { sheets } = await readSav(await file.arrayBuffer());
    setSheet(applyUserMissing(sheets[0]));
  }

  return (
    <>
      <input type="file" accept=".sav" onChange={onFile} />
      {sheet && (
        <p>
          {sheet.rows.length} rows × {sheet.variables.length} variables
        </p>
      )}
    </>
  );
}

Node — a file on disk

import { readFile } from "node:fs/promises";
import { readSav } from "datareader-spss";

const nodeBuf = await readFile("survey.sav");

// Slice an exact ArrayBuffer (Node pools Buffers behind a shared store):
const bytes = nodeBuf.buffer.slice(
  nodeBuf.byteOffset,
  nodeBuf.byteOffset + nodeBuf.byteLength,
);

const parsed = await readSav(bytes);
for (const v of parsed.sheets[0].variables) {
  console.log(v.name, v.type, v.label ?? "");
}

Reading variables and rows

import { readSav, type CellValue } from "datareader-spss";

const { sheets, encoding } = await readSav(buffer);
const { variables, rows } = sheets[0];

// `rows` is CellValue[][], column-aligned to `variables` by index.
const genderCol = variables.findIndex((v) => v.name === "gender");
const genderValues: CellValue[] = rows.map((row) => row[genderCol]);

// A variable's value labels (e.g. 1 -> "Male", 2 -> "Female"):
const gender = variables[genderCol];
console.log(gender.valueLabels); // [{ value: 1, label: "Male" }, ...]
console.log(encoding); // e.g. "utf-8" or "windows-1252"

API

readSav(buf, opts?)

function readSav(
  buf: ArrayBuffer,
  opts?: Partial<SavLimits>,
): Promise<ParsedFile>;

Reads a .sav/.zsav file end-to-end (header → dictionary → extensions → variables → data cases, inflating ZSAV blocks as needed) and resolves to a ParsedFile. opts tightens the resource ceilings (see Security & limits). Rejects with a SavError on malformed or hostile input.

applyUserMissing(sheet)

function applyUserMissing(sheet: Sheet): Sheet;

Returns a new Sheet in which every cell matching its variable's declared MissingSpec is replaced by null (user-missing → system-missing). Non-mutating — the input sheet is untouched. Use it when you want declared missing codes treated as missing; skip it when you need the literal codes.

SavError

class SavError extends Error {}

Thrown when a resource bound rejects a malformed or hostile file (a bad file, not a reader bug). Because it extends Error, a plain catch still catches it; check err instanceof SavError to distinguish a rejected/hostile file from a genuinely unsupported-but-valid construct (which throws a plain Error).

DEFAULT_LIMITS

const DEFAULT_LIMITS: SavLimits = {
  // ncases × nvars ceiling (~500k rows × 100 vars)
  maxCells: 50_000_000,
  // ZSAV inflate output ceiling (512 MiB)
  maxInflatedBytes: 512 * 1024 * 1024,
};

The generous defaults that readSav merges your opts over. No well-formed file is affected.

Types

// null = system-missing
type CellValue = string | number | Date | null;

type Measure = "nominal" | "ordinal" | "scale" | "unknown";

type MissingSpec =
  | { kind: "none" }
  | { kind: "discrete"; values: number[] }
  | { kind: "range"; lo: number; hi: number }
  | { kind: "range+discrete"; lo: number; hi: number; value: number }
  | { kind: "strings"; values: string[] };

type SpssFormat = {
  type: number;
  width: number;
  decimals: number;
  isDate: boolean;
};

type Variable = {
  name: string;
  label?: string;
  type: "numeric" | "string";
  // string byte width, or the merged width for very-long strings
  width?: number;
  missing: MissingSpec;
  valueLabels?: Array<{ value: CellValue; label: string }>;
  format: SpssFormat;
  measure: Measure;
};

type Sheet = {
  name: string;
  variables: Variable[];
  rows: CellValue[][];
};

type ParsedFile = {
  format: "sav";
  encoding?: string;
  sheets: Sheet[];
};

type SavLimits = {
  maxCells: number;
  maxInflatedBytes: number;
};

For a .sav file, sheets always contains exactly one sheet (named "data"); rows is row-major and column-aligned to variables.

Missing values. null in a cell is always system-missing. Declared user-missing values (a survey's 99 = "no answer", or a range) are kept as their literal number/string so you can see them — call applyUserMissing(sheet) to fold them to null. A variable's declaration is on variable.missing:

  • { kind: "none" } — no declared missing values.
  • { kind: "discrete", values } — up to three discrete codes.
  • { kind: "range", lo, hi } — a missing range [lo, hi].
  • { kind: "range+discrete", lo, hi, value } — a range plus one discrete code.
  • { kind: "strings", values } — discrete codes for a string variable.

Format coverage

| Area | Supported | | -------------- | ------------------------------------------------------------------- | | Magic | $FL2 (uncompressed / RLE) and $FL3 (ZSAV) | | Compression | uncompressed, RLE (SPSS bytecode), ZSAV (zlib deflate blocks) | | Variables | numeric, string, and very-long strings (> 255 bytes, segment-merged) | | Dates | SPSS date/time formats decoded to JavaScript Date | | Value labels | numeric and string value-label sets, resolved per variable | | Missing values | discrete, range, range + discrete, and string missing specs | | Encodings | file-declared charset passed straight to the platform TextDecoder; unsupported labels rejected, never mis-decoded (see Security & limits) | | Endianness | little-endian |

Correctness

Every supported construct is validated value-for-value against R haven (which wraps the ReadStat C library) across a fixture matrix covering compression modes, numeric/string/very-long variables, dates, value labels, each missing-value kind, and encodings — plus a real-world 2.78 MB survey file. The oracle fixtures live alongside the source and are asserted on every test run.

Security & limits

datareader-spss is designed to parse untrusted files safely. A hostile .sav cannot make it OOM or hang: every attacker-controlled allocation and loop is bounded, and any bound violation throws a catchable SavError rather than exhausting memory or spinning.

  • Reads never allocate beyond the remaining file bytes.
  • The data-cell budget (maxCells) and the ZSAV inflate budget (maxInflatedBytes) are enforced while streaming, so a decompression bomb aborts before it materializes.
  • Malformed dictionaries, out-of-spec record counts, and truncated headers are rejected up front.
  • The file's declared charset (subtype-20) is handed straight to the runtime's TextDecoder. Node (full ICU) decodes every WHATWG label; Bun decodes a subset — UTF-8/16, windows-1252/latin1, and the CJK pages (shift_jis, gbk, big5, euc-jp), but not single-byte non-Latin pages such as windows-1251/koi8-r. A charset the runtime can't decode throws a SavError rather than silently producing mojibake — an unreadable codepage fails loud (a zero-dep reader ships no codepage tables of its own).

Tune the ceilings for memory-constrained environments:

await readSav(buf, {
  maxCells: 5_000_000,
  maxInflatedBytes: 64 * 1024 * 1024,
});

Recommendations for consumers:

  • Still bound the input size before you hand a buffer to readSav (reject files larger than you expect for your use case).
  • Treat Variable.name (and other file-derived strings) as untrusted — don't use a variable name as a plain-object key without care; prefer a Map or a null-prototype object to avoid prototype-pollution surprises from adversarial names.

Roadmap

datareader-spss reads modern little-endian .sav/.zsav files correctly today, and is intentionally read-only. On the map for future releases:

  • Big-endian system files — the header already detects byte order; big-endian data reading and haven validation are not yet implemented.
  • Portable .por format — the older text-based SPSS portable format.
  • Writing .sav files — the reader is read-only; a writer is a possible future addition.

Issues and contributions are welcome.

License

MIT © 2026 everydaydesign

Credits