npm package discovery and stats viewer.

Discover Tips

  • General search

    [free text search, go nuts!]

  • Package details

    pkg:[package-name]

  • User packages

    @[username]

Sponsor

Optimize Toolset

I’ve always been into building performant and accessible sites, but lately I’ve been taking it extremely seriously. So much so that I’ve been building a tool to help me optimize and monitor the sites that I build to make sure that I’m making an attempt to offer the best experience to those who visit them. If you’re into performant, accessible and SEO friendly sites, you might like it too! You can check it out at Optimize Toolset.

About

Hi, 👋, I’m Ryan Hefner  and I built this site for me, and you! The goal of this site was to provide an easy way for me to check the stats on my npm packages, both for prioritizing issues and updates, and to give me a little kick in the pants to keep up on stuff.

As I was building it, I realized that I was actually using the tool to build the tool, and figured I might as well put this out there and hopefully others will find it to be a fast and useful way to search and browse npm packages as I have.

If you’re interested in other things I’m working on, follow me on Twitter or check out the open source projects I’ve been publishing on GitHub.

I am also working on a Twitter bot for this site to tweet the most popular, newest, random packages from npm. Please follow that account now and it will start sending out packages soon–ish.

Open Software & Tools

This site wouldn’t be possible without the immense generosity and tireless efforts from the people who make contributions to the world and share their work via open source initiatives. Thank you 🙏

© 2026 – Pkg Stats / Ryan Hefner

@thinletterio/vqweb

v0.1.0

Published

Vector-quantised query encoders (.vqw containers) in the browser on WebGPU: load a 105-240 MiB file, embed a query in ~100 ms, keep your document index unchanged.

Readme

@thinletterio/vqweb

Vector-quantised query encoders in the browser on WebGPU. Load a released .vqw file (105–240 MiB), embed a query in about 100 ms on an integrated GPU, and search your existing document index — the index was built with the full-precision model and does not change. Apache-2.0.

  • Models: huggingface.co/thinletter — harrier-0.6b (English, 2.1 / 1.8 bits per weight, 105–129 MiB) and Qwen3-Embedding-0.6B (Czech-calibrated, 3.6 bits, 237 MiB). Each model card says which document encoder and settings the file is a client of, and what it keeps of the full-precision retrieval quality (paired bootstrap on the public test split; the browser reproduces the simulation).
  • Demo: thinletter.io/demo. Method and numbers: the technical report.

Install

npm install @thinletterio/vqweb

or without a bundler, straight from a CDN:

<script type="module">
  import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js';
  // or: https://cdn.jsdelivr.net/npm/@thinletterio/vqweb/dist/vqweb.js
</script>

Requirements: WebGPU with the shader-f16 feature (Chrome / Edge on desktop and Android; Safari 26 and Firefox where the adapter reports shader-f16; probeWebGPU() tells you). There is no CPU fallback: check probeWebGPU() first and use the scalar GGUF files from the same Hugging Face organisation with wllama where WebGPU is missing.

Use

import { loadVqwClient, probeWebGPU } from '@thinletterio/vqweb';

const gpu = await probeWebGPU();
if (!gpu.ok) throw new Error(gpu.reason);

const client = await loadVqwClient(
  'https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw',
  { onProgress: ({ phase, loaded, total }) => console.log(phase, loaded, total) },
);

const q = await client.embed('Kolik stojí parkování v centru Prahy?');   // Float32Array(1024), L2-normalised
// score = dot(q, docVector) for every document vector of your fp32 / int8 index built with Qwen/Qwen3-Embedding-0.6B

console.log(client.info.model.teacher, client.info.quant.bpwBlocks, client.info.timing);
client.dispose();

What loadVqwClient does: downloads the container (with progress; a .chunks.json byte-chunk manifest as on thinletter.io works too), keeps it in the origin's private file system (OPFS) so the next visit loads from disk, uploads the codebooks, indices and scales to the GPU, fetches the tokenizer named in the container header, compiles the WebGPU pipelines and runs one warm-up query. The container header carries the query prompt, the EOS rule, the byte-fallback token table and, for the *p files, the prompt's precomputed keys and values — embed(text) takes the raw query text.

Options: cache: 'none' (memory only), device (bring your own GPUDevice with shader-f16), powerPreference ('low-power' default = the integrated GPU where there is a choice), tokenizerUrl / tokenizerConfigUrl (override the files next to the container), warmup: false. client.embedMany(texts) runs queries one after another (the runtime is batch 1). client.embedWithStats(text) adds the token count and the milliseconds. clearCache(url) removes a cached file.

In a Web Worker

The package touches no DOM. Load the client in a worker to keep the page responsive during the ~1–3 s of upload and pipeline compilation:

// worker.js
import { loadVqwClient } from '@thinletterio/vqweb';
let client;
onmessage = async ({ data }) => {
  if (data.load) { client = await loadVqwClient(data.load, { onProgress: (p) => postMessage({ progress: p }) }); postMessage({ ready: client.info }); }
  if (data.embed) postMessage({ id: data.id, vec: await client.embed(data.embed) });
};

React

const [client, setClient] = useState(null);
useEffect(() => { let c; loadVqwClient(URL).then((x) => setClient(c = x)); return () => c?.dispose(); }, []);

Which file

| file | base model (the index must be built with it) | bits / weight | size | keeps of fp32 nDCG@10 | |---|---|---|---|---| | harrier-0.6b-vq2.0p-scifact.vqw | microsoft/harrier-oss-v1-0.6b | 2.10 + prompt K/V | 122 MiB | SciFact 98.1 % | | harrier-0.6b-vq2.0-scidocs.vqw | microsoft/harrier-oss-v1-0.6b | 2.10 | 127 MiB | SciDocs 94.7 % | | harrier-0.6b-vq1.75p-scifact.vqw | microsoft/harrier-oss-v1-0.6b | 1.83 + prompt K/V | 107 MiB | SciFact 95.8 % | | qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw | Qwen/Qwen3-Embedding-0.6B | 3.57 | 237 MiB | WebFAQ-cs 96.8 % |

All files: thinletter/harrier-0.6b-vq-clients and thinletter/qwen3-embedding-0.6b-vq-clients. A file calibrated on one corpus family transfers to similar text; for your own index the repository has the recipe to verify (paired interval over your queries) and to compile a client calibrated on your corpus.

Numbers to expect

Integrated Intel GPU (Xe-LPG), Chrome: 103 ms per query for the 2.1-bit harrier file, 112 ms for the 3.6-bit Qwen3 file (p50, ~30 tokens); load 2–3 s after the download; GPU memory = the file size plus ~40 MiB of activations. Discrete GPUs and phones are not measured yet. The browser output equals the exporter's simulation of the same file to ±0.003 nDCG@10.

Advanced

The bundle exports the runtime modules unchanged (parseVqw, VqwContainer, VqwRuntime, VqwTokenizer, requestDevice, KERNELS), so a bench or a layer test can drive the GPU pass directly; client.embedTokens(ids, { profile: true }) returns per-kernel GPU timestamps where the adapter supports them. Source, format description and verification ladder: client/vqweb in the repository; this package is built from it by packages/vqweb/build.mjs.

Licence

Apache-2.0 (see LICENSE and NOTICE: bundles @huggingface/tokenizers 0.2.0, Apache-2.0, and a Q2_K decoder ported from llama.cpp's gguf-py, MIT). The model files carry the licence of their base model (MIT for harrier, Apache-2.0 for Qwen3-Embedding). Contact: [email protected]