@thinletterio/vqweb
v0.1.0
Published
Vector-quantised query encoders (.vqw containers) in the browser on WebGPU: load a 105-240 MiB file, embed a query in ~100 ms, keep your document index unchanged.
Maintainers
Readme
@thinletterio/vqweb
Vector-quantised query encoders in the browser on WebGPU. Load a released .vqw file (105–240 MiB), embed a query in
about 100 ms on an integrated GPU, and search your existing document index — the index was built with the full-precision
model and does not change. Apache-2.0.
- Models: huggingface.co/thinletter — harrier-0.6b (English, 2.1 / 1.8 bits per weight, 105–129 MiB) and Qwen3-Embedding-0.6B (Czech-calibrated, 3.6 bits, 237 MiB). Each model card says which document encoder and settings the file is a client of, and what it keeps of the full-precision retrieval quality (paired bootstrap on the public test split; the browser reproduces the simulation).
- Demo: thinletter.io/demo. Method and numbers: the technical report.
Install
npm install @thinletterio/vqwebor without a bundler, straight from a CDN:
<script type="module">
import { loadVqwClient } from 'https://thinletter.io/vqweb/vqweb.js';
// or: https://cdn.jsdelivr.net/npm/@thinletterio/vqweb/dist/vqweb.js
</script>Requirements: WebGPU with the shader-f16 feature (Chrome / Edge on desktop and Android; Safari 26 and Firefox where the
adapter reports shader-f16; probeWebGPU() tells you). There is no CPU fallback: check probeWebGPU() first and use the scalar GGUF files from the same Hugging
Face organisation with wllama where WebGPU is missing.
Use
import { loadVqwClient, probeWebGPU } from '@thinletterio/vqweb';
const gpu = await probeWebGPU();
if (!gpu.ok) throw new Error(gpu.reason);
const client = await loadVqwClient(
'https://huggingface.co/thinletter/qwen3-embedding-0.6b-vq-clients/resolve/main/qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw',
{ onProgress: ({ phase, loaded, total }) => console.log(phase, loaded, total) },
);
const q = await client.embed('Kolik stojí parkování v centru Prahy?'); // Float32Array(1024), L2-normalised
// score = dot(q, docVector) for every document vector of your fp32 / int8 index built with Qwen/Qwen3-Embedding-0.6B
console.log(client.info.model.teacher, client.info.quant.bpwBlocks, client.info.timing);
client.dispose();What loadVqwClient does: downloads the container (with progress; a .chunks.json byte-chunk manifest as on thinletter.io
works too), keeps it in the origin's private file system (OPFS) so the next visit loads from disk, uploads the codebooks,
indices and scales to the GPU, fetches the tokenizer named in the container header, compiles the WebGPU pipelines and runs
one warm-up query. The container header carries the query prompt, the EOS rule, the byte-fallback token table and, for the
*p files, the prompt's precomputed keys and values — embed(text) takes the raw query text.
Options: cache: 'none' (memory only), device (bring your own GPUDevice with shader-f16), powerPreference
('low-power' default = the integrated GPU where there is a choice), tokenizerUrl / tokenizerConfigUrl (override the
files next to the container), warmup: false. client.embedMany(texts) runs queries one after another (the runtime is
batch 1). client.embedWithStats(text) adds the token count and the milliseconds. clearCache(url) removes a cached file.
In a Web Worker
The package touches no DOM. Load the client in a worker to keep the page responsive during the ~1–3 s of upload and pipeline compilation:
// worker.js
import { loadVqwClient } from '@thinletterio/vqweb';
let client;
onmessage = async ({ data }) => {
if (data.load) { client = await loadVqwClient(data.load, { onProgress: (p) => postMessage({ progress: p }) }); postMessage({ ready: client.info }); }
if (data.embed) postMessage({ id: data.id, vec: await client.embed(data.embed) });
};React
const [client, setClient] = useState(null);
useEffect(() => { let c; loadVqwClient(URL).then((x) => setClient(c = x)); return () => c?.dispose(); }, []);Which file
| file | base model (the index must be built with it) | bits / weight | size | keeps of fp32 nDCG@10 |
|---|---|---|---|---|
| harrier-0.6b-vq2.0p-scifact.vqw | microsoft/harrier-oss-v1-0.6b | 2.10 + prompt K/V | 122 MiB | SciFact 98.1 % |
| harrier-0.6b-vq2.0-scidocs.vqw | microsoft/harrier-oss-v1-0.6b | 2.10 | 127 MiB | SciDocs 94.7 % |
| harrier-0.6b-vq1.75p-scifact.vqw | microsoft/harrier-oss-v1-0.6b | 1.83 + prompt K/V | 107 MiB | SciFact 95.8 % |
| qwen3-0.6b-vq3.5d2-webfaq_q-vocfull.vqw | Qwen/Qwen3-Embedding-0.6B | 3.57 | 237 MiB | WebFAQ-cs 96.8 % |
All files: thinletter/harrier-0.6b-vq-clients and thinletter/qwen3-embedding-0.6b-vq-clients. A file calibrated on one corpus family transfers to similar text; for your own index the repository has the recipe to verify (paired interval over your queries) and to compile a client calibrated on your corpus.
Numbers to expect
Integrated Intel GPU (Xe-LPG), Chrome: 103 ms per query for the 2.1-bit harrier file, 112 ms for the 3.6-bit Qwen3 file (p50, ~30 tokens); load 2–3 s after the download; GPU memory = the file size plus ~40 MiB of activations. Discrete GPUs and phones are not measured yet. The browser output equals the exporter's simulation of the same file to ±0.003 nDCG@10.
Advanced
The bundle exports the runtime modules unchanged (parseVqw, VqwContainer, VqwRuntime, VqwTokenizer,
requestDevice, KERNELS), so a bench or a layer test can drive the GPU pass directly; client.embedTokens(ids,
{ profile: true }) returns per-kernel GPU timestamps where the adapter supports them. Source, format description and
verification ladder: client/vqweb
in the repository; this package is built from it by packages/vqweb/build.mjs.
Licence
Apache-2.0 (see LICENSE and NOTICE: bundles @huggingface/tokenizers 0.2.0, Apache-2.0, and a Q2_K decoder ported from
llama.cpp's gguf-py, MIT). The model files carry the licence of their base model (MIT for harrier, Apache-2.0 for
Qwen3-Embedding). Contact: [email protected]
