qvdjs
v4.0.0
Published
Library for reading/writing Qlik Sense/QlikView QVD files in Javascript.
Maintainers
Readme
qvdjs
Read and write Qlik Sense and QlikView (QVD) files from Node.js
npm install qvdjsimport {QvdDataFrame} from 'qvdjs';
const df = await QvdDataFrame.fromQvd('sales.qvd');
console.log(df.shape); // [ 1705805, 20 ]
console.log(df.head(5));No Qlik installation, no ODBC driver, no running engine — just the file. Node 22 or newer, published as a dual ESM/CommonJS package, MIT licensed, one runtime dependency.
📖 Full documentation at qvdjs.ptarmiganlabs.com — guides, API reference, the QVD format explained, measured performance, and troubleshooting.
⚠️ This library is based on reverse engineering
Qlik does not publish the QVD format. Everything qvdjs knows about it was worked out by reading files Qlik produced, and what it writes is compared with what Qlik Sense writes from the same data - below.
That works well enough to be useful, and is not the same as being exact. Test with your own files, and validate output in your own Qlik environment before relying on it. Reading any bundled Qlik file and writing it back reproduces its symbol table byte for byte, whichever way it was read; a value you change is written as Qlik stores that kind of value. Where that boundary lies is the first thing worth reading: What a round trip preserves.
Checked against Qlik Sense
qvdjs is compared with Qlik Sense itself, not only with its own reading of the format. 48 fields in 8 datasets - integers, doubles, text (non-ASCII, empty and number-like included), booleans, dates, timestamps, and NULLs beside numbers, text, dates and timestamps - are defined twice, as Qlik load-script expressions and as the JavaScript values a caller would write. Qlik Sense and qvdjs each write them to a QVD, and three things are checked:
- qvdjs writes the content Qlik writes - the same symbols, of the same types, with the same text, on the same rows.
- Given Qlik's tags and number formats, qvdjs writes them unchanged - the bit layout and a tag the values contradict, below, are the only parts of a field's header left to differ.
- Qlik reads qvdjs's file as it reads its own - loaded into Qlik and stored again, with
Text(),Num(),IsNum(),IsText(),IsNull()andYear()the same on every row.
The live comparison runs on demand against a Qlik Sense server; its last run, against Qlik Sense Enterprise on Windows build 50699, passed for all 48 fields. Qlik's files from that run are kept in the test suite, so every CI run compares qvdjs's output with them again - on Linux, Windows and macOS, under Node 22, 24 and 26 - along with reference files Qlik wrote for booleans, dates and timestamps.
Three differences are by design. Qlik often gives a field more bits in the index than its values need,
where qvdjs uses the fewest. qvdjs adds no tags or number formats of its own - Qlik works tags out when
it loads a file, and keeps any the file carries. And qvdjs drops a tag the values contradict, even one
Qlik wrote: Qlik tags dates outside 1980 to 2080 as text when their field has no date format, although
its own IsNum() reads every one as a number. QVD is also QlikView's format, but QlikView is not part
of the comparison.
Six ways to open a file
fromQvd is the general one, and often not the one you want. A QVD is an XML header, then a symbol
table holding every distinct value, then a bit-packed index table of one code per cell. Two things
separate these calls: how far into the file a read has to go, and how much of it the read then
holds.
| You want | Call | How far it reads, and what it holds |
| ---------------------------------------- | ----------------------------------------- | -------------------------------------------------------------------- |
| Rows, to index, iterate or write back | QvdDataFrame.fromQvd(path) | Everything, and materialises every row |
| Rows from a file too large to hold | QvdDataFrame.iterate(path, {chunkSize}) | Everything, but holds two chunks — a 96 MB heap against 512 MB |
| A few columns of a large file | QvdColumnTable.fromQvd(path) | Everything, but stops before building rows — 141 MiB against 385 MiB |
| Only the schema: names, row count, types | QvdDataFrame.readMetadata(path) | The header alone. Constant cost, whatever the file's size |
| To know whether a read will fit first | QvdDataFrame.checkRead(path, options) | The header, and a walk of the symbols the read would decode. Answers rather than reads; a read it approves is not refused for memory |
| Pages, by position and in any order | QvdDataFrame.open(path, options) | The header, then a page per call — rows(), columns(), await check(), close(). A column is walked once, and a page decodes only the values its rows use |
import {QvdDataFrame, QvdColumnTable} from 'qvdjs';
// What is in this file? Costs the same whether it is 20 KB or 20 GB.
const {columns, rowCount} = await QvdDataFrame.readMetadata('sales.qvd');
// Will reading it fit? It walks the symbols the read would decode, never a record, and every
// suggestion it gives has been checked.
const {fits, suggestions} = await QvdDataFrame.checkRead('sales.qvd');
// Paging: a page can be asked for by position, forwards or back. Its first touch of a column walks the
// column once; after that it decodes only the values its rows use. check() takes the page's own steps.
const qvd = await QvdDataFrame.open('sales.qvd');
if ((await qvd.check({offset: 5_000_000, limit: 100})).fits) {
const page = await qvd.rows({offset: 5_000_000, limit: 100});
}
await qvd.close();
// Sum one column without ever building a row.
const table = await QvdColumnTable.fromQvd('sales.qvd');
let total = 0;
for (const value of table.column('amount')) {
if (typeof value === 'number') total += value;
}
// Every row of a file that will not fit, a chunk at a time.
for await (const chunk of QvdDataFrame.iterate('sales.qvd', {chunkSize: 50_000})) {
process(chunk.data);
}
// Rows, when you want rows.
const df = await QvdDataFrame.fromQvd('sales.qvd');Reaching for fromQvd when you wanted one of the others is the common mistake, and
fromQvd(path, {maxRows: 0}) is not a substitute for readMetadata: it loads no rows but still
reads the whole symbol table, which grows with the data, and can be refused for its size where
readMetadata never is.
What a cell holds
| The file stores | The cell holds |
| ------------------------------------------------------------------- | ------------------------------------------------------- |
| A number | A number |
| A string | A string, never parsed - '007' stays '007' |
| A dual: a number with the text Qlik displays, such as a date's text | Its number. The text is kept, and textAt returns it |
| NULL | null |
{duals: 'text'} reads a dual as its text instead, and {duals: 'both'} as a frozen QvdDual holding
both halves. {coerceNumericStrings: true} reads a numeric-looking string as a number. Whichever way a
frame was read, it writes back the symbols it was read from.
import {QvdDataFrame, qlikSerialToDate} from 'qvdjs';
const df = await QvdDataFrame.fromQvd('rides.qvd');
df.data[0][0]; // 42382.260416666664 - the date is its number, as Qlik sums and sorts it
df.textAt(0, 'started'); // '2016-01-13 06:15:00' - the text Qlik displays
qlikSerialToDate(df.data[0][0]).toISOString(); // '2016-01-13T06:15:00.000Z'Writing
const df = await QvdDataFrame.fromDict({
columns: ['Region', 'Year', 'Amount'],
data: [
['North', 2025, 1200.0],
['South', 2025, 980.5],
],
});
await df.toQvd('out.qvd', {
onProgress: ({stage, percent}) => console.log(`${stage}: ${percent}%`),
});The QVD is built beside the destination under a temporary name and renamed over it, so a write that
fails — a full disk, a killed process, an error from the disk — leaves the previous file exactly as it
was, and anything reading the path meanwhile gets the old file or the new one and never a part of
either. toQvd() resolves once the contents have reached the disk. The rename that puts the file in
place is not flushed, so a power loss moments after the call can still lose the replacement, or a QVD
that did not exist before — what it cannot leave is a damaged one.
Two options turn those off, each on its own:
| Option | What false does |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| atomic: false | Rewrites the file where it stands, as versions before 2.0.0 did. It keeps the file's identity — hard links, its owner, permissions set on the file itself — and needs neither room for two copies at once nor permission to create files in the directory. It also destroys the previous file as the write begins. A QVD that does not exist yet is renamed into place in either mode, since there is nothing to rewrite. |
| fsync: false | Resolves as soon as the operating system has accepted the bytes, without waiting for the disk. Faster, and a power loss just after the write can then leave a damaged file rather than the old one. |
A destination that exists and is not a regular file — a device such as /dev/null, or a FIFO — is
written in place whatever you ask for, because replacing it would replace the device itself.
Files it produces load in Qlik Sense as Qlik's own do - checked for
each kind of value it writes. A number is written as a pure number - an integer
within 32 bits as an integer, any other number as a double - a string as a string, and null as
NULL. A dual value - a QvdDual, or an object with exactly the keys number and text - is written
as a dual, which is how a number gets the text Qlik displays. As in Qlik, a field holds one text per
number, the first written. Anything else is refused with a QvdValidationError.
Reading part of a file
The three calls that read data share the same options — offset, limit (or maxRows), fields,
duals, coerceNumericStrings, onProgress and signal — because they are three answers about the
same file rather than three features.
// A window, and only the columns you need.
const df = await QvdDataFrame.fromQvd('large.qvd', {
offset: 1_000_000,
limit: 1000,
fields: ['Region', 'Amount'],
});
console.log(df.loadStats);
// { symbolTableBytes, totalRows, rowsLoaded, offset, symbolFiltering, symbolsKept }The limit is memory, not file size. A read holds the symbols of the fields it reads and what it
builds — the rows, or the columns — and reads both the symbol areas and the records a slice of at most
16 MiB at a time rather than holding the file. So a QVD of any size can be read whole when its result
fits, a window of it when the window's values fit — 100 rows of a column of 705 MB of distinct texts read
at Node's default heap — and iterate() over one holds those symbols and two chunks of rows. A read with
fields never reads the other fields' symbols at all.
Before it decodes anything, a read walks the symbols it will decode, a slice at a time, to see what each one
is — an integer, a text of so many characters, a dual — and is charged for what they decode to rather than
for their bytes on disk. The walk adds a quarter to a third to a read of every row of a column of short
values, and can nearly double the time of a small window of one, while long texts cost it little.
memorySafetyFactor: 0, which switches the check off, skips it. A page from QvdDataFrame.open() takes its
census from the walk its first touch of each column makes, which it cannot skip, and is charged for the values its
rows use, from the stretches of the column they fall in: 100 rows of that 705 MB column cost a few megabytes, not
the column. Its check() takes the page's own steps, so the two agree, and returns a promise since 4.0.0.
A read too large to fit throws a catchable QvdValidationError instead of a fatal Reached heap limit
that no try/catch can intercept. Its context.reason is 'memory', and context.check carries the
answer checkRead() would have given, bar a fields suggestion only checkRead() makes. Each of its
suggestions gives a read that fits:
try {
const df = await QvdDataFrame.fromQvd('huge.qvd');
} catch (error) {
const limit = error.context?.check?.suggestions.find((suggestion) => suggestion.option === 'limit');
if (error.context?.reason === 'memory' && limit) {
await QvdDataFrame.fromQvd('huge.qvd', {limit: limit.value});
}
}There is no limit suggestion when no row count helps, because the symbols alone exceed the budget.
context.recommendedMaxRows is 0 then, so never retry with it unchecked: {maxRows: undefined} reads
the whole file. Raise the heap as the --max-old-space-size suggestion says, or read fewer
columns: checkRead() with the same options suggests which. When it is iterate() that overflowed, the suggestion is chunkSize rather than limit,
and the context carries recommendedChunkSize as well.
Metadata
Table name, comments, number formats and field tags survive a read, can be changed on any frame, and
are written back out - less a tag or number format the written values contradict, such as $numeric
on a field that now holds text.
const df = await QvdDataFrame.fromQvd('products.qvd');
df.fileMetadata.tableName; // 'Products'
df.getFieldMetadata('ProductKey'); // { comment, numberFormat, tags, ... }
df.setFileMetadata({tableName: 'UpdatedProducts'});
await df.toQvd('products-v2.qvd');File paths are sandboxed
Every path is checked against an allowed directory before it is opened. Leaving the option out does not mean "anywhere" — it means the current working directory.
await QvdDataFrame.fromQvd('sales.qvd', {allowedDir: '/var/data/qvd'});
await QvdDataFrame.fromQvd(anyPath, {allowedDir: '/'}); // deliberately unrestrictedContainment is decided by the filesystem — both paths are resolved through symlinks and compared by device and inode — so a link inside the allowed directory pointing out of it is refused, whether or not the file it points at exists.
A check is only as good as the open that follows it, so the open is held to what was checked. The
path is checked again immediately before the file is opened, and the file opened is the one the
check approved: a symlink swapped in at its name afterwards is refused rather than followed, and so
is a different file at that name. Either is a QvdSecurityError whose context.reason is
'changed_after_check'. What cannot be covered from Node — a directory further up the path
swapped for a symlink in the moment between the check and the open — is the one gap left, so keep
the allowed directory's own subdirectories out of untrusted hands.
Errors
QvdError and five subclasses, each carrying a code and a context object:
QvdValidationError (QVD_VALIDATION_ERROR), QvdCorruptedError (QVD_CORRUPTED_ERROR),
QvdParseError (QVD_PARSE_ERROR), QvdIOError (QVD_IO_ERROR) and QvdSecurityError
(QVD_SECURITY_ERROR). A file-system failure — a missing file, a directory, a permission error, a
full disk, a file another program has locked on Windows — is a QvdIOError: the system code is
error.context.code (ENOENT, EACCES, ENOSPC, …) and Node's own error is error.cause. On
Windows an EPERM or EBUSY from error.context.operation === 'rename' means the destination is open
in another program — Qlik Sense reading it, most often — and the file it holds is unchanged. A
QvdSecurityError says why in error.context.reason: 'outside_allowed_directory', 'null_byte',
or 'changed_after_check' for a file that was not the one the check approved by the time it was
opened. The one
QvdIOError without error.context.code or error.cause is a write that the system accepted none of
without reporting an error; error.context.filePosition says where the file stops.
error.name is always the class name, in the CommonJS and ES module builds alike, so it works across
a module boundary where instanceof may not.
Known limitations
Honest boundaries rather than an issue list — these are the ones that change what you can do.
| Limitation | What it means in practice |
| ------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| A selected column's symbols are always walked whole | {offset, limit}, iterate() and QvdColumnTable all avoid materialising rows, but each walks the whole symbol area of every column it selects, whatever its window, a slice of at most 16 MiB at a time: a stored index in any row can address any value, and only a walk to the end of an area can check it holds the values its header claims. A read of every row decodes every value in those areas. The first page of an open file to touch a column walks it the same way, once, and then decodes only the values its rows use, from checkpoints the walk recorded, keeping them for later pages; a window over a symbol table larger than symbolFilteringThreshold (50 MiB by default) decodes only the values its rows use, though it still holds a slot for each value it walks past, so on a column of short values it saves little. fields bounds both, since the other columns' areas are never read. No read is constant-memory in the number of values in a high-cardinality column. |
| A killed process can leave a temporary file | A write builds <name>.qvdjs-<hex>.tmp beside the QVD and removes it however the write ends — unless the process is killed outright, or the removal is itself refused, in which case the error names the file in context.temporaryFile. The QVD is untouched either way, and such a file can be deleted. {atomic: false} writes none when it rewrites a QVD that exists, and damages the QVD instead; a QVD that does not exist yet is renamed into place in either mode. |
| The writer takes numbers, strings, dual values and null | Anything else - a Date, a boolean, an array, a QvdSymbol - is refused with a QvdValidationError naming the field and the row. Convert first: a Date to new QvdDual(dateToQlikSerial(date), text), the serial and the text Qlik shows, which is how Qlik stores a date; a boolean to -1 and 0, as a Qlik comparison stores it. |
Documentation
qvdjs.ptarmiganlabs.com is the documentation site, and it is where everything above is covered properly:
| | | | ------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------- | | Getting started | Install, choose an entry point, first read and write | | Guides | One page per task, runnable sample code | | Concepts | The QVD format, symbols and duals, bit stuffing, the memory model | | Reference | Every class, method, option and error | | Performance | Measured baselines, and how to read them | | Testing | How qvdjs is tested, against what, and how often | | Troubleshooting | What each failure means, and what to do about it |
Benchmarks run weekly and publish to ptarmiganlabs.github.io/qvdjs.
Reporting bugs
Email [email protected]. Include the qvdjs and Node versions,
the call that failed, and the whole error, with its code and context.
Contributors
- Göran Sander and Ptarmigan Labs — creator of qvdjs. Added lazy loading, columnar and chunked reads, header-only metadata reads, exposed QVD header metadata, ESM/CJS support, memory safety limits, security hardening, multi-platform testing and the automated release process.
- Constantin Müller — author of qvd4js, whose code qvdjs started from. Parts of it are still in qvdjs, used under qvd4js's MIT licence.
Licence
MIT © Ptarmigan Labs AB, and © Constantin Müller for the code that
comes from qvd4js. The full text is in the LICENSE file shipped with the package.
