x11-dri
v0.9.1
Published
Optional native companion to the pure-JS x11 package: direct OpenGL ES rendering (GBM/EGL, dlopen'd at runtime) and dma-buf plumbing (udmabuf, DMA_BUF sync, dup) for the DRI3 + Present path on Linux — and on macOS, the client half of XQuartz's Apple-DRI d
Maintainers
Readme
x11-dri — native companion for node-x11's DRI3 + Present path
The x11 package is pure JavaScript
and stays that way: its
DRI3 extension
can pass dma-buf descriptors to the X server over the ordinary unix-socket
connection with no native code at all. What JavaScript cannot do is produce
those dma-bufs — that takes a GPU driver or an ioctl. This optional addon
fills exactly that hole:
Gpu/Surface— an OpenGL ES 2.0 or 3.0 rendering context on a DRM render node (GBM + EGL) whose finished frames are exportable as dma-buf fds: render →swap()→{fd, stride, modifier}→DRI3.PixmapFromBuffer→Present.Pixmap, with a swapchain that resizes in place with the window. (Linux)apple.Context— on macOS, the client half of XQuartz'sApple-DRIdirect rendering: import the WindowServer surface the X server exported for a window and bind a real-GPU CGL context to it, driven by the sameglbelow — see the macOS section.gl— a WebGL-flavored subset of GL ES 2.0 driving that context from JS: shaders and programs, buffers and vertex attributes, draws, textures (including compressed uploads), blending, framebuffer objects for rendering to a texture, the uniform setters, program introspection, andreadPixels— plus vertex array objects, instanced drawing, multiple render targets, 3D/array textures, multisampling, fences and GPU timer queries where the driver has them.createUdmabuf(size)— CPU memory turned into a dma-buf by the kernel's/dev/udmabuf, with the pixels mapped into JS as anArrayBuffer: the GPU-less way to feed DRI3 (the same trick Xwayland uses), where supported by the server's driver.gl.importDmabuf({...})— the same traffic in reverse: a dma-buf someone else allocated becomes a GL texture with no copy. That is the last step of the compositor path —Composite.NameWindowPixmap→DRI3.BuffersFromPixmap→ here — and how a decoder's or another client's buffer gets sampled. (Linux)mapDmabuf(fd, size?)— the CPU-read counterpart: a descriptor the caller did not allocate, mapped into JS as anArrayBuffer. WhatcreateUdmabufalready gives for one it did. (Linux)dup(fd),dmabufSync(fd, flags)— descriptor plumbing (DRI3 sends consume their fds;dupkeeps a copy) and CPU-access bracketing.
Complete samples live in the main repo — a self-contained folder you can npm install && npm start:
examples/dri3/cube.js
(spinning GPU cube) and
examples/dri3/software.js.
The examples/ folder here is the other half of that: no X server
and no DRI3, one GL feature per file, rendering off-screen and writing a PNG
you can open.
Installing
npm install x11-dri # no toolchain needed on linux x64/arm64, macOS arm64The npm tarball bundles prebuilt binaries for linux-x64,
linux-arm64 (glibc ≥ 2.31 — Debian 11 / Ubuntu 20.04 and everything
newer), darwin-arm64 and darwin-x64 (macOS 11+, Apple Silicon and
Intel), built in CI from the released tag. The install script just verifies
the matching one loads, so a box with no build tools installs from the
tarball alone — and because the loader also resolves the prebuild for
process.platform-process.arch at require() time, the package keeps
working under npm install --ignore-scripts. The addon is Node-API, so one
binary per platform/arch covers every supported Node (and Electron)
version.
Anything else (musl/Alpine, armv7, riscv64, forced rebuilds
with --build-from-source) compiles automatically with node-gyp, and that
needs only a C toolchain: the addon has no build-time dependency on
gbm/EGL/GLES — libgbm.so.1, libEGL.so.1 and libGLESv2.so.2 are
dlopen()ed at runtime (Mesa's ABI is stable) and it degrades with clear
errors where a library or device is missing. probe() reports what is
available; npm test runs a self-check that skips whatever this machine
lacks.
The dma-buf/DRI3 features are Linux-only — that is a property of the
platform, not a missing port (see the macOS section for why). On macOS the
package instead carries the native half of XQuartz's own direct-rendering
path: the apple namespace binds a real-GPU CGL context to an XQuartz
window through the Apple-DRI extension, driven by the same gl object and
the same ES2 shaders. Everything Linux-specific still loads there and
reports itself unavailable, so cross-platform code can require() the
package unconditionally and branch on probe().
API sketch
const dri = require('x11-dri');
dri.probe(); // { gbm, egl, gles, udmabuf }
dri.listRenderNodes(); // ['/dev/dri/renderD128', ...]
const gpu = new dri.Gpu({ // opens a render node (no X auth
devicePath: undefined, // needed), gbm + EGL + a context
format: dri.FORMAT.XRGB8888, // must match the window depth
depthSize: 16, // EGL depth buffer bits
stencilSize: 0, // 8 to stencil — see below
glVersion: 'auto' // 'auto' | 3 | 2 — see below
});
const surface = gpu.createSurface(w, h); // GBM swapchain (add
gpu.makeCurrent(surface); // dri.GBM_USE.LINEAR for
// cross-device consumers)
const gl = gpu.gl; // clearColor, shaders, drawElements…
// ... draw ...
const out = surface.swap();
// out: { key, generation, isNew, width, height } and, the first time a buffer
// appears, { fd, stride, offset, modifier (BigInt) } — hand fd to
// DRI3.PixmapFromBuffer (it is consumed), cache the pixmap by
// out.generation + out.key (see below).
// out === null: every buffer still held — wait for PresentIdleNotify.
surface.release(out); // when PresentIdleNotify says so
surface.resize(w2, h2); // window resized: same Surface, new
// swapchain, generation moves on
surface.destroy(); gpu.destroy();One EGL context per Gpu, one thread, GL calls valid between makeCurrent
and destroy — deliberately no more machinery than a renderer needs.
Resizing, and what a key is unique within
A window that resizes does not need a new surface:
surface.resize(width, height); // the same Surface, a new swapchain
surface.generation; // 0, 1, 2, … — every resize moves it onThe handle stays valid, the context keeps its GL objects and stays current if
it was, the buffers still locked are released, and the next swap() reports a
fresh set — every one of them isNew, with a new dma-buf fd to import.
Resizing to the size it already has does nothing at all, so a drag can call
this on every ConfigureNotify and pay only for the sizes that differ. What
it will not do is decide when: rounding the window size up to a granularity,
so that a continuous drag reallocates a handful of times instead of once per
frame, is policy, and policy belongs to the caller — this is the mechanism
under it.
generation is the other half, and it earns its place whether or not anything
ever resizes. key is a GEM handle: per-DRM-fd, and recycled as soon as the
buffer behind it is freed. It is unique among the buffers of one swapchain
and no further. Here is the same driver handing the same handle to a buffer of
a different size two resizes later:
| generation | size | key | stride | | --- | --- | --- | --- | | 0 | 64×64 | 1 | 256 | | 1 | 128×96 | 3 | 512 | | 2 | 256×192 | 1 | 1024 |
So key the pixmap cache by the pair, and give the swap result itself back to
release() rather than the bare key:
const out = surface.swap();
cache.set(`${out.generation}:${out.key}`, pixmap); // not out.key alone
...
surface.release(out); // false, and harmless, if a resize got there firstA late release(out.key) — an IdleNotify for a buffer of the swapchain that
just went away — would free whichever live buffer inherited that handle.
release(out) reads the generation off the result, sees it is not the current
one, and answers false instead. The pixmaps of a generation you have left
behind are yours to free: the buffers under them are already gone.
TypeScript
Declarations ship with the package (index.d.ts), so there is nothing to
install and no @types entry to look for. The runtime is CommonJS, so they
are named exports — import { Gpu } from 'x11-dri' — with no default export,
because there is no .default at runtime.
They describe the binding rather than WebGL, and the differences are the useful part:
import { Gpu, GLContext } from 'x11-dri';
const gpu = new Gpu({ glVersion: 3 });
const surface = gpu.createSurface(1024, 768);
gpu.makeCurrent(surface);
const out = surface.swap();
if (out?.isNew) {
const fd: number = out.fd; // only in scope because isNew narrowed it
}
if (gpu.features?.instancedArrays)
gpu.gl.drawArraysInstanced(gpu.gl.TRIANGLE_STRIP, 0, 4, 1000);swap() returns a union discriminated on isNew, so the dma-buf fields are
reachable exactly where they exist. features and glVersion are optional
until makeCurrent has run, which is when they become knowable. GL objects
are plain numbers, getUniformLocation answers -1 rather than null, and
getShaderParameter answers a number — all as the binding does, not as
WebGL's typings do.
Two checks keep the declarations honest: npm test compares every declared
name against the addon's actual exports in both directions (no GPU and no
TypeScript needed, so it runs everywhere), and npm run test:types compiles
test-types.ts against them under strict, where a row of
@ts-expect-error lines fail the build if the mistakes below them stop being
mistakes.
Which ES version
glVersion defaults to 'auto': ask EGL for ES 3.0, and fall back to ES 2.0
when the display has no ES 3.0-capable config or refuses the context. Pass
3 to insist — you get an error naming the version rather than a silent
downgrade — or 2 to pin.
Two numbers come back, and they are not the same one:
gpu.contextVersion // 2 or 3: what EGL was asked for and granted
gpu.makeCurrent(surface);
gpu.glVersion // { major, minor, string } — what the driver reportsA version request is a floor, not a ceiling. ES 3.0 is backward
compatible with ES 2.0 and EGL is allowed to hand back more than was asked
for: Mesa answers glVersion: 2 with an ES 3.0 context, so
contextVersion === 2 with glVersion.major === 3 is normal and not a bug.
glVersion is the one to branch on — it is what the driver will actually
honour, and it is what gates the optional entry points above.
gpu.glVersion needs a current context to exist, so like features it
appears on makeCurrent rather than in the constructor.
Depth and stencil bits
depthSize and stencilSize are part of the EGL config query, not of the
context, so the constructor is the only place they can be asked for — there is
nothing to turn on afterwards:
const gpu = new dri.Gpu({ depthSize: 24, stencilSize: 8 });
gpu.depthSize // what the chosen config carries, e.g. 24
gpu.stencilSize // 8 — 0 when it was not asked forstencilSize defaults to 0, which is why it has to be asked for: a
default framebuffer with no stencil bits passes every stencil test, so
stencilFunc/stencilOp/clearStencil run without error and change nothing
at all. That is the failure mode behind a stencil-then-cover fill — the
standard way to fill an arbitrary vector path, and what
NVG_STENCIL_STROKES-style overlapping strokes need — painting everywhere
instead of inside the path.
Like a version request, a bit count is a floor, not an exact size: EGL may
hand back a config with more (a 16-bit depth request usually lands on 24, and
asking for stencil usually brings a packed depth24/stencil8 along with it).
gpu.depthSize and gpu.stencilSize read back what the config actually
carries — the same relationship contextVersion has with glVersion, except
that these are known from the constructor, no current context required.
The alternative, when the surface has no stencil bits, is to render to a
framebuffer of your own with a DEPTH24_STENCIL8 renderbuffer and blit — an
extra full-surface pass every frame, and still no way to stencil to the
default framebuffer.
macOS takes the same two options on dri.apple.Context, spelled identically
(kCGLPFADepthSize / kCGLPFAStencilSize there), so a renderer asks for a
stencil buffer the same way on both. There is no read-back on that side: the
request is all AppleContext records.
What gl covers
Enough of ES 2.0 to drive a real renderer, in WebGL's spelling and argument order, so code and tutorials carry over:
| area | entry points |
| --- | --- |
| programs | createShader, shaderSource, compileShader, getShaderParameter, getShaderInfoLog, createProgram, attachShader, linkProgram, getProgramParameter, getProgramInfoLog, useProgram, bindAttribLocation, deletes |
| geometry | createBuffer, bindBuffer, bufferData, bufferSubData, vertexAttribPointer, enableVertexAttribArray, disableVertexAttribArray, vertexAttrib1f–4f, drawArrays, drawElements |
| uniforms | getUniformLocation, uniform1f/2f/3f/4f, uniform1i/2i/3i/4i, uniform1fv–4fv, uniform1iv, uniformMatrix2fv/3fv/4fv |
| textures | createTexture, bindTexture, activeTexture, texImage2D, texSubImage2D, compressedTexImage2D, compressedTexSubImage2D, texParameteri/f, generateMipmap, deleteTexture |
| framebuffers | createFramebuffer, bindFramebuffer, framebufferTexture2D, framebufferRenderbuffer, checkFramebufferStatus, createRenderbuffer, bindRenderbuffer, renderbufferStorage, deletes |
| per-fragment state | blendFunc, blendFuncSeparate, blendEquation, blendEquationSeparate, blendColor, depthFunc, depthMask, depthRange, colorMask, scissor, polygonOffset, stencilFunc, stencilOp, stencilMask, stencilFuncSeparate, stencilOpSeparate, stencilMaskSeparate, clearStencil, cullFace, frontFace |
| introspection | getActiveUniform, getActiveAttrib, getUniform, getAttachedShaders, getShaderSource, getShaderPrecisionFormat, getVertexAttrib, getVertexAttribOffset, getBufferParameter, getTexParameter, getFramebufferAttachmentParameter, getRenderbufferParameter, getSupportedExtensions, validateProgram, isBuffer/isProgram/isShader/isTexture/isFramebuffer/isRenderbuffer/isEnabled |
| optional (see features) | createVertexArray, bindVertexArray, deleteVertexArray, isVertexArray; drawArraysInstanced, drawElementsInstanced, vertexAttribDivisor; drawBuffers; texImage3D, texSubImage3D, copyTexSubImage3D, compressedTexImage3D, compressedTexSubImage3D, framebufferTextureLayer; texStorage2D, texStorage3D; renderbufferStorageMultisample, blitFramebuffer; readBuffer; fenceSync, clientWaitSync, waitSync, getSyncParameter, deleteSync, isSync; createQuery, deleteQuery, isQuery, beginQuery, endQuery, getQuery, getQueryParameter, getQueryObjectui64v; queryCounter; importDmabuf |
| the rest | clear, clearColor, clearDepthf, viewport, enable, disable, lineWidth, pixelStorei, getParameter, getIntegerv, getFloatv, getBooleanv, getError, getString, readPixels, finish, flush |
texImage2D accepts null pixels, which is how a texture is allocated to be
rendered into. getParameter answers in the type the parameter has — a
number, a boolean, or an array for VIEWPORT, SCISSOR_BOX,
COLOR_CLEAR_VALUE, COLOR_WRITEMASK and COMPRESSED_TEXTURE_FORMATS;
getIntegerv/getFloatv/getBooleanv are the raw single-value escape hatch.
Introspection is how code that did not write the shader drives it anyway.
getActiveUniform(program, i) walks the uniforms a linked program actually
kept, answering { name, size, type } (and null past the end, so a loop
can stop without disturbing getError()); arrays appear once, as u[0] with
their length. getUniform(program, location) reads a value back in its own
type — a Float32Array for a vec, a boolean for a bool — by asking the
program what that uniform is; getUniformfv/getUniformiv are the raw form
that takes a component count instead.
Compressed uploads pass their bytes to the driver untouched: the block layout
belongs to the format, and the byte count comes from the TypedArray. Which
internalformat values are legal is per-driver, so ask first —
getSupportedExtensions() names the formats, getParameter(gl.COMPRESSED_TEXTURE_FORMATS)
enumerates the enums. examples/compressed-texture.js encodes DXT1 blocks by
hand and renders the result beside the uncompressed original.
What is optional, and how to ask
Vertex array objects, instanced drawing, multiple render targets, 3D and array textures, immutable storage, multisampling, the read buffer and sync objects are core in ES 3.0 and extensions before it; GPU timer queries are an extension on every ES version. So whether they exist at all is a property of the driver and the context rather than of this build. They are resolved separately from everything else, against a live context, and reported per feature:
gpu.makeCurrent(surface);
gpu.features // { vertexArrayObject, instancedArrays, drawBuffers,
// texture3D, textureStorage, multisample,
// readBuffer, sync, timerQuery, timestampQuery,
// dmabufImport, dmabufImportModifiers,
// dmabufImportExternal }
if (gpu.features.instancedArrays)
gl.drawArraysInstanced(gl.TRIANGLE_STRIP, 0, 4, count);features appears on makeCurrent because that is the first moment the
answer is knowable, and it is refreshed on each call; apple.Context sets it
on attach and makeCurrent, and gl.getFeatures() answers for whichever
context is current — the one to ask from code that holds only gl. It is the
probe, and typeof gl.beginQuery === 'function' is not: gl is one object
shared by every context of both flavors, so every function is always there. A
feature is true only when every entry point it needs resolved, so a driver
offering half an extension reports it absent rather than throwing partway
through a frame. Calling one that is missing throws a message naming the
feature — the wrappers never call through a null pointer.
Two details this hides. Mesa exports the whole ES 3.2 symbol set from
libGLESv2.so.2 whatever the context supports, so finding
glDrawArraysInstanced there says nothing about being allowed to call it —
the core spellings are gated on the context reporting ES 3.0. And drivers
that have the extension but not the core function often export neither,
offering glDrawArraysInstancedEXT (or …ANGLE, or …NV) through
eglGetProcAddress alone — so each feature carries a list of candidate
spellings and takes the first that both resolves and is advertised.
The newer groups, per flavor. On the CGL flavor "always" means Apple's GL 4.1 core profile, the default; the legacy profile's route is in parentheses.
| feature | entry points | GLES flavor (Gpu) | CGL flavor (apple.Context) |
| --- | --- | --- | --- |
| multisample | renderbufferStorageMultisample, blitFramebuffer | ES 3.0; on ES 2.0 the ANGLE or NV framebuffer_multisample + framebuffer_blit pair | always (ARB_framebuffer_object) |
| readBuffer | readBuffer | ES 3.0, or NV_read_buffer | always (core since GL 1.0) |
| sync | fenceSync, clientWaitSync, waitSync, getSyncParameter, deleteSync, isSync | ES 3.0, or APPLE_sync | always (ARB_sync) |
| timerQuery | createQuery, deleteQuery, isQuery, beginQuery, endQuery, getQuery, getQueryParameter, getQueryObjectui64v | only with GL_EXT_disjoint_timer_query, on ES 3.0 as on 2.0 | always (EXT_timer_query) |
| timestampQuery | queryCounter | only with GL_EXT_disjoint_timer_query | always, over a clock of 0 bits (absent) |
The separate stencil setters are core ES 2.0 and need no flag.
3D and array textures
Same call, different target: TEXTURE_3D filters across the third axis,
TEXTURE_2D_ARRAY keeps its layers independent — a fractional layer index
rounds to one of them rather than blending two, which is what makes an array
texture the right home for an atlas and a 3D texture the right home for a
volume.
gl.bindTexture(gl.TEXTURE_2D_ARRAY, tex);
gl.texStorage3D(gl.TEXTURE_2D_ARRAY, 1, gl.RGBA8, w, h, layers); // allocate once
gl.texSubImage3D(gl.TEXTURE_2D_ARRAY, 0, 0, 0, 0, w, h, layers,
gl.RGBA, gl.UNSIGNED_BYTE, pixels); // then fill
gl.framebufferTextureLayer(gl.FRAMEBUFFER, gl.COLOR_ATTACHMENT0, tex, 0, 2);texStorage2D/texStorage3D allocate the whole mipmap pyramid once in a
sized format and refuse to be called twice; only the contents change
afterwards, through texSubImage. Sampling either target needs GLSL ES 3.00
(sampler3D, sampler2DArray), so it needs an ES 3.0 context — see
glVersion above. examples/texture-3d.js ray-marches a 64³ volume beside
the array texture and the slices it interpolates between.
One-pass non-zero fills, and multisampled edges
stencilFuncSeparate, stencilOpSeparate and stencilMaskSeparate set one
facing's stencil state and leave the other's alone, which is what counts a
winding number in a single pass — front faces up, back faces down:
gl.colorMask(false, false, false, false);
gl.stencilFunc(gl.ALWAYS, 0, 0xff);
gl.stencilOpSeparate(gl.FRONT, gl.KEEP, gl.KEEP, gl.INCR_WRAP);
gl.stencilOpSeparate(gl.BACK, gl.KEEP, gl.KEEP, gl.DECR_WRAP);
drawFans(); // both windings, one draw
gl.colorMask(true, true, true, true);
gl.stencilFunc(gl.NOTEQUAL, 0, 0xff); // the non-zero rule
gl.stencilOp(gl.ZERO, gl.ZERO, gl.ZERO); // clean for the next layer
drawCover();A framebuffer with no stencil buffer passes every stencil test, so make sure
there is one. Neither backend's default framebuffer has one unless the context
asked for it: new dri.Gpu({ stencilSize: 8 }) on Linux (gpu.stencilSize
reads back what the config granted — see
Depth and stencil bits),
new dri.apple.Context({ stencilSize: 8 }) on macOS. A framebuffer you build
yourself needs a DEPTH24_STENCIL8 renderbuffer attached; createTarget's
already has one. The test is to draw with NOTEQUAL 0 over a cleared
stencil — which must draw nothing.
Multisampling (features.multisample) is render, then resolve: draw into a
framebuffer of renderbufferStorageMultisample renderbuffers, and
blitFramebuffer it into a single-sample one, which is the only way to read
or sample what was drawn:
const samples = Math.min(4, gl.getParameter(gl.MAX_SAMPLES));
gl.renderbufferStorageMultisample(gl.RENDERBUFFER, samples, gl.RGBA8, w, h);
// ... attach it to msaaFbo, draw ...
gl.bindFramebuffer(gl.READ_FRAMEBUFFER, msaaFbo);
gl.bindFramebuffer(gl.DRAW_FRAMEBUFFER, target.fbo); // an IOSurface target will do
gl.blitFramebuffer(0, 0, w, h, 0, 0, w, h, gl.COLOR_BUFFER_BIT, gl.NEAREST);ES 3.0 resolves only between identical formats and identical rectangles.
readBuffer (features.readBuffer) chooses which attachment a blit, or
readPixels, reads.
Timing a frame without stalling it
finish() times a frame by waiting for it, stalling the thread — on the
Cocoa path until the surface is done too. A GPU timer brackets the frame
instead, and its reading is there a frame or two later with nobody having
waited for it:
// each frame
const q = gl.createQuery();
gl.beginQuery(gl.TIME_ELAPSED, q);
drawFrame();
gl.endQuery(gl.TIME_ELAPSED);
pending.push(q);
// then, each frame, whatever has come back — asking never waits
while (pending.length && gl.getQueryParameter(pending[0], gl.QUERY_RESULT_AVAILABLE)) {
const done = pending.shift();
const ns = gl.getQueryParameter(done, gl.QUERY_RESULT);
if (!gl.getParameter(gl.GPU_DISJOINT_EXT)) // false on desktop GL
learn(ns);
gl.deleteQuery(done);
}- Asking whether there is one.
features.timerQuery, thengl.getQuery(gl.TIME_ELAPSED, gl.QUERY_COUNTER_BITS) > 0: a driver may offer the entry points over a counter of no bits. On the CGL flavor timers are core GL 3.3 (Apple Silicon gives a 32-bit counter; 64 blended quads on an M1 Pro read about 0.6 ms). On the GLES flavor they needGL_EXT_disjoint_timer_query— on ES 3.0 as much as 2.0, since ES's own query objects have no timer. - Disjoint readings. The ES extension reports when something (a power
state change, say) disturbed the GPU clock, voiding the readings taken since
the last ask.
getParameter(gl.GPU_DISJOINT_EXT)answers — and resets — it there, and answersfalseon desktop GL, which has no such notion, without raising theINVALID_ENUMGL would; so the loop above is the same on both flavors. - Precision. Results are 64-bit nanosecond counts.
getQueryParameteranswers a Number, exact below 2^53 ns — 104 days, far past any frame.getQueryObjectui64vanswers the same reading as a BigInt with all 64 bits, for the one kind of value that can pass 2^53: an absoluteTIMESTAMP. - Timestamps.
queryCounter(q, gl.TIMESTAMP)(features.timestampQuery) needsgetQuery(gl.TIMESTAMP, gl.QUERY_COUNTER_BITS) > 0as well, and Apple's GL answers 0 — on macOS, time spans withTIME_ELAPSED. That width is a claim rather than a promise: virgl answers 64 bits and then reads back 0 every time, so a clock counts as ticking only once two readings differ. - The stall is still there if asked for.
QUERY_RESULTfor a result not yet available waits for it. AskQUERY_RESULT_AVAILABLEfirst.
Fences (features.sync) answer the coarser question — has the GPU finished
everything up to here — in the same non-blocking way. fenceSync(gl.SYNC_GPU_COMMANDS_COMPLETE, 0)
after a frame, then clientWaitSync(fence, 0, 0) polls: TIMEOUT_EXPIRED
until it has passed, ALREADY_SIGNALED after. flush() first, or pass
SYNC_FLUSH_COMMANDS_BIT, or the fence may never reach the GPU to signal. A
timeout other than 0 is nanoseconds, as a Number or a BigInt, and waitSync
takes gl.TIMEOUT_IGNORED (WebGL 2's -1, since a Number cannot hold GL's
2^64 - 1). The handle fenceSync returns is a number like every other object
here, not the driver's pointer: the binding keeps the pointer, and refuses a
handle that is deleted, made up, or another context's rather than hand GL
something it cannot validate.
Importing a dma-buf
Everything above produces buffers; this consumes one. A descriptor that
arrives from anywhere else — DRI3.BuffersFromPixmap on a redirected
window's pixmap, a video decoder, another process — becomes an EGLImage
and then an ordinary texture, sampled where it lies:
// X side (pure JS, node-x11 >= 4.1 under Bun's receiveFds): the pixmap of a
// redirected window, described as planes
const buf = await dri3.BuffersFromPixmap(pixmap);
// { width, height, modifier, depth, bpp, planes: [{ fd, stride, offset }] }
const image = gl.importDmabuf({
width: buf.width, height: buf.height,
fourcc: dri.FORMAT.XRGB8888, // any DRM fourcc, not just these two
modifier: buf.modifier, // omit for the implicit-modifier form
planes: buf.planes
});
gl.bindTexture(image.target, image.texture); // then sample it like any other
// ... draw ...
image.destroy(); // glDeleteTextures + eglDestroyImageKHRThe plane shape is exactly what BuffersFromPixmap replies with, and exactly
what PixmapFromBuffers takes, because it is the same description of the same
buffer travelling the other way.
Four things worth knowing:
- The descriptors are consumed. On success
importDmabufcloses them, the same ownership rule DRI3's fd-carrying requests andswap()'s exported fd follow — EGL holds its own reference to the buffer from then on. On failure it throws and leaves them open, so a caller can report or retry.dup()first if you need to keep one. targetis in the answer, not assumed. A single RGB plane lands onTEXTURE_2D; multi-planar and YUV imports come back asTEXTURE_EXTERNAL_OES, which GLSL reads through asamplerExternalOESrather than asampler2D. Passtargetexplicitly to insist.- Ask before you try.
gpu.features.dmabufImportsays whether the driver hasEGL_EXT_image_dma_buf_importat all,features.dmabufImportModifierswhether an explicitmodifiermay be passed (without it, only the implicit form works), andfeatures.dmabufImportExternalwhether the external target exists. All three arefalseon the macOS/CGL backend, which has no EGL.probe()cannot answer these — they are extensions of an initialized EGL display, andprobe()opens no devices — so it returns a string saying to readfeaturesaftermakeCurrent()instead. destroy()belongs to the importing context. A texture name means something else in every other context, so destroying an image while a differentGpuis current throws rather than deleting whatever that context happens to call by the same number —makeCurrentthe right one first. Destroying theGputakes its images with it, so adestroy()after that does nothing, as does a second one.
examples/dmabuf-import.js renders both sides of that: one udmabuf shown
twice, imported on the left and mapped-then-uploaded on the right.
For a buffer the CPU should read rather than the GPU, mapDmabuf(fd, size?)
returns { buffer, size, writable, sync(flags), close() } — the same shape
createUdmabuf returns, minus the allocation. The descriptor is borrowed
there, not consumed: it stays yours to send on. Only exporters that implement
mmap can be mapped at all (udmabuf and linear/dumb buffers do, tiled GPU
allocations generally do not), and close() releases the mapping at once
rather than waiting for the collector.
writable is the one to check before writing through buffer. A dma-buf
descriptor carries an access mode like any other fd, and a read-only one is
ordinary rather than exotic: DRM_RDWR is opt-in when a buffer is exported,
and gbm_bo_get_fd does not pass it — so even this package's own swap()
descriptor maps read-only on most drivers. Mapping still succeeds, for
reading; it is writing through such a mapping that faults.
Still not covered
The rest of ES 3.0: sampler objects, uniform buffer objects, transform
feedback, primitive restart, and getUniformuiv for unsigned-integer uniforms
(getUniform answers null for a type it cannot read). Query objects are
here for the timers; their occlusion targets would go through the same calls
but have no constants in GL. Adding one is still a small wrapper per entry
point in src/x11dri.c plus a line in the EXPORT block — the JS name is
derived from the glFoo export automatically, and anything past ES 2.0
belongs in the optional table beside the features above.
How it fits together
GPU (render node) X server
----------------- --------
render into a buffer
export -> dma-buf fd --- fd over unix socket (DRI3) ---> pixmap
Present.Pixmap(window, pixmap) ------ vsync'd flip/copy -> on screen
<--- PresentCompleteNotify (pace the next frame)
<--- PresentIdleNotify (buffer reusable)The protocol side — DRI3, Present, and the descriptor-passing socket — is implemented in pure JS by the main package; see docs/ext/dri3.md and docs/ext/present.md.
macOS / XQuartz
Two separate facts, and the second is the interesting one:
- The DRI3 path does not work under XQuartz, and cannot be made to.
- Accelerated direct rendering into an XQuartz window still works —
through XQuartz's own mechanism, the
Apple-DRIextension, whose native half this package now carries (dri.apple, macOS only).
Why DRI3 is out
Measured against XQuartz 21.1.23 (X.Org 21.1.23) on macOS 15.2, Apple Silicon:
| Piece | On XQuartz | Consequence |
| --- | --- | --- |
| Present | v1.2, works — Present.Pixmap accepted, PresentCompleteNotify and PresentIdleNotify both delivered | the pacing/buffer-recycling loop from the samples runs unchanged |
| DRI3 | not implemented — QueryExtension says absent, X.require('dri3') fails | no PixmapFromBuffer, so no way to turn a buffer fd into a pixmap |
| dma-buf | no such kernel object on Darwin | nothing to export, and nothing to send |
| GBM | no libgbm on macOS | Gpu cannot allocate exportable buffers |
| EGL | XQuartz ships Mesa's libGLESv2/libOSMesa but no libEGL | no way to create the ES context, even ignoring the above |
| udmabuf | /dev/udmabuf is a Linux driver | the CPU-memory fallback is out too |
| dma-buf import | needs EGL and a dma-buf to import | gl.importDmabuf and mapDmabuf report unavailable |
The gaps compound: even with a GPU context, there is no dma-buf to export;
even with a dma-buf, there is no DRI3 request to hand it to. The pure-JS
fallback that always works is CreatePixmap + PutImage + Present.Pixmap
— a copy per frame, no native code.
What works instead: Apple-DRI
XQuartz's direct rendering runs the buffer handoff in the opposite direction from DRI3. The client does not produce a buffer and send it; the server exports the window's own WindowServer surface to the client, and the client renders straight into it:
this process X server (XQuartz)
------------ ------------------
apple.clientId() --- AppleDRICreateSurface(win, cid) ---> exports the
<-------------- key[2] ---------------- window's surface
ctx.attach(key) (xp_import_surface + CGL context)
glDraw... straight into the window's backing store
ctx.flush() (CGLFlushDrawable — WindowServer composites)
<--- AppleDRISurfaceNotify ------------ moved/resized:
ctx.update()After the attach, nothing crosses the X socket per frame — no pixels, no
requests. This is exactly the machinery XQuartz's libGL gives GLX clients
(Apple's private-but-ABI-stable Xplugin library plus CGL), minus GLX; the
context is the real GPU (OpenGL-on-Metal — glGetString(RENDERER) answers
"Apple M1 Pro" and the like).
The division of labor mirrors the DRI3 path exactly. The X protocol side —
AppleDRICreateSurface, the SurfaceNotify event — is plain protocol for
the x11 package (see
examples/xquartz/appledri.js, written in
node-x11's extension style). This addon supplies what JavaScript cannot: the
WindowServer handshake, the surface import, and the CGL context.
const cid = dri.apple.clientId(); // WindowServer handshake
// X side: AppleDRI.CreateSurface(screen, wid, cid) -> { key, uid }
const ctx = new dri.apple.Context({ depthSize: 16 });
ctx.attach(key); // import + bind; context is current
// ... render with dri.gl — same object, same calls as the Linux path ...
ctx.flush(); // present (this backend's swap)
// on ConfigureNotify / SurfaceNotify(changed): ctx.update()
// on SurfaceNotify(destroyed): CreateSurface again, then ctx.attach(newKey)examples/xquartz/cube.js is the complete program — the
DRI3 cube sample, ported to this path.
What makes the same gl object work on macOS: the system OpenGL framework
exports the whole ES 2.0 name set, and the context is created as a core
profile (GL 4.1 on Metal hardware), where ARB_ES2_compatibility compiles
GLSL ES 1.00 — so ES2/WebGL1 shaders run unchanged, with two core-profile
potholes smoothed over inside the addon: source with no #version gets
#version 100 prepended (as its own source string, so your line numbers
survive in info logs), and the context carries a default vertex array object
the way browsers do. On Apple's GL 4.1 every optional features entry —
VAOs, instancing, drawBuffers, 3D textures, texStorage (via
GL_ARB_texture_storage) — resolves true. Shader-language versions are the
one visible difference: GLSL ES 1.00 or desktop GLSL 4.10 compile,
#version 300 es does not (Apple never shipped ARB_ES3_compatibility).
Constraints to know about: the WindowServer only talks to processes in a
logged-in GUI session (an SSH login gets a clear error from
apple.clientId()); pacing is yours — flush() returns immediately by
default and a timer sets the frame rate (setSwapInterval(1) gives real
vsync at the price of blocking the event loop up to a frame); and Present
plays no part in this path. A timer needs a rate to aim for, and XQuartz's
RandR advertises modes with no timing data — so apple.refreshRate()
asks macOS directly: the fastest rate any connected display is running at, in
Hz (max across displays, the same pacing-ceiling semantics as taking the
fastest CRTC from RandR on Linux), or null where there is no answer, e.g.
over SSH.
probe() reports the whole picture at runtime:
require('x11-dri').probe();
// {
// platform: 'darwin',
// dmabuf: false, // false => no library will fix DRI3 here
// gbm: 'GBM needs Linux DRM/dma-buf — no equivalent on darwin',
// egl: 'dlopen(libEGL.dylib) ... no such file',
// gles: true,
// appledri: true, // libXplugin + OpenGL.framework load
// udmabuf: 'dma-buf is a Linux kernel facility with no darwin equivalent ...',
// dmabufImport: 'dma-buf is a Linux kernel facility with no darwin equivalent ...'
// }Every capability is true when usable or a string explaining why not:
dmabuf: false means the DRI3 pipeline is off the table on this host, and
appledri: true means the XQuartz pipeline is on it.
Building a cross-platform GL surface on top of this (ntk's rendering
contexts, react-x11's <glarea>, scene graphs above them):
docs/glarea-portability.md is the integration
map — the two paths side by side, the exact work items per repo, and the
portability rules for consumer code.
Why not DRI3.Open?
The protocol's own way to get a DRM device fd is a reply carrying a
descriptor, and receiving descriptors is the one thing node-x11's pure-JS
transport must never do (it aborts the Node process — see
lib/fdpass.js
in the main package). Render nodes make it unnecessary: they exist precisely
so applications can render without asking a display server for permission.
On multi-GPU machines, probe: import a test buffer per node with
DRI3.PixmapFromBuffer(..., cb) and keep the node the server accepts.
