wurld-core
v1.3.0
Published
Reader, writer and live-streaming primitives for the wurld posed sensor-video format
Maintainers
Readme
wurld
wurld — the World's Unbroken Record of Localization & Depth Posed sensor video in one playable WebM — RGB video + per-frame camera pose + intrinsics + timestamps + bit-exact metric depth (and confidence / object IDs), in a single file that an ordinary video player treats as plain RGB.
wurld is the missing interchange layer between the things that produce posed
RGBD video (phone capture apps, SLAM systems, robots, synthetic renderers) and the
things that consume it (NeRF/splat training, robot learning, world-model pipelines,
SLAM evaluation). Today every producer invents a directory layout — COLMAP dirs,
transforms.json, TUM text files, per-app zip formats — and every consumer maintains a
zoo of parsers. wurld replaces the zoo with one canonical-convention container:
- Payload: chromapakz tracks — VP9 RGB + lossless, bit-exact uint16 planes for depth / confidence / IDs. Decodes natively in the browser via WebCodecs.
- Metadata: one Matroska
WURLDtag holding cameras (COLMAP-style models), per-frame poses and sensor timestamps, signal semantics (how uint16 maps to meters), and world info (metric scale, gravity). See SPEC.md. - Conventions are fixed, not declared: RDF camera axes (OpenCV/COLMAP), camera-to-world, wxyz quaternions, meters, seconds. Converters normalize on write; consumers never branch on axis flags.
Install
pip install wurld # Python reader/writer, converters, CLI
npm install wurld-core # JavaScript: same format, byte-identical recordsOptional Python extras: wurld[record3d] for .r3d imports, wurld[mcap] for
Foxglove export, wurld[dev] for the test suite.
Quickstart
# working from a checkout instead: pip install -e ".[dev]"
# write a synthetic demo sequence (analytic RGBD + orbit poses)
wurld demo demo.wurld.webm
# inspect
wurld info demo.wurld.webm
# convert real datasets (auto-detects TUM / transforms.json / COLMAP / Stray Scanner)
wurld convert path/to/rgbd_dataset_freiburg1_desk desk.wurld.webm
wurld convert path/to/transforms.json scene.wurld.webm
wurld convert path/to/colmap_project scene.wurld.webm --images path/to/images
wurld convert path/to/stray_capture scan.wurld.webm # needs ffmpeg on PATH
# extract back out
wurld extract demo.wurld.webm out/ --format tum # or transforms | colmap
# check a file against the spec (exit 1 on a MUST violation)
wurld validate scene.wurld.webmvalidate turns SPEC's normative requirements into executable checks, so anyone
writing a producer can confirm their output before shipping it. Findings name the
section they come from — pose-table ordering and precedence, camera/video
resolution agreement, quaternion normalisation, timestamp monotonicity, value-map
sanity, rig and IMU references — with MUST violations as errors and SHOULD as
warnings.
Python API:
import wurld as wl
seq = wl.read("demo.wurld.webm")
seq.rgb # (T, H, W, 4) uint8, lazily decoded
seq.depth_meters(0) # (H, W) float, NaN = invalid
seq.c2w(0) # 4x4 camera-to-world (RDF, meters)
seq.K("0") # 3x3 intrinsics (K("0", frame_index=i) honors overrides)
seq.frames[0].t # sensor timestamp (authoritative, may be non-uniform)
seq.rigs # camera-to-rig calibration; seq.rig_c2w(i, "1") derives poses
seq.imu["imu0"].samples # (N, 7) [t, gyro xyz, accel xyz]Long sequences (>10k frames) automatically pack poses into a binary frame table
(45 bytes/frame; SPEC §7) — wl.write(..., frames_format="binary") forces it.
JavaScript API (wurld-core, no dependencies — works in browsers and Node):
import { readDocument, WurldRecorder, StreamSplitter } from 'wurld-core';
const { doc, imu, rgbStreams } = readDocument(bytes);
doc.cameras; // calibration, keyed by camera id
doc.frames[0].t; // seconds; .q_wxyz, .tr, .pose_valid
imu.imu0?.[0]; // { t, gyro, accel }
rgbStreams; // display streams present, primary firstreadDocument resolves poses through the SPEC §9 precedence chain —
WURLD_FRAMES, then concatenated WURLD_POSES chunks, then the JSON array —
which is the part worth not hand-rolling: a reader that honours only the JSON
array sees zero poses on any phone recording, and one that assumes a binary
table throws on everything else. readWurldTags and unpackFrames are still
exported for callers who want the pieces.
The 45-byte records are byte-identical to the Python writer's, and all three readers are checked against a shared conformance corpus rather than against their own expectations.
Writing live (WurldRecorder) needs a chromapakz encoder, which you supply — it
is an optional peer dependency, so reading costs you no native install.
Browser viewer
npm install # fetches the chromapakz decoder for the demo
python3 -m http.server 8000 # from the repo root
# open http://localhost:8000/viewer/index.html and drop a .wurld.webm
# (or .../viewer/index.html?src=../demo.wurld.webm)Zero-install, WebCodecs-decoded: seekable playback with per-frame point-cloud
reprojection, camera trajectory + frustum, RGB and metric-depth panes, and the same
file playing in a plain <video> element.
Hosted, with a sample loaded — nothing to clone or install: kmatzen.com/wurld
No install at all
The metadata lives in standard Matroska tags, so standard tools read it.
ffprobe prints the whole document — cameras, conventions, signal semantics,
and (for batch-written files) every pose:
ffprobe -v error -show_entries format_tags=WURLD -of default=nw=1:nk=1 scene.wurld.webmEXTRACTING.md is the cookbook: poses to CSV, intrinsics, RGB frames, and depth in metres using nothing but ffmpeg and arithmetic — including the exact triangle-fold and inverse-depth formulas, checked against the reference reader to 1.5 µm. It is also honest about the two things that route does not reach (binary pose tables and IMU streams) and what to do instead.
Use cases
USE_CASES.md surveys what people actually build with posed RGBD
— feed-forward reconstruction, splat and NeRF training, SLAM benchmarking, robot
rigs, dataset distribution — maps each to the format, and is explicit about the
cases where wurld is the wrong tool (multi-agent scenes, camera-less LiDAR
sweeps, geospatial survey, non-rigid capture). Seven scenarios ship as runnable,
tested examples under examples/.
It also carries the measured playback matrix and the HDR plans. In short: VLC, Chrome and ffmpeg play these files and pick the right track; QuickTime and iOS Photos cannot open them at all, because AVFoundation has no WebM demuxer. That is a deliberate trade — VP9 lossless is what makes bit-exact depth possible — and desktop viewing is VLC or IINA.
ROS 2
wurld ros2 export writes a real rosbag2 — CDR-encoded sensor_msgs and
tf2_msgs, so ros2 bag play works and ROS nodes can subscribe. (The older
wurld convert --to mcap writes Foxglove jsonschema channels, which Foxglove
Studio reads but no ROS node can.)
wurld ros2 export scene.wurld.webm ./bag # mcap storage; --storage sqlite3 also
wurld ros2 import ./bag out.wurld.webm/camera/<id>/image_raw sensor_msgs/msg/Image rgb8, one per stream
/camera/<id>/camera_info sensor_msgs/msg/CameraInfo
/camera/<id>/depth/image_raw sensor_msgs/msg/Image 32FC1, metres
/tf tf2_msgs/msg/TFMessage world -> *_optical_frame
/imu/<id> sensor_msgs/msg/ImuTwo conventions this gets right, both of which fail silently when got wrong:
ROS quaternions are xyzw where wurld's are wxyz, and ROS optical frames are
RDF (REP 145) — the same convention wurld uses — so world -> <cam>_optical_frame
is c2w with no axis conversion. The _optical_frame suffix is what tells a
consumer not to apply one.
Depth goes out as 32FC1 metres rather than 16UC1 millimetres so NaN survives;
in the 16-bit convention 0 means both "no return" and "at the sensor".
An HDR display track goes out as rgb16 (10-bit PQ codes, with a warning that
they are display-referred rather than linear — ROS has no field for the transfer
function). Every display stream is exported, not just the primary, and /tf
carries a frame per camera — the second derived from the rig, so a stereo pair arrives with its
calibrated baseline intact. Images are streamed rather than decoded whole, since
seq.rgb on a real EuRoC sequence is 4.2 GB per stream.
Fidelity is one-way. wurld → rosbag2 is exact. The return leg is not: a bag carries no quantization range, so depth is requantized (measured under 5 µm on a 1.5–2 m scene) and images re-encode through lossy VP9 (mean |Δ| ≈ 3/255). Good for moving between ecosystems, not for archival round trips.
Needs pip install 'wurld[ros2]', which does not require ROS itself.
Conformance corpus
Three readers — Python, JavaScript, C++ — is three chances to disagree.
conformance/vectors/ is 88 KiB of small files paired with the parse a reader
must produce, and one harness judges all three against it:
pytest tests/test_conformance.py # python, javascript and c++Expectations are generated from intent rather than captured from a reader
(conformance/generate.py), so the corpus cannot enshrine a bug all three
happen to share — which a golden file dumped from the Python reader would.
The vectors cover JSON and binary pose tables, unposed frames, stereo streams,
rigs, IMU, distorted camera models, float16_bits, unicode, and a single-frame
file. If you write a fourth reader, this is the definition of correct.
C++ reader
cpp/include/wurld.hpp is a single-header C++17 reader with no dependencies
— calibration, per-frame poses, timestamps, signal descriptors, rigs and IMU:
#include "wurld.hpp"
auto doc = wurld::read("scene.wurld.webm");
for (const auto& f : doc.frames)
if (f.pose_valid) use(f.c2w()); // row-major 4x4, camera-to-worldIt does not decode pixels or depth; that needs libvpx and chromapakz, and
the point of the header is that it drops onto a robot without dragging a codec
stack along. doc.cluster_start reports where the Clusters begin for consumers
that do want to hand bytes to a decoder. Traversal seeks over payloads and reads
only Tags, so opening a 10 GB file costs a handful of seeks.
Agreement with the Python reader is the actual requirement, so it is tested that
way: tests/test_cpp_reader.py diffs both readers field-by-field over JSON and
binary pose tables, rigs, IMU, unposed frames and the example files. A second
implementation that only agrees with itself is not a second implementation.
cmake -S cpp -B cpp/build && cmake --build cpp/build
./cpp/build/wurld_info scene.wurld.webmWriting from C++
cpp/include/wurld_write.hpp attaches a metadata layer to an already-encoded
WebM. That split is the point: a robot already has chromapakz for encoding, and
pulling libvpx into this header would destroy the zero-dependency property that
made the reader worth having.
#include "wurld_write.hpp"
wurld::WriteDoc doc;
doc.cameras["0"] = {"PINHOLE", 640, 480, {525, 525, 320, 240}};
doc.frames.push_back({0, 0.0, "0", true, {1,0,0,0}, {0,0,0}});
doc.world_json = R"({"metric_scale":true})";
wurld::write_file("encoded.webm", "out.wurld.webm", doc); // or wurld::attach(bytes, doc)Inserting tags shifts every Cluster, so Cues are rebuilt at the new offsets and the SeekHead is regenerated — a carried-over Cues element seeks into the middle of a Cluster, which presents as corrupt video rather than as a muxing bug. Re-attaching replaces the previous wurld tags instead of accumulating them, while leaving foreign tags (chromapakz's own) untouched.
The bar is equality with Python, not merely readability: pack_frames and
pack_imu must emit byte-identical buffers, and a C++-written file must satisfy
wurld validate and read identically through all three readers
(tests/test_cpp_writer.py).
Recording from C++
attach() rebuilds a whole file at once — fine for finalising a clip, useless
to a robot recording for hours. cpp/include/wurld_stream.hpp is the
incremental form: metadata woven between the encoder's Clusters as they arrive,
so peak memory is one Cluster plus 45 bytes per pose.
wurld::StreamWriter w([&](const std::string& b){ out.write(b.data(), b.size()); }, doc);
encoder.on_chunk = [&](const std::string& c){ w.on_encoder_chunk(c); };
for (...) { w.add_frame(pose); encoder.add_frame(pixels); }
encoder.finish();
w.finish();Same dependency split as everything else here: chromapakz encodes, wurld weaves.
The layout is SPEC §9's live form and is byte-compatible with the Python
StreamWriter — poses go out ahead of the Cluster carrying them, so a
recording killed mid-flight still reads back everything written before the
interruption. That is asserted, by truncating a file and reading it.
Collections: a corpus as a dataset
One file is one sequence; training is ten thousand of them. A collection (SPEC §14) is a manifest plus the files it names — a sidecar, not a new container, so every member stays an ordinary playable wurld file.
wurld index captures/ -o collection.json # headers only; no pixels decoded
wurld collection collection.json --membersfrom wurld import Collection
c = Collection.read("collection.json")
c.locate(12345) # global frame -> (member, frame within it)
for item in c.iter_frames(fields=("rgb", "depth"), shard=(0, 4), shuffle=7):
...Indexing reads each member's header and stops at the first Cluster: a file with 100x the pixels of another, at the same frame count, indexes for the same ~8 KiB. Sharding keeps whole files together so each member decodes once, and yields every frame exactly once — asserted across ranks x DataLoader workers, because a bad split does not crash, it quietly trains on duplicates.
With PyTorch (pip install 'wurld[torch]'):
from wurld.integrations.torch_data import WurldIterableDataset
ds = WurldIterableDataset("collection.json", fields=("rgb", "depth"))
loader = DataLoader(ds, batch_size=8, num_workers=4) # shards itselfVerify before you train — the manifest caches frame counts, and global indexing
is computed from them, so a member that was re-exported one frame shorter shifts
every index after it while locate() keeps returning an answer:
wurld collection collection.json --verify # exits 1 on driftIt re-reads headers, which is the same cheap read that built the manifest.
--checksum additionally compares recorded hashes, which is the only way to see
a change that leaves the header identical.
Measured at 10,000 members (scripts/bench_collection.py): index 1.9 s, manifest
3.4 MiB, locate() 0.46 µs, verify 1.0 s, streaming 31,700 frames/s for metadata
and 8,400 with pixels — peak RSS 58.7 MB, flat. Streaming holds a Cluster, not a
file: see USE_CASES scenario 7 for what measuring that turned up.
A collection asserts nothing about how members relate in space or time; that is the separate scene-manifest concern reserved in SPEC §11.
Layout
SPEC.md— the format (v1.2)wurld/— Python reference implementation (container, conventions, EBML tag layer, converters, CLI, synthetic test scene)tests/— round-trip suite (pytest): bit-exact depth, pose fidelity, converter round trips, COLMAP binary parsing, validation errorsviewer/— single-file browser viewercpp/— dependency-free single-header C++17 metadata reader +wurld_infoconformance/— cross-implementation test vectors and their generatorwurld/converters/ros2.py— rosbag2 bridge (real ROS 2 messages)wurld/collection.py— manifests, global indexing, sharded streamingwurld/integrations/— nerfstudio DataParser, PyTorch datasetsexamples/— runnable scenario walkthroughs (see USE_CASES.md)LANDSCAPE.md— the market research that motivated this projectUSE_CASES.md— pipelines, scenarios, and the format's limits
Status / verified
- Python, JavaScript and C++ suites green on Linux and macOS (see the
testworkflow for the current counts — a number written here goes stale within a day). Depth round-trips bit-exactly through VP9; poses, timestamps and intrinsics survive every converter round trip; binary frame tables, rigs, IMU streams and both Stray resampling policies are covered. - Three readers agree on a shared conformance corpus whose expectations are
generated from intent rather than captured from any one of them
(
conformance/), so Python, JavaScript and C++ cannot drift apart quietly. - Checked against real downloads, not only fixtures: TUM
freiburg1_desk(572/573 poses at 0.000000000 mm, depth within 2e-3 of the source PNGs) and EuRoCV1_01_easy(poses recomputed independently from the raw ground-truth csv agree to 3.2e-13; stereo baseline 11.01 cm from the real calibration). - Every exporter against every shape of legal file — unposed frames, no
display track, stereo, HDR — must succeed or refuse with a reason, never an
internal error (
tests/test_export_matrix.py). ffmpeg/ffproberead and fully decode wurld files with zero warnings — three VP9 tracks for a mono capture, four for a stereo one.ffprobesurfaces theWURLDdocument itself as a format tag, so calibration and conventions are readable with no wurld install; only the binary pose table and IMU are invisible, because ffmpeg's Matroska demuxer dropsTagBinary(see EXTRACTING.md).- Viewer verified in Chrome (native WebCodecs decode, no WASM), including binary frame tables.
1.0
SPEC declared 1.0 (frozen semantics; 1.x additions are additive-only, and
§11 codifies file scope: one rig, one clock). chromapakz pinned to the released
0.4.0 — no more git installs. LICENSE (MIT), CHANGELOG, CI (Python 3.11/3.13
on Linux+macOS plus a viewer parity job), and packaging validated (sdist+wheel
build clean, twine-checked, fresh-venv install + CLI smoke). Phone importers (Stray, Polycam, Record3D) and
WurldCam remain fixture/harness-validated pending a real capture. TUM is the
exception and is checked against the real download by tests/test_real_tum.py
(scripts/fetch_tum.sh fetches it; the test skips without it, and a weekly CI
job runs it).
v0.9: bounded-memory iteration, trim, auto-lazy
Sequence.iter_frames(start, stop): memory-bounded streaming decode — one Cluster (~1 s) decoded at a time, so hour-long files iterate without holding the take in RAM. Pre-cadence files fall back to a full decode with a logged warning.wurld trim in out --frames 30:60: cut a range into a new file — poses rebased, signals sliced bit-exactly, quantization specs preserved, in-range IMU kept, provenance noted in the world description.- Viewer auto-lazy: the probe's
Content-Rangetotal switches large files (>32 MB) to on-demand cluster scrubbing automatically;?lazy=1|0forces either way. - Multi-RGB design proposal filed as ChromaPakZ #47 — the schema/ABI sketch for true stereo pixel storage, awaiting maintainer review before implementation.
v0.8: partial decode, EuRoC stereo+IMU, release plumbing
Sequence.fetch_frames(indices): local partial decode — reading 3 frames of a 10k-frame file no longer decodes the other clusters (same cluster-splice machinery as remote access, bit-exact parity tested).- EuRoC MAV importer (
wurld convert MH_01_easy out.wurld.webm): the first real exercise of rigs + IMU together — cam0 posed video viaT_WB @ T_BS, both cameras' OPENCV calibration plus abodyrig (camera-to-body extrinsics;rig_c2wderives cam1 poses), imu0 with imu-to-cam0 extrinsics and measured rate. Dependency-free mini-YAML parser forsensor.yaml. (Video carried cam0 only at 0.8; both eyes ship as display streams now — see Unreleased.) - MCAP
/tf: FrameTransform world→camera per posed frame, so Foxglove's 3D panel places the moving camera correctly. - ChromaPakZ #45 merged (signal keyframe cadence) and release PR
#46 (0.4.0) staged — once
published, wurld pins
chromapakz>=0.4.0instead of git installs. iOS vendor libs rebuilt from merged main.
v0.7: random video access over ranges
- Cluster-independent decode (enabled by
ChromaPakZ #45, pending merge):
signal tracks now keyframe at the RGB cadence, so
[header + Cluster k]decodes that second of footage bit-exactly in isolation (+1.0% file size). wurld.remote.fetch_frames(fetch, indices): random access to any frames of a remote file — SeekHead → Cues → only the touched Clusters are fetched and spliced-decoded. Verified: bit-exact vs full decode, untouched clusters never downloaded, and a clear error for files written before the keyframe cadence.- Viewer lazy scrubbing (
?lazy=1&src=…): poses render from the header instantly; video clusters fetch on demand as you scrub. Verified in Chrome: jumping frame 1 → 81 of the demo took 4 ranged requests total with the middle cluster never downloaded. - WurldCam confidence: on-device recordings now carry ARKit confidence
as a second lossless signal (0/1/2 labels), verified device-free through the
macOS harness. (Note:
build-native.shneeds a chromapakz checkout with #45 for cluster-independent app recordings.)
v0.6: on-device wurld recording, ffmpeg-native metadata
- WurldCam v2: the iOS app now records wurld directly on-device —
libvpx + the chromapakz C core cross-compiled for iOS
(
ios/scripts/build-native.sh), Swift bindings over thedc_stream_*streaming ABI, and a Swift port of the pose-weavingStreamWriter(.wurld.webm/.r3dtoggle in the UI). The recording pipeline is verified without a device:ios/scripts/verify-pipeline.shcompiles the app's own writer+encoder for macOS, records a synthetic take, and validates it with the Python reader (poses at f32 precision vs analytic ground truth, exact f64 timestamps, NaN-invalid depth preserved, chunked layout + consolidated table, ffmpeg-clean). ARKit-on-hardware remains the one untested link. - ffmpeg needs no patch: ffprobe already surfaces the complete WURLD
JSON document as a format tag —
ffprobe -show_entries format_tags=WURLD -of json scene.wurld.webmyields cameras, conventions, and (for JSON-frame files) every pose, in any ffmpeg-based tool. Binary pose tables/IMU tags don't surface; everything else does. The roadmap's "ffmpeg demuxer patch" is retired as unnecessary.
v0.5: Record3D, MCAP/Foxglove, ranged viewer, nerfstudio
- Record3D importer (
wurld convert capture.r3d out.wurld.webm): the .r3d zip layout confirmed against the app author's own snippets and four community parsers — column-major K, scalar-last ARKit quaternions, LZFSE float32-meters depth (NaN invalid, float16 export variant handled), 0/1/2 confidence. Needspip install wurld[record3d]. - MCAP export (
wurld extract scene.wurld.webm out.mcap --format mcap): Foxglove-ready jsonschema channels (/camera/pose,/camera/imagejpeg,/camera/depth16UC1 bit-exact codes,/camera/calibration,/imu/<id>) plus the full WURLD document as an MCAP metadata record. Needs[mcap]extra. - Progressive ranged viewer:
viewer/index.html?src=...now loads via Range requests — trajectory and all poses render from the ~18KB header before any video byte, then video streams through the network decoder. Falls back to full download when the server lacks ranges.scripts/range_server.pyis a range-capable dev server (python's builtin lacks Range). - nerfstudio DataParser (
wurld.integrations.nerfstudio_parser): reads a .wurld.webm directly (frames extracted to a<file>.cache/beside it), poses converted to nerfstudio's convention, metric depth viadepth_unit_scale_factor— verified against a live nerfstudio 1.1.5 install.
v0.4: range-request access + Polycam
- SeekHead (SPEC §9.1): batch files begin with a fixed-width SeekHead covering the header elements, the first Cluster, and Cues. Combined with the metadata-first layout, a static file on S3/CDN serves calibration and every pose in at most two ranged reads.
wurld.remote:fetch_header(http_fetcher(url))pulls all metadata + poses without downloading video — verified <2% of file bytes on a 60-frame sequence;cues_offset/header_extentexpose what a video-seeking client needs next.- Polycam raw importer: keyframes/cameras JSON (
fx/fy/cx/cy,t_00..t_23ARKit c2w), microsecond-timestamp stems (sorted numerically), 16-bit mm depth, 0/127/255 confidence labels,corrected_cameras/corrected_imagespreferred withcorrected=Falseopt-out, same--at depth|rgbresolution policy as Stray. Like Stray: validated against synthetic fixtures built from polyform's published schema — a real capture to confirm conventions is welcome.
v0.3: fully streamable playback and recording
Streaming layout (SPEC §9): all wurld metadata — calibration and every pose — now lands before the first Cluster; Cues are rebuilt for the new offsets. A progressive reader has full pose data before the first video byte; ffmpeg still decodes and seeks cleanly.
Live recording:
viewer/wurld.jsprovidesWurldRecorder(wraps chromapakz's streaming encoder, weavesWURLD_POSESchunks before each Cluster, consolidates a pose table on finish) andWurldLivePlayer(extracts pose tags from the byte stream, forwards clean video bytes to the network decoder). A crash-truncated recording is still a valid, fully-posed file up to its last flushed chunk.viewer/live.html: record → stream → play in one page — the player consumes only the byte stream and reconstructs the point cloud live (verified in Chrome: 46/46 frames + poses received during recording; finalized file passes buffered decode).Python
wurld.stream.StreamReader: incremental parser for live/growing streams; batchwl.read()also accepts crash-truncated live files via chunk concatenation.Python live recording —
wurld.StreamWriter(chromapakz ≥ 0.4.0):w = wl.StreamWriter(out.write, cameras={"0": cam}, has_rgb=True, signal_meta=[wl.SignalMeta("depth", "depth", {"type": "inverse_depth", "near": 0.4, "far": 12.0})]) w.add_frame(frame, rgb=rgba, signals={"depth": {"float": z}}) # per capture tick w.add_imu("imu0", samples) # any cadence w.finish()Verified: bit-exact depth through the live path, IMU chunk concatenation, progressive parse parity, crash-truncation survival, and ffmpeg-clean output. Pose-only takes (RGB + poses, no depth) work too — enabled upstream by ChromaPakZ #44.
v0.2
- Binary frame tables (
WURLD_FRAMESTagBinary, 45 B/frame) for 10^6+-frame sequences; automatic beyond 10k frames. - Multi-camera rigs: camera-to-rig calibration block +
rig_c2w()derivation. - Per-frame intrinsics overrides (zoom / autofocus drift).
- IMU streams: packed 32 B/sample gyro+accel tags with extrinsics.
- Stray Scanner importer (first real capture-app source; ffmpeg-based, ARKit RUB→RDF pose conversion, honest RGB/depth resolution policy). Validated against synthetic fixtures — a real capture to confirm conventions is welcome.
Roadmap (v1.0)
Real-device validation (one WurldCam capture + one third-party-app capture confirms every convention); LeRobot depth-backend PR upstream (staged privately); ChromaPakZ #46 merge + 0.4.0 publish, then pin here. Multi-RGB tracks per the #47 design have since landed, and the EuRoC importer now stores both cameras' pixels with SPEC §4.4 binding frame camera ids to RGB streams.
