@dcl/catalyst-storage
v5.0.0
Published
Storage component for catalysts (WKC)
Downloads
856
Readme
catalyst-storage
The Catalyst Storage Library provides multiple implementations to handle file storage for Catalyst servers. This allows users to store and retrieve content through different backends like S3, folder-based storage, or in-memory solutions. It abstracts the complexity of interacting with these systems, offering a unified API for managing file storage.
Installation
npm install @dcl/catalyst-storage
Requires Node.js 24 or newer (engines.node: >=24). The package is published as CommonJS and
targets ES2023, which Node 24 implements in full.
Id validation
The folder-based and in-memory backends enforce the same id-shape rules, so an id one accepts is one the other accepts — previously the in-memory backend took '', 'foo.gzip', '../evil' and './x', all of which the folder-based backend rejects, so a service whose id handling was exercised in tests failed in production. Rejections are a PathNotContainedError (exported, so instanceof is stable), which callers should treat as a bad request rather than a storage fault. retrieve() is the one exception: it answers undefined for an unaddressable id, because there is nothing to serve. The validators themselves are internal — see Public surface.
The .gzip reservation is case-insensitive, and so is the reserved staging directory. The collision they prevent is a filesystem one, and half the filesystems this library runs on fold case: on APFS, NTFS or an SMB/CIFS mount <id>.GZIP is <id>.gzip, so storing it overwrote the compressed representation of another id — that id's reads then failed to inflate, its contentSize came out of the wrong file's last four bytes, and allFileIds() reported it while never listing the id that had clobbered it. Rejecting every case gives one id namespace on every filesystem rather than one that quietly widens on ext4. The same applies to a case variant of tempDirectoryName in flat mode, which otherwise landed a file inside the staging directory where enumeration cannot see it and which the next construction refuses to start over.
Two divergences remain, and both are structural — the folder-based backend knows things about its filesystem that an id-shape rule cannot:
- The reserved staging-directory name is a folder-based configuration value, so the in-memory backend cannot know it. In flat mode
'.tmp-writes/foo'is rejected by the folder-based backend and accepted by the in-memory one. A service that may run against either should not use ids under its configuredtempDirectoryName. - Two ids where one is a path prefix of the other cannot both be stored on a filesystem — nothing can hold a file and a directory at one path. Storing
aand thena/b(or the reverse) makes the second store reject withPathNotContainedError, while the in-memory and S3 backends take both. Hash prefixes usually hide this by putting the two ids in different shards, so it is reachable in flat mode and, via a 4-hex shard collision, rarely in hash mode too. Both directions are checked before anything is staged — the parent path forathena/b, and the commit target itself for the reverse — because each used to surface as a bareENOTDIR/EISDIRfrom the commit rename, several awaits after the id was accepted. A compressed store is refused in the reverse case too, rather than succeeding: it would otherwise leave an id whose whole reads work but whose byte ranges can never be served, since the range path has to publish its decompressed copy at exactly the occupied path. Reads of such an id report it absent —false/undefined— which is what makes them agree with the store refusing it and withdeleteresolving for it. Nothing can exist beneath a file, so absence is provable rather than assumed; this previously rejected with a bareENOTDIR, so one such id also destroyed the answers for every other id in anexistMultiplebatch.
A store refuses a source it cannot read, on every backend. What it rejects with depends on why:
- A source that has already errored rejects with that stream's own error, unchanged. The stream already knows what went wrong and replacing that with a generic code would discard the only useful diagnosis.
- A source that has been consumed, destroyed or closed — in whole or in part — rejects with
ERR_STREAM_PREMATURE_CLOSE. - A source in a non-utf8 encoding mode (
latin1,hex, …) rejects with a message naming the encoding: every backend re-encodes string chunks as utf8, which does not round-trip for those, so the bytes stored would not be the bytes read.
Piping a consumed stream writes zero bytes and resolves, so this used to commit an empty object under the id and report success, which in a content-addressed store is permanent: exist(id) answers true, so nothing ever re-fetches and the real content never lands. It is reachable from any caller that hashes, sniffs or measures a body before storing it, and from a retry that reuses its source. unshift-ing bytes back does not make a source storable again — hand over a fresh one. A rejected store consumes its source: it is destroyed on the way out, so a caller retrying after correcting the id must supply a new stream. The in-memory backend always rejected the fully-consumed case; the folder-based and S3 ones stored 0 bytes, and all three accepted a partially consumed source.
The S3 backend does not validate ids: keys are opaque to S3, an existing bucket may already hold keys these rules would reject, and refusing them would make that content unreadable. A service that may run against more than one backend should apply the same id rules itself before storage is reached.
Supported storage types
- S3 Storage: Store and retrieve content from AWS S3 buckets.
- Folder-based Storage: Local file storage on disk.
- In-memory Storage: Temporary storage for testing or lightweight operations.
Public surface
Everything importable from this package is listed in src/index.ts, and package.json declares an
exports map, so @dcl/catalyst-storage/dist/... paths are no longer reachable. The surface is:
| | |
|---|---|
| factories | createFolderBasedFileSystemContentStorage, createS3BasedFileSystemContentStorage, createAwsS3BasedFileSystemContentStorage, createInMemoryStorage, createFsComponent |
| types | IContentStorageComponent, ContentItem, FileInfo, AppComponents, IFileSystemComponent, FolderStorageOptions, S3ContentStorageOptions, FileTypeLoader |
| helpers | bufferToStream, streamToBuffer, SimpleContentItem, assertStorableStream |
| errors | PathNotContainedError, RangeNotSupportedError, UncommittedIntentSurvivedError |
The errors are exported as runtime values precisely so instanceof is stable, and the list is decided
by ESCAPE, not by existence: an error a public method can actually reach a caller with is contract, and
one that is only thrown and caught inside the package is not — exporting that would be inventing an API
nobody can use.
PathNotContainedErrorescapesexist,fileInfo,delete,storeStreamandstoreStreamAndCompress.retrieve()converts it toundefinedinstead.RangeNotSupportedErrorescapes the S3retrieve()for a range over an encoded object — the read contract asks callers to answer 416 for it, which is impossible without this export.UncommittedIntentSurvivedErrorescapesstoreStream/storeStreamAndCompresswhen a commit failed and its journal could not be cleared. ItsstagedPathis the actionable part.DecompressionLimitExceededErroris not exported: every path that raises it is caught byretrieve(), which answersundefined. A cap breach is therefore observed as "nothing to serve", not as a typed error.
assertStorableStream is the one validator that is public, because the rule it enforces is one a caller can
otherwise only discover by failing: a body that has been read from — to hash it, sniff it, or measure it —
cannot supply the content any more, and every backend refuses it. A caller that must inspect a body before
storing can check the source it is about to hand over instead of learning from a rejected store. The id
validators stay internal: they encode invariants that only make sense applied at a specific point, and the
backends apply them on the caller's behalf.
A source that still has a 'readable' listener attached is refused by every backend for the same reason,
even though nothing has been read from it yet. That listener is a competing consumer: a backend that pipes
never sees the body flow at all (a store of such a source used to hang forever, with no timeout in this
package), and a backend that reads explicitly — S3, via the MIME head peek — races it and stores only the
chunks it happens to win. Measured on the S3 backend with a reader attached to a 20-chunk source: 0 of 2000
bytes committed, and the store resolved. Hand over a fresh source rather than one someone else is reading.
streamToBuffer(stream, maxBytes?) takes an optional cap, and it is opt-in on purpose: callers know
their own content sizes, and a default ceiling would start rejecting bodies that store and read correctly
today. Pass one when the stream is a decoded one — streamToBuffer(await item.asStream()) over
attacker-supplied compressed content inflates without limit, because the folder-based backend's
decompressMaxFileSize bounds only the range-request inflation path, not a full read.
Previously every module was re-exported wholesale, which made a lot of internals public by accident —
the MIME detector and its ESM loader memo, the id validators, the bounded-map helper, the stream
teardown utilities, the range validators, compressContentFile. None of them were imported by any
consumer, and each was something this package would otherwise have to keep working forever. They all
still exist and are still used internally; they are simply no longer contract.
Migrating. If you import from a dist/ path, switch to the package root — every symbol that was
reachable that way and is still public is exported from it:
-import { bufferToStream } from '@dcl/catalyst-storage/dist/content-item'
-import { createInMemoryStorage } from '@dcl/catalyst-storage/dist/in-memory-storage-component'
+import { bufferToStream, createInMemoryStorage } from '@dcl/catalyst-storage'If you relied on a withdrawn symbol, open an issue rather than deep-importing: several of them (the id validators in particular) encode invariants that only make sense applied at a specific point, and exporting them again is cheap if there is a real use for it.
Cancellation scope
storeStream and storeStreamAndCompress accept an AbortSignal. Cancelling stops the work and rejects with the caller's reason, and a store that already completed before observing the abort is allowed to succeed. What "nothing was stored" guarantees differs by backend, because the two commit through different machinery:
- Folder-based — absolute. The commit is a local
renamethis storage fully controls, with a checkpoint at every phase boundary, so a cancelled store never leaves content at a canonical path. The previous version of the id stays intact. - S3 — one bounded window. The abort tears down the in-flight request itself (
PutObject,UploadPart,CompleteMultipartUpload), so the key does not appear; a partially uploaded multipart upload is also cleaned up, and that cleanup deliberately ignores the signal that triggered it. What cannot be covered is a request S3 has already received in full when the abort fires — tearing down the connection does not un-send those bytes, and the service may still apply them. The residue is bounded: S3 object writes are atomic, so the key is either absent or holds the complete content (never partial or mixed), and because content is addressed by its own hash the worst outcome is the correct bytes existing under their own id after a store reported as cancelled.
S3 storage: operational contract
Reads distinguish "not here" from "cannot be read".
retrieve(),exist()andfileInfo()report absence —undefined/false— only for a definitive not-found:NotFound/NoSuchKeyor a 404. Everything else — 500, 503SlowDown, a network fault, expired credentials — rejects. Reporting those as absent made a throttled or unreachable bucket look like an empty node and stopped it being retried. Callers mapping to HTTP should return 404 forundefined/falseand 5xx for a rejection.A 403 rejects unless you opt in, and that opt-in is
report403AsAbsent. A principal withouts3:ListBucketgets 403 instead of 404 for a missing key, and nothing in the response separates that from a real denial. So for such a principal every read of absent content rejects — passreport403AsAbsent: trueto answer those reads as absence instead.It is a configuration flag, not something construction infers. The startup
ListObjectsV2probe cannot decide it: a 403AccessDeniedon that probe is also what an IAM policy scoped toarn:aws:s3:::my-bucketreturns for any other bucket — including a misspelled one — because the implicit deny is evaluated before S3 checks whether the bucket exists, so theNoSuchBucketguard below never fires. Inferring the lenient mode from it turned a one-character typo inBucketinto a node that answered "I hold nothing" for every id, for its whole lifetime, while its writes rejected loudly: precisely the silently-empty node the probe exists to prevent, produced by the probe.The probe still runs, and still refuses to start over a missing bucket, but it now only reports what it saw:
- it succeeds → the principal has
s3:ListBucket, so a missing key answers 404 and every 403 is a genuine authorization failure. Ifreport403AsAbsentis set anyway, a warning says to remove it. - it fails with a 403 (
AccessDenied, or a credential/clock one such asInvalidAccessKeyId,SignatureDoesNotMatch,RequestTimeTooSkewed) → a warning names the ambiguity and tells you to grants3:ListBucketor setreport403AsAbsentdeliberately. Reads follow the flag. - it fails any other way — a network blip, a throttle → a warning, and reads follow the flag.
Defaulting to strict is deliberate, because the two readings are not symmetric: answering "cannot read" for content that is genuinely absent costs a retry, while answering "absent" for content that is present and unreadable is a silent data-loss report.
ListObjectsV2is a GET with an XML error body, so the real code survives — unlikeHeadObject, whose empty body is why the ambiguity exists at all.A missing or misnamed bucket fails construction.
HeadObjectagainst one answers 404NotFound, byte-identical to a missing key, so every id reported absent and nothing was logged. The same startup probe detects it (NoSuchBucket/404) and refuses to start, because a silently empty node is worse than one that will not boot.- it succeeds → the principal has
The reported
encodingis only what is still applied to the bytes.fileInfo()andretrieve()run the rawContent-Encodingthrough the same predicate that decides whetherasStream()decodes, so the two can never disagree about one id. Codings that describe nothing still applied are dropped, andnullis reported when none remains:| stored
Content-Encoding| reportedencoding|contentSize| rangeable | |---|---|---|---| | (absent) /identity|null| known | yes | |aws-chunked|null| known | yes | |gzip, aws-chunked|gzip|null| no (416) | |gzip|gzip|null| no (416) |identityand an empty header mean "not encoded", so reporting them verbatim forces every caller to special-case a value carrying no information.aws-chunked— which S3 writes itself on flexible-checksum uploads — describes the transfer of the bytes and is already undone by the time a body reaches us, so a bareaws-chunkedobject streams plain bytes and has a known size; reporting it anyway handed callers aContent-Encodingto forward for content that is not encoded at all. Surviving codings keep their original spelling and order, so what is reported stays a valid header value. Previously the metadata surfaces tested raw truthiness and reportedcontentSize: null— "size unknown" — for content whose size was plainly known.aws-chunkedis the only transfer coding treated this way. Barechunkedis aTransfer-Encodingvalue rather than a content coding, so an object whoseContent-Encodingsayschunkedcarries metadata that is simply wrong; it is left unrecognised, which meansasStream()refuses it and namesasRawStream()instead of quietly serving the bytes as though the header agreed.More than one real coding is refused, not partially undone.
asStream()applies at most one decoder, so an object stored asgzip, brwould have Brotli undone and be handed back still gzipped — under a contract that says the stream yields decompressed content. Such an item reportsencoding: 'gzip, br'and anullcontentSize, refuses a range, and throws fromasStream()namingasRawStream(); the stored bytes stay reachable. No backend here writes more than one coding — the folder-based one writesgzip, S3 writes none — so this only arises for an object an operator or a migration put there.The content-type detector is resolved at construction, and can be supplied.
file-typeis ESM-only, so the default reaches it through a dynamic import; that import is awaited while the component is being built rather than left in flight, so a resolution problem is reported before any traffic arrives instead of degrading the first stores that race it. PassfileTypeLoaderto use your own detector, or to avoid the dynamic import entirely.The loader is called once, and the module it returns is what every store uses — it is not re-invoked per store, so an injected loader never ends up on the hot path. The one exception is a loader that rejects: construction then only warns (detection is metadata, and a store must not fail because the detector is unavailable) and stores retry it, so a transient failure does not permanently downgrade every later store to
application/octet-stream.A single store is capped at roughly 50 GiB unless you set
partSize. Bodies reach@aws-sdk/lib-storagewrapped by the MIME head-peek, and that wrapper exposes no length, so lib-storage cannot right-size parts the way it does for aBufferor anfs.ReadStream: it falls back to its 5 MiB minimum, and with S3's 10,000-part limit that is the ceiling — reported asExceeded 10000 partsonly after 50 GiB has crossed the wire. PasspartSize(bytes) to raise it: 64 MiB lifts the ceiling to about 640 GiB, at the cost of proportionally more memory per concurrent upload. Budget that aspartSize× 4: lib-storage uploads four parts concurrently (itsqueueSize, which this pins explicitly so an SDK bump cannot change the arithmetic), each holding a full part, so 64 MiB parts cost roughly 256 MiB per in-flight store. This previously said "buffers a part at a time", which understated the footprint fourfold — enough to have a container sized from it OOM-killed mid-upload, after the bytes had already crossed the wire. It is validated at construction against S3's own 5 MiB–5 GiB bounds and must be an integer, because every way of getting it wrong otherwise fails late — below the minimum every store rejects with a bareEntityTooSmall,0andNaNare silently ignored by lib-storage, and a non-integer passes its own guard and then breaks any upload that needs a second part. Available on both S3 factories.Metadata and bytes can come from different versions of a key — a documented window, not a closed one, exactly as on the folder-based backend.
retrieve()readssize/encoding/contentSizewith aHeadObject, while the bytes come from aGetObjectissued when the consumer opens the lazy stream. An id overwritten with different content in between is therefore served under the previous version's advertised length: a 100-byte object re-stored at 95 bytes answers a{start: 90, end: 99}range with 5 bytes undersize: 10, and a shrink paststartsurfaces as a raw SDKInvalidRange.This only arises when an id is overwritten with different content, which the content-addressed model this storage exists for does not do. A caller that both allows it and forwards
sizeas an HTTPContent-Lengthmust re-check after streaming rather than trust the advertised value.An
IfMatchprecondition on theGetObjectwas implemented and withdrawn. It does close the window, but it fires on any ETag change rather than any content change, and the two are not the same: where the ETag is not a digest of the body (SSE-KMS, SSE-C), or when two writers pick different multipart part boundaries — whichpartSizemakes reachable across a rolling deploy — re-storing identical bytes rotates it. That traded a rare wrong answer for usage this storage says does not happen against a routine 412 for usage that is entirely correct.Ranges are not served for encoded objects. The
Rangeheader addresses the STORED bytes, so a range over an object with aContent-Encodingwould hand back a fragment of a compressed stream — whichasStream()then fails to inflate, or (for a range starting at 0) inflates into the whole object while advertising the requested length. S3 keeps no uncompressed-size metadata, so the logical bounds cannot be applied service-side; such a request rejects withRangeNotSupportedError(exported), which callers should map to 416 rather than the 5xx the rest of the read contract prescribes.identity(and an empty header) mean not encoded, so those objects range normally. This storage never writesContentEncodingitself. The folder-based backend answers the same call by inflating into its decompression cache, which has no S3 equivalent.A failed or cancelled store cleans up its own multipart parts.
lib-storageissues itsAbortMultipartUploadonly on the paths that run beforeCompleteMultipartUpload, so a complete request that is cancelled or fails leaves every uploaded part in the bucket — invisible, and billed until a lifecycle rule reaps it. This storage therefore captures the upload id as it is issued and aborts the upload itself on any failure path, on the real client so the signal that caused the teardown cannot cancel the cleanup too. It is best-effort and idempotent: a failure to clean up is logged, never substituted for the caller's error.CompleteMultipartUploaddeliberately stays cancellable — exempting it would close the same window, but only by letting a cancelled store run the commit to completion, widening the residue described above from a last-packet race into the full duration of that request.delete()verifies what it removed.DeleteObjectsreports per-key failures inside a 200 response, so a resolved request is not a completed delete; any reported failure rejects and names the keys that survived. The id list is also chunked to S3's 1000-key per-request limit (the SDK does not split it), and deleting nothing is a no-op rather than theMalformedXMLan empty request would return.Those chunks are issued concurrently (8 in flight, so up to 8000 keys at once): serialized, a million-id GC sweep was 1000 sequential round trips of pure latency for work that shares no state. The consequence is that a rejection says nothing about the ordering of what was removed — earlier chunks are not necessarily complete and later ones not necessarily untouched.
deleteis idempotent, so retrying the whole list is the recovery, and the error message states exactly that rather than the ordering guarantee the serialized version could make.Enumeration follows the continuation token, not just
IsTruncated. A page is followed whenever it carries aNextContinuationTokenand does not explicitly sayIsTruncated: false. AWS always sets the flag, but an S3-compatible endpoint (MinIO, Ceph, a gateway) may return the token and omit it — which ended the walk after one page and let a GC or sync sweep conclude the bucket held only its first 1000 keys. A token is still required to continue: a page that explicitly saysIsTruncated: truewithout one fails the enumeration instead of completing with a silently partial bucket view, and so does a token the endpoint has already issued, which would otherwise cycle the walk over pages it has seen.The cursor is checked before the page is yielded. Both of those rejections mean the enumeration cannot be completed, so deciding after yielding let a caller streaming ids consume the page first — and for a repeated token that page is a duplicate of one already delivered, so a GC or sync sweep acted on repeats before the iterator rejected. Nothing is emitted from a page the walk is about to refuse to continue past. A final page (
IsTruncated: false) is exempt: its token, stale or repeated, no longer means anything, and rejecting it would fail a listing that is in fact complete.Content type is detected from the head of the stream, never from the whole object: only the first 4100 bytes are buffered, and the body streams straight through to the managed upload. If detection is unavailable the store still succeeds as
application/octet-stream, and says so in the logs — a silent fallback is indistinguishable from content that genuinely has no signature.Batch surfaces bound their concurrency.
existMultiple()andfileInfoMultiple()cap in-flightHeadObjectrequests instead of issuing one per id at once, which drew 503SlowDownon large batches.Whoever creates the client owns it.
createAwsS3BasedFileSystemContentStorageconstructs its ownS3Client, so it exposes astop()that destroys it — the SDK's guidance is to shut a client down explicitly in Node, or its sockets stay open long after the last request.createS3BasedFileSystemContentStoragetakes an injected client and deliberately does not destroy it: the caller owns it and may share it with other components.
Folder-based storage: operational contract
The folder-based storage stages writes through a reserved directory and renames them into place, which is what makes them crash-atomic and makes a read concurrent with a write observe one complete version rather than a half-written file. rename is therefore a required capability of the filesystem component; the bundled createFsComponent provides it.
This comes with these explicit rules:
One live instance per storage root, with exclusive ownership of the tree — this is a hard requirement. The root, its shard directories and everything beneath them must be created and managed only by this storage under the service user; no other writer and no pre-existing symlinks anywhere under the root (reads, writes and deletes resolve paths through the OS and would follow a planted symlink outside the root — only the reserved temp path is actively checked, via
lstat, because staging is where files are created most often). In-memory coordination (path locks, decompress-cache tracking, staged-write ownership) is also per-instance; two instances sharing a root can delete each other's staged files and race their caches.A
ContentItem's metadata and its bytes can come from different versions of an id.sizeandcontentSizeare measured whenretrieve()runs; the stream is opened later, whenasStream()is called. A store landing in between can unlink the file (soasStream()fails — treat it as a retryable miss) or replace it, in which case the stream yields the new version's bytes under the previous version's advertised size, and a gzip item'scontentSizeis read from what is no longer the file's trailer and may be an arbitrary number. This only arises when an id is overwritten with different content, which the content-addressed model this storage exists for does not do. A caller that both allows it and forwardssizeas an HTTPContent-Lengthmust verify after streaming instead of trusting the advertised value. Closing the window entirely requires opening the stream eagerly atretrieve()and reading through that descriptor, which the filesystem component has no capability for today.Atomicity covers process crashes, not power loss. Staged data is deliberately not
fsync'd before the commit rename. Against process death — OOM-kill, eviction, crash — this is airtight: a canonical path only ever holds the previous file or the complete new one. Against a power loss or kernel panic it is not, and the gap is worth stating precisely:renameorders metadata, so the directory entry can survive while the staged file's data blocks never reached the disk. The file may then be missing, zero-length or partial at its canonical path. (ext4'sauto_da_allocheuristic forces the flush when renaming over an existing file, but the common case here is a fresh content id whose target does not exist.) This is a deliberate trade — content is content-addressed and re-downloadable, so durability past process death does not justify an fsync per write — but it means a host that lost power can hold a truncated file thatexist()reports as present, and consumers must be able to detect and discard unreadable content rather than trusting presence. It is the one case where this storage cannot promise complete-or-absent.Interrupted commits are reconciled at construction. A stored id spans two possible representations (
<id>and<id>.gzip); commits that transition between them journal an intent first, and a crash between the commit and its cleanup is resolved at the next construction in favor of the committed representation — reads can never prefer a stale counterpart. A repair that cannot be completed fails construction (reads do not consult intents, so running over an unreconciled state would serve the stale representation for the process lifetime).A symlinked reserved path is rejected at construction when the filesystem component provides
lstat(the bundled one does); withoutlstat, the exclusive-root model is the guarantee that no symlinks exist under the root.Compression goes through the injected filesystem component.
storeStreamAndCompresspipes throughcompressContentFile, which reads and writes via the same component as every other operation, so a custom adapter that virtualizes paths gets compressed stores as well as atomic raw writes. It is not part of the public surface: it was previously reachable only because every module was re-exported wholesale, no consumer imported it, and calling it outside this storage means reasoning about the staging and commit rules it assumes. UsestoreStreamAndCompress.retrieve()distinguishes "not here" from "cannot be read". It resolves toundefinedwhen there is nothing to serve — the id is absent, it does not resolve to a servable path, it exceededdecompressMaxFileSize, or the content file was deleted while being read — and rejects when storage work done byretrieve()itself fails: EACCES/EIO/ENOSPC on its own directories, a corrupt gzip inflated for a range request, a failed decompression commit, a missing staging directory, or an id whose raw/gzip state is mixed and could not be repaired (a quarantined id is present on disk and still enumerated byallFileIds(), so refusing it is a "cannot be read", never a 404). Full non-range gzip reads are different because they return a lazyContentItem:retrieve(id)can resolve before the gzip is opened or decompressed, and the same storage/corruption failures may surface later fromContentItem.asStream()or while consuming that stream. Callers mapping this to HTTP should return 404 forundefinedand 5xx for both aretrieve()rejection and a returned stream that fails; answering 404 for an unreadable disk makes a broken node look like an empty one and stops it being retried.RangeErroris the caller's fault, not the storage's, and belongs on 416: it means bounds this storage rejects outright, including astartpast the end of the object, and the folder-basedretrieve()re-raises it ahead of its failure logging for exactly that reason. Read the 5xx rule as applying to everything else; taking it literally answered 500, and paged an operator, for a malformedRangeheader. (RangeNotSupportedErroris also a 416, but only the S3 backend raises it — see its own section; the folder-based backend serves an encoded range by inflating instead.)The distinction is drawn by on-disk state, not by the error: an
ENOENTcounts as a miss only when the content is provably gone — checked by re-probing it — so the same code arriving from a rename, a staged write or a damaged shard directory surfaces as the storage fault it is. The error itself cannot be used as evidence here:pipelinedestroys upstream streams with the downstream error, so a failing staged write reaches the content stream as the very same object a concurrently deleted file would produce.statfollows the same rule. OnlyENOENT/ENOTDIRcan mean absent, and even then absence is reported only once the id's parent directory is proven to still be an intact directory — a shard that was removed, or replaced by a file, makes the whole shard unreadable and rejects instead. A present-but-unreadable file is likewise never reported as missing. A damaged shard also drops its cached entry, so writes recreate the tree as soon as whatever is occupying the path is gone; this storage never removes that obstruction itself, on the same principle as the reserved-namespace checks — it does not destroy what it cannot prove it owns.exist()andfileInfo()reject on a non-containable id rather than reportingfalse/undefined.fileInfo().contentSizeis a hint, but a truthful one. For a gzip item it comes from the format's trailer, sonullmeans the size could not be read out of the format — the file is too short to hold a trailer, or a concurrent overwrite moved it — never a number this storage guessed. The trailer read is verified against a secondstat: the read opens the path after the size that positioned it was measured, so an overwrite in between would otherwise address the wrong file — a smaller replacement short-reads, and a larger one returns four bytes from the middle of the new compressed body as if they were the trailer. Only an unchanged size is trusted; anything else re-reads once against the fresh size, and reportsnullrather than a guess if the file is still moving. It is not a signal for content above 4GB: gzip's ISIZE field is only accurate mod 2³², so such an original reports the wrapped value rather thannull, and callers must not treat a smallcontentSizeas proof the content is small. A trailer read that fails for a storage reason rejects instead, because callers cannot distinguish the two and at least one bounds range requests withcontentSize ?? size, where a masked failure would silently substitute the compressed size. The trailer is stored, possibly attacker-controlled data, so it is never used to bound decompression.retrieve()reports the same number. A gzipContentItemstreams decompressed bytes fromasStream(), so itscontentSizeis read from the trailer exactly asfileInfo()does, andsizeremains the stored (compressed) length. The two surfaces never disagree about one id.exist()answers about presence, not readability. Only a file provably gone is absent —ENOENT,ENOTDIR, orENAMETOOLONG, which means no such name can exist anywhere. A probe failing for any other reason — EIO, or a shard directory that was removed or replaced — rejects, matchingfileInfo()andretrieve()rather than reportingfalsefor a broken store.existMultiple()inherits this, and both cap how many ids they probe concurrently so a large batch cannot exhaust the process file-descriptor limit.Note the precise limit of this: presence is decided by
stat, which succeeds on a file whose own mode denies reading. Achmod 000file is therefore reported present, and fails later when its stream is opened. What is detected here is a damaged path, not an unreadable file.An id must name a file of its own. An id is used verbatim as a path under the directory ids resolve against, and
path.joinnormalizes what it builds, so different id strings can land on one file.a/../victim,./victim,/victimanda//../victimall resolve to the path ofvictim;a//victimresolves to that ofa/victim;''and.resolve to the containment directory itself (the storage root in flat mode). Containment does not catch any of this — the result is still inside the root, it is just somebody else's file — so a caller accepting untrusted ids could overwrite, read or delete another id's content. In flat mode that is direct; with hash prefixes the shard comes from the unnormalized id, so it also needs a collision on the first four SHA-1 hex digits, which is only ~2¹⁶ work by varying a prefix.The rule enforced is the invariant itself rather than a list of the bad forms: an id must resolve to exactly its own path, i.e.
path.relative(<containment dir>, <resolved path>) === id. Every aliasing form fails that equality by construction — normalizing is precisely what makes the resolved path differ from the id that produced it — so the guarantee does not depend on having enumerated the forms correctly. It is also the inverse of howallFileIds()recovers an id from a path. For that inverse to hold, one more name is reserved: an id may not end in.gzip, because that is this storage's own suffix for the compressed representation of another id — storing bothfooandfoo.gzipmaderetrieve('foo')serve the other id's bytes andallFileIds()report a phantom. Ids are likewise rejected if they contain a NUL byte, which no filename can hold. An empty id is rejected separately, being the one input where the equality holds trivially.This is orthogonal to the containment check:
../evilresolves to exactly its own path and round-trips cleanly — it is simply outside the root, and containment is what refuses it.Because
path.join/path.relativefollow the platform's separator rules, the check is platform-aware: on POSIX a backslash is an ordinary filename character, soa\..\victimnames its own file and is accepted; on Windows the same equality rejects it — but it also rejects forward-slash nested ids there, so separator-containing ids are POSIX-only. Ids containing separators are otherwise fully supported: they nest into subdirectories and round-trip throughallFileIds(). Two limits apply. One is the prefix rule above —aanda/bcannot both exist on a filesystem. The other is a total id length of 1024 bytes, enforced by the shared id rules on stores only, so that the assembled path stays insidePATH_MAX(1024 on macOS/BSD, 4096 on Linux) instead of failing late with a bareENAMETOOLONG; already-stored longer ids stay readable, enumerable and deletable. A root long enough to push a legal id pastPATH_MAXanyway is reported as a storage fault, not a bad id, because the root is the deployment's choice and no caller can correct it.The comparison is lexical, so the storage root must be on a case-sensitive, normalization-preserving filesystem — treat that as part of the exclusive-ownership requirement above, not as advice. On a case-insensitive volume (macOS by default, Windows, any SMB/CIFS mount)
victimandVICTIMboth satisfy the equality yet open the same file, so the aliasing this rule prevents returns: the second store overwrites the first id's content,exist()answers true for both, andallFileIds()yields only one — so a GC pass diffing enumeration againstexist()is served the wrong bytes. An id differing from the reserved directory only by case can likewise reach inside it. Hash prefixes hide this unless the two spellings land in one 4-hex shard; flat mode has no shards, so it is reachable there directly.A per-store guard for this was tried and reverted as worse than the hole: it compared only the raw basename, so two
storeStreamAndCompresscalls still corrupted silently (and whether a pair was refused or corrupted depended on an invisible compression decision); it ignored every path segment but the last; the fold it used strips trailing dots and spaces, which APFS preserves, so it refused provably distinct names; and it cost areaddirper store — 39 ms/store at 50k entries in flat mode against ~0.25 ms. Doing it properly means probing case folding and trailing-tail stripping independently, comparing both representations of an id, checking every segment, and not reading a directory per store. Until then: run production on ext4 or XFS.Reads never create directories.
exist(),retrieve()andfileInfo()resolve an id to its path without touching the filesystem. Previously every call created the parent: with hash prefixes that merely pre-created shards over time, but in flat mode ids nest, so probing a never-stored nested id (exist('a/b/c/missing')) lefta/b/c/behind permanently — a caller passing through untrusted ids grew the inode count without limit, andallFileIds()then walked those empty trees on every enumeration. A consequence: a shard directory that does not exist now means "nothing was ever stored here", so a read of an id in it is a miss. A directory this instance created or observed that then disappeared, or stopped being a directory, or that cannot be read at all, still rejects — the "a fault is not a miss" rule is unchanged for content that is actually present. Every successful stat and every classified miss records the directory, so damage under a shard the instance is actually serving is always loud.The answer is derived from the tree on every read, not from remembered damage, so it does not depend on how deep the id is, how many times it is asked, or what this instance happened to observe earlier. A read that finds nothing walks up to the first real directory and classifies what it finds: everything missing below that directory is a fault if any of those paths is one this instance saw exist, and an ordinary miss otherwise.
One case reads as absence even though a directory was destroyed: a path that now holds a regular file where content legitimately lives. This storage serves that file as an id, so ids nested under it are unstorable —
storeStreamrefuses them withPathNotContainedErroranddeleteresolves for them — and reads reporting absence is what makes those three agree. The transition is logged, since it destroyed whatever was under that path."Where content legitimately lives" is decided by rebuilding the path from the id the file's name spells, not by how deep it sits. With hash prefixes an id's shard is
sha1(the whole id), so a file at<root>/<shard>/ais content foraonly when that shard issha1('a')— otherwise nothing resolves to it, and crediting it would report every id beneath it absent on the strength of something this storage cannot serve. A file where a shard belongs is the reachable case, and it makes 1/65,536 of the keyspace read as an empty node. Such a file, a file whose name spells no valid id, a fifo, a socket, a device node, and the root itself are all faults — on reads, and on stores, where they raise a plain storage error rather thanPathNotContainedError, because no id of the caller's can be corrected to fix them.A read decides that per id, not per path. An id has two representation paths, and a foreign node at one says nothing about the other, so
<id>.gzipbeing a socket while the raw file is intact reads normally: the fault is deferred and raised only once no representation has served the id. Reporting it any earlier would answer "cannot be read" for content this storage can read — the same defect as answering absence for a fault, pointing the other way.deleteis deliberately the exception: it unlinks and resolves, which is what lets an operator clear a foreign node through the API instead of only on the filesystem.A compressed commit claims the prefix too, by a different mechanism. It leaves the raw path free, so the filesystem raises no objection to a directory being created there — but that path is exactly where a byte RANGE of the compressed id has to publish its decompressed copy, so allowing it left an id that serves whole reads and can never serve a range, with no way back: the path cannot be freed while the nested id lives there, and neither re-storing the compressed id (both commit targets are refused) nor deleting it recovers. The same end state reached in the other order — directory first, then a compressed store — was already refused for precisely this reason, so allowing this one made the verdict depend on arrival order rather than on the state.
storeStream('a2/b')afterstoreStreamAndCompress('a2')is therefore refused withPathNotContainedError, at any depth between the two, whether or not the directory already exists, and before themkdirrather than undone after it — an empty directory left behind by the rejection would be the same breakage, and this component exposes normdir. The check costs no syscall where it cannot fire: a shard name never spells its own id, so hash-mode stores settle it in CPU alone, and hash mode is unaffected in any case (a2anda2/bland in different shards). The two stores are also serialized against each other, since each check alone is a preflight: a compressed store probes its commit targets while the raw path is free and commits only after consuming and compressing its whole source, so a nested store creating the directory inside that window used to leave the unrangeable pair with both checks having been correct when they ran. The nested store takes the path lock of every prefix level it inspects — shallowest first, which is the ordering that makes the nesting deadlock-free — and the compressed commit re-probes its raw path under that same lock before publishing, so whichever store arrives second is refused. Either store shape can therefore fail late, after its source is consumed, when another id claims a path it commits to mid-flight — with the samePathNotContainedErrorthe preflight gives, never a bareEEXISTorEISDIR. The gzip commit re-probes the raw path because renaming onto<id>.gzipwould otherwise succeed and leave the unrangeable pair; every other commit target reports the conflict by failing, so the failure is classified instead, which costs nothing when there is no conflict.The read follows the store: an id that can never be created reads as absent, so a destroyed
a2/bwhose prefix is now owned by a compresseda2reports nothing rather than a fault — the same answer, and for the same reason, as a prefix owned by a raw file. The transition is logged either way, since content was destroyed.An ancestor that is a regular file is a miss, not a fault, when it is a directory this instance never observed. A filesystem cannot hold a file and a directory at one path, so no file can exist at
a/b/c/dwhilea/bis another id's content — the id is provably absent, exactly as an over-long name is, and this is now the answer at every depth and in both shard modes. It previously rejected with a bareENOTDIR, which made the surfaces contradict each other for one id:storeStreamrefused it with the typedPathNotContainedError,deleteresolved, andexist/fileInfo/retrieveproduced an untyped 5xx that also destroyed the answers for every other id in anexistMultiplebatch. Reachable in flat mode without any corruption —store('a')then a read ofa/b, ora.gzip/bafter a compressedstore('a'). The store still refuses such an id, which is the right asymmetry: a read asks a question whose true answer is "nothing", while a store asks to create something no filesystem can hold.Enumeration only yields ids the point lookups accept. An id recovered by slicing a file's path is the inverse of storing it only if the file is where this storage would have put it, and with hash prefixes it often is not: a foreign file at
<root>/3ec6/aspells the ida, whose own shard issha1('a')— a different directory — soexist('a')answeredfalsefor an id enumeration had just reported. A GC consumer acting on that pair would delete the realafrom its own shard and leave the foreign file behind. Each entry's id is now required to round-trip to the file being yielded (the raw path, orgzipPathOffor a.gzipname), which also drops names no id can spell at all, such asx.gzip.gzip— previously yielded asx.gzip, whoseexist()then threw. It costs CPU on the walk and no syscalls: measured +3.6% per id in hash mode and +21% in flat mode, against a phantom that can drive a sweep into deleting live content. Entries that are not regular files are skipped for the same reason: a directory is another id's nesting, and a fifo, socket or device node is foreign state a read now rejects, so yielding either would hand a sweep an id whoseexist()throws — failing that batch on every retry, forever. Types the listing does not report (UNKNOWN, or a symlink, or anfsadapter that answers onlyisDirectory()) stay enumerable, so the filter never costs astatper entry and never silently empties a listing.Enumeration retains a bounded amount, whatever the directory holds.
allFileIds()reads a large directory twice — once to learn which entries have a.gzipsibling, once to yield — instead of probing that sibling with a syscall per raw file, which doubled the syscalls of a full walk. Both the buffered listing and the set of compressed names are capped atMAX_BUFFERED_DIRECTORY_ENTRIES, and a large directory descends into subdirectories as it meets them rather than collecting their names first — a flat root ofd000001/x,d000002/x, … otherwise retained one string per top-level directory, which is the same unbounded shape as holding one per entry (measured 1.6MB at 50k directories, now 0.4MB flat, with the first id arriving in 62ms instead of 99ms). So retention is flat rather than proportional to the directory: measured on a flat root of compressed entries, 1.0MB retained at the first id for 50k, 150k and 300k entries alike, against 5.3MB / 15.9MB / 31.2MB when the name set was uncapped (~104 bytes per entry, so ~104MB for a million). Ids are yielded as they arrive in the second read, but the first id does wait for the first read to drain — the sibling question is answered from a listing, so the listing has to complete. Paying that is what keeps the common flat-mode shape, a root of raw-only content, free of astatper id. This matters only in flat mode: with hash prefixes a shard holdstotal/65,536entries, so any approach is cheap.allFileIds()is not guaranteed to yield a set. Every id present throughout an enumeration comes out at least once, and only ids the point lookups accept come out at all — but one case can yield an id twice, so a consumer acting on the output must be idempotent: a single flat-mode directory overMAX_BUFFERED_DIRECTORY_ENTRIESentries, where a raw file whose.gzipsibling fell outside the capped snapshot is yielded rather than probed (the reasoning is under Enumeration retains a bounded amount below). With hash prefixes it cannot arise.allFileIds()yields ids, not filenames. An id is the path of its file relative to the directory ids resolve against — the shard directory with hash prefixes, the root in flat mode — so an id containing path separators (which nests it into subdirectories) round-trips through enumeration instead of collapsing onto its last segment. The optionalprefixfilters those ids, never the on-disk filename, so it cannot match a.gzipextension.An id present throughout an enumeration is always yielded. Deciding whether a raw file is an id of its own or the decompression cache of its
.gzipsibling needs the directory's contents, and reading the directory twice to answer it — once for the compressed names, once to yield — made that answer stale between the two: a raw↔gzip transition landing in the gap made the second read skip a raw whose gzip no longer existed, so an id holding a complete representation for the whole enumeration was yielded by neither read whileexist()answeredtruefor it.A directory small enough to hold in memory (up to
MAX_BUFFERED_DIRECTORY_ENTRIES, 4096 entries) is now decided from a single read, which removes the staleness entirely: a skip is justified by an entry from the same read, so the id is yielded exactly once — never zero times, never twice. With hash prefixes this is every directory, since a shard holds total/65,536 entries and a root of 268 million ids still fits. It also halves thegetdentstraffic of a full walk: measured 34,510opendircalls for 20,000 ids down to 17,256, and 117.9 → 64.5 µs/id.A larger directory — a flat-mode root with hundreds of thousands of ids in one place — falls back to the two-read walk, because buffering that listing retained ~300 bytes per entry (47MB before the first id came out for 200k ids, ~290MB for a million). There the skip is confirmed with a
statbefore it is taken, and only for a raw that actually has a gzip sibling, which closes the id-hiding case. What remains in that path alone is the opposite, benign window: a transition can still let the second read see both representations and yield the id twice. Enumerating an id twice costs an idempotent repeat; failing to enumerate one that is present under-reports what the node holds, so that is the direction to fail in.The compressed-name set is capped as well, so a directory holding more than 4096 gzips cannot make it grow. Past the cap it is used one-sidedly: membership still suppresses a raw entry (after the confirming
stat), while absence never does — absence from a partial snapshot is not evidence, so a raw entry the snapshot does not hold is yielded rather than probed. That costs no syscalls, and what it can cost is a duplicate — the id comes out via its.gzipentry as well — never an omission, which is the direction stated above. Duplicates are bounded by the raw/gzip pairs whose gzip fell outside the snapshot, i.e. decompression-cache copies, not by the directory size.Probing those unrecorded names instead was tried and withdrawn. It cost one
statper raw id in any directory whose gzip count crossed the cap — including a flat root with a handful of compressed ids and hundreds of thousands of raw-only ones, measured at 20,000 stats for 20,000 raw ids against 0 for the same directory with no gzips — which is exactly the per-entry cost this two-pass walk exists to avoid. It also read a directory named<id>.gzipas a compressed representation and suppressed the valid raw id<id>outright, so the two paths disagreed:acame out zero times from the two-read walk and once from the buffered one. A directory can legitimately have that name, sincestoreStream('a.gzip/b')is a valid store — only an id ending in.gzipis reserved — so the confirming probe now asks whether the sibling is a regular file, not merely whether astatsucceeds.Decompressed range-cache copies are NOT reclaimed after an unclean shutdown. Cache tracking is in-memory, so a process killed without reaching
stop()leaves the decompressed copies it wrote on disk with no record of them: they are invisible to eviction, and toallFileIds()(which skips a raw whose.gzipexists). They are reclaimed only when the id is written again or deleted, so repeated hard kills accumulate disk usage bounded bydecompressCacheMaxSizeper run.decompressCacheMaxSizeis enforced on admission, not only on a timer. Recording a new decompressed copy that pushes the total past the budget triggers an eviction immediately; previously the budget was consulted only by the periodic sweep (default every 5 minutes), so a burst of range requests over distinct gzip-only ids — the shape of a sync or backfill pass — could write far past it before the first tick. Two things keep the cache usable rather than hostile to the reads that fill it. The most recently used entry is never size-evicted, because a budget smaller than a single decompressed file would otherwise delete the copy the request that just created it is about to read; and an entry is pinned from the moment it is recorded until the consumer's stream has its file descriptor, becauseretrieve()returns a lazyContentItemand protecting only the single most-recent entry is a guarantee for one reader and no more — with concurrent inflations of distinct ids, 20 parallel reads of present content produced a spuriousENOENT. A pin expires after a few seconds so an item that is never read cannot exempt its entry indefinitely, and releasing one re-checks the budget, since that is the moment a candidate becomes evictable.decompressMaxConcurrentInflations(default 4) is what bounds the transient overshoot. Admission can only be checked after an inflated file has been committed, and the eviction it triggers cannot be awaited — it needs the path lock the committing read still holds — so the cache settles at up todecompressMaxConcurrentInflations × decompressMaxFileSizeabove the budget before eviction catches up (1 GB at the defaults). Unbounded, that multiplier was the caller's own request concurrency: 50 concurrent cold range reads measured 36x over budget. Excess range reads queue rather than fail.A startup sweep that adopted them into the new instance's tracker was implemented and withdrawn: "a raw next to its own
.gzip" is not sufficient evidence that the raw is derived. A runtime quarantine, a crash mid-compression, and simply two ids differing by the suffix all produce that shape, and adopting on that premise deleted live content. Reclaiming these safely needs positive proof — inflating the gzip and comparing — which is not worth paying on every start.One directory name under the root is reserved (default
.tmp-writes, configurable viatempDirectoryName). Ids resolving into it are rejected. WithdisablePrefixHash(flat mode) the root is the content namespace, so the factory refuses to start if the reserved directory pre-exists with content it cannot prove it owns — pre-existing ids there would otherwise become silently unreachable after an upgrade. To resolve: migrate those files out, configure a differenttempDirectoryName, or restore the ownership marker if they are staging leftovers from a previous run.
