@cavsnode/dedup-sdk
v0.2.0
Published
Official TypeScript SDK for the CAVS Deduplication Service — content-defined chunking, tenant-scoped deduplication and versioned objects over your own object storage.
Maintainers
Readme
@cavsnode/dedup-sdk
Official TypeScript SDK for the CAVS Deduplication Service — the storage data plane that sits between your application and your object storage, splitting content into content-defined chunks, storing each distinct chunk once, and keeping the manifests, packs, indexes and versions that make that reversible.
It speaks the same /api/v1/dedup API the CAVS console uses, and covers both planes:
- Control plane — domains, namespaces, packs, index generations, garbage collection, encryption keys, S3 gateway credentials, tenant exports, usage.
- Data plane —
put,get,head,delete,list,listVersions. This is the part the console never exposed.
No node: imports. The transport is the global fetch and the data plane is Web
Streams, so the same build runs on Node 18+, on an edge runtime and in a
browser — though a browser should not be holding a service-account token. Ships
ESM + CJS with type declarations.
Install
npm install @cavsnode/dedup-sdkAuthentication
export CAVS_TOKEN=cavs_ci_... # an access token from the dashboard
export CAVS_API=https://cavsnode.com # or your self-hosted Hub
export CAVS_ORG=<organization slug or id>The token is sent as Authorization: Bearer … and never appears in a log line, a
message or an error.
Issue one from Dedup → Connect → Generate token, or from Organization
settings → Tokens. Tokens start with cavs_pat_, cavs_ci_ or cavs_repo_.
Which scope each call needs
| Scope | Covers |
|---|---|
| dedup:read | objects.get/head/exists/list/listVersions, domains.get/usage/listNamespaces/listPacks/listIndexes/listGCRuns/listExports, organizations.entitlement/usage/listDomains, namespaces.get |
| dedup:write | Everything above, plus objects.put/delete, batchCheck and the proof endpoints |
| dedup:admin | Everything above, plus domains.update/delete/createNamespace, namespaces.update/delete, publishIndex, startGC, rotateKeys, listKeys, the credential endpoints and createExport |
Most applications want dedup:write and nothing else. dedup:admin can delete
bytes and spend money; give it to the automation that runs collections, not to the
service that stores objects.
Token kinds
A PAT (cavs_pat_) acts as the person who created it, so its effective
permission is their role ∩ its scopes — it loses access the moment they are
demoted. A CI token (cavs_ci_) has no user behind it and is authorized by its
scopes alone, which is what you want for a deployment that must keep working when
people change roles.
Restricting a token to one domain
A token can be bound to a single deduplication domain, and optionally to a single
namespace inside it. Anything outside the binding answers 404, and a bound token
cannot create a new domain or namespace to escape it. Bind in the token dialog, or
over the API with dedup_domain_id / dedup_namespace_id.
This is the containment worth having: the credential your build agent holds should reach the one bucket it writes to, not everything your organization has stored.
Quick start
import { CAVSDedup } from "@cavsnode/dedup-sdk";
const cavs = CAVSDedup.fromEnv();
// A domain is the deduplication boundary. Chunks are shared inside one, never across.
const domain = await cavs.organizations.createDomain(cavs.requireOrg(), {
name: "Build artifacts",
slug: "build-artifacts",
});
// A namespace is a logical bucket. It bounds naming and versioning, not dedup.
const ns = await cavs.domains.createNamespace(domain.id, {
name: "assets",
versioning: true,
});
const result = await cavs.objects.put({
namespaceId: ns.id,
key: "models/resnet50.bin",
body: fileStream, // string | Uint8Array | ArrayBuffer | Blob | ReadableStream
size: fileSize,
contentType: "application/octet-stream",
metadata: { framework: "pytorch" },
});
console.log(result.logical_bytes); // what you sent
console.log(result.unique_bytes); // what the organization did not already have
console.log(result.stored_bytes); // what landed in the bucket, after compression
console.log(result.deduped_bytes); // what deduplication removedThose four numbers are deliberately separate. One collapsed "savings" figure cannot answer the question that actually matters — whether the data shrank because it repeats or because it compresses — and only the first is deduplication's doing.
Run the same put twice and the second returns unique_bytes: 0: the same bytes were
sent, and none were stored.
Reading
const object = await cavs.objects.get({ namespaceId: ns.id, key: "models/resnet50.bin" });
object.size; // bytes this response carries
object.readAmplification; // backend bytes fetched per byte served
object.packsTouched; // distinct packs this read had to open
// The body is a ReadableStream, still streaming. Consume it once.
await pipeline(object.body, createWriteStream("./resnet50.bin"));readAmplification is the honest cost of the restore. Above ~2 the profile is
fetching much more than it delivers, and domains.usage() will say so in
chunker_advice — the number exists so a fragmented profile shows up here rather
than only on the storage provider's invoice.
Byte ranges work, and only touch the packs the range needs:
const head = await cavs.objects.get({
namespaceId: ns.id,
key: "video.mp4",
range: { start: 0, end: 1023 },
});
head.partial; // true
head.totalSize; // the object's full sizeDeleting
await cavs.objects.delete({ namespaceId: ns.id, key: "models/resnet50.bin" });
await cavs.objects.delete({ namespaceId: ns.id, key: "a.bin", versionId: "…" });A delete removes a logical reference, never chunks. Whether those bytes are still needed is garbage collection's decision, because another object, another version, or another member of the same domain may reference exactly the same content. On a versioned namespace, deleting a key writes a delete marker and the history survives.
Garbage collection
await cavs.domains.startGC(domain.id); // dry run
await cavs.domains.startGC(domain.id, { dryRun: false, compact: true }); // for realdryRun defaults to true, here and on the server. Collection deletes data and
compaction spends real money, so the destructive version has to be asked for.
Client-side deduplication
const answer = await cavs.batchCheck(domain.id, [
{ hash: blake3Hex, length: 1048576 },
]);
// answer.results[i].state === "present" | "upload-required"Scoped to your own domain, with a constant-shape response and a metered hash budget: a tenant is only ever told about content it already holds, so the endpoint cannot be used to discover whether somebody else stored a particular file.
What else is on the client
| | |
|---|---|
| cavs.catalog | services, plans, profiles (public — no token needed) |
| cavs.organizations | entitlement, usage, listDomains, createDomain |
| cavs.domains | get, update, delete, usage, usageSeries, listNamespaces, createNamespace, listPacks, listUploads, listIndexes, publishIndex, listGCRuns, startGC, listKeys, rotateKeys, listCredentials, createCredential, revokeCredential, listExports, createExport, cancelExport, issueProof, verifyProof |
| cavs.namespaces | get, update, delete |
| cavs.objects | put, get, head, exists, delete, list, listVersions |
Errors
Every failure is a typed subclass of DedupError, carrying the server's code,
request_id and details:
AuthenticationError (401), AuthorizationError (403 — also legal hold and
retention), NotFoundError (404), ConflictError (409 — a concurrent write to the
same key, safe to retry), PlanLimitError (402 — details carries the limit),
RateLimitError (429), RangeNotSatisfiableError (416), TemporaryServiceError
(5xx, network, timeout).
import { PlanLimitError } from "@cavsnode/dedup-sdk";
try {
await cavs.objects.put({ /* … */ });
} catch (err) {
if (err instanceof PlanLimitError) {
console.error(err.message, err.details); // { limit, size } or { limit, used, plan }
}
}Deleting a domain
domains.delete() is refused with a 409 export_required unless the domain has a
complete, verified export — once garbage collection releases the chunks, the data
cannot be reconstructed outside CAVS. { force: true } skips that check; it is the
"there is no copy and I know it" flag, not a retry.
Retries, timeouts and cancellation
Requests retry up to 4 times on 429/5xx/network errors with exponential backoff
and jitter, honouring Retry-After. An upload whose body is a ReadableStream is
not retried — a stream can only be read once, so a retry would resend a drained
body. Hand over a Blob or a Uint8Array instead when you want retries.
Pass an AbortSignal to cancel anything:
const ac = new AbortController();
const p = cavs.objects.get({ namespaceId, key, signal: ac.signal });
ac.abort();Development
npm ci
npm run lint && npm run typecheck && npm test
npm run build # tsup → dist/ (ESM + CJS + d.ts)A working client built on this SDK lives in ../playground.
License
Apache 2.0 — see ../../LICENSE.
