@leettools/vault-core
v0.2.0
Published
Immutable, schema-bound JSONL artifact vault core
Readme
@leettools/vault-core
Standalone, database-free storage and recovery for immutable, schema-bound JSONL artifacts.
ArtifactVault registers JSON Schema documents, validates each JSONL record
against its exact vaulted writer schema, and commits content-addressed blobs plus
immutable manifests. Append-only logs commit verified, hash-chained byte ranges
for data that only grows. It needs no database or running service. An optional
daemon archives inactive blobs and closed logs with transparent gzip storage and
can copy additions to another vault through a filesystem path or rsync over SSH.
Install
npm install @leettools/vault-coreNode.js 24 or newer is required.
Example
import { ArtifactVault, JSON_SCHEMA_DIALECT } from '@leettools/vault-core'
const vault = new ArtifactVault('.app-vault')
const schema = await vault.registerSchema({
name: 'example.event',
version: '1.0.0',
schema: {
$schema: JSON_SCHEMA_DIALECT,
$id: 'urn:example:schema:event:1.0.0',
'x-leet-schema-name': 'example.event',
'x-leet-schema-version': '1.0.0',
type: 'object',
properties: {
id: { type: 'string' },
status: { enum: ['created', 'processed'] },
},
required: ['id', 'status'],
additionalProperties: false,
},
producer: { system: 'example-app' },
})
const artifact = await vault.putJsonl({
content: '{"id":"evt_1","status":"created"}\n',
schema,
kind: 'example.event',
producer: { system: 'example-app' },
})
await vault.validateArtifact(artifact.artifactId)
const bytes = await vault.readArtifactBytes(artifact.artifactId)
for await (const item of vault.artifacts({ kind: 'example.event' })) {
console.log(item.artifact_id, item.records.record_count)
}
const result = await vault.verify({ deep: true, records: true })
console.log(result.valid)API Shape
registerSchemastores a schema as an immutable schema artifact.putJsonlvalidates records and commits a new immutable data artifact.readArtifactreturns the manifest for an artifact ID.readArtifactBytesreturns verified JSONL bytes, records consumer access, and transparently restores a cold gzip blob to the raw hot representation.artifactsiterates immutable v3 manifests with optional filters.validateArtifactvalidates stored records against their vaulted writer schema.verifyaudits manifests, blobs, schemas, records, document revisions, and logs.openArtifactStream/openLegacySnapshotStreamstream verified existing content from raw or gzip storage without rehydrating it.createLog,appendLog,closeLog, andabandonLogmanage private logs;linkLog,observeLogSource, andcommitLogcapture an external writer's file without ever writing it. Every mutation takes theexpectedHeadit builds on.openLogStream,logCommits,latestLogCommit,logStatus,logs, andverifyLogread and audit logs at any published head.importLogconverts a historical copy into a closed, deep-verified log;publishMigrationReceipt,retireArtifact, andcollectGarbageretire the superseded copies and reclaim unreferenced blobs.applyLogCommitsis the replica receiver contract for verbatim log commits;readLogManifestBytesandcreateLogReplicalet any transport set a replica up.withPublicationLockcoordinates external garbage-collection roots withcollectGarbage, whoseextraRootsmay be a function evaluated under its lock.
const log = await vault.createLog({ kind: 'example.session-log', schema, producer: { system: 'example-app' } })
const commit = await vault.appendLog(log.logId, '{"id":"evt_2","status":"created"}\n', { expectedHead: log.head })
await vault.closeLog(log.logId, {
expectedHead: { seq: commit.seq, byteEnd: commit.byte_end, commitSha256: commit.commit_sha256 },
reason: 'done',
})
for await (const chunk of await vault.openLogStream(log.logId)) process.stdout.write(chunk)Strict logs (the default) reject invalid appends as a whole; report logs preserve malformed input and record per-line diagnostics. Closed logs are archived like blobs; open and abandoned logs stay raw.
The package still exports deprecated legacy manifest v2 helpers for migration compatibility. New application code should use immutable manifest v3 through the generic APIs above.
Cold Storage
# One pass using the default seven-day inactivity window.
leet-vault archive /path/to/vault --json
# Reconcile immediately and then hourly, without replication.
leet-vault daemon /path/to/vault --inactive-days 7 --archive-interval-ms 3600000The archiver stores verified gzip data alongside the content-addressed raw path, then removes only the raw representation. Artifact IDs, manifests, logical byte counts, and SHA-256 digests remain unchanged. A later consumer read verifies and rehydrates the raw bytes automatically. Maintenance validation and replication neither record access nor leave cold source blobs rehydrated.
Access is recorded at most once per digest per UTC day in retained versioned
JSONL files under artifacts/.state/access/YYYY-MM-DD.jsonl. Archive passes
consult only the current and preceding retention buckets, so the default is
conservatively seven to eight days. Journals are local operational state and are
not replicated. Concurrent writers use a transient per-day lock, and an archive
pass refreshes access recorded during its run before removing raw data. Blobs
referenced by legacy v2 manifests are not compressed.
One archive pass is available in application code:
import { VaultArchiver } from '@leettools/vault-core'
const archiver = new VaultArchiver('.app-vault')
const result = await archiver.archiveOnce({ inactiveDays: 7 })Use the unified daemon for continuous maintenance and optional replication:
import { VaultDaemon, VaultReplicator } from '@leettools/vault-core'
const controller = new AbortController()
const root = '.app-vault'
const daemon = new VaultDaemon(root, {
replicator: new VaultReplicator(root, '.replica-vault'),
})
await daemon.run({
inactiveDays: 7,
archiveIntervalMs: 3_600_000,
replicationIntervalMs: 5_000,
signal: controller.signal,
onArchive: (result) => console.log(result),
onSync: (result) => console.log(result),
onError: (task, error) => console.error(`${task} failed; retrying:`, error),
})Physical v3 blob paths are omitted from the published TypeScript API. Use
readArtifactBytes or artifact get --content; enforce this at the OS boundary
with a dedicated service account when clients must not read the vault directory.
Replication
# Copy committed artifacts once.
leet-vault sync /path/to/vault --target backup@archive:/srv/vault --json
# Run archiving and replication in one foreground process.
leet-vault daemon /path/to/vault --target backup@archive:/srv/vault --interval-ms 5000
# Avoid duplicate blob storage when a local target shares the source filesystem.
leet-vault daemon /path/to/vault --target /srv/local-replica --hard-link
# Log timestamped scan, validation, transfer, and publication progress to stderr.
leet-vault daemon /path/to/vault --target backup@archive:/srv/vault --verboseThe source must exist; the target is created on demand. SSH targets use
[user@]host:/absolute/path, including SSH config aliases. Both hosts need
rsync 3+, and the sender needs SSH. The receiver needs SSH access, a POSIX shell,
standard Unix tools (cmp, ln, wc, sync, uname, etc.), and a filesystem supporting
hard links. One of sha256sum, shasum, or openssl is required for remote
SHA-256 verification. gzip is required only when validating an independently
encoded compressed representation already at the target. No remote Node
installation, vault package, mount, or daemon is required. Configure unattended
SSH authentication and host-key trust; connections use batch mode with host-key
checking. Optional --ssh-port and --ssh-identity flags apply to both SSH
commands and rsync transfers.
Filesystem targets still work with --target /mnt/archive/vault. Those roots
must be separate, non-nested directories, including through symlinks. Direct
S3 and HTTP transports are not implemented.
For roots on the same filesystem, --hard-link publishes missing content blobs
as hard links rather than copying their bytes. Manifests remain independent,
and existing target blobs are never replaced. Because linked blobs share an
inode, directly modifying either path would corrupt both vaults; direct artifact
mutation is unsupported in every replication mode.
Run the daemon under systemd, launchd, or a container supervisor for automatic startup and process restarts. The first pass copies existing artifacts; subsequent passes discover additions from any writer. Polling runs every five seconds by default, measured after each pass. Archive and replication schedules are independent, but the daemon serializes their source-vault passes so archiving cannot change a blob during replication. SIGINT and SIGTERM cancel active SSH/rsync commands, stop filesystem replication after its current artifact, or stop immediately while idle. Local writers do not depend on the daemon or target availability.
Add --status-port 9467 to expose GET /healthz and versioned
GET /v1/status on 127.0.0.1. The status includes archive and replication
pass counts, last results/errors, and replication progress, but no artifact
content. The listener is disabled by default and the CLI restricts it to
loopback. VaultDaemon.status() and serveStatus() provide the same
observability to embedded applications.
Filesystem replication also replicates logs, re-verifying and publishing each
source commit verbatim; --hard-link shares log data with the source.
--max-log-bytes (default 256 MiB) is a soft per-pass budget, and a damaged source
history is counted in logsBlocked while other logs continue. SSH replication
sends only the committed bytes each log gained since the target's last pass. The
sender verifies them first and compresses them for the wire (an archived closed log
travels as its whole gzip); the remote shell helper checks the target's committed
tail and appends. This additionally needs gzip, tail, tr, dd, mv,
and ps remotely. Each pass resumes from the previous pass's verified journal
checkpoint, so log work tracks new commits. The remote vault opens received logs read-only and verifies newly
received journal records on access.
Replication preserves exact v3 manifest bytes and transfers schemas before dependent artifacts, publishing each manifest only after its blob is verified. It preserves verified raw or gzip representations and excludes access journals. Every existing raw or gzip sibling is verified against the manifest's logical content. Existing content is never overwritten and target-only artifacts are retained. Conflicts and corruption stop replication. The CLI reports transient I/O errors with their stacks and retries on the next interval.
Each pass scans manifests without a timestamp cursor. Restarts catch up from the
source and destination files, including replayed artifacts with old dates.
Only missing files are added to the live target. Local validation is cached
while file metadata is unchanged. Remote passes inventory the target first and
transfer only blobs it holds in neither representation, then checksum the
transferred files on both hosts and reuse identical remote manifests without
retransmitting payloads; consider a longer interval for large vaults. Staging files and unreferenced
blobs are ignored; an interrupted transfer is safe to retry. Mutable legacy v2
manifests are skipped and reported as legacyArtifactsSkipped.
Rsync transfers the missing part of the committed file list to a fresh private
staging directory, using --checksum and --link-dest. A shell helper verifies
every staged blob and manifest's expected digest and byte count, then validates
every existing raw or gzip destination against the manifest's logical content —
including for blobs the inventory already found there — before publishing
with atomic hard links. Missing SHA-256 tooling is
a configuration failure that stops replication. The helper flushes blobs
before schemas, schemas before data manifests, and all commits before returning
success. Published files are never replaced or deleted. Successful and failed
passes both clean their own staging directory; only a cancelled pass may leave an
artifacts/.staging/rsync-* directory for manual cleanup. The next pass safely
reconciles the live target. See the
rsync manual for transfer details.
The same engine is available in the library:
import { RsyncVaultReplicator } from '@leettools/vault-core'
const replicator = new RsyncVaultReplicator('.app-vault', 'backup@archive:/srv/vault')
const result = await replicator.syncOnce()
// artifactsCopied, artifactsPresent, blobsCopied, bytesCopied, legacyArtifactsSkipped, logs* counters
const controller = new AbortController()
await replicator.run({
intervalMs: 5_000,
signal: controller.signal,
onSync: (result) => console.log(result),
onError: (error) => console.error('Replication I/O failed; retrying:', error),
onProgress: ({ phase, completed, total, done }) => {},
})onProgress is optional and also accepted by syncOnce. It is called
synchronously, possibly once per file, with { phase, completed, total?,
bytesCompleted?, bytesTotal?, done }. Phases run in order and each ends with
done: true: scan and replicate (then logs when the vault has logs) for filesystem targets; scan, validate (then logs),
transfer, confirm, verify, and publish for SSH targets. transfer
totals only the files offered to rsync, excluding blobs the target inventory
already found, and a finished transfer stays below that total when the receiver
already has a file; verify and publish total every plan entry instead. An exception thrown by the listener fails the pass.
The target can also be { host, path, port?, identityFile? }. Optional third
argument { sshExecutable?, rsyncExecutable? } selects local executable paths.
For filesystem targets use new VaultReplicator(sourceRoot, targetRoot) with the
same methods, or pass { hardLink: true } as the third argument for same-filesystem
blob linking. Abort the controller to stop. Without onError, errors propagate
immediately; with it, retryable I/O and transport failures retry. A
VaultTransportError exposes retryable, exitCode, stderr in its message, and
the original cause for spawn failures. Conflicts, corrupt data, invalid options,
and missing executables stop replication. A replicator runs one pass at a time; independent replicators
can safely publish to the same target concurrently. No process starts
automatically when constructing a vault or replicator.
In the source repository, npm run test:ssh exercises the CLI daemon with a
temporary loopback SSH server, generated keys, and pinned host-key trust. It
checks recovery from host outages and process restarts, then cleans up. Run as
a regular user with rsync 3+ and OpenSSH client/server tools installed. Set
LEET_VAULT_TEST_SSHD for a server executable outside /usr/sbin/sshd. This test
runs separately from the default test suite and is required in CI.
CLI
The package installs a leet-vault command.
leet-vault init /path/to/vault --json
leet-vault schema register /path/to/vault --name example.event --version 1.0.0 --schema schema.json --json
leet-vault schema get /path/to/vault schema_... --json
leet-vault schema list /path/to/vault --json
leet-vault artifact put /path/to/vault --kind example.event --schema-ref schema-ref.json --content records.jsonl --producer-system example-app --json
leet-vault artifact get /path/to/vault art_... --content
leet-vault artifact list /path/to/vault --kind example.event --json
leet-vault artifact validate /path/to/vault art_... --json
leet-vault artifact retire /path/to/vault art_... --superseded-by art_... --json
leet-vault log create /path/to/vault --kind example.session-log --schema-ref schema-ref.json --json
leet-vault log append /path/to/vault log_... --content records.jsonl --expected-seq 0 --expected-byte-end 0 --json
leet-vault log link /path/to/vault /path/to/session.jsonl --kind example.session-log --schema-ref schema-ref.json --json
leet-vault log commit /path/to/vault log_... --current-head --allow-unterminated --json
leet-vault log get /path/to/vault log_... --content
leet-vault log import /path/to/vault --kind example.session-log --schema-ref schema-ref.json --artifact-id art_... --json
leet-vault gc /path/to/vault --json
leet-vault verify /path/to/vault --json
leet-vault archive /path/to/vault --inactive-days 7 --json
leet-vault daemon /path/to/vault --inactive-days 7 --archive-interval-ms 3600000 --status-port 9467 --jsonThe CLI mirrors the library API for scripts and CI. Prefer the library for
application code. schema register returns a VaultSchemaReference; pass it
back to artifact put via --schema-ref. Most commands accept --json for
machine-readable output.
demo writes a self-contained example vault. inspect lists manifests, prints
artifact content previews, and runs deep record verification. verify runs the
same integrity checks without printing artifact content; pass --quick for
manifest/blob-presence checks only.
