@finbheara/idfs
v1.0.0
Published
Stores large files by hash in a remote store and leaves a text pointer in git.
Maintainers
Readme
idfs
idfs stores large files by hash in a remote store and leaves a text pointer (version idfs/1) in git.
It replaces git-lfs for repositories whose large files are not needed on every checkout,
that outgrow Github's LFS storage and transfer quotas. The goal is an experience similar
to git-annex but with the worktree usability and object store layout of lfs
It includes migration from git-lfs and can interoperate with it. Pointers are always exchanged explicitly. git
behaves consistently regardless of how many or few files have been hydrated.
Install
idfs is TypeScript run directly by Bun (1.4.2 or later). Add it to a consuming repo as a dev dependency, pinned to an exact version:
bun add -d --exact @finbheara/[email protected]{
"devDependencies": {
"@finbheara/idfs": "1.0.0"
}
}Pin the exact version, not a range, so local hooks, CI and every workstation run the same idfs.
Commit bun.lock, and install with bun install --frozen-lockfile in CI. The package installs
an idfs bin, so bun idfs <verb> (or bunx idfs <verb>) works from anywhere in the repo.
Wire the hooks and the clean filter as shown under
Configuration in a consuming repo.
Upgrading is a deliberate change: bump the pinned version, read CHANGELOG.md,
and commit the new lockfile. A major version changes an invariant or the pointer format.
Library entry point
@finbheara/idfs/pointer is the resolver: parse, inspect, format, objectKey and the
version constants. A tool that reads pointers imports it instead of parsing pointer text itself.
It is the only supported import, covered by semver like the CLI. Paths under src/ and the bare
package name are not importable. The subpath is TypeScript source, so it needs Bun or a bundler.
import { parse } from "@finbheara/idfs/pointer";Pointer
A tracked path holds this text instead of its bytes:
version idfs/1
oid sha256:<64 lowercase hex>
size <bytes>- It is the git-lfs v1 pointer with a different version line. The oid is the same sha256 git-lfs uses. GitHub does not treat it as an LFS pointer, so pushes are not checked against GitHub's LFS store.
- The reader also accepts legacy git-lfs v1 pointers, so history keeps resolving.
- A pointer carries no date, tag or path. Two copies of the same bytes have identical pointers.
Commands
bun idfs <verb>:
| verb | does | network |
|---|---|---|
| add <file> | hash the file into the cache, hardlink it back, stage the pointer | none |
| pull <glob> | fetch missing objects into the cache, verify them, hardlink them into the worktree. --cached-only relinks without network. Records each glob for this worktree. A bare pull uses the recorded globs, and with none links nothing; '**' asks for everything. --forget <glob> drops a recorded glob and removes no file. See What a worktree hydrates. Reads a remote over its publicUrl when its credentials are not set (see Pulling without credentials) | GET |
| push | upload every object the presence log does not show on each remote that accepts its use. Refuses, sending nothing, when no remote accepts an object's use | conditional PUT |
| whereis <path> | list the remotes holding a path's object, from the presence log | none |
| check | fail if any path in the idfs class holds bytes instead of a pointer. Read-only | none |
| fsck | re-hash the cache against the pointers and report. Read-only | none |
| link <path> | print a presigned GET URL for a path's object. --expires (default 1d, at most 7d), --remote. Refuses restricted content unless --restricted | none |
| manifest | print path, oid, size, use for every pointer in a tree. --url adds a public URL column | none |
| migrate <glob> | move tracked git-lfs files into idfs: cache from the local LFS store, upload, rewrite pointers to idfs/1, hand the glob from filter=lfs to filter=idfs in .gitattributes. Uploads each object only to remotes that accept its use, and fails the paths of an object no remote accepts. Stages, never commits. --dry-run, --report, --fetch-lfs, --no-archive. See Migrating from git-lfs | conditional PUT |
There is deliberately no drop, sync or get.
When the reader of stdout goes away, as in idfs manifest | head, a read-only verb (whereis,
check, fsck, link, manifest, migrate --dry-run) stops at that write and exits 0 with
nothing on stderr. A verb with side effects never stops part way: it drops the rest of its output,
runs to completion and keeps its own exit status, so hook post-checkout still exits 0 and
hook pre-push still lets the push through. It notes the lost output on stderr if stderr is still
open. Any other write error, such as ENOSPC or EIO, fails the verb loudly, possibly part way.
How it fits together
- Resolver. The only code that parses pointers.
- Cache.
<git-common-dir>/idfs/objects/aa/bb/<oid>, shared by every linked worktree. An object enters it only after its sha256 and size match, and is stored mode 0444. - Remotes. One interface:
putIfAbsent,head,get,list(prefix), and an optionalpresign. There is no delete. Each implementation maps its provider's responses topresent / absent / created / refused.- R2: the S3 API, region
auto, path-style. Under a retention lock, R2 answers a conditional PUT on an existing key with 409, not 412. That 409 counts as present only after a HEAD confirms the size. - S3: AWS, including
DEEP_ARCHIVEstorage for an archive remote. - SSH: a directory on a host reached over one multiplexed
sshconnection.putIfAbsentuploads to a temp name, then hardlinks to the final name, which fails if the name exists. - Public HTTP: read-only, over a remote's
publicUrl, with plain GET and no credentials.pulluses it in place of an r2 or aws remote whose credentials are not set. It has theread-onlycapability:putIfAbsentis refused andlistthrows, both without a request. - Every remote uses the key layout
lfs/objects/aa/bb/<oid>, the same as git-lfs's local store. - A remote has a role,
primaryorarchive, and a cost.pullnever reads an archive remote. - Requests are SigV4-signed over
fetch. Bun's S3 client cannot sendIf-None-Matchor a storage class.
- R2: the S3 API, region
- Presence log. An orphan branch,
idfs-log, written with git plumbing. It holds one file per storage location, one sorted line per object,oid size confirmed-at, and merges by union. A line is added only after the remote confirms the object. A clone that has fetched the branch knows what every remote holds without a request. A remote is listed once, to seed its file, and never again. - CLI and git hooks. pre-commit runs
check. pre-push runspush. post-checkout runspull --cached-onlyon the globs this worktree has recorded, and does nothing when there are none.
What a worktree hydrates
Nothing, until it asks. Content arrives on an explicit pull, never on checkout and never by
default. Linked worktrees share one cache that can hold every object the repository has ever
had, while each worktree usually needs a few of them.
pull <glob>..., with or without--cached-only, appends its globs to this worktree's hydrate list,<git-dir>/idfs/hydrate.<git-dir>isgit rev-parse --git-dir: the worktree's own git dir (.git/worktrees/<name>for a linked worktree), not the shared common dir, so every worktree keeps its own list.- The file holds one git pathspec per line, relative to the top of the tree, in the order first asked for. Duplicates are not added and blank lines are ignored. It may be edited by hand.
- The list only ever narrows what is linked.
pullrefuses, recording and sending nothing, a pathspec git rejects (../x,:(bogus)x), one that looks narrow but matches everything (.,*,:/,:(top)), and excludes alone (:!a, which git reads as everything buta). An exclude beside what it narrows is fine:pull 'a/**' ':!a/big.pdf'.'**'is the one way to ask for everything. hook post-checkoutrelinks from the cache exactly the recorded globs, with no request. With none recorded it does nothing and prints nothing, so a freshgit worktree addlinks no file. A hand-edited line that fails the checks above is skipped with a warning naming it, the valid lines still relink, and an unreadable list is a warning. The hook never fails a checkout.- A bare
pulluses the recorded globs. With none, it links nothing, sends nothing, prints a hint and exits 0.pull '**'is the explicit way to ask for everything, and records**. pull --forget <glob>...drops exact lines from the list and exits 1 if one was not recorded. It is local state, not content: forgetting a glob unlinks, removes and changes no file or cache object. Files already linked stay linked until git next checks out their path.addandmigratelink only the paths they handle and record nothing.
Worktree files are hardlinks
Tools such as pdftotext need real paths, so pull hardlinks cache objects into the worktree.
- A clean filter turns the bytes back into the pointer, so
git statusstays clean. The index entry is re-staged with zeroed stat after linking. - Pointer paths need
merge=binary. Otherwise a conflict writes markers into a pointer. - Mode 0444 blocks appends, truncation and in-place writes.
sed -ireplaces the link and leaves the cache alone.chmod u+wand vim's:wq!still write through, sofsckexists to catch that, andpullre-fetches anything that fails verification. - Containers mount the cache read-only (
-v <cache>:/idfs:ro).
Configuration in a consuming repo
.gitattributes:idfsmarks the paths that must be pointers.idfs-use=restrictedmarks restricted content. The pointer paths also carrymerge=binaryandfilter=idfs.- Remote definitions: name, provider, bucket or host path, role, cost, and the uses it accepts. Credentials come from the environment, or from an env file that fills only what the environment leaves unset.
- Hooks: the consuming repo's hook scripts call the idfs entry points.
Remote definitions live in .idfs.json at the top of the worktree. It names environment
variables and never holds credential values. Without credentialsEnv, a remote named cold reads
IDFS_COLD_ACCESS_KEY_ID and IDFS_COLD_SECRET_ACCESS_KEY.
{
"gitRemote": "origin",
"remotes": [
{ "name": "r2", "provider": "r2", "bucket": "example-bucket", "endpointEnv": "R2_ENDPOINT",
"role": "primary", "cost": 10,
"credentialsEnv": { "accessKeyId": "R2_ACCESS_KEY_ID", "secretAccessKey": "R2_SECRET_ACCESS_KEY" } },
{ "name": "cold", "provider": "aws", "bucket": "example-archive", "region": "us-east-2",
"storageClass": "DEEP_ARCHIVE", "role": "archive", "cost": 100 },
{ "name": "nas", "provider": "ssh", "host": "host.example", "path": "/srv/idfs-store", "role": "primary", "cost": 5 }
]
}Env file
Linked worktrees have no environment of their own. To avoid exporting credentials before every call, point idfs at an env file once. Git config is shared by every worktree of a repository:
git config idfs.envFile ~/.config/idfs/creds.env
chmod 600 ~/.config/idfs/creds.envLookup order for the path: $IDFS_ENV_FILE if set (the override), else git config
idfs.envFile (local, then global). A leading ~/ is expanded, and a relative path is resolved
against the repository root.
Lookup order for each variable: the process environment, then the env file. The environment always wins, and the file only fills a variable that is not set at all (an exported empty string is left alone).
- Only the variables a configured remote names are read:
endpointEnv,credentialsEnv(or the defaultIDFS_<NAME>_...pair) andpresignCredentialsEnv. Every other key in the file is ignored. The values are held in idfs's own copy of the environment, never exported toprocess.envor to a child process. - Lines are
KEY=VALUE, with#comment lines, an optionalexportprefix, and optional'single'or"double"quotes (\",\\,\n,\tin double quotes). There is no variable expansion. An unquoted value ends at#. - The file is read only when a needed variable is unset. If it is missing or unreadable then, the "is not set" error also names its path. Otherwise idfs stays silent.
- A file readable by group or world gets a warning on stderr naming the path and mode, and is still read. Keep it at 0600. No message ever prints a value.
Permitted use
Restricted content is content that must never be served or listed publicly: no public URL in
manifest --url, no URL from link without --restricted, never fetched from a public URL by pull, and never
offered by a public host. Most repositories have few such paths. It does not limit which remotes
hold the bytes unless a remote opts out with accepts.
A path's use is its idfs-use attribute. Only an unspecified or unset attribute, or the value
public, is public. Any other value, including a bare idfs-use or a misspelled value, is
restricted. manifest prints the use read the same way.
acceptslists the uses whose content a remote may hold:"public","restricted", or both. It defaults to both. A remote that must not hold restricted content, for example a bucket served directly at a public URL, says"accepts": ["public"]. A public bucket serves any key to whoever knows the oid, and oids are in the pointers in git, so restricted objects belong only on remotes that are not served publicly, or behind a host that refuses them.An object's use is the strictest use of every path that holds it. One restricted path makes the object restricted on every path.
push,migrate,pullover a public URL,manifest --urlandlinkshare this rule. An object whose path cannot be found counts as restricted.pushand the pre-push hook look at the commits being pushed. An object is restricted if, in one of those commits, a path holds it and that commit's attributes mark the path non-public. A path that was restricted only before its content changed does not restrict the later content.migratereads the index (worktree attributes) and HEAD's history the same way.Each guard reads one view of the repository:
| guard | paths read | attributes used | |---|---|---| |
manifest,manifest --url| the tree at--rev(HEAD by default) | that tree's, plus--attributes| |link| the path itself, the index, and HEAD | the worktree's for the path and the index; HEAD's and the worktree's for HEAD | |pullover a public URL | the index | the worktree's | |push, pre-push hook | every commit being pushed | each commit's own | |migrate| the index, and HEAD's history | the worktree's for the index; each commit's own for history |So
manifestandlinkdo not read history: an object restricted only in an older commit gets a URL frommanifest --urland fromlink, whilepushstill keeps it off a remote that accepts onlypublic.Paths are judged from the top of the tree, from whatever directory a command runs in. A path git gives no answer for counts as restricted.
Known gap: only the commits being pushed are read. An object that is restricted only on a local branch that is not being pushed does not count.
pushrefuses, before any request, when any object's use is accepted by no configured remote. The pre-push hook then refuses the git push.migratefails the paths of such an object and uploads nothing for it. With--no-archive, archive remotes do not count.push --remote <name>sends that remote only what it accepts. The rest is left to the other remotes, and is not refused as long as one of them accepts it.
The presence log has one file per storage location, not per remote name. The file is named by
a location id: r2:<endpoint host>/<bucket>, aws:s3.<region>.amazonaws.com/<bucket> or
ssh:<host>[:<port>]<absolute path>, with every byte outside A-Z a-z 0-9 . _ - escaped as
%XX. For example, nas above logs to ssh%3Ahost.example%2Fsrv%2Fidfs-store.log. Two definitions of one
location share a log, and renaming a remote keeps its log.
Share links
Two optional keys on a remote definition:
publicUrl: an https base under which the remote's keys are served without credentials, for example a public bucket domain. It must end with/and carry no query, fragment or credentials.manifest --urlprintspublicUrl+ key, e.g.https://files.example/lfs/objects/aa/bb/<oid>, for content whoseidfs-useispublic, taken from the cheapest non-archive remote with apublicUrlthat the presence log shows holding the object. Restricted content, and content no such remote holds, gets an empty column. An object reachable from any restricted path counts as restricted on every path, public ones included. An empty column hides nothing: anyone who knows an oid can fetch it from a public host, so keep restricted content off remotes that have one.presignCredentialsEnv(r2 and aws only): the environment variables holding the key pair thatidfs linksigns with. Production must sign with a read-only key. A read-only pair can be given to whatever issues links without granting writes, and rotating it revokes every outstanding link without touching uploads. Without this key,linksigns with the remote's own key and warns on stderr.
idfs link <path> signs locally and sends no request. It picks the cheapest non-archive r2 or
aws remote that the presence log shows holding the object, or the remote named by --remote.
SigV4 caps a presigned URL at seven days. Presigning uses long-lived access keys: temporary
credentials that need a session token are not supported.
A presigned URL is a working public link until it expires, so link refuses restricted content
by default: it exits 1, prints nothing on stdout, and names on stderr the path that restricts the
object (the path itself, or another path holding the same object) and the override. With
--restricted it signs anyway and notes on stderr that the URL exposes restricted content.
--restricted on public content changes nothing. Neither case sends a request. With --remote, it also warns when the
presence log does not show the object on that remote, because the URL may answer 404. The URL carries the signing key's access key id and a
signature, never the secret. That is inherent to SigV4 query signing: treat the URL itself as a
credential for that one object until it expires.
{ "name": "r2", "provider": "r2", "bucket": "example-bucket", "endpointEnv": "R2_ENDPOINT",
"role": "primary", "cost": 10, "publicUrl": "https://files.example/",
"presignCredentialsEnv": { "accessKeyId": "R2_LINK_ACCESS_KEY_ID", "secretAccessKey": "R2_LINK_SECRET_ACCESS_KEY" } }Pulling without credentials
CI and release jobs can hydrate objects with no credentials at all. pull reads an r2 or aws
remote over its publicUrl when neither of the remote's credential variables is set. Nothing
else needs configuring: the publicUrl already in .idfs.json is enough, and the endpoint
variable may be unset too.
bun idfs pull 'data/**' # no IDFS_*/R2_* credentials in the environment- Each such remote is announced on stderr, naming the unset variables and the URL. With one
variable set and the other missing,
pullstops with the usual error instead of quietly reading without credentials. An ssh remote is always read over ssh. - Requests are
GET <publicUrl>lfs/objects/aa/bb/<oid>with no credential header. Redirects are followed by hand, at most 5, never from https to http, and every hop is logged. - The body streams to disk, capped at exactly the pointer's size, with no fixed maximum. A
Content-Lengthabove it is refused before reading, a longer body is cut off, and a shorter one fails. The partial file is removed. A network error fails that remote for that object, andpulltries the next remote. - Every object is verified (sha256 and size) before it enters the cache, exactly as with a signed read. A wrong body is rejected and nothing is cached.
- Listing is never public; fetching a blob is. The public path only GETs (and can HEAD)
one object by its oid. The read-only remote has no way to list:
listthrows without a request. A presence log is seeded only bylog seedwith credentials, never from a public URL. - An unauthenticated read never writes the presence log, and never PUTs.
pullnever reads an archive remote, public or not. - Restricted content is never requested from a public URL, and is never expected there. An
object reachable from any path in the index whose
idfs-useis notpublicis restricted on every path, as withmanifest --url.pullfetches it only from a remote read with credentials. Otherwise it names each such path on stderr, counts the object as missing and exits non-zero. It never skips one silently. - That rule only keeps
pullfrom asking. Keeping restricted content off a remote that has apublicUrlis up to routing: give such a remote"accepts": ["public"], so restricted content is pushed only to remotes without one.
Hook and filter wiring:
.githooks/pre-commit: bun idfs hook pre-commit
.githooks/pre-push: bun idfs hook pre-push "$@" # stdin passed through
.githooks/post-checkout: bun idfs hook post-checkout "$@" # relinks recorded globs only
git config filter.idfs.process "bun idfs hook filter-process" # clean only, no smudge
git config filter.idfs.clean "bun idfs hook clean %f" # fallbackhook filter-processspeaks git's long-running filter protocol, so one process cleans every file of a git command. It advertisescapability=cleanonly: no smudge and no delay.hook clean %fstarts one process per file, at roughly 60-130 ms each. The firstgit statusafterpullcleans every relinked path, so with 3,000 hydrated paths that costs minutes. Only recorded globs are relinked, so this scales with what a worktree asked for, not with the cache.- git uses
processwhen both are set. Keepcleanfor a setup that lacksprocess. - Both give the same output, byte for byte, from the same code. Content past 64 MiB streams through a temp file in the cache instead of memory.
- A file the filter process fails on gets
status=error, and git keeps that file's bytes for that command, as whenhook cleanexits non-zero. pre-commit'scheckstill refuses bytes. - Do not set
filter.idfs.required=true. There is no smudge, so a required filter makes every checkout of a pointer path fail withsmudge filter idfs failed. - A temp file of content past 64 MiB is removed on every failure and on SIGTERM, SIGINT and
SIGPIPE. A process killed with SIGKILL can leave one in
<cache>/tmp.fsckreports those older than an hour, and never removes them.
Migrating from git-lfs
migrate keeps every oid: the idfs/1 pointer carries the same oid sha256: and size as the
git-lfs pointer. It rewrites a git-lfs pointer only when the version line is the only change.
That means the pointer has no ext-* lines and its lines are in the canonical order (version,
oid, size). Any other git-lfs pointer fails its path. It is not cached, uploaded or rewritten,
and migrate exits 1.
--report prints one line per matching path to stdout as each path is classified. The summary
moves to stderr. The line is tab-separated:
path oid size status| status | meaning | migrated |
|---|---|---|
| version-only | git-lfs pointer. Its idfs/1 rewrite differs only in the version line | yes |
| ext-lines | git-lfs pointer with ext-* lines, which idfs/1 cannot carry | no, fails the path |
| non-canonical | git-lfs pointer in another form, such as size before oid | no, fails the path |
| already-idfs | idfs/1 pointer | unchanged |
| not-pointer | not a pointer. oid and size are - | unchanged |
The dry run and the real run print the same report. Neither makes a request to build it. To gate a migration on it:
bun idfs migrate --dry-run --report '*.pdf' 'data/**' > migrate-report.tsv
awk -F'\t' '$4 != "version-only" && $4 != "already-idfs"' migrate-report.tsvThe status is the path's classification, not its final outcome. A version-only path can still
fail later, for example when its object is missing from the local git-lfs store or a required
remote does not confirm it, or be skipped for local changes. Those outcomes are in the summary,
which --report sends to stderr. The exit status is 1 if any path fails, including a path whose
pointer would change more than the version line.
A path holding a control character (such as a tab or newline), a double quote or a backslash is
C-quoted the way git quotes paths: wrapped in ", with \t, \n, \", \\ and the like,
and other control characters as three-digit octal (\001). Non-ASCII is printed as is, as with
core.quotePath=false. Any other path is printed unquoted, so a line always has four fields.
Later
A web service that serves a published manifest as a browsable layout with share links.
Development
bun install
bun run typecheck
bun run lint
bun test
bun run test:ssh # SSH contract suite against throwaway OpenSSH containers (needs docker)
bun run test:live # R2 and AWS contract runs against scratch buckets, credentials from the envReleasing
- On a branch, set
versioninpackage.jsonand move the[Unreleased]notes inCHANGELOG.mdunder a## [X.Y.Z] - <date>heading. Merge the PR. - Tag the merge commit on
mainand push the tag:git tag vX.Y.Z && git push origin vX.Y.Z. .github/workflows/release.ymlchecks that the tag equalsv+ thepackage.jsonversion and that the changelog has the entry, runs typecheck, lint and tests, and publishes withnpm publish --provenance --access public.
The workflow authenticates with npm trusted publishing (GitHub OIDC); the repo stores no npm
token. Trusted publishing is configured on npmjs.com once the package exists, so the first
version is published by hand from a clean checkout of the tag:
npm publish --access public --provenance=false (provenance can only be generated in CI).
Never publish a version twice; fix forward with a new patch version.
License
MIT. See LICENSE.
