@patronage/alchemy-d1-state
v0.4.0
Published
Alchemy v2 custom state store backed by one shared Cloudflare D1 database, one table per project
Readme
@patronage/alchemy-d1-state
An Alchemy v2 custom state store backed by one shared Cloudflare D1 database, one table per project.
It replaces Alchemy's built-in Cloudflare state-store worker. Instead of each project running its own state worker behind a rotatable shared credential, every project writes to its own table in a single D1 database that is provisioned once by hand and never managed by an Alchemy stack.
0.2.0 adds exclusive stage ownership, atomic empty-stage cleanup, and child lease adoption (childEnvironment(), ownership.adoptFrom). There is no schema migration. The ownership row is a second reserved sentinel FQN in the existing (stack, stage, fqn, data) shape.
pnpm add @patronage/alchemy-d1-statealchemy and effect are peer dependencies — your stack supplies both.
Runtime
Published engines.node is ^24.0.0. This repository develops and tests on Node 24.20.0 (.nvmrc).
Node 24.14.1 node:sqlite returns an empty string for TEXT that starts with NUL. This store's reserved sentinel FQNs start with NUL. Local SQLite through the injected client option, and this package's own tests, need Node 24.20.0 or later. Cloudflare D1 HTTP is not node:sqlite and is unaffected.
Usage
import { d1State } from "@patronage/alchemy-d1-state";
import * as Alchemy from "alchemy";
import * as Cloudflare from "alchemy/Cloudflare";
import * as Effect from "effect/Effect";
export default Alchemy.Stack(
"MyApp",
{
providers: Cloudflare.providers(),
state: d1State({
databaseId: process.env.ALCHEMY_STATE_DB_ID!,
table: "my_app", // this project's table
}),
},
Effect.gen(function* () {
// ... resources ...
})
);d1State options:
| Option | Required | Notes |
| --- | --- | --- |
| databaseId | yes | The shared state database's D1 id. |
| table | yes | A bare SQL identifier ([A-Za-z_][A-Za-z0-9_]*) interpolated into DDL/DML — this project's table. |
| accountId | no | Defaults to CLOUDFLARE_ACCOUNT_ID. |
| apiToken | no | Defaults to CLOUDFLARE_API_TOKEN; needs D1 edit on the account. |
| baseUrl | no | Overrides the Cloudflare API root. |
| client | no | An injected D1Client — used by tests to run the real store code against local SQLite with no network. |
| clock | no | Local epoch-millis clock; defaults to Date.now. Decides only when an owner renews as a side effect of a write. No fence trusts it. |
| ownership | no | { stack, stage, holder, ttlMillis?, adoptFrom?, requireAdoption? }. The run takes exclusive ownership of that stage on its first state access and releases it when the run's scope closes. With adoptFrom: process.env, a run launched by a parent that holds the stage adopts the parent's lease instead. See Stage ownership. |
The service is built once per stack run and memoised with Effect.cached, mirroring Alchemy's own inMemoryState/localState layers, so the D1 backend is only exercised on real state access, never at layer-construction time.
Choosing the store for a stage
A stack that also runs against a local emulator wants Alchemy's own on-disk state, not D1. stateForStage takes that decision as an input and returns the matching layer:
import { stateForStage } from "@patronage/alchemy-d1-state";
import { isLocalEmulationStage } from "@patronage/factory-ci/alchemy";
state: stateForStage({
local: isLocalEmulationStage(stage),
props: {
databaseId: process.env.ALCHEMY_STATE_DB_ID!,
table: "my_app",
},
}),local: true returns Alchemy's localState(), which keeps the state tree under .alchemy/ in the working directory. local: false returns d1State(props).
The decision is an input, never derived here. This package inspects no stage name, reads no environment variable, and does not depend on @patronage/factory-ci; the contract runs one way only. props stays required so one call site carries both branches. On the local branch the props are unused and no D1 client is built, so a local run needs no Cloudflare credentials.
Storage model
Each project's table is keyed (stack, stage, fqn) with upsert semantics: a write to an existing key replaces the row rather than erroring or appending. Two reserved sentinel FQNs share the table with the resources, and every resource-facing query and the stack/stage inventory exclude them:
- the stack's resolved output (
getOutput/setOutput), one row per(stack, stage); - the stage's ownership row (see Stage ownership), one row per
(stack, stage)while a lease exists.
Both sentinels are rows in the existing (stack, stage, fqn, data) shape. No migration adds them, and get/set/delete refuse a sentinel FQN, so a resource can never be confused with one.
Stage ownership
One (stack, stage) is owned by at most one holder at a time. The store enforces this in the database, not in the calling process: every write is one SQL statement whose WHERE clause re-checks the ownership row, so a stale holder's write changes zero rows and fails.
const state = buildService(client, { databaseId, table: "my_app" });
const cleaned = await Effect.runPromise(
Effect.gen(function* () {
const owner = yield* state.acquireStageOwnership({
stack: "MyApp",
stage: "pr-123",
holder: "cleanup-run-42",
});
const outcome = yield* owner.deleteEmptyStage(); // "removed" | "absent" | "occupied"
yield* owner.release();
return outcome;
})
);acquireStageOwnership fails with StageOwnershipError (reason: "held", plus the holder label and lease expiry) while another holder's lease is unexpired. Distinct stages are independent. The returned StageOwnership gives you:
| Member | What it does |
| --- | --- |
| state | A StateService whose writes to the owned stage are fenced by the lease. Reads are unfenced and may target any stack or stage. Writes to any other stage are refused. |
| renew() | Extends the lease by its TTL. Fails with reason: "lost" once the lease expired, was released, or was taken over. Writes past the lease's half-life renew it as a side effect, so an active run keeps its stage. |
| release() | Drops the lease. Idempotent. On an adopted handle it drops nothing. |
| deleteEmptyStage() | Atomically removes the stack output row if, in the same statement, it is {}, no resource row exists, and the lease still holds. It never deletes a resource row. |
| childEnvironment() | The environment entries a child process needs to adopt this lease. They carry the bearer token: spread them into the child's environment and nowhere else. |
The default TTL is ten minutes (DEFAULT_OWNERSHIP_TTL_MILLIS). A holder that stops renewing — a killed process — loses the stage after the TTL, and the next acquireStageOwnership takes it over. The token that proves ownership is generated in the holder's process and leaves it only through childEnvironment(), into a child the holder launches; the table stores its SHA-256, and neither StageOwnership nor StageSnapshot nor any error carries either.
Unowned writes. A client that writes without ownership (the plain d1State layer, or buildService used directly) is refused with a StateStoreError whose cause is a StageOwnershipError while any unexpired owner exists. Two unowned clients still race each other; give every mutating entry point an ownership to get mutual exclusion.
Running Alchemy under a lease. Pass ownership to d1State. The stack run acquires the stage on its first state access — a plan, a deploy or a destroy — and releases it when the run's scope closes. A second run on the same stage fails at its first state call with the holder's label and lease expiry.
Handing a lease to a child process. A parent that holds the stage — a CI lifecycle that plans, applies and checks around an alchemy child — keeps holding it across the child. The parent spreads owner.childEnvironment() into the child's environment; the child passes its own environment as ownership.adoptFrom:
// The child's alchemy.run.ts
state: d1State({
databaseId,
table: "my_app",
ownership: {
stack: "MyApp",
stage,
holder: runId,
adoptFrom: process.env,
requireAdoption: true,
},
}),Adoption is fenced like every other ownership operation: one statement, on the token hash and an unexpired row by the database clock. It renews the lease the way a write past half-life does and changes nothing else about the row; the holder label stays the parent's. An expired, released or superseded lease cannot be adopted (reason: "lost"), so a child never revives a lease its parent lost. An adopted handle's release() drops nothing: the lease ends when the parent releases it. With requireAdoption: true, a missing or empty handoff refuses with reason: "lost" before any database access. This prevents a direct invocation of the child entry from minting its own lease and bypassing its parent lifecycle. A supplied handoff must still pass the existing token and scope fence; presence alone grants nothing. Invalid non-boolean requireAdoption values fail as StateStoreError before database access.
Keep requireAdoption omitted or false for standalone/dev paths that are allowed to acquire their own lease. In that default mode, a missing or empty handoff acquires as usual. The consumer chooses which entries require a parent; the store knows no stage names or deploy policy. stateForStage({ local: true, ... }) still chooses local disk state and does not apply D1 ownership options.
Adoption resets the lease window to the adopting process's TTL. Adoption is the renewal statement, and that statement assigns expiresAt = <database now> + ttl where ttl is the adopting call's own — its ttlMillis if it supplies one, the store default otherwise. It does not take the maximum against the stored expiry, so adoption can shorten the remaining window as easily as lengthen it: a parent holding a thirty-minute custom lease that is adopted by a child taking the ten-minute default has its expiry moved twenty minutes earlier.
Exclusivity is unaffected — a competing acquisition still fails with held — and a shortened window fails closed rather than silently. If the parent's lease lapses earlier than it expected, the fenced renewal on the way out of the leased window changes no rows, and the operation reports OWNERSHIP_LOST instead of proceeding on a lease it no longer holds. The cost is a refused operation and a reconciling retry, not a double mutation. A consumer that sets ownershipTtlMillis should therefore assume an adopting child resets the window to its own TTL, and size the child's TTL for the work rather than only the parent's.
The parent's continuous lease is what makes an unowned mutating child impossible: while the parent holds the stage, a child that takes its own lease is refused at its first state call, and a child that takes none is refused at its first write. The store's fence decides; nothing in the child is trusted to opt in.
What fencing does and does not cover. Ownership fences the state database only. An Alchemy provider call that already reached Cloudflare before the lease was lost is not undone; the refused state write is how the loss surfaces. The run fails at that resource, and the remote resource it created or changed exists without a matching state row. Recovering that drift — re-running under a fresh lease so Alchemy adopts or replaces it — is the consumer's decision.
Lease timestamps are written from, and compared against, the database's own clock inside each statement (julianday('now'), exposed as DB_NOW). No fence reads the caller's clock, so a client whose clock is fast, slow, or stale sees the same expiry as every other client, and a write delayed past its lease is refused when it executes, not when it was issued. The clock prop only decides when an owner renews as a side effect of a write. A release that changes no rows while this holder's row is still present fails instead of reporting success; under d1State({ ownership }) that failure surfaces from the run's scope close.
Stage snapshot
snapshotStage({ stack, stage }) reads resources, stack output and ownership in one SQL statement, so the three agree with each other:
const snapshot =
yield * state.snapshotStage({ stack: "MyApp", stage: "pr-123" });
// { stack, stage, resources: ["MyApp/Worker"],
// output: { present: true, empty: false, value: {...} },
// ownership: { held: true, holder: "deploy-run-41", expiresAt: 1725480000000 },
// empty: false }empty is true when no resource row exists and the output row is absent or {}. An ownership row alone does not make a stage occupied. An unreadable store is a StateStoreError failure, never an empty snapshot, so cleanup code cannot mistake an outage for absence.
Provisioning
The store creates and writes rows; it never creates the database. Provision that once, by hand, with wrangler, and do not let an Alchemy stack manage it — an Alchemy-managed store database would let one project's destroy delete every project's state.
1. Create the state database
wrangler d1 create my-alchemy-stateKeep the printed database id; that is what each stack passes as databaseId (Patronage projects read it from ALCHEMY_STATE_DB_ID). One database serves every project — adding a project adds a table, not a database.
2. Create this project's table
The table shape is part of the package's contract, so the migrations ship in the tarball:
node_modules/@patronage/alchemy-d1-state/provisioning/migrations/
0001_create_project_table.sql template — replace `project_table` with your table name
0002_create_hq_table.sql a worked example of that templateCopy 0001_create_project_table.sql into your own migrations/ directory, rename project_table to the name you pass to d1State({ table }), and apply it:
wrangler d1 migrations apply my-alchemy-state --remoteAdding a project is purely additive — a new table, no other project's data or credentials touched. If the migration is skipped the store will CREATE TABLE IF NOT EXISTS its own table on first write; running the migration is still preferable, because it makes the schema reviewable and diffable before any state depends on it.
3. Enable durability
The state database is the only record of what your stacks own, so back it up independently of D1:
- Time Travel is on by default (30-day point-in-time restore). Confirm with
wrangler d1 time-travel info my-alchemy-state. - Add a scheduled export to R2 — e.g.
wrangler d1 export my-alchemy-state --remote --output=state-$(date +%F).sqlon a cron.
Schema
CREATE TABLE IF NOT EXISTS "<table>" (
stack TEXT NOT NULL,
stage TEXT NOT NULL,
fqn TEXT NOT NULL,
data TEXT NOT NULL,
PRIMARY KEY (stack, stage, fqn)
);FQNs are stored via Alchemy's encodeFqn (/ → __). The output and ownership sentinel rows use the same columns; an adopter on this schema needs no further migration to use stage ownership.
Fail-closed secret guard
This store ships with no encryption key by design. Built-in Alchemy stores rely on at-rest encryption to make persisting an Effect Redacted<T> marker safe; this store has nothing to rely on, so it refuses instead.
set() and setOutput() scan the encoded value before writing and fail with a StateStoreError if a Redacted marker appears anywhere in it. The message names the JSON path of the leak plus the resource FQN (set) or the stack/stage the output belongs to (setOutput). It is an Effect failure, not a thrown exception — catch it with Effect.catchTag, not try/catch. Nothing is written when the guard fires: the scan runs before the upsert.
Keep secrets out of persisted props/attr entirely. Store a reference or id and read the secret from Secrets Store or the environment at runtime.
Trust boundary
A Cloudflare API token with D1 access is account-wide: any token that can write one project's table can read and write every project's table in the shared database. Table-per-project is operational isolation, not a security boundary. What this store removes is a shared rotatable credential and secrets at rest — not cross-project inaccessibility. Treat every project's CLOUDFLARE_API_TOKEN as able to reach all state in the shared database, and scope and rotate tokens accordingly.
Qualification
The store's own suite runs the real store code against local SQLite through the injected D1Client, so every fence, every meta.changes count and every DB_NOW comparison is executed by a real SQL engine rather than modelled.
Two further suites run this package as a consumer installs it. In this repository, software-factory-hq/alchemy/lifecycle-integration.test.ts drives runAlchemyLifecycle against this store in a file-backed SQLite database while a second process adopts the parent's lease and writes through it — a real contending process, not a stub. And factory-ci/tests/packed-lifecycle-smoke.test.ts packs this package and @patronage/factory-ci together and repeats the exercise from the tarballs, over one installed copy of alchemy and one of effect. That last one detects a second copy of alchemy resolved by this package's tarball: an error this store raises must be instanceof the StateStoreError class the consumer imported, and a duplicate names its class the same, so only instanceof tells them apart. That detection is all it claims, and the claim stops at this package. A duplicate under @patronage/factory-ci's own tarball leaves the whole suite green — the check never compares a class that package raised — and is not asserted anywhere. Neither is what a duplicate would go on to break: Alchemy's services are keyed by string, so a faithful duplicate still satisfies the same State key and nothing else in the suite notices. It proves nothing about effect, which is unproven rather than proven.
Exports
| Export | Kind | Notes |
| --- | --- | --- |
| d1State | function | The state: layer factory for Alchemy.Stack. Takes D1StateLayerProps. |
| buildService | function | The underlying D1StateService builder; d1State composes it with a lazily-resolved client. |
| STATE_STORE_ID | constant | Telemetry slug ("cloudflare-d1") reported as alchemy.state_store.id. |
| DEFAULT_OWNERSHIP_TTL_MILLIS | constant | Default lease length, ten minutes. |
| DB_NOW | constant | The SQL expression for the database clock in epoch milliseconds; every lease timestamp and expiry predicate uses it. |
| StageOwnershipError | class | Tagged error (_tag: "StageOwnershipError") with reason: "held" \| "lost", target, and the current holder/expiresAt when known. |
| D1StateProps | type | { databaseId, table, accountId?, apiToken?, fetch?, baseUrl?, client?, clock? }. |
| D1StateLayerProps | type | D1StateProps plus optional ownership: { stack, stage, holder, ttlMillis?, adoptFrom?, requireAdoption? }. |
| D1StateService | type | StateService plus snapshotStage and acquireStageOwnership, which also adopts a parent's lease from adoptFrom. |
| StageOwnership | type | A held lease: state, renew, release, deleteEmptyStage, childEnvironment. |
| StageOwnershipProps | type | { holder, ttlMillis?, adoptFrom? }. |
| StageSnapshot | type | { stack, stage, resources, output, ownership, empty }. |
| StageTarget, StageOwnershipView, StackOutputOccupancy, StageCleanupOutcome, StageOwnershipRefusal | types | The parts of the two above. |
| createHttpD1Client | function | Builds a D1Client that calls the real Cloudflare D1 HTTP query API. |
| lazyHttpD1Client | function | Defers client construction (and credential reads) until first use. |
| D1Client | type | Minimal client interface: run one parameterised SQL statement, get rows back. |
| D1QueryResult | type | Row set returned by a single D1Client statement. |
| HttpD1ClientProps | type | Props for createHttpD1Client. |
| findRedactedPath | function | Scans an encoded value for an Effect Redacted marker and returns its JSON path; backs the secret guard. |
License
MIT
