Full-tree review before task decomposition found five composition
defects (mechanisms specced correctly in isolation, composition
unruled) and a set of caller-facing gaps. ADR-012 rules each:
- Sweeps are never errors: aborts are GcAbortCause report data in
Ok(SweepReport) (ProtectFailed / SweeperLock / NoLivenessSources);
GcAborted retired from the error enum; direct-delete refusal is
the GcRefuse error covering the full protection set
- Pin token gains its wire shape: blobs/put response {token, digest};
blobs/have's token-renewal form (one digest + token); token
validity domain = the minting serving node, process-lifetime
mapping
- Fleet mode is an explicit constructor declaration (fleet: true),
never inferred from engine choice
- All fleet GC state hosts on the fleet's kv engine (postgres) — one
arbitration domain; large=pg-lo fleet nodes required onto the same
pg instance (composite predicate's new clause); large=local fleet
puts pin-row-first, publish-second
- Engine-state seam reduced: sqlite pin/sweep-lock bodies dropped
(dead machinery); non-SQL engines stage delete-window candidates
in-process; one-window-host rule per instance
- Facade clarifications: fall-through for all key-addressed ops,
kv-only put-time rejection, mem+local dual-tier valid, error-model
member/return-shape ruling (trait/facade family split), has ->
bool, PinState variants, fleet liveness-table registration form,
window executor = the next sweep
Alignment edits across all specs and ADR-005/008/009/010/011
(bracketed corrections per the established pattern); OQ-11 (pg-only-kv
feature graph) added to the parked index for auditability.
Verification: two independent review passes; all findings resolved;
verdict READY for task decomposition.
19 KiB
status, last_updated
| status | last_updated |
|---|---|
| draft | 2026-10-03 (ADR-012 — composite predicate gains the pg-lo same-instance clause; kv-only rejection is put-time; GC state hosts on the kv engine; window wording) |
Backends, tiers, and dispatch
What this is
The physical storage layer: the Backend trait contract every engine
implements, the two shipped tiers (kv, large — the tier formerly named
fs; renamed in ADR-011), and the size-threshold dispatch that routes
between them. The trait is the crate's second stable seam (with the
key encoding, ADR-002): engines must survive digest-layer evolution,
so the boundary stays dumb.
Vocabulary is pinned in requirements.md (ADR-011): tier = contract
- routing class (exactly two; ADR-004), engine = concrete impl of one tier's contract (exactly five), and this document's table is the spec home of the constructor table (ADR-011 §4).
What the Backend trait is (and is not)
The trait is a CAS-tier contract: it accepts, serves, and enumerates
immutable content-addressed entries (hash → bytes), and its list()/
delete() shapes exist to feed GC (ADR-005). That is its whole job, and
the contract assumes it: complete-list-equals-sweep-safety only makes
sense for a store where every entry is a standalone immutable blob.
It follows that the trait is not a general-purpose storage abstraction,
and a downstream must never implement it to host data that is not
content-addressed blobs — a manifest table, a queryable index, or any
mutable/structured state. Such storage is consumer-owned, beside the
pool (a downstream's own sqlite/whatever, ADR-004): it needs no
digest-byte addressing, no list()-completeness, and no sweep
participation, so the CAS-tier contract fits it worse than it fits any
purpose-built option. The wrong turn to rule out is "my data isn't
blobs, therefore I implement a Backend for it"; nothing a consumer needs
is ever reachable through that door.
The Backend trait contract (ADR-003; I/O seams per ADR-010)
- Opaque byte keys, opaque byte-valued entries. Engines never learn
what a digest is; typed
Keyconverts at the store boundary (ADR-002). This is also the persistence boundary — a kv row, alocalfilename, or an LO companion row has no schema-migration story, so the byte layout must be self-describing (algorithm-tagged keys). - Methods:
has/get/put/delete/list/name/size(the length probe added by ADR-008).size(key)returns the entry's length orNoneif absent — metadata only, never a content fetch: kv engines probe the indexedsizecolumn both row shapes carry (POC #5 B5);localstats the file; pg-lo reads the companion row. Store-levelstatand pool accounting ride this method. - The I/O seams are crate-internal abstractions, not engine shapes
(ADR-010 §1 — the gap that would otherwise force guessing):
getreturns a read cursor the store core drives. kv-engine cursors materialize the whole value (bounded by tier policy); large-engine cursors stream —local's is a pread loop over the committed file, pg-lo's drives descriptorlesslo_get(oid, off, len)windows (ADR-009 §3, the C4 posture). Both serveread_range(slice + slice digest, ADR-006) at one seam. The cursor is not public API: store consumers see the facade's(len, stream)shapes with the cursor behind them.puthas two forms: the whole-value put (kv tier — the settled CAS row write) and the staged put (large tier — ADR-003's stage-then-commit-rename as a named per-engine contract: the engine hands out a staging cursor; commit publishes; every failure path converges on discard with zero residue). Both land in the CAS entry shape; both participate in pin-before-publish (ADR-005/010 §2 — the store core holds the joint entry+pin transaction via the engine-state seam).- Per-engine conformance: an engine is admitted only if its cursor/stage/list bodies pass the uniform contract suite — the POC #7 gate (10/10 exact-count sweep-outcome tests) is the norm; suite details in store-api.md's invariants.
list()complete by contract. A malformed list is as deadly as an incomplete one: list correctness is only observable through GC (POC #1 finding 2 — the redb key-vs-value trap produced a sweep that deleted the wrong blobs, silently). Therefore:- every engine implementation must prove list correctness through sweep-outcome tests (the invariant's test gate lives in store-api.md);
- an implementation that cannot enumerate (the rudolfs S3 anti-lesson) must not ship — GC is a structural requirement, not an optional extra (ADR-005).
- Virgin-store reads are no-ops: read paths treat a missing table as absent/empty (POC #1 finding 2).
- Namespace-blind: engines never see namespaces, tenants, or
reference structure (ADR-005 — the rudolfs inversion; physical
storage is
hash → bytesflat). - GC state never rides this trait (ADR-010 §2): pins, window rows,
and sweep locks are store-core-owned via a crate-internal
engine-state companion seam on the SQL-backed engines (sqlite,
postgres, pg-lo) — invisible at the public trait;
localandmemcannot express it and are never fleet engines. All fleet GC state (pin rows, window rows, the sweep lock) hosts on the fleet's kv-tier engine — postgres — regardless of the large tier's engine; pg-lo hosts no fleet state (ADR-012 §4). Per-engine clause applicability: ADR-012 §5's table (sqlite and pg-lo carry window-family clauses only; pin/sweep-lock clauses are the postgres host's).
Shipped tiers and engines
Two tiers, exactly — this counts trait implementations this crate ships as CAS tiers (the contracts dispatch routes between); it is the complete set the problem requires, and the "third backend" fear is a category error fixed in ADR-004 (tier count is not engine count). It says nothing about other storage existing in a deployment: a downstream runs whatever else it needs on its own media, beside the pool, above the store's seams (see "What the Backend trait is").
| Tier | Engine | Feature | Default | Fleet-valid |
|---|---|---|---|---|
| kv | sqlite | — (tier default) | yes (tier default-on) | no |
| kv | postgres | postgres |
no | yes |
| kv | mem | mem |
no | no (test/reference engine) |
| large | local | — (tier default) | yes (tier default-on) | shared media only, explicit assertion |
| large | pg-lo | pg-lo |
no | yes |
kv tier (tier default-on, feature kv); engines sqlite (default) + postgres (feature postgres, default-off) + mem (feature mem, default-off)
Small blobs — kv is the tier name; the engine is a per-node
constructor choice between the shipped engines (ADR-007; ADR-003
carries sqlite's on-disk format trade: file compat is upstream
sqlite's guarantee). The shipped trait impls write plain
content-addressed rows into the chosen engine, nothing else (ADR-010:
the fleet state rides the companion seam, never the public trait). A
third non-mem engine remains the substitution door as amended: a new
ADR with its own sweep-safety proof (the door is exercised — ADR-007
in-repo); and a downstream's other storage — schemas, manifests,
whatever it likes including some other sqlite file — never comes near
this tier at all, per "What the Backend trait is".
The collapse ADR-007 buys: the classical downstream stack (git server, vfs node) needed three storage systems — kv + large content + relational. Here the kv tier rides the relational engine itself, so a node provisions one relational engine + large-tier storage; consumer tables (refs, manifests, queues) sit beside the pool on the same engine or file by the deployment's own choice (co-tenancy note: same- file consumer writes share sqlite's single-writer ceiling — that sharing is the node's trade; separate files remain available).
Evidence (POC #3 finding A4, first-party measured): sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime — most git objects, workspace files, manifests), with the crossover at ~128-256 KiB where fs stops paying the B-tree row rewrite and wins.
sqlite engine (the default): WAL + synchronous=NORMAL as the
shipped durability tier (matching what the benchmark measured and what
iroh's store ships); bounded reads; has as an EXISTS probe; prepared
statements; read paths tolerate the no-tables-yet database (see
contract). Solo economics win the single-machine case by measurement;
the WAL single-writer ceiling (~1.2k objects/s under contention,
POC #5 B3) is an order above single-node push rates.
postgres engine (feature postgres, default-off; ADR-007): the
cross-machine-writers / many-client-node / already-running-pg engine.
Contract-identical row shape (key bytea PK / value bytea / size),
ON CONFLICT DO NOTHING CAS, complete list() as a stable cursor
over a vacuuming table, batch-delete via key = ANY($1). Impl
requirements per POC #5 finding B5: configurable synchronous_commit
posture (the parity knob with the sqlite shipped tier), pooled
connections with per-connection prepared-statement discipline, stated
autovacuum + fillfactor tuning, ~40× small-tier storage overhead
accepted (B6). Sweep-safety proof via the same exact-count
sweep-outcome CI gate, run against dockerized postgres under
--all-features. Concurrency evidence (POC #5 B3/B4): gets/puts scale
near-linearly to ~37k puts/s at 24 conns — the only engine whose
throughput increases under concurrency.
mem engine (feature mem, default-off; ADR-011 §2 — formerly
described as the "mem backend"/"testing tier" before the vocabulary
was pinned): a BTreeMap-shaped ephemeral implementation of the kv
tier's contract, including size (trivially). For tests and
in-process ephemerality — never a production story, never
fleet-valid, and the contract-reference engine: every other
engine's conformance tests mirror its suite (the role it already
played in POC #1's miniature).
large tier (tier default-on, feature large; the tier formerly named fs — renamed in ADR-011); engines local (default) + pg-lo (feature pg-lo, default-off)
Large blobs. Like the kv tier (ADR-007), the large tier is one contract, multiple engines — the medium is a per-node constructor choice (ADR-008; REQ-2 in requirements.md makes the engine concept load-bearing for fleets):
localengine (the default; the ADR-003 backend as originally specced): flat sharded layout:{hex-prefix}/{hex-prefix}/{hash}sharding survives from iroh's conclusion (limits directory size on huge pools); stage-then-commit- rename for the two-pass unknown-length path (the staged-put form, ADR-010); pread-based range reads (POC #3 finding A2 — local range serving is sound, e.g. for packfiles).pg-loengine (shipped, featurepg-lo, default-off; ADR-009): postgres Large Objects as the large tier's storage — the same engine-behind-one-trait move as ADR-007, one tier over, admitted on POC #7's measured evidence (contract 10/10 exact-count tests; durable put ≈ fs durable put at ≥1 MiB, beats it <1 MiB on the POC disk; gets ride descriptorlesslo_get(oid, off, len)windows over the companion table as contract authority; LO creation transactional — zero crash orphans). Named deltas: catalog space reused-but-never-returned (monitor rel size;VACUUM FULLis the operator's shrink path), autovacuum inherited, cached-get 20–50× behind page-cache fs single-stream (~700 MB/s aggregate at 16 readers — the fleet serving picture).- Fleet locality contract (ADR-008, extended by ADR-009/010): over
one shared pool, the large tier is one of: shared media (every
node's
localroot on the same fleet-shared media — soundness caveats and the required media guarantees in ADR-008), re-routed (one storage node serves pool large-blob content via the ops surface; client nodes are kv-only over pool content), orpg-lofor the whole fleet (ADR-009 — one engine, no shared media, no re-routing; pool content lives in the same pg instance the kv tier rides; the joint entry+pin tx is store-core-held, ADR-010 §2). Mixed per-node-local large tiers over one pool is the partitioning failure this contract exists to prevent — a documented deployment invariant the constructor cannot fully prove but must be declared against (ADR-008 names the enforceable seam and the detection symptom). The kv tier has the mirrored rule (ADR-010 §3): a fleet constructor requires a pool-shared kv engine (postgres); two sqlite engines are two pools, never one. - The composite fleet-validity rule (one sentence, both tiers): a
constructor configuration is fleet-valid iff kv = postgres AND
(large = pg-lo on the same pg instance as the kv engine OR
large =
local-on-declared-shared-media OR large = none with the re-routing posture carrying pool large content) — plus an explicitfleet: trueconstructor declaration (ADR-012 §3: fleet mode is never inferred). Any other combination over one shared pool is invalid — the constructor requires the declarations that make this predicate checkable per node; cross-node truth remains the deployment's verified invariant (ADR-008's seam). The pg-lo same-instance clause formalizes what ADR-009's consolidation posture already assumed (pool content lives in the pg instance the kv tier rides) — it is what makes the joint entry+pin tx expressible; cross-instance kv=postgres + large=pg-lo is valid only as a non-fleet configuration. This predicate is backends-and-dispatch.md's and ADR-011 §4's table combined; it is stated here once so no reader composes it by inference.
Size-threshold dispatch (ADR-003)
- Routing is a pure function of content length. Same content ⇒ same length ⇒ same tier; re-puts are deterministic. No content ever migrates between tiers — the migration question existed only to patch the unknown-length asymmetry, which ADR-003's pre-threshold buffering eliminates by construction.
- Default threshold: 128 KiB (constructor-tunable) — the bottom of the measured crossover zone (~128-256 KiB, POC #3 A4; chosen at the zone's conservative edge, not its midpoint). Re-tuning per deployment media is a constructor parameter, not an API change.
- Get fall-through: small-tier miss queries the large tier (deterministic, cheap — a stat probe).
- Per-namespace or per-tenant engine configuration: rejected (ADR-003 §Consequences — it would re-weld namespacing into the physical layer, the rudolfs anti-pattern ADR-005 inverts).
Constructor modes (ADR-011 §4)
- dual-tier (the default): kv + large, size-threshold dispatch
(ADR-003). Any engine pair the tables allow — including
kv = mem+large = local, the dispatch-coverage test shape (ADR-012 §6.3; both tiers are load-bearing there too). - kv-only (ADR-008's re-routing client posture, now named): no large tier; over-threshold puts are rejected at put time — known-length puts reject immediately, unknown-length puts reject at mid-stream threshold overflow (ADR-012 §6.2: lengths are not known at construction, so "constructor-time error" was strictly impossible). The re-routing posture handles over-threshold content via the ops surface — the consumer's composition, not a store mode.
- mem-only: the mem engine alone, no dispatch threshold in effect; a testing/embedder-ephemeral posture, never production.
- Single-tier SQL modes (postgres-only, pg-lo-only tiers) do not exist: both tiers are load-bearing (ADR-004). "Postgres-only" as in one SQL instance serving both tiers exists and is the ADR-009 consolidation: dual-tier mode with kv=postgres + large=pg-lo over one pool — under fleet mode, required to be literally the same instance (ADR-012 §4).
Where a new engine could come from
The trait is open to future implementations (network stores, S3-like
tiers), but nothing in the current consumer set requires one, and the
contract is deliberately hostile to half-implementations (complete
list(), GC-participating delete). Any future engine — per tier —
is a new ADR carrying its own sweep-safety story (ADR-007 is the
precedent for a kv engine addition; ADR-008 named pg-lo and ADR-009
admitted it with POC #7's evidence — the door is exercised twice).
This crate's roadmap is not blocked on one (see open-questions.md —
alkfs intake may name needs externally; OQ-08).
Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| 003 | Backend contract & dispatch | opaque keys, complete list, pure-function routing, no migration |
| 004 | Two tiers, no third | scope boundary against manifest-layer absorption |
| 005 | Namespace-blindness | engines see hashes only |
| 007 | Two kv engines | sqlite (default) + postgres behind one trait; one relational engine + large tier per node |
| 008 | size probe + fleet GC + large-tier engines |
trait length probe; DB-backed pins/advisory-locked sweeper for fleets; large tier engine-selectable (local default) |
| 009 | pg-lo admitted | postgres Large Objects as the large tier's second engine, on POC #7 |
| 010 | I/O seams + GC-state home | read cursor / staged put; store-core-owned GC state via the engine-state seam; the kv fleet-validity rule |
| 011 | Vocabulary | tier/engine/instance/node/fleet; mem is a kv engine; fs → large; constructor modes |
| 012 | Fleet activation & GC-state host | fleet: constructor declaration; all fleet GC state on the kv engine; pg-lo same-instance clause; kv-only put-time rejection |
Open Questions
- OQ-08: alkfs requirement intake may name storage requirements (e.g., durability tiers, sync-friendly layouts) that touch this layer — open, external owner (alkfs Phase 0).
References
docs/research/poc-trait-dispatch-findings.mdfindings 1/2/6docs/research/poc-largeblob-findings.mdfindings A2/A4 (+ the re-runnable benchmark harness)docs/research/poc-postgres-kv-findings.md— the pg arm's measured curves + engine-posture deltas (ADR-007's evidence base)docs/research/poc-redb-kv-findings.md— redb ruled out at a durability-tier mismatch (POC #6; recorded in ADR-007)docs/research/poc-pglo-findings.md— pg-lo's admission evidence (POC #7; ADR-009)docs/research/iroh-blobs-eval.md— the fs-layout conclusions borrowed (sharding, crash ordering, inline thresholds rejected as weld)- rudolfs notes —
list()anti-lesson, decorator alternative noted and not adopted (threshold dispatch chose the simpler policy; ADR-003 §Context) - ADR-003, ADR-004, ADR-005; ADR-012 (fleet activation, GC-state host, the composite predicate's same-instance clause); store-api.md (the invariants engines must satisfy); requirements.md (the pinned vocabulary)