Files
alkblobs/docs/architecture/backends-and-dispatch.md
T
glm-5.3-flash 7b9d904b8a docs(architecture): ADR-012 — pre-decomposition consistency rulings; tree verified single-valued
Full-tree review before task decomposition found five composition
defects (mechanisms specced correctly in isolation, composition
unruled) and a set of caller-facing gaps. ADR-012 rules each:

- Sweeps are never errors: aborts are GcAbortCause report data in
  Ok(SweepReport) (ProtectFailed / SweeperLock / NoLivenessSources);
  GcAborted retired from the error enum; direct-delete refusal is
  the GcRefuse error covering the full protection set
- Pin token gains its wire shape: blobs/put response {token, digest};
  blobs/have's token-renewal form (one digest + token); token
  validity domain = the minting serving node, process-lifetime
  mapping
- Fleet mode is an explicit constructor declaration (fleet: true),
  never inferred from engine choice
- All fleet GC state hosts on the fleet's kv engine (postgres) — one
  arbitration domain; large=pg-lo fleet nodes required onto the same
  pg instance (composite predicate's new clause); large=local fleet
  puts pin-row-first, publish-second
- Engine-state seam reduced: sqlite pin/sweep-lock bodies dropped
  (dead machinery); non-SQL engines stage delete-window candidates
  in-process; one-window-host rule per instance
- Facade clarifications: fall-through for all key-addressed ops,
  kv-only put-time rejection, mem+local dual-tier valid, error-model
  member/return-shape ruling (trait/facade family split), has ->
  bool, PinState variants, fleet liveness-table registration form,
  window executor = the next sweep

Alignment edits across all specs and ADR-005/008/009/010/011
(bracketed corrections per the established pattern); OQ-11 (pg-only-kv
feature graph) added to the parked index for auditability.

Verification: two independent review passes; all findings resolved;
verdict READY for task decomposition.
2026-10-03 07:14:45 +00:00

19 KiB
Raw Blame History

status, last_updated
status last_updated
draft 2026-10-03 (ADR-012 — composite predicate gains the pg-lo same-instance clause; kv-only rejection is put-time; GC state hosts on the kv engine; window wording)

Backends, tiers, and dispatch

What this is

The physical storage layer: the Backend trait contract every engine implements, the two shipped tiers (kv, large — the tier formerly named fs; renamed in ADR-011), and the size-threshold dispatch that routes between them. The trait is the crate's second stable seam (with the key encoding, ADR-002): engines must survive digest-layer evolution, so the boundary stays dumb.

Vocabulary is pinned in requirements.md (ADR-011): tier = contract

  • routing class (exactly two; ADR-004), engine = concrete impl of one tier's contract (exactly five), and this document's table is the spec home of the constructor table (ADR-011 §4).

What the Backend trait is (and is not)

The trait is a CAS-tier contract: it accepts, serves, and enumerates immutable content-addressed entries (hash → bytes), and its list()/ delete() shapes exist to feed GC (ADR-005). That is its whole job, and the contract assumes it: complete-list-equals-sweep-safety only makes sense for a store where every entry is a standalone immutable blob.

It follows that the trait is not a general-purpose storage abstraction, and a downstream must never implement it to host data that is not content-addressed blobs — a manifest table, a queryable index, or any mutable/structured state. Such storage is consumer-owned, beside the pool (a downstream's own sqlite/whatever, ADR-004): it needs no digest-byte addressing, no list()-completeness, and no sweep participation, so the CAS-tier contract fits it worse than it fits any purpose-built option. The wrong turn to rule out is "my data isn't blobs, therefore I implement a Backend for it"; nothing a consumer needs is ever reachable through that door.

The Backend trait contract (ADR-003; I/O seams per ADR-010)

  • Opaque byte keys, opaque byte-valued entries. Engines never learn what a digest is; typed Key converts at the store boundary (ADR-002). This is also the persistence boundary — a kv row, a local filename, or an LO companion row has no schema-migration story, so the byte layout must be self-describing (algorithm-tagged keys).
  • Methods: has / get / put / delete / list / name / size (the length probe added by ADR-008). size(key) returns the entry's length or None if absent — metadata only, never a content fetch: kv engines probe the indexed size column both row shapes carry (POC #5 B5); local stats the file; pg-lo reads the companion row. Store-level stat and pool accounting ride this method.
  • The I/O seams are crate-internal abstractions, not engine shapes (ADR-010 §1 — the gap that would otherwise force guessing):
    • get returns a read cursor the store core drives. kv-engine cursors materialize the whole value (bounded by tier policy); large-engine cursors stream — local's is a pread loop over the committed file, pg-lo's drives descriptorless lo_get(oid, off, len) windows (ADR-009 §3, the C4 posture). Both serve read_range (slice + slice digest, ADR-006) at one seam. The cursor is not public API: store consumers see the facade's (len, stream) shapes with the cursor behind them.
    • put has two forms: the whole-value put (kv tier — the settled CAS row write) and the staged put (large tier — ADR-003's stage-then-commit-rename as a named per-engine contract: the engine hands out a staging cursor; commit publishes; every failure path converges on discard with zero residue). Both land in the CAS entry shape; both participate in pin-before-publish (ADR-005/010 §2 — the store core holds the joint entry+pin transaction via the engine-state seam).
    • Per-engine conformance: an engine is admitted only if its cursor/stage/list bodies pass the uniform contract suite — the POC #7 gate (10/10 exact-count sweep-outcome tests) is the norm; suite details in store-api.md's invariants.
  • list() complete by contract. A malformed list is as deadly as an incomplete one: list correctness is only observable through GC (POC #1 finding 2 — the redb key-vs-value trap produced a sweep that deleted the wrong blobs, silently). Therefore:
    • every engine implementation must prove list correctness through sweep-outcome tests (the invariant's test gate lives in store-api.md);
    • an implementation that cannot enumerate (the rudolfs S3 anti-lesson) must not ship — GC is a structural requirement, not an optional extra (ADR-005).
  • Virgin-store reads are no-ops: read paths treat a missing table as absent/empty (POC #1 finding 2).
  • Namespace-blind: engines never see namespaces, tenants, or reference structure (ADR-005 — the rudolfs inversion; physical storage is hash → bytes flat).
  • GC state never rides this trait (ADR-010 §2): pins, window rows, and sweep locks are store-core-owned via a crate-internal engine-state companion seam on the SQL-backed engines (sqlite, postgres, pg-lo) — invisible at the public trait; local and mem cannot express it and are never fleet engines. All fleet GC state (pin rows, window rows, the sweep lock) hosts on the fleet's kv-tier engine — postgres — regardless of the large tier's engine; pg-lo hosts no fleet state (ADR-012 §4). Per-engine clause applicability: ADR-012 §5's table (sqlite and pg-lo carry window-family clauses only; pin/sweep-lock clauses are the postgres host's).

Shipped tiers and engines

Two tiers, exactly — this counts trait implementations this crate ships as CAS tiers (the contracts dispatch routes between); it is the complete set the problem requires, and the "third backend" fear is a category error fixed in ADR-004 (tier count is not engine count). It says nothing about other storage existing in a deployment: a downstream runs whatever else it needs on its own media, beside the pool, above the store's seams (see "What the Backend trait is").

Tier Engine Feature Default Fleet-valid
kv sqlite — (tier default) yes (tier default-on) no
kv postgres postgres no yes
kv mem mem no no (test/reference engine)
large local — (tier default) yes (tier default-on) shared media only, explicit assertion
large pg-lo pg-lo no yes

kv tier (tier default-on, feature kv); engines sqlite (default) + postgres (feature postgres, default-off) + mem (feature mem, default-off)

Small blobs — kv is the tier name; the engine is a per-node constructor choice between the shipped engines (ADR-007; ADR-003 carries sqlite's on-disk format trade: file compat is upstream sqlite's guarantee). The shipped trait impls write plain content-addressed rows into the chosen engine, nothing else (ADR-010: the fleet state rides the companion seam, never the public trait). A third non-mem engine remains the substitution door as amended: a new ADR with its own sweep-safety proof (the door is exercised — ADR-007 in-repo); and a downstream's other storage — schemas, manifests, whatever it likes including some other sqlite file — never comes near this tier at all, per "What the Backend trait is".

The collapse ADR-007 buys: the classical downstream stack (git server, vfs node) needed three storage systems — kv + large content + relational. Here the kv tier rides the relational engine itself, so a node provisions one relational engine + large-tier storage; consumer tables (refs, manifests, queues) sit beside the pool on the same engine or file by the deployment's own choice (co-tenancy note: same- file consumer writes share sqlite's single-writer ceiling — that sharing is the node's trade; separate files remain available).

Evidence (POC #3 finding A4, first-party measured): sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime — most git objects, workspace files, manifests), with the crossover at ~128-256 KiB where fs stops paying the B-tree row rewrite and wins.

sqlite engine (the default): WAL + synchronous=NORMAL as the shipped durability tier (matching what the benchmark measured and what iroh's store ships); bounded reads; has as an EXISTS probe; prepared statements; read paths tolerate the no-tables-yet database (see contract). Solo economics win the single-machine case by measurement; the WAL single-writer ceiling (~1.2k objects/s under contention, POC #5 B3) is an order above single-node push rates.

postgres engine (feature postgres, default-off; ADR-007): the cross-machine-writers / many-client-node / already-running-pg engine. Contract-identical row shape (key bytea PK / value bytea / size), ON CONFLICT DO NOTHING CAS, complete list() as a stable cursor over a vacuuming table, batch-delete via key = ANY($1). Impl requirements per POC #5 finding B5: configurable synchronous_commit posture (the parity knob with the sqlite shipped tier), pooled connections with per-connection prepared-statement discipline, stated autovacuum + fillfactor tuning, ~40× small-tier storage overhead accepted (B6). Sweep-safety proof via the same exact-count sweep-outcome CI gate, run against dockerized postgres under --all-features. Concurrency evidence (POC #5 B3/B4): gets/puts scale near-linearly to ~37k puts/s at 24 conns — the only engine whose throughput increases under concurrency.

mem engine (feature mem, default-off; ADR-011 §2 — formerly described as the "mem backend"/"testing tier" before the vocabulary was pinned): a BTreeMap-shaped ephemeral implementation of the kv tier's contract, including size (trivially). For tests and in-process ephemerality — never a production story, never fleet-valid, and the contract-reference engine: every other engine's conformance tests mirror its suite (the role it already played in POC #1's miniature).

large tier (tier default-on, feature large; the tier formerly named fs — renamed in ADR-011); engines local (default) + pg-lo (feature pg-lo, default-off)

Large blobs. Like the kv tier (ADR-007), the large tier is one contract, multiple engines — the medium is a per-node constructor choice (ADR-008; REQ-2 in requirements.md makes the engine concept load-bearing for fleets):

  • local engine (the default; the ADR-003 backend as originally specced): flat sharded layout: {hex-prefix}/{hex-prefix}/{hash} sharding survives from iroh's conclusion (limits directory size on huge pools); stage-then-commit- rename for the two-pass unknown-length path (the staged-put form, ADR-010); pread-based range reads (POC #3 finding A2 — local range serving is sound, e.g. for packfiles).
  • pg-lo engine (shipped, feature pg-lo, default-off; ADR-009): postgres Large Objects as the large tier's storage — the same engine-behind-one-trait move as ADR-007, one tier over, admitted on POC #7's measured evidence (contract 10/10 exact-count tests; durable put ≈ fs durable put at ≥1 MiB, beats it <1 MiB on the POC disk; gets ride descriptorless lo_get(oid, off, len) windows over the companion table as contract authority; LO creation transactional — zero crash orphans). Named deltas: catalog space reused-but-never-returned (monitor rel size; VACUUM FULL is the operator's shrink path), autovacuum inherited, cached-get 20–50× behind page-cache fs single-stream (~700 MB/s aggregate at 16 readers — the fleet serving picture).
  • Fleet locality contract (ADR-008, extended by ADR-009/010): over one shared pool, the large tier is one of: shared media (every node's local root on the same fleet-shared media — soundness caveats and the required media guarantees in ADR-008), re-routed (one storage node serves pool large-blob content via the ops surface; client nodes are kv-only over pool content), or pg-lo for the whole fleet (ADR-009 — one engine, no shared media, no re-routing; pool content lives in the same pg instance the kv tier rides; the joint entry+pin tx is store-core-held, ADR-010 §2). Mixed per-node-local large tiers over one pool is the partitioning failure this contract exists to prevent — a documented deployment invariant the constructor cannot fully prove but must be declared against (ADR-008 names the enforceable seam and the detection symptom). The kv tier has the mirrored rule (ADR-010 §3): a fleet constructor requires a pool-shared kv engine (postgres); two sqlite engines are two pools, never one.
  • The composite fleet-validity rule (one sentence, both tiers): a constructor configuration is fleet-valid iff kv = postgres AND (large = pg-lo on the same pg instance as the kv engine OR large = local-on-declared-shared-media OR large = none with the re-routing posture carrying pool large content) — plus an explicit fleet: true constructor declaration (ADR-012 §3: fleet mode is never inferred). Any other combination over one shared pool is invalid — the constructor requires the declarations that make this predicate checkable per node; cross-node truth remains the deployment's verified invariant (ADR-008's seam). The pg-lo same-instance clause formalizes what ADR-009's consolidation posture already assumed (pool content lives in the pg instance the kv tier rides) — it is what makes the joint entry+pin tx expressible; cross-instance kv=postgres + large=pg-lo is valid only as a non-fleet configuration. This predicate is backends-and-dispatch.md's and ADR-011 §4's table combined; it is stated here once so no reader composes it by inference.

Size-threshold dispatch (ADR-003)

  • Routing is a pure function of content length. Same content ⇒ same length ⇒ same tier; re-puts are deterministic. No content ever migrates between tiers — the migration question existed only to patch the unknown-length asymmetry, which ADR-003's pre-threshold buffering eliminates by construction.
  • Default threshold: 128 KiB (constructor-tunable) — the bottom of the measured crossover zone (~128-256 KiB, POC #3 A4; chosen at the zone's conservative edge, not its midpoint). Re-tuning per deployment media is a constructor parameter, not an API change.
  • Get fall-through: small-tier miss queries the large tier (deterministic, cheap — a stat probe).
  • Per-namespace or per-tenant engine configuration: rejected (ADR-003 §Consequences — it would re-weld namespacing into the physical layer, the rudolfs anti-pattern ADR-005 inverts).

Constructor modes (ADR-011 §4)

  • dual-tier (the default): kv + large, size-threshold dispatch (ADR-003). Any engine pair the tables allow — including kv = mem + large = local, the dispatch-coverage test shape (ADR-012 §6.3; both tiers are load-bearing there too).
  • kv-only (ADR-008's re-routing client posture, now named): no large tier; over-threshold puts are rejected at put time — known-length puts reject immediately, unknown-length puts reject at mid-stream threshold overflow (ADR-012 §6.2: lengths are not known at construction, so "constructor-time error" was strictly impossible). The re-routing posture handles over-threshold content via the ops surface — the consumer's composition, not a store mode.
  • mem-only: the mem engine alone, no dispatch threshold in effect; a testing/embedder-ephemeral posture, never production.
  • Single-tier SQL modes (postgres-only, pg-lo-only tiers) do not exist: both tiers are load-bearing (ADR-004). "Postgres-only" as in one SQL instance serving both tiers exists and is the ADR-009 consolidation: dual-tier mode with kv=postgres + large=pg-lo over one pool — under fleet mode, required to be literally the same instance (ADR-012 §4).

Where a new engine could come from

The trait is open to future implementations (network stores, S3-like tiers), but nothing in the current consumer set requires one, and the contract is deliberately hostile to half-implementations (complete list(), GC-participating delete). Any future engine — per tier — is a new ADR carrying its own sweep-safety story (ADR-007 is the precedent for a kv engine addition; ADR-008 named pg-lo and ADR-009 admitted it with POC #7's evidence — the door is exercised twice). This crate's roadmap is not blocked on one (see open-questions.md — alkfs intake may name needs externally; OQ-08).

Design Decisions

ADR Decision Summary
003 Backend contract & dispatch opaque keys, complete list, pure-function routing, no migration
004 Two tiers, no third scope boundary against manifest-layer absorption
005 Namespace-blindness engines see hashes only
007 Two kv engines sqlite (default) + postgres behind one trait; one relational engine + large tier per node
008 size probe + fleet GC + large-tier engines trait length probe; DB-backed pins/advisory-locked sweeper for fleets; large tier engine-selectable (local default)
009 pg-lo admitted postgres Large Objects as the large tier's second engine, on POC #7
010 I/O seams + GC-state home read cursor / staged put; store-core-owned GC state via the engine-state seam; the kv fleet-validity rule
011 Vocabulary tier/engine/instance/node/fleet; mem is a kv engine; fs → large; constructor modes
012 Fleet activation & GC-state host fleet: constructor declaration; all fleet GC state on the kv engine; pg-lo same-instance clause; kv-only put-time rejection

Open Questions

  • OQ-08: alkfs requirement intake may name storage requirements (e.g., durability tiers, sync-friendly layouts) that touch this layer — open, external owner (alkfs Phase 0).

References

  • docs/research/poc-trait-dispatch-findings.md findings 1/2/6
  • docs/research/poc-largeblob-findings.md findings A2/A4 (+ the re-runnable benchmark harness)
  • docs/research/poc-postgres-kv-findings.md — the pg arm's measured curves + engine-posture deltas (ADR-007's evidence base)
  • docs/research/poc-redb-kv-findings.md — redb ruled out at a durability-tier mismatch (POC #6; recorded in ADR-007)
  • docs/research/poc-pglo-findings.md — pg-lo's admission evidence (POC #7; ADR-009)
  • docs/research/iroh-blobs-eval.md — the fs-layout conclusions borrowed (sharding, crash ordering, inline thresholds rejected as weld)
  • rudolfs notes — list() anti-lesson, decorator alternative noted and not adopted (threshold dispatch chose the simpler policy; ADR-003 §Context)
  • ADR-003, ADR-004, ADR-005; ADR-012 (fleet activation, GC-state host, the composite predicate's same-instance clause); store-api.md (the invariants engines must satisfy); requirements.md (the pinned vocabulary)