- requirements.md (new): the three use cases pinned as REQ-1..4 with the node/pool/fleet/engine vocabulary defined once — ends per-session re-derivation of consumer facts. REQ-2 (replicator fleet over one shared pg pool, incl. large-blob serving) is recorded as a planning fact predating all POCs, on the operator's authority. - ADR-008: Backend trait gains size(key) length probe (before any backend ships data — ADR-002 one-way-door discipline); fleet GC mechanism (DB-backed pin rows committed atomically with entries, TTL+renewal semantics, liveness = embedder-owned table, protect callback is single-node-only, advisory-locked single sweeper, staged re-arbitrated delete window on both engines); fs tier becomes engine-selectable (local default; pg-lo named candidate) with the fleet locality contract (shared media or re-routing; mixed tiers are a documented deployment invariant, not a constructor-provable one). - poc-pglo-spec.md (new): POC #7 spec — pg Large Objects as the fs tier's pg-lo engine; instruments, decision gate, registered in phase-0.md OQ-BL-06. - ADR-003/005 status amendments point to ADR-008's extensions; specs ripple (backends/store-api/gc/ops/overview/README). - open-questions.md: deferral-policy header gains the decisions-vs- sequenced-work distinction; pg-lo's why-not-parked audit trail recorded. - research fixes: postgres POC renumbered #4->#5 to the canonical register (phase-0 OQ-BL-06), redb cross-refs fixed, thinking- artifact sentence in B1 replaced with the honest reading. Verification: docs-only change; reference-integrity sweep across the tree (ADR/REQ/POC refs resolve); architecture-reviewer pass on the delta — original 3 criticals addressed, its follow-up (fleet liveness form, pin TTL, staged-delete semantics, enforceability, shared-media caveats) fixed in this commit.
7.5 KiB
ADR-003: Backend contract, shipped backends, and dispatch policy
Status
Accepted (engine pin amended by ADR-007 — the kv tier ships two
engines, sqlite and postgres, behind this contract; the dispatch layer
and trait are unchanged. The trait was extended by ADR-008 with a
size length probe — before any backend ships data — and the fs
tier became engine-selectable (local default; pg-lo candidate), so
this ADR's "engines"/"backends" counts are read through ADR-004's
tier≠engine scope note)
Context
OQ-BL-02 settled the base contract in Phase 0 (lean Backend trait,
opaque keys, complete list(), size-threshold dual dispatch) with
named residuals: migration-between-backends policy, per-namespace
backend config, and the unknown-length put asymmetry (POC #3 finding
A6: an unknown-length re-put of a small blob lands on fs, duplicating
content the same digest's known-length put would put in kv).
One framing correction also lands here: the "two backends is ugly" / "three backends" framing that influenced early rounds — corrected to its real content as ADR-004 (two backends are required, and the manifest layer is downstream, not a backend). This ADR records the physical layer; ADR-004 records the boundary.
Evidence in hand (no further POC or consumer-waiting needed):
- POC #1 findings 1/2/6 (trait seam, list traps, dispatch thinness),
- POC #3 findings A1/A4/A5/A6 (put shapes, benchmark, stage hygiene, the asymmetry),
- rudolfs anti-lesson (S3 backend with no list can never GC).
Decision
Backend trait contract
- Opaque byte keys (
&[u8]), opaque byte values. TypedKeyconverts at the store boundary only (ADR-002). Backends must survive digest-layer evolution — a kv row or fs filename has no migration story, so opaque, self-describing bytes are the durable choice. (The generic-over-Kalternative was considered and rejected in POC #1 finding 1: with one key type, generics buy nothing and break the persistence boundary.) - Methods:
has/get/put/delete/list/name/size(the length probe added by ADR-008 before any backend shipped data). Streaming and range shapes ride the store layer above (fsgetyields a file handle, not a loaded buffer). list()complete by contract — and well-formed: a malformed list (wrong destructure, values-as-keys) is as deadly as an incomplete one, and silent until a sweep deletes the wrong things (POC #1 finding 2). Consequence: backend implementations are accepted only with sweep-outcome tests proving list correctness.- Virgin-store reads are no-ops (missing table = absent/empty).
- Namespace-blind (ADR-005): backends see hashes only.
Shipped backends (two, default-on; plus a testing tier)
kv(default-on): sqlite for small blobs — measured ~9-10× faster than fs at 1-16 KiB (POC #3 A4, first-party benchmark), the regime holding most git objects, workspace files, and manifests. WAL +synchronous=NORMALas shipped durability; prepared statements; no-tables-yet tolerance. (The tier's engine set was widened to include postgres by ADR-007.)fs(default-on) for large blobs — flat{hex-prefix}/{hex-prefix}/{hash}sharding (iroh's surviving conclusion); stage-then-commit-rename; pread range reads (sound locally per ADR-006).mem(default-off): ephemeralBTreeMaptier for tests and in-process ephemerality. Never a production story.
Both production backends are default-on (batteries-included, the
alkgit feature-model precedent); mem is not.
Dispatch policy
- Size-threshold routing; default 128 KiB, constructor-tunable (midpoint of the measured ~128-256 KiB crossover zone; POC #3 A4).
- Unknown-length puts buffer in memory up to the threshold; the
buffer's fate has three pinned semantics (finding A6's option (d)):
- Threshold exceeded mid-stream (non-restartable sources are normal — network streams can't be replayed): the buffered prefix flushes into a fresh stage file, and the stream continues appending into it. There is no "restart the stream" requirement anywhere.
- Verification runs over the eventual stored artifact as a whole: the memory buffer for the sub-threshold kv commit; the complete staged bytes for the fs commit. "Two passes over staged bytes" and "hash-check against the canonical derivation" are both literally true under this — there is no stitched-input case left ambiguous (the buffer is in the stage file by the time the second pass runs).
- Stream error before commit: no entry lands and no stage file remains (the stage-hygiene invariant extends to the buffered case: buffer discard on the failure path, exactly like stage discard). This is finding A6's option (d), decided over the alternatives:
- (a) tier-correction sweeps are moving machinery that re-creates migration, and are unnecessary if routing never misplaces a blob;
- (b) rejecting unknown-length small puts is hostile to streaming producers for no safety gain (dedup still holds under (d));
- (c) accepting duplicated content across tiers makes
list()lie about the pool (two entries for one digest) and complicates GC's live-set diffs for no benefit. Option (d) makes routing a pure function of content length — the property that eliminates the asymmetry and the migration question with it. Batch-scope pinning (POC #1 finding 3) is orthogonal — it protects in-flight puts, it does not route them — and ships anyway (ADR-005).
- No migration between backends. Content is immutable; its length never changes; routing is pure. The migration question existed only to patch the asymmetry and is resolved by elimination: there is nothing to migrate. A future backend needing migration writes a new ADR.
- No per-namespace backend configuration. It would re-weld namespacing into the physical layer — the exact rudolfs anti-pattern ADR-005 inverts. Tenancy differences that ever require physical separation (a different pool per tenant) are a deployment/topology concern (two stores), not a backend-config concern.
Consequences
Positive
- The dispatch is deterministic and test-invariant: same content ⇒ same backend, always — including across mixed put paths.
- GC's enumerate-everything substrate (
list()unions) is contract- enforced upfront rather than discovered against the rudolfs failure. - The shipped default matches the measured economics on the git/ workspace regime, re-tunable per deployment without API change.
Negative
- Unknown-length puts of large content pay two passes (forced by the preamble — ADR-002 §Consequences) plus staging I/O; bounded-memory buffering adds a copy for sub-threshold unknown-length puts (cheap, per the benchmark).
- sqlite's on-disk format enters the crate's surface durability story (file compatibility across rusqlite/sqlite versions is upstream sqlite's guarantee, not ours).
Neutral
- A network-backed backend (S3-like) remains possible via the same trait but ships nothing here; its list/GC obligations are the acceptance gate.
References
docs/research/poc-trait-dispatch-findings.mdfindings 1/2/6docs/research/poc-largeblob-findings.mdfindings A1/A4/A5/A6docs/research/phase-0.mdOQ-BL-02- backends-and-dispatch.md; store-api.md
- ADR-002 (key bytes), ADR-004 (boundary vs manifest layer), ADR-005 (namespace-blindness, pins)