Files
alkblobs/docs/architecture/decisions/003-backend-contract-and-dispatch.md
T
glm-5.3-flash 281c37876e docs(architecture): ADR-008 — trait size probe, fleet GC, fs-tier engines; requirements anchor
- requirements.md (new): the three use cases pinned as REQ-1..4 with
  the node/pool/fleet/engine vocabulary defined once — ends per-session
  re-derivation of consumer facts. REQ-2 (replicator fleet over one
  shared pg pool, incl. large-blob serving) is recorded as a planning
  fact predating all POCs, on the operator's authority.
- ADR-008: Backend trait gains size(key) length probe (before any
  backend ships data — ADR-002 one-way-door discipline); fleet GC
  mechanism (DB-backed pin rows committed atomically with entries,
  TTL+renewal semantics, liveness = embedder-owned table, protect
  callback is single-node-only, advisory-locked single sweeper,
  staged re-arbitrated delete window on both engines); fs tier becomes
  engine-selectable (local default; pg-lo named candidate) with the
  fleet locality contract (shared media or re-routing; mixed tiers are
  a documented deployment invariant, not a constructor-provable one).
- poc-pglo-spec.md (new): POC #7 spec — pg Large Objects as the fs
  tier's pg-lo engine; instruments, decision gate, registered in
  phase-0.md OQ-BL-06.
- ADR-003/005 status amendments point to ADR-008's extensions; specs
  ripple (backends/store-api/gc/ops/overview/README).
- open-questions.md: deferral-policy header gains the decisions-vs-
  sequenced-work distinction; pg-lo's why-not-parked audit trail
  recorded.
- research fixes: postgres POC renumbered #4->#5 to the canonical
  register (phase-0 OQ-BL-06), redb cross-refs fixed, thinking-
  artifact sentence in B1 replaced with the honest reading.

Verification: docs-only change; reference-integrity sweep across the
tree (ADR/REQ/POC refs resolve); architecture-reviewer pass on the
delta — original 3 criticals addressed, its follow-up (fleet liveness
form, pin TTL, staged-delete semantics, enforceability, shared-media
caveats) fixed in this commit.
2026-10-03 03:41:49 +00:00

7.5 KiB
Raw Blame History

ADR-003: Backend contract, shipped backends, and dispatch policy

Status

Accepted (engine pin amended by ADR-007 — the kv tier ships two engines, sqlite and postgres, behind this contract; the dispatch layer and trait are unchanged. The trait was extended by ADR-008 with a size length probe — before any backend ships data — and the fs tier became engine-selectable (local default; pg-lo candidate), so this ADR's "engines"/"backends" counts are read through ADR-004's tier≠engine scope note)

Context

OQ-BL-02 settled the base contract in Phase 0 (lean Backend trait, opaque keys, complete list(), size-threshold dual dispatch) with named residuals: migration-between-backends policy, per-namespace backend config, and the unknown-length put asymmetry (POC #3 finding A6: an unknown-length re-put of a small blob lands on fs, duplicating content the same digest's known-length put would put in kv).

One framing correction also lands here: the "two backends is ugly" / "three backends" framing that influenced early rounds — corrected to its real content as ADR-004 (two backends are required, and the manifest layer is downstream, not a backend). This ADR records the physical layer; ADR-004 records the boundary.

Evidence in hand (no further POC or consumer-waiting needed):

  • POC #1 findings 1/2/6 (trait seam, list traps, dispatch thinness),
  • POC #3 findings A1/A4/A5/A6 (put shapes, benchmark, stage hygiene, the asymmetry),
  • rudolfs anti-lesson (S3 backend with no list can never GC).

Decision

Backend trait contract

  • Opaque byte keys (&[u8]), opaque byte values. Typed Key converts at the store boundary only (ADR-002). Backends must survive digest-layer evolution — a kv row or fs filename has no migration story, so opaque, self-describing bytes are the durable choice. (The generic-over-K alternative was considered and rejected in POC #1 finding 1: with one key type, generics buy nothing and break the persistence boundary.)
  • Methods: has/get/put/delete/list/name/size (the length probe added by ADR-008 before any backend shipped data). Streaming and range shapes ride the store layer above (fs get yields a file handle, not a loaded buffer).
  • list() complete by contract — and well-formed: a malformed list (wrong destructure, values-as-keys) is as deadly as an incomplete one, and silent until a sweep deletes the wrong things (POC #1 finding 2). Consequence: backend implementations are accepted only with sweep-outcome tests proving list correctness.
  • Virgin-store reads are no-ops (missing table = absent/empty).
  • Namespace-blind (ADR-005): backends see hashes only.

Shipped backends (two, default-on; plus a testing tier)

  • kv (default-on): sqlite for small blobs — measured ~9-10× faster than fs at 1-16 KiB (POC #3 A4, first-party benchmark), the regime holding most git objects, workspace files, and manifests. WAL + synchronous=NORMAL as shipped durability; prepared statements; no-tables-yet tolerance. (The tier's engine set was widened to include postgres by ADR-007.)
  • fs (default-on) for large blobs — flat {hex-prefix}/{hex-prefix}/{hash} sharding (iroh's surviving conclusion); stage-then-commit-rename; pread range reads (sound locally per ADR-006).
  • mem (default-off): ephemeral BTreeMap tier for tests and in-process ephemerality. Never a production story.

Both production backends are default-on (batteries-included, the alkgit feature-model precedent); mem is not.

Dispatch policy

  • Size-threshold routing; default 128 KiB, constructor-tunable (midpoint of the measured ~128-256 KiB crossover zone; POC #3 A4).
  • Unknown-length puts buffer in memory up to the threshold; the buffer's fate has three pinned semantics (finding A6's option (d)):
    1. Threshold exceeded mid-stream (non-restartable sources are normal — network streams can't be replayed): the buffered prefix flushes into a fresh stage file, and the stream continues appending into it. There is no "restart the stream" requirement anywhere.
    2. Verification runs over the eventual stored artifact as a whole: the memory buffer for the sub-threshold kv commit; the complete staged bytes for the fs commit. "Two passes over staged bytes" and "hash-check against the canonical derivation" are both literally true under this — there is no stitched-input case left ambiguous (the buffer is in the stage file by the time the second pass runs).
    3. Stream error before commit: no entry lands and no stage file remains (the stage-hygiene invariant extends to the buffered case: buffer discard on the failure path, exactly like stage discard). This is finding A6's option (d), decided over the alternatives:
    • (a) tier-correction sweeps are moving machinery that re-creates migration, and are unnecessary if routing never misplaces a blob;
    • (b) rejecting unknown-length small puts is hostile to streaming producers for no safety gain (dedup still holds under (d));
    • (c) accepting duplicated content across tiers makes list() lie about the pool (two entries for one digest) and complicates GC's live-set diffs for no benefit. Option (d) makes routing a pure function of content length — the property that eliminates the asymmetry and the migration question with it. Batch-scope pinning (POC #1 finding 3) is orthogonal — it protects in-flight puts, it does not route them — and ships anyway (ADR-005).
  • No migration between backends. Content is immutable; its length never changes; routing is pure. The migration question existed only to patch the asymmetry and is resolved by elimination: there is nothing to migrate. A future backend needing migration writes a new ADR.
  • No per-namespace backend configuration. It would re-weld namespacing into the physical layer — the exact rudolfs anti-pattern ADR-005 inverts. Tenancy differences that ever require physical separation (a different pool per tenant) are a deployment/topology concern (two stores), not a backend-config concern.

Consequences

Positive

  • The dispatch is deterministic and test-invariant: same content ⇒ same backend, always — including across mixed put paths.
  • GC's enumerate-everything substrate (list() unions) is contract- enforced upfront rather than discovered against the rudolfs failure.
  • The shipped default matches the measured economics on the git/ workspace regime, re-tunable per deployment without API change.

Negative

  • Unknown-length puts of large content pay two passes (forced by the preamble — ADR-002 §Consequences) plus staging I/O; bounded-memory buffering adds a copy for sub-threshold unknown-length puts (cheap, per the benchmark).
  • sqlite's on-disk format enters the crate's surface durability story (file compatibility across rusqlite/sqlite versions is upstream sqlite's guarantee, not ours).

Neutral

  • A network-backed backend (S3-like) remains possible via the same trait but ships nothing here; its list/GC obligations are the acceptance gate.

References

  • docs/research/poc-trait-dispatch-findings.md findings 1/2/6
  • docs/research/poc-largeblob-findings.md findings A1/A4/A5/A6
  • docs/research/phase-0.md OQ-BL-02
  • backends-and-dispatch.md; store-api.md
  • ADR-002 (key bytes), ADR-004 (boundary vs manifest layer), ADR-005 (namespace-blindness, pins)