Files
alkblobs/docs/architecture/decisions/003-backend-contract-and-dispatch.md
glm-5.3-flash bee86b0dbd docs(architecture): ADR-013 — family composition, reactive boundary, sweep lease; OQ-11 resolved
- No sibling (alkstore) dependency in-core; reactivity is
  deployment-layer composition above the facade — cadence stays
  embedder-owned (ADR-005), with alkstore the named substrate for it;
  in-crate notify/scheduler/queues/invalidation hooks are non-goals.
- Fleet sweeper lock pinned as 'named exclusive sweep lease'
  semantics; shipped provider unchanged (pg advisory clause);
  provider swaps are ADR-007-shaped substitution doors.
- Composition contract for downstream ADRs: co-tenancy of CAS/GC
  state + reactive machinery + consumer/liveness tables on one fleet
  host; no-cross-store-atomicity ruled; compensation rests on
  idempotent re-puts, windowed deletes, sweeps-as-report-data.
- OQ-11 resolved (the register's one live circular hedge): engine
  features pull their own driver clusters — sqlite becomes feature
  'sqlite' (default-on, carries rusqlite); 'kv'/'large' tier features
  retired; 'local' ungated; pg-pure builds are supported,
  non-default. Deciding fact already in-tree (REQ-2/ADR-009).
- requirements.md: UC-4/REQ-5 records the composition facts
  (operator record 2026-10-11, corroborated by alkstore's
  consumer-inventory).
- Publish guard stated pre-code: facade ships before publish;
  semver-is-contract from first publish (alkstore ADR-017 discipline).
- Amended accepted ADRs' feature wording (ADR-010 §1, ADR-011 §3/§4,
  ADR-003) with supersession notes; README/overview/open-questions/
  backends/store-api/gc aligned; parked index now empty.

Verification: docs-only round (no Rust surface touched); full-tree
single-value check on the feature graph and sweep-lease wording.
2026-10-11 16:47:22 +00:00

8.4 KiB
Raw Permalink Blame History

ADR-003: Backend contract, shipped backends, and dispatch policy

Status

Accepted (engine pin amended by ADR-007 — the kv tier ships two engines, sqlite and postgres, behind this contract; the dispatch layer and trait are unchanged. The trait was extended by ADR-008 with a size length probe — before any backend ships data — and the second tier became engine-selectable (local default; pg-lo admitted by ADR-009), so this ADR's "engines"/"backends" counts are read through ADR-004's tier≠engine scope note. The I/O shapes behind get/put are pinned by ADR-010 (read cursor / staged put). The tier's name changed: fs → large (ADR-011 — its routing criterion is length, not medium); this ADR's per-engine counts also gained the mem kv engine as ADR-011's vocabulary resolves — mem is a kv engine, not a tier)

Context

OQ-BL-02 settled the base contract in Phase 0 (lean Backend trait, opaque keys, complete list(), size-threshold dual dispatch) with named residuals: migration-between-backends policy, per-namespace backend config, and the unknown-length put asymmetry (POC #3 finding A6: an unknown-length re-put of a small blob lands on fs, duplicating content the same digest's known-length put would put in kv).

One framing correction also lands here: the "two backends is ugly" / "three backends" framing that influenced early rounds — corrected to its real content as ADR-004 (two backends are required, and the manifest layer is downstream, not a backend). This ADR records the physical layer; ADR-004 records the boundary.

Evidence in hand (no further POC or consumer-waiting needed):

  • POC #1 findings 1/2/6 (trait seam, list traps, dispatch thinness),
  • POC #3 findings A1/A4/A5/A6 (put shapes, benchmark, stage hygiene, the asymmetry),
  • rudolfs anti-lesson (S3 backend with no list can never GC).

Decision

Backend trait contract

  • Opaque byte keys (&[u8]), opaque byte values. Typed Key converts at the store boundary only (ADR-002). Backends must survive digest-layer evolution — a kv row or fs filename has no migration story, so opaque, self-describing bytes are the durable choice. (The generic-over-K alternative was considered and rejected in POC #1 finding 1: with one key type, generics buy nothing and break the persistence boundary.)
  • Methods: has/get/put/delete/list/name/size (the length probe added by ADR-008 before any backend shipped data). Streaming and range shapes ride the store layer above (fs get yields a file handle, not a loaded buffer).
  • list() complete by contract — and well-formed: a malformed list (wrong destructure, values-as-keys) is as deadly as an incomplete one, and silent until a sweep deletes the wrong things (POC #1 finding 2). Consequence: backend implementations are accepted only with sweep-outcome tests proving list correctness.
  • Virgin-store reads are no-ops (missing table = absent/empty).
  • Namespace-blind (ADR-005): backends see hashes only.

Shipped tiers (two, default-on; the kv tier also carries the mem engine)

  • kv (default-on): sqlite for small blobs — measured ~9-10× faster than fs at 1-16 KiB (POC #3 A4, first-party benchmark), the regime holding most git objects, workspace files, and manifests. WAL + synchronous=NORMAL as shipped durability; prepared statements; no-tables-yet tolerance. (The tier's engine set was widened to include postgres by ADR-007, and mem — the ephemeral reference engine — by ADR-011's vocabulary resolution.)
  • fs (default-on; renamed large by ADR-011) for large blobs — flat {hex-prefix}/{hex-prefix}/{hash} sharding (iroh's surviving conclusion); stage-then-commit-rename; pread range reads (sound locally per ADR-006). (The tier's engine set was widened by ADR-008/ 009.)
  • mem (default-off): ephemeral BTreeMap kv engine for tests and in-process ephemerality — an engine of the kv tier, not a third tier (ADR-011 §2; the pre-ADR-011 "testing tier" phrasing is superseded). Never a production story; the contract-reference engine.

Both production tiers are default-on (batteries-included, the alkgit feature-model precedent); mem is not. (Feature shape superseded by ADR-013 §4: "default-on" now reads as constructor-presence default — the sqlite and local engines — not tier features; engine features pull their own drivers.)

Dispatch policy

  • Size-threshold routing; default 128 KiB, constructor-tunable (chosen at the bottom of the measured ~128-256 KiB crossover zone — the conservative edge of the flat zone; POC #3 A4).
  • Unknown-length puts buffer in memory up to the threshold; the buffer's fate has three pinned semantics (finding A6's option (d)):
    1. Threshold exceeded mid-stream (non-restartable sources are normal — network streams can't be replayed): the buffered prefix flushes into a fresh stage file, and the stream continues appending into it. There is no "restart the stream" requirement anywhere.
    2. Verification runs over the eventual stored artifact as a whole: the memory buffer for the sub-threshold kv commit; the complete staged bytes for the fs commit. "Two passes over staged bytes" and "hash-check against the canonical derivation" are both literally true under this — there is no stitched-input case left ambiguous (the buffer is in the stage file by the time the second pass runs).
    3. Stream error before commit: no entry lands and no stage file remains (the stage-hygiene invariant extends to the buffered case: buffer discard on the failure path, exactly like stage discard). This is finding A6's option (d), decided over the alternatives:
    • (a) tier-correction sweeps are moving machinery that re-creates migration, and are unnecessary if routing never misplaces a blob;
    • (b) rejecting unknown-length small puts is hostile to streaming producers for no safety gain (dedup still holds under (d));
    • (c) accepting duplicated content across tiers makes list() lie about the pool (two entries for one digest) and complicates GC's live-set diffs for no benefit. Option (d) makes routing a pure function of content length — the property that eliminates the asymmetry and the migration question with it. Batch-scope pinning (POC #1 finding 3) is orthogonal — it protects in-flight puts, it does not route them — and ships anyway (ADR-005).
  • No migration between backends. Content is immutable; its length never changes; routing is pure. The migration question existed only to patch the asymmetry and is resolved by elimination: there is nothing to migrate. A future backend needing migration writes a new ADR.
  • No per-namespace backend configuration. It would re-weld namespacing into the physical layer — the exact rudolfs anti-pattern ADR-005 inverts. Tenancy differences that ever require physical separation (a different pool per tenant) are a deployment/topology concern (two stores), not a backend-config concern.

Consequences

Positive

  • The dispatch is deterministic and test-invariant: same content ⇒ same backend, always — including across mixed put paths.
  • GC's enumerate-everything substrate (list() unions) is contract- enforced upfront rather than discovered against the rudolfs failure.
  • The shipped default matches the measured economics on the git/ workspace regime, re-tunable per deployment without API change.

Negative

  • Unknown-length puts of large content pay two passes (forced by the preamble — ADR-002 §Consequences) plus staging I/O; bounded-memory buffering adds a copy for sub-threshold unknown-length puts (cheap, per the benchmark).
  • sqlite's on-disk format enters the crate's surface durability story (file compatibility across rusqlite/sqlite versions is upstream sqlite's guarantee, not ours).

Neutral

  • A network-backed engine (S3-like) remains possible via the same trait but ships nothing here; its list/GC obligations are the acceptance gate; per-tier, via a new ADR (ADR-011's vocabulary: an engine addition, not a tier addition).

References

  • docs/research/poc-trait-dispatch-findings.md findings 1/2/6
  • docs/research/poc-largeblob-findings.md findings A1/A4/A5/A6
  • docs/research/phase-0.md OQ-BL-02
  • backends-and-dispatch.md; store-api.md
  • ADR-002 (key bytes), ADR-004 (boundary vs manifest layer), ADR-005 (namespace-blindness, pins)