Files
alkblobs/docs/architecture/decisions/008-trait-size-probe-fleet-gc-and-large-engines.md
glm-5.3-flash 7b9d904b8a docs(architecture): ADR-012 — pre-decomposition consistency rulings; tree verified single-valued
Full-tree review before task decomposition found five composition
defects (mechanisms specced correctly in isolation, composition
unruled) and a set of caller-facing gaps. ADR-012 rules each:

- Sweeps are never errors: aborts are GcAbortCause report data in
  Ok(SweepReport) (ProtectFailed / SweeperLock / NoLivenessSources);
  GcAborted retired from the error enum; direct-delete refusal is
  the GcRefuse error covering the full protection set
- Pin token gains its wire shape: blobs/put response {token, digest};
  blobs/have's token-renewal form (one digest + token); token
  validity domain = the minting serving node, process-lifetime
  mapping
- Fleet mode is an explicit constructor declaration (fleet: true),
  never inferred from engine choice
- All fleet GC state hosts on the fleet's kv engine (postgres) — one
  arbitration domain; large=pg-lo fleet nodes required onto the same
  pg instance (composite predicate's new clause); large=local fleet
  puts pin-row-first, publish-second
- Engine-state seam reduced: sqlite pin/sweep-lock bodies dropped
  (dead machinery); non-SQL engines stage delete-window candidates
  in-process; one-window-host rule per instance
- Facade clarifications: fall-through for all key-addressed ops,
  kv-only put-time rejection, mem+local dual-tier valid, error-model
  member/return-shape ruling (trait/facade family split), has ->
  bool, PinState variants, fleet liveness-table registration form,
  window executor = the next sweep

Alignment edits across all specs and ADR-005/008/009/010/011
(bracketed corrections per the established pattern); OQ-11 (pg-only-kv
feature graph) added to the parked index for auditability.

Verification: two independent review passes; all findings resolved;
verdict READY for task decomposition.
2026-10-03 07:14:45 +00:00

16 KiB

ADR-008: Trait length probe, fleet GC, and large-tier engines

Status

Accepted (the §3 pg-lo bullet's "candidate, not shipped" posture was superseded 2026-10-03 by ADR-009, which admitted pg-lo on POC #7's passed gate — the two §Consequences passages that said pg-lo "remains unbuilt" are resolved by that same supersession; the tier this ADR's title and text call "fs" is renamed large by ADR-011; the fleet GC mechanism's state home and the joint-tx ownership are ruled by ADR-010; fleet-mode activation, the GC-state host, the sweep-abort return shape, and the window executor are ruled by ADR-012; every other part of this ADR stands as written)

Context

Three defects surfaced in the post-ADR-007 architecture review, all concentrated at the seams ADR-007 touched and all load-bearing for REQ-2 (a replicator node may run multiple store instances — a fleet — sharing one pool via postgres; requirements.md):

  1. The Backend trait has no length probe. The store surface promises stat(key) → Option<EntryMeta> ("cheap length/type probe"; store-api.md) and pool accounting rides list + stat (ADR-005), but the trait's methods are has/get/put/delete/list/ name with no value-metadata access. The kv engines' row shapes both carry a size column (POC #5 B5), and the large tier can stat a file — but with no trait method, kv stat degenerates to a full value load. The trait is the crate's second stable seam (ADR-003); discovering this during implementation would be a breaking trait change after backends exist.
  2. ADR-005's sweep-safety protocol is per-node by construction. Pins, the pin/liveness arbitration lock, and delete-window arbitration all live inside one process (ADR-005 §Decision). That is sound for sqlite and single-instance postgres (REQ-1, REQ-3 topologies) — but under a fleet (REQ-2), node A's sweep enumerates the shared pool and arbitrates against A's in-process pin map; node B's in-flight put is invisible. Data loss, structurally: the topology REQ-2 requires does not satisfy the protocol ADR-005 specifies.
  3. The large tier is node-local, silently partitioning a fleet. The large tier stores to a sharded layout on the node's own filesystem (ADR-003). Under a fleet sharing one kv pool, a >threshold blob put via node A exists only on A's disk; node B's get fall-through queries its own fs and misses, and list() unions disagree per node. REQ-2 makes this topology a shipping requirement, not a hypothetical.

These three travel together: the fleet topology forces #2 and #3, and #3's resolution constrains what the large tier must support, while #1 is a trait-level gap any topology hits.

Decision

1. size(key) → Option<u64> joins the Backend trait

The trait gains a length probe: size(key) returns the entry's byte length or None if absent. Required of all engines; semantics per tier:

  • kv engines: an indexed column probe (SELECT size …WHERE key), never a value load. The size column both POC row shapes carry becomes contract-relevant.
  • large engine (local): stat/metadata read, no content fetch.
  • The store facade's stat (store-api.md) and the accounting path (ADR-005) ride this method; the has method remains a pure existence probe.

size is a trait addition before any backend ships data — the one-way-door discipline of ADR-002/003 applied to the trait itself.

2. Fleet GC: the delete window is durable and re-arbitrated in the database; pins are DB-backed under fleet engines

The store facade's GC invariant is unchanged — a visible pool entry is never deleted while liveness protects it, and an in-flight put is never deleted by a sweep started before it committed (ADR-005). Under a fleet, the mechanism moves into the shared engine:

  • Pins are rows. put's pin-before-publish (ADR-005) becomes two inserts in one transaction on the fleet engine: the entry row and its pin row commit together (or neither does). The pin row carries an owner (a store-instance id — ADR-011's vocabulary; "node id" is this ADR's original wording, superseded because a node may host several instances) and expiry: the pin guard renews its row every TTL/3 (default TTL 60 s, constructor-tunable — fleet constructor parameters, not API). Protection = an un-expired pin row: at delete-time arbitration, only expiry > now() pins count; an expired pin is abandoned (its holder failed to renew — the same liveness-failure signal as an in-process guard's drop) and is reaped by the sweeper's maintenance step (the GC-of-GC cleanup). A node that pauses past TTL loses its pins; delete-then-recover (ADR-005) remains the correctness backstop — content is re-put byte-identically under the same key.
  • Registered liveness sources have one fleet form: a liveness table. A fleet embedder registers its liveness as a named table of (key, …) rows in the shared engine (schema pinned by the embedder, shape contractually (key bytea, …) — any columns it needs beyond the key). In-process callback views are not a fleet liveness form: delete-time arbitration queries/finger-locks tables per key, so a source must be a table for the SQL protocol to see it. Single-node embedders keep callback sources unchanged (ADR-005).
  • The protect callback is single-node-only. A fleet sweeper cannot consult another node's in-process callback, so under fleet engines the callback's fleet expression is: the embedder materializes its protect set into its liveness table (or pin rows) before invoking sweep — the callback's abort semantics remain for single-node/topology-local sweeps that cannot express their protection as rows.
  • One sweeper per pool. A fleet node acquires a pg advisory lock before sweeping and holds it for the whole sweep. Two nodes sweeping concurrently is not a state this crate tolerates — the lock is the admission gate, and a second-node sweep returns the typed GcAborted (no deletion) rather than racing. GcAborted thus carries two cause classes (protection-source failure — ADR-005; sweeper lock contention — this ADR); the typed error's fleet variant identifies which. (Return carrier corrected by ADR-012 §1: the abort is GcAbortCause::SweeperLock report data — Ok(report) — the typed sweep error is retired; the two cause classes and the no-deletion semantics stand.)
  • Delete-time arbitration is SQL-level. The delete window's re-check (ADR-005: "arbitrate under the pin/liveness lock at delete time") executes as SELECT … FOR UPDATE against pin rows and liveness tables at delete time, then batch-deletes. The staged re-arbitrated delete: during mark, candidate-dead keys are staged as rows carrying visible_after = mark_end + window (the delete window's delay — default 5 min, constructor-tunable per deployment); the executor (the sweeper, or the next sweep's owner) re-arbitrates each staged candidate at execution time against pins and liveness tables before deleting. This is the durable delete window on every SQL-backed engine — the window table is engine rows on sqlite as well; what differs by topology is only where the liveness state lives (in-process maps and callbacks for single nodes; tables and SQL arbitration for fleets), not whether staging exists. (ADR-012 §5/§6.7: "every SQL-backed engine" is the final wording — local/mem stage in-process, window-inline; the executor of overdue staged rows is the next sweep.) No external queue dependency is required (the pattern is engine-native transactions; consumer-side queues like honker or pg-boss-lineage tools remain consumer-layer options, per ADR-004's boundary).
  • Single-node postures are unchanged. sqlite and single-instance postgres keep ADR-005's in-process protocol verbatim: the pin map, RAII guards, and arbitration lock stay in-process; only the persistence of the delete window (the staged-candidate table) is shared machinery. The pin-row expiry/ttl machinery exists only where fleet engines are active. (Activation is explicit: ADR-012 §3's fleet: true constructor declaration turns the fleet machinery on; it is never inferred from engine choice alone — a single-instance postgres declared fleet: false keeps the in-process protocol.)

3. The large tier is engine-selectable; fleet topology dictates which engines are permissible

The large tier is a tier, not a medium. It ships engines the way the kv tier does (ADR-007), and a deployment's topology determines which are valid:

  • local (default): the sharded-layout filesystem backend, unchanged (ADR-003). Valid for any single-node topology (REQ-1, REQ-4) and for fleet nodes using shared media (below).
  • Fleet + shared media (deployment requirement): multi-node pools may keep the local engine by pointing every node's large tier at shared/replicated media. The soundness claim is qualified, not assumed: immutable-once-committed files plus stage-then-commit- rename are sound on media that provides (a) atomic rename within a directory, (b) close-to-open consistency (a written+closed file is complete to any later reader — NFSv3+ with the close-to-open discipline; NFSv4 and EFS-class stores qualify), and (c) no cross-node POSIX locking reliance (this store takes none — its cross-node coordination is the fleet engine's locks, never file locks; honker's NFS warning is about sqlite's lock protocol, not this shape). Verifying (a)-(c) on the deployment's actual media is the deployment's responsibility; a verification pass is recorded ops work, not crate code.
  • Fleet re-routing (topology alternative): a fleet whose nodes do not share media designates one storage node; large-blob writes and reads over the pool route through it via the ops surface (verified fetch/put; ADR-001). Client nodes' constructors are configured with kv-only mode (ADR-011 §4's name for what this bullet first called "kv-tier-only dispatch") — no local large tier over pool content at all; over-threshold puts reject at put time (ADR-012 §6.2), and the re-routing posture handles them via the ops surface; a client node's node-private, consumer-owned large-tier content is outside the pool and outside this crate's GC entirely: it is the consumer's own storage beside the pool per ADR-004, not a constructor mode over pool content, and the composite fleet-validity rule (backends-and-dispatch.md) therefore has no "local on node-private media over one pool" entry. The facade shape on a client node is its own store for sub-threshold content plus its registered ops-put path for large content — a consumer-side composition, not a store-level mode.
  • pg-lo (shipped; admitted by ADR-009): postgres Large Objects as the large tier's storage — the same "engine behind one contract" move as ADR-007, one tier over. Originally named here as a candidate gated on measured evidence, POC #7 ran and passed that gate (docs/research/poc-pglo-findings.md: companion-table authority, lo_get window gets, durable put ≈ fs's durable put, cached gets behind page-cache fs by a measured 20-50x single-stream, ~700 MB/s aggregate at 16 readers), and ADR-009 admitted the engine the same day — this bullet's original "candidate, not shipped / REQ-2's immediate answers are shared media or re-routing; pg-lo the consolidation option" posture is historical; the engine is shipped, feature pg-lo, and it is the fleet consolidation option as a first-class engine (ADR-012 §4 adds the same-instance requirement under fleet mode).
  • The no-mixed-large-tiers rule is a deployment invariant, enforced at the enforceable seam. Cross-node configuration cannot be validated by any one constructor (it sees only its own node). What the constructor can enforce: a fleet-mode constructor (fleet engines active) requires an explicit large-tier declaration — local (with the operator's assertion that its root is fleet-shared media), routed (kv-only mode, ADR-011 §4; no local filesystem over pool content), or pg-lo. The cross-node truth ("is node B's local actually the same media as node A's?") is a deployment invariant this crate documents and ops verifies, not something a constructor can prove; a partitioned tier's observable signature (get fall-through misses for entries another node wrote) is the documented detection symptom. The composite fleet-validity predicate over both tiers is pinned in backends-and-dispatch.md ("The composite fleet-validity rule").

In sum: a fleet's large tier over pool content is either shared media (all nodes' local roots on the same fleet-shared media) or routed (one storage node); mixing per-node-local media over one pool is the partitioning failure this decision exists to prevent.

Consequences

Positive

  • The trait now expressibly supports stat and pool accounting on every engine; the "cheap probe" claim of store-api.md is true rather than aspirational.
  • REQ-2 (fleet over one pool) has a complete, specced GC story: pins move to the shared engine atomically with put-commits, one sweeper holds the lock, deletes re-arbitrate in-transaction. The ADR-005 invariant extends to fleets as a stated mechanism, not an accident of single-process luck.
  • The large tier's fleet problem has named answers: shipped-code-free for shared media and re-routing, shipped-code for the consolidation option (pg-lo — admitted by ADR-009, the same POC-first door as ADR-007).

Negative

  • The Backend trait gains a method before any backend ships — the correct moment, but every engine's task list grows by one method.
  • Fleet GC is real state machinery: pin rows with expiry (GC-of-GC surface for orphaned pins at node crash — the expiry is the cleanup handle), advisory-lock coordination, and SQL-level arbitration added on top of the already-intricate delete-window state machine.
  • pg-lo was built under ADR-009 after POC #7 passed — the measured, deliberate sequence of the ADR-006 pattern (admission gate first, engine second), not a hedge; resolved same-day.
  • (ADR-012 §4: fleet GC state hosts on the fleet's kv engine — the large tier's local engine never joins the pin transaction; large=local fleet puts commit the pin row first, publish (commit-rename) second.)

Neutral

  • The sqlite and single-node-postgres postures are unchanged; this ADR costs existing paths nothing at runtime.
  • pg-lo, if admitted, rides the same driver stack ADR-007 already ships (tokio-postgres + deadpool) via SQL lo_* functions — no new driver dependency (the postgres_large_object crate is a dead 0.15-era io-trait glue; unnecessary; POC #7 confirms: ~400 lines of SQL-statement shapes over the existing stack).

References

  • requirements.md — REQ-2 (the fleet fact this ADR binds), REQ-1/3/4 (the topologies left unchanged)
  • POC #5 (poc-postgres-kv-findings.md B5 — row shapes with size columns; the driver stack pg-lo would ride)
  • ADR-002 (the one-way-door discipline this trait change applies), ADR-003 (the trait contract and admission gate size and pg-lo amend/enter), ADR-004 (tier≠engine, extended to the large tier), ADR-005 (the GC invariant and single-node protocol extended here, not replaced), ADR-006 (the deferred-cost pattern pg-lo's admission follows), ADR-007 (the engine-behind-one-trait precedent)
  • backends-and-dispatch.md (the trait and tier surface this amends); store-api.md (stat's contract); gc-and-namespaces.md (the fleet GC mechanism); ops-surface.md (the re-routing topology's transport); ADR-010 (the fleet mechanism's state home and the joint-tx ownership this ADR's §2 implied but did not rule); ADR-011 (the vocabulary — instance/fleet — this ADR's topology language resolves to; the large tier renamed large)