# ADR-008: Trait length probe, fleet GC, and large-tier engines ## Status Accepted (the §3 pg-lo bullet's "candidate, not shipped" posture was superseded 2026-10-03 by ADR-009, which admitted pg-lo on POC #7's passed gate — the two §Consequences passages that said pg-lo "remains unbuilt" are resolved by that same supersession; the tier this ADR's title and text call "fs" is renamed `large` by ADR-011; the fleet GC mechanism's state home and the joint-tx ownership are ruled by ADR-010; fleet-mode activation, the GC-state host, the sweep-abort return shape, and the window executor are ruled by ADR-012; every other part of this ADR stands as written) ## Context Three defects surfaced in the post-ADR-007 architecture review, all concentrated at the seams ADR-007 touched and all load-bearing for REQ-2 (a replicator node may run multiple store instances — a fleet — sharing one pool via postgres; requirements.md): 1. **The `Backend` trait has no length probe.** The store surface promises `stat(key) → Option` ("cheap length/type probe"; store-api.md) and pool accounting rides `list` + `stat` (ADR-005), but the trait's methods are `has/get/put/delete/list/ name` with no value-metadata access. The kv engines' row shapes both carry a `size` column (POC #5 B5), and the large tier can stat a file — but with no trait method, kv `stat` degenerates to a full value load. The trait is the crate's second stable seam (ADR-003); discovering this during implementation would be a breaking trait change after backends exist. 2. **ADR-005's sweep-safety protocol is per-node by construction.** Pins, the pin/liveness arbitration lock, and delete-window arbitration all live inside one process (ADR-005 §Decision). That is sound for sqlite and single-instance postgres (REQ-1, REQ-3 topologies) — but under a fleet (REQ-2), node A's sweep enumerates the *shared* pool and arbitrates against A's in-process pin map; node B's in-flight put is invisible. Data loss, structurally: the topology REQ-2 requires does not satisfy the protocol ADR-005 specifies. 3. **The large tier is node-local, silently partitioning a fleet.** The large tier stores to a sharded layout on the node's own filesystem (ADR-003). Under a fleet sharing one kv pool, a >threshold blob put via node A exists only on A's disk; node B's get fall-through queries its own fs and misses, and `list()` unions disagree per node. REQ-2 makes this topology a shipping requirement, not a hypothetical. These three travel together: the fleet topology forces #2 and #3, and #3's resolution constrains what the large tier must support, while #1 is a trait-level gap any topology hits. ## Decision ### 1. `size(key) → Option` joins the `Backend` trait The trait gains a length probe: `size(key)` returns the entry's byte length or `None` if absent. Required of all engines; semantics per tier: - **kv engines**: an indexed column probe (`SELECT size …WHERE key`), never a value load. The `size` column both POC row shapes carry becomes contract-relevant. - **large engine (`local`)**: stat/metadata read, no content fetch. - The store facade's `stat` (store-api.md) and the accounting path (ADR-005) ride this method; the `has` method remains a pure existence probe. `size` is a trait addition *before* any backend ships data — the one-way-door discipline of ADR-002/003 applied to the trait itself. ### 2. Fleet GC: the delete window is durable and re-arbitrated in the database; pins are DB-backed under fleet engines The store facade's GC invariant is unchanged — **a visible pool entry is never deleted while liveness protects it, and an in-flight put is never deleted by a sweep started before it committed** (ADR-005). Under a fleet, the mechanism moves into the shared engine: - **Pins are rows.** `put`'s pin-before-publish (ADR-005) becomes two inserts in one transaction on the fleet engine: the entry row and its pin row commit together (or neither does). The pin row carries an owner (a store-instance id — ADR-011's vocabulary; "node id" is this ADR's original wording, superseded because a node may host several instances) and expiry: the pin guard renews its row every TTL/3 (default TTL 60 s, constructor-tunable — fleet constructor parameters, not API). **Protection = an un-expired pin row**: at delete-time arbitration, only `expiry > now()` pins count; an expired pin is abandoned (its holder failed to renew — the same liveness-failure signal as an in-process guard's drop) and is reaped by the sweeper's maintenance step (the GC-of-GC cleanup). A node that pauses past TTL loses its pins; delete-then-recover (ADR-005) remains the correctness backstop — content is re-put byte-identically under the same key. - **Registered liveness sources have one fleet form: a liveness table.** A fleet embedder registers its liveness as a named table of `(key, …)` rows in the shared engine (schema pinned by the embedder, shape contractually `(key bytea, …)` — any columns it needs beyond the key). In-process callback *views* are not a fleet liveness form: delete-time arbitration queries/finger-locks tables per key, so a source must be a table for the SQL protocol to see it. Single-node embedders keep callback sources unchanged (ADR-005). - **The protect callback is single-node-only.** A fleet sweeper cannot consult another node's in-process callback, so under fleet engines the callback's fleet expression is: the embedder materializes its protect set into its liveness table (or pin rows) *before* invoking sweep — the callback's abort semantics remain for single-node/topology-local sweeps that cannot express their protection as rows. - **One sweeper per pool.** A fleet node acquires a pg advisory lock before sweeping and holds it for the whole sweep. Two nodes sweeping concurrently is not a state this crate tolerates — the lock is the admission gate, and a second-node sweep returns the typed `GcAborted` (no deletion) rather than racing. `GcAborted` thus carries two cause classes (protection-source failure — ADR-005; sweeper lock contention — this ADR); the typed error's fleet variant identifies which. *(Return carrier corrected by ADR-012 §1: the abort is `GcAbortCause::SweeperLock` report data — `Ok(report)` — the typed sweep error is retired; the two cause classes and the no-deletion semantics stand.)* - **Delete-time arbitration is SQL-level.** The delete window's re-check (ADR-005: "arbitrate under the pin/liveness lock at delete time") executes as `SELECT … FOR UPDATE` against pin rows and liveness tables at delete time, then batch-deletes. The **staged re-arbitrated delete**: during mark, candidate-dead keys are staged as rows carrying `visible_after = mark_end + window` (the delete window's delay — default 5 min, constructor-tunable per deployment); the executor (the sweeper, or the next sweep's owner) re-arbitrates each staged candidate *at execution time* against pins and liveness tables before deleting. This is the durable delete window on every SQL-backed engine — the window table is engine rows on sqlite as well; what differs by topology is only where the liveness state lives (in-process maps and callbacks for single nodes; tables and SQL arbitration for fleets), not whether staging exists. *(ADR-012 §5/§6.7: "every SQL-backed engine" is the final wording — `local`/`mem` stage in-process, window-inline; the executor of overdue staged rows is the next sweep.)* No external queue dependency is required (the pattern is engine-native transactions; consumer-side queues like honker or pg-boss-lineage tools remain consumer-layer options, per ADR-004's boundary). - **Single-node postures are unchanged.** sqlite and single-instance postgres keep ADR-005's in-process protocol verbatim: the pin map, RAII guards, and arbitration lock stay in-process; only the *persistence* of the delete window (the staged-candidate table) is shared machinery. The pin-row expiry/ttl machinery exists only where fleet engines are active. *(Activation is explicit: ADR-012 §3's `fleet: true` constructor declaration turns the fleet machinery on; it is never inferred from engine choice alone — a single-instance postgres declared `fleet: false` keeps the in-process protocol.)* ### 3. The large tier is engine-selectable; fleet topology dictates which engines are permissible The large tier is a tier, not a medium. It ships engines the way the kv tier does (ADR-007), and a deployment's topology determines which are valid: - **`local` (default):** the sharded-layout filesystem backend, unchanged (ADR-003). Valid for any single-node topology (REQ-1, REQ-4) and for fleet nodes using **shared media** (below). - **Fleet + shared media (deployment requirement):** multi-node pools may keep the `local` engine by pointing every node's large tier at shared/replicated media. The soundness claim is *qualified*, not assumed: immutable-once-committed files plus stage-then-commit- rename are sound on media that provides (a) atomic rename within a directory, (b) close-to-open consistency (a written+closed file is complete to any later reader — NFSv3+ with the close-to-open discipline; NFSv4 and EFS-class stores qualify), and (c) no cross-node POSIX locking reliance (this store takes none — its cross-node coordination is the fleet engine's locks, never file locks; honker's NFS warning is about sqlite's lock protocol, not this shape). Verifying (a)-(c) on the deployment's actual media is the deployment's responsibility; a verification pass is recorded ops work, not crate code. - **Fleet re-routing (topology alternative):** a fleet whose nodes do not share media designates **one storage node**; large-blob writes and reads over the pool route through it via the ops surface (verified fetch/put; ADR-001). Client nodes' constructors are configured with **kv-only mode** (ADR-011 §4's name for what this bullet first called "kv-tier-only dispatch") — no local large tier over pool content at all; over-threshold puts reject at put time (ADR-012 §6.2), and the re-routing posture handles them via the ops surface; a client node's *node-private, consumer-owned* large-tier content is outside the pool and outside this crate's GC entirely: it is the consumer's own storage beside the pool per ADR-004, not a constructor mode over pool content, and the composite fleet-validity rule (backends-and-dispatch.md) therefore has no "local on node-private media over one pool" entry. The facade shape on a client node is its own store for sub-threshold content plus its registered ops-put path for large content — a consumer-side composition, not a store-level mode. - **`pg-lo` (shipped; admitted by ADR-009):** postgres Large Objects as the large tier's storage — the same "engine behind one contract" move as ADR-007, one tier over. Originally named here as a candidate gated on measured evidence, POC #7 ran and passed that gate (`docs/research/poc-pglo-findings.md`: companion-table authority, `lo_get` window gets, durable put ≈ fs's durable put, cached gets behind page-cache fs by a measured 20-50x single-stream, ~700 MB/s aggregate at 16 readers), and ADR-009 admitted the engine the same day — this bullet's original "candidate, not shipped / REQ-2's immediate answers are shared media or re-routing; pg-lo the consolidation option" posture is historical; the engine is shipped, feature `pg-lo`, and it is the fleet consolidation option as a first-class engine (ADR-012 §4 adds the same-instance requirement under fleet mode). - **The no-mixed-large-tiers rule is a deployment invariant, enforced at the enforceable seam.** Cross-node configuration cannot be validated by any one constructor (it sees only its own node). What the constructor *can* enforce: a fleet-mode constructor (fleet engines active) requires an explicit large-tier declaration — `local` (with the operator's assertion that its root is fleet-shared media), routed (kv-only mode, ADR-011 §4; no local filesystem over pool content), or `pg-lo`. The cross-node truth ("is node B's `local` actually the same media as node A's?") is a deployment invariant this crate documents and ops verifies, not something a constructor can prove; a partitioned tier's observable signature (get fall-through misses for entries another node wrote) is the documented detection symptom. The composite fleet-validity predicate over both tiers is pinned in backends-and-dispatch.md ("The composite fleet-validity rule"). In sum: a fleet's large tier over pool content is either shared media (all nodes' `local` roots on the same fleet-shared media) or routed (one storage node); mixing per-node-local media over one pool is the partitioning failure this decision exists to prevent. ## Consequences **Positive** - The trait now expressibly supports `stat` and pool accounting on every engine; the "cheap probe" claim of store-api.md is true rather than aspirational. - REQ-2 (fleet over one pool) has a complete, specced GC story: pins move to the shared engine atomically with put-commits, one sweeper holds the lock, deletes re-arbitrate in-transaction. The ADR-005 invariant extends to fleets as a *stated* mechanism, not an accident of single-process luck. - The large tier's fleet problem has named answers: shipped-code-free for shared media and re-routing, shipped-code for the consolidation option (pg-lo — admitted by ADR-009, the same POC-first door as ADR-007). **Negative** - The `Backend` trait gains a method before any backend ships — the correct moment, but every engine's task list grows by one method. - Fleet GC is real state machinery: pin rows with expiry (GC-of-GC surface for orphaned pins at node crash — the expiry is the cleanup handle), advisory-lock coordination, and SQL-level arbitration added on top of the already-intricate delete-window state machine. - `pg-lo` was built under ADR-009 after POC #7 passed — the measured, deliberate sequence of the ADR-006 pattern (admission gate first, engine second), not a hedge; resolved same-day. - *(ADR-012 §4: fleet GC state hosts on the fleet's kv engine — the large tier's `local` engine never joins the pin transaction; large=local fleet puts commit the pin row first, publish (commit-rename) second.)* **Neutral** - The sqlite and single-node-postgres postures are unchanged; this ADR costs existing paths nothing at runtime. - `pg-lo`, if admitted, rides the same driver stack ADR-007 already ships (tokio-postgres + deadpool) via SQL `lo_*` functions — no new driver dependency (the `postgres_large_object` crate is a dead 0.15-era io-trait glue; unnecessary; POC #7 confirms: ~400 lines of SQL-statement shapes over the existing stack). ## References - [requirements.md](../requirements.md) — REQ-2 (the fleet fact this ADR binds), REQ-1/3/4 (the topologies left unchanged) - POC #5 (`poc-postgres-kv-findings.md` B5 — row shapes with `size` columns; the driver stack pg-lo would ride) - ADR-002 (the one-way-door discipline this trait change applies), ADR-003 (the trait contract and admission gate `size` and pg-lo amend/enter), ADR-004 (tier≠engine, extended to the large tier), ADR-005 (the GC invariant and single-node protocol extended here, not replaced), ADR-006 (the deferred-cost pattern pg-lo's admission follows), ADR-007 (the engine-behind-one-trait precedent) - [backends-and-dispatch.md](../backends-and-dispatch.md) (the trait and tier surface this amends); [store-api.md](../store-api.md) (`stat`'s contract); [gc-and-namespaces.md](../gc-and-namespaces.md) (the fleet GC mechanism); [ops-surface.md](../ops-surface.md) (the re-routing topology's transport); ADR-010 (the fleet mechanism's state home and the joint-tx ownership this ADR's §2 implied but did not rule); ADR-011 (the vocabulary — instance/fleet — this ADR's topology language resolves to; the large tier renamed `large`)