Files
alkblobs/docs/architecture/decisions/008-trait-size-probe-fleet-gc-and-large-engines.md
glm-5.3-flash 7b9d904b8a docs(architecture): ADR-012 — pre-decomposition consistency rulings; tree verified single-valued
Full-tree review before task decomposition found five composition
defects (mechanisms specced correctly in isolation, composition
unruled) and a set of caller-facing gaps. ADR-012 rules each:

- Sweeps are never errors: aborts are GcAbortCause report data in
  Ok(SweepReport) (ProtectFailed / SweeperLock / NoLivenessSources);
  GcAborted retired from the error enum; direct-delete refusal is
  the GcRefuse error covering the full protection set
- Pin token gains its wire shape: blobs/put response {token, digest};
  blobs/have's token-renewal form (one digest + token); token
  validity domain = the minting serving node, process-lifetime
  mapping
- Fleet mode is an explicit constructor declaration (fleet: true),
  never inferred from engine choice
- All fleet GC state hosts on the fleet's kv engine (postgres) — one
  arbitration domain; large=pg-lo fleet nodes required onto the same
  pg instance (composite predicate's new clause); large=local fleet
  puts pin-row-first, publish-second
- Engine-state seam reduced: sqlite pin/sweep-lock bodies dropped
  (dead machinery); non-SQL engines stage delete-window candidates
  in-process; one-window-host rule per instance
- Facade clarifications: fall-through for all key-addressed ops,
  kv-only put-time rejection, mem+local dual-tier valid, error-model
  member/return-shape ruling (trait/facade family split), has ->
  bool, PinState variants, fleet liveness-table registration form,
  window executor = the next sweep

Alignment edits across all specs and ADR-005/008/009/010/011
(bracketed corrections per the established pattern); OQ-11 (pg-only-kv
feature graph) added to the parked index for auditability.

Verification: two independent review passes; all findings resolved;
verdict READY for task decomposition.
2026-10-03 07:14:45 +00:00

293 lines
16 KiB
Markdown

# ADR-008: Trait length probe, fleet GC, and large-tier engines
## Status
Accepted (the §3 pg-lo bullet's "candidate, not shipped" posture was
superseded 2026-10-03 by ADR-009, which admitted pg-lo on POC #7's
passed gate — the two §Consequences passages that said pg-lo "remains
unbuilt" are resolved by that same supersession; the tier this ADR's
title and text call "fs" is renamed `large` by ADR-011; the fleet GC
mechanism's state home and the joint-tx ownership are ruled by
ADR-010; fleet-mode activation, the GC-state host, the sweep-abort
return shape, and the window executor are ruled by ADR-012; every
other part of this ADR stands as written)
## Context
Three defects surfaced in the post-ADR-007 architecture review, all
concentrated at the seams ADR-007 touched and all load-bearing for
REQ-2 (a replicator node may run multiple store instances — a fleet —
sharing one pool via postgres; requirements.md):
1. **The `Backend` trait has no length probe.** The store surface
promises `stat(key) → Option<EntryMeta>` ("cheap length/type
probe"; store-api.md) and pool accounting rides `list` + `stat`
(ADR-005), but the trait's methods are `has/get/put/delete/list/
name` with no value-metadata access. The kv engines' row shapes
both carry a `size` column (POC #5 B5), and the large tier can stat a
file — but with no trait method, kv `stat` degenerates to a full
value load. The trait is the crate's second stable seam (ADR-003);
discovering this during implementation would be a breaking trait
change after backends exist.
2. **ADR-005's sweep-safety protocol is per-node by construction.**
Pins, the pin/liveness arbitration lock, and delete-window
arbitration all live inside one process (ADR-005 §Decision). That
is sound for sqlite and single-instance postgres (REQ-1, REQ-3
topologies) — but under a fleet (REQ-2), node A's sweep enumerates
the *shared* pool and arbitrates against A's in-process pin map;
node B's in-flight put is invisible. Data loss, structurally: the
topology REQ-2 requires does not satisfy the protocol ADR-005
specifies.
3. **The large tier is node-local, silently partitioning a fleet.** The
large tier stores to a sharded layout on the node's own filesystem
(ADR-003). Under a fleet sharing one kv pool, a >threshold blob
put via node A exists only on A's disk; node B's get fall-through
queries its own fs and misses, and `list()` unions disagree per
node. REQ-2 makes this topology a shipping requirement, not a
hypothetical.
These three travel together: the fleet topology forces #2 and #3, and
#3's resolution constrains what the large tier must support, while #1 is
a trait-level gap any topology hits.
## Decision
### 1. `size(key) → Option<u64>` joins the `Backend` trait
The trait gains a length probe: `size(key)` returns the entry's byte
length or `None` if absent. Required of all engines; semantics per
tier:
- **kv engines**: an indexed column probe (`SELECT size …WHERE key`),
never a value load. The `size` column both POC row shapes carry
becomes contract-relevant.
- **large engine (`local`)**: stat/metadata read, no content fetch.
- The store facade's `stat` (store-api.md) and the accounting path
(ADR-005) ride this method; the `has` method remains a pure
existence probe.
`size` is a trait addition *before* any backend ships data — the
one-way-door discipline of ADR-002/003 applied to the trait itself.
### 2. Fleet GC: the delete window is durable and re-arbitrated in the database; pins are DB-backed under fleet engines
The store facade's GC invariant is unchanged — **a visible pool entry
is never deleted while liveness protects it, and an in-flight put is
never deleted by a sweep started before it committed** (ADR-005).
Under a fleet, the mechanism moves into the shared engine:
- **Pins are rows.** `put`'s pin-before-publish (ADR-005) becomes two
inserts in one transaction on the fleet engine: the entry row and
its pin row commit together (or neither does). The pin row carries
an owner (a store-instance id — ADR-011's vocabulary; "node id" is
this ADR's original wording, superseded because a node may host
several instances) and expiry: the pin guard renews its row every
TTL/3 (default TTL 60 s, constructor-tunable — fleet constructor
parameters, not API). **Protection = an un-expired pin row**: at
delete-time arbitration, only `expiry > now()` pins count; an
expired pin is abandoned (its holder failed to renew — the same
liveness-failure signal as an in-process guard's drop) and is reaped
by the sweeper's maintenance step (the GC-of-GC cleanup). A node
that pauses past TTL loses its pins; delete-then-recover (ADR-005)
remains the correctness backstop — content is re-put byte-identically
under the same key.
- **Registered liveness sources have one fleet form: a liveness
table.** A fleet embedder registers its liveness as a named table of
`(key, …)` rows in the shared engine (schema pinned by the embedder,
shape contractually `(key bytea, …)` — any columns it needs beyond
the key). In-process callback *views* are not a fleet liveness form:
delete-time arbitration queries/finger-locks tables per key, so a
source must be a table for the SQL protocol to see it. Single-node
embedders keep callback sources unchanged (ADR-005).
- **The protect callback is single-node-only.** A fleet sweeper cannot
consult another node's in-process callback, so under fleet engines
the callback's fleet expression is: the embedder materializes its
protect set into its liveness table (or pin rows) *before* invoking
sweep — the callback's abort semantics remain for
single-node/topology-local sweeps that cannot express their
protection as rows.
- **One sweeper per pool.** A fleet node acquires a pg advisory lock
before sweeping and holds it for the whole sweep. Two nodes
sweeping concurrently is not a state this crate tolerates — the
lock is the admission gate, and a second-node sweep returns the
typed `GcAborted` (no deletion) rather than racing. `GcAborted`
thus carries two cause classes (protection-source failure —
ADR-005; sweeper lock contention — this ADR); the typed error's
fleet variant identifies which. *(Return carrier corrected by
ADR-012 §1: the abort is `GcAbortCause::SweeperLock` report data —
`Ok(report)` — the typed sweep error is retired; the two cause
classes and the no-deletion semantics stand.)*
- **Delete-time arbitration is SQL-level.** The delete window's
re-check (ADR-005: "arbitrate under the pin/liveness lock at delete
time") executes as `SELECT … FOR UPDATE` against pin rows and
liveness tables at delete time, then batch-deletes. The **staged
re-arbitrated delete**: during mark, candidate-dead keys are staged
as rows carrying `visible_after = mark_end + window` (the delete
window's delay — default 5 min, constructor-tunable per
deployment); the executor (the sweeper, or the next sweep's owner)
re-arbitrates each staged candidate *at execution time* against
pins and liveness tables before deleting. This is the durable
delete window on every SQL-backed engine — the window table is
engine rows on
sqlite as well; what differs by topology is only where the liveness
state lives (in-process maps and callbacks for single nodes;
tables and SQL arbitration for fleets), not whether staging exists.
*(ADR-012 §5/§6.7: "every SQL-backed engine" is the final wording
— `local`/`mem` stage in-process, window-inline; the executor of
overdue staged rows is the next sweep.)* No external queue
dependency is required (the pattern is
engine-native transactions; consumer-side queues like honker or
pg-boss-lineage tools remain consumer-layer options, per ADR-004's
boundary).
- **Single-node postures are unchanged.** sqlite and single-instance
postgres keep ADR-005's in-process protocol verbatim: the pin map,
RAII guards, and arbitration lock stay in-process; only the
*persistence* of the delete window (the staged-candidate table) is
shared machinery. The pin-row expiry/ttl machinery exists only
where fleet engines are active. *(Activation is explicit: ADR-012
§3's `fleet: true` constructor declaration turns the fleet
machinery on; it is never inferred from engine choice alone — a
single-instance postgres declared `fleet: false` keeps the
in-process protocol.)*
### 3. The large tier is engine-selectable; fleet topology dictates which engines are permissible
The large tier is a tier, not a medium. It ships engines the way the kv
tier does (ADR-007), and a deployment's topology determines which are
valid:
- **`local` (default):** the sharded-layout filesystem backend,
unchanged (ADR-003). Valid for any single-node topology (REQ-1,
REQ-4) and for fleet nodes using **shared media** (below).
- **Fleet + shared media (deployment requirement):** multi-node pools
may keep the `local` engine by pointing every node's large tier at
shared/replicated media. The soundness claim is *qualified*, not
assumed: immutable-once-committed files plus stage-then-commit-
rename are sound on media that provides (a) atomic rename within a
directory, (b) close-to-open consistency (a written+closed file is
complete to any later reader — NFSv3+ with the close-to-open
discipline; NFSv4 and EFS-class stores qualify), and (c) no
cross-node POSIX locking reliance (this store takes none — its
cross-node coordination is the fleet engine's locks, never file
locks; honker's NFS warning is about sqlite's lock protocol, not
this shape). Verifying (a)-(c) on the deployment's actual media is
the deployment's responsibility; a verification pass is recorded
ops work, not crate code.
- **Fleet re-routing (topology alternative):** a fleet whose nodes do
not share media designates **one storage node**; large-blob writes
and reads over the pool route through it via the ops surface
(verified fetch/put; ADR-001). Client nodes' constructors are
configured with **kv-only mode** (ADR-011 §4's name for what this
bullet first called "kv-tier-only dispatch") — no local large tier
over pool content at all; over-threshold puts reject at put time
(ADR-012 §6.2), and the re-routing posture handles them via the ops
surface; a client node's *node-private,
consumer-owned* large-tier content is outside the pool and outside
this crate's GC entirely: it is the consumer's own storage beside
the pool per ADR-004, not a constructor mode over pool content, and
the composite fleet-validity rule (backends-and-dispatch.md)
therefore has no "local on node-private media over one pool" entry.
The facade shape on
a client node is its own store for sub-threshold content plus its
registered ops-put path for large content — a consumer-side
composition, not a store-level mode.
- **`pg-lo` (shipped; admitted by ADR-009):** postgres Large Objects
as the large tier's storage —
the same "engine behind one contract"
move as ADR-007, one tier over. Originally named here as a candidate
gated on measured evidence, POC #7 ran and passed that gate
(`docs/research/poc-pglo-findings.md`: companion-table authority,
`lo_get` window gets, durable put ≈ fs's durable put, cached gets
behind page-cache fs by a measured 20-50x single-stream, ~700 MB/s
aggregate at 16 readers), and ADR-009 admitted the engine the same
day — this bullet's original "candidate, not shipped / REQ-2's
immediate answers are shared media or re-routing; pg-lo the
consolidation option" posture is historical; the engine is shipped,
feature `pg-lo`, and it is the fleet consolidation option as a
first-class engine (ADR-012 §4 adds the same-instance requirement
under fleet mode).
- **The no-mixed-large-tiers rule is a deployment invariant, enforced at
the enforceable seam.** Cross-node configuration cannot be validated
by any one constructor (it sees only its own node). What the
constructor *can* enforce: a fleet-mode constructor (fleet engines
active) requires an explicit large-tier declaration — `local` (with the
operator's assertion that its root is fleet-shared media), routed
(kv-only mode, ADR-011 §4; no local filesystem over pool content), or
`pg-lo`. The
cross-node truth ("is node B's `local` actually the same media as
node A's?") is a deployment invariant this crate documents and ops
verifies, not something a constructor can prove; a partitioned
tier's observable signature (get fall-through misses for entries
another node wrote) is the documented detection symptom. The
composite fleet-validity predicate over both tiers is pinned in
backends-and-dispatch.md ("The composite fleet-validity rule").
In sum: a fleet's large tier over pool content is either shared media (all
nodes' `local` roots on the same fleet-shared media) or routed (one
storage node); mixing per-node-local media over one pool is the
partitioning failure this decision exists to prevent.
## Consequences
**Positive**
- The trait now expressibly supports `stat` and pool accounting on
every engine; the "cheap probe" claim of store-api.md is true rather
than aspirational.
- REQ-2 (fleet over one pool) has a complete, specced GC story: pins
move to the shared engine atomically with put-commits, one sweeper
holds the lock, deletes re-arbitrate in-transaction. The ADR-005
invariant extends to fleets as a *stated* mechanism, not an
accident of single-process luck.
- The large tier's fleet problem has named answers: shipped-code-free
for shared media and re-routing, shipped-code for the consolidation
option (pg-lo — admitted by ADR-009, the same POC-first door as
ADR-007).
**Negative**
- The `Backend` trait gains a method before any backend ships — the
correct moment, but every engine's task list grows by one method.
- Fleet GC is real state machinery: pin rows with expiry (GC-of-GC
surface for orphaned pins at node crash — the expiry is the cleanup
handle), advisory-lock coordination, and SQL-level arbitration
added on top of the already-intricate delete-window state machine.
- `pg-lo` was built under ADR-009 after POC #7 passed — the
measured, deliberate sequence of the ADR-006 pattern (admission
gate first, engine second), not a hedge; resolved same-day.
- *(ADR-012 §4: fleet GC state hosts on the fleet's kv engine — the
large tier's `local` engine never joins the pin transaction;
large=local fleet puts commit the pin row first, publish
(commit-rename) second.)*
**Neutral**
- The sqlite and single-node-postgres postures are unchanged; this ADR
costs existing paths nothing at runtime.
- `pg-lo`, if admitted, rides the same driver stack ADR-007 already
ships (tokio-postgres + deadpool) via SQL `lo_*` functions — no
new driver dependency (the `postgres_large_object` crate is a dead
0.15-era io-trait glue; unnecessary; POC #7 confirms: ~400 lines of
SQL-statement shapes over the existing stack).
## References
- [requirements.md](../requirements.md) — REQ-2 (the fleet fact this
ADR binds), REQ-1/3/4 (the topologies left unchanged)
- POC #5 (`poc-postgres-kv-findings.md` B5 — row shapes with `size`
columns; the driver stack pg-lo would ride)
- ADR-002 (the one-way-door discipline this trait change applies),
ADR-003 (the trait contract and admission gate `size` and pg-lo
amend/enter), ADR-004 (tier≠engine, extended to the large tier),
ADR-005 (the GC invariant and single-node protocol extended here,
not replaced), ADR-006 (the deferred-cost pattern pg-lo's admission
follows), ADR-007 (the engine-behind-one-trait precedent)
- [backends-and-dispatch.md](../backends-and-dispatch.md) (the trait
and tier surface this amends);
[store-api.md](../store-api.md) (`stat`'s contract);
[gc-and-namespaces.md](../gc-and-namespaces.md) (the fleet GC
mechanism); [ops-surface.md](../ops-surface.md) (the re-routing
topology's transport); ADR-010 (the fleet mechanism's state home
and the joint-tx ownership this ADR's §2 implied but did not rule);
ADR-011 (the vocabulary — instance/fleet — this ADR's topology
language resolves to; the large tier renamed `large`)