Full-tree review before task decomposition found five composition
defects (mechanisms specced correctly in isolation, composition
unruled) and a set of caller-facing gaps. ADR-012 rules each:
- Sweeps are never errors: aborts are GcAbortCause report data in
Ok(SweepReport) (ProtectFailed / SweeperLock / NoLivenessSources);
GcAborted retired from the error enum; direct-delete refusal is
the GcRefuse error covering the full protection set
- Pin token gains its wire shape: blobs/put response {token, digest};
blobs/have's token-renewal form (one digest + token); token
validity domain = the minting serving node, process-lifetime
mapping
- Fleet mode is an explicit constructor declaration (fleet: true),
never inferred from engine choice
- All fleet GC state hosts on the fleet's kv engine (postgres) — one
arbitration domain; large=pg-lo fleet nodes required onto the same
pg instance (composite predicate's new clause); large=local fleet
puts pin-row-first, publish-second
- Engine-state seam reduced: sqlite pin/sweep-lock bodies dropped
(dead machinery); non-SQL engines stage delete-window candidates
in-process; one-window-host rule per instance
- Facade clarifications: fall-through for all key-addressed ops,
kv-only put-time rejection, mem+local dual-tier valid, error-model
member/return-shape ruling (trait/facade family split), has ->
bool, PinState variants, fleet liveness-table registration form,
window executor = the next sweep
Alignment edits across all specs and ADR-005/008/009/010/011
(bracketed corrections per the established pattern); OQ-11 (pg-only-kv
feature graph) added to the parked index for auditability.
Verification: two independent review passes; all findings resolved;
verdict READY for task decomposition.
8.6 KiB
ADR-009: pg-lo — postgres Large Objects as the large tier's second engine
Status
Accepted (admission evidence: POC #7, passed 2026-10-03 —
docs/research/poc-pglo-findings.md; the ADR-008-named candidate's
gate. The tier this ADR calls "fs" is renamed large by ADR-011);
§2's "put = one transaction" tx-ownership statement is corrected by
ADR-010 §2 (the store core holds the tx; the engine contributes its
staged-put mechanics into it — engine content unchanged); ADR-012 §4
adds the same-pg-instance requirement for fleet-mode pg-lo nodes
(the consolidation posture this ADR described is the fleet-valid
shape).
Context
ADR-008 named pg-lo as the candidate large-tier engine for REQ-2's
fleet topology (multiple store instances sharing one pg pool, serving
large/beyond-threshold pool content) — expressly gated on the ADR-003
admission door: measured evidence first, in the ADR-007 shape. POC #7
ran that door (poc-pglo-spec.md → poc-pglo-findings.md,
findings C1–C7): a ~400-line PgLoBackend over the ADR-003/008 trait
contract (including size), 10/10 exact-count sweep-outcome contract
tests, on the POC #5 driver stack (tokio-postgres + deadpool-postgres
via SQL lo_* functions — no new dependency, as ADR-008 §Neutral
projected), against dockerized postgres:16.
The gate's three legs all passed:
- Contract — complete
list()(= companion table == LO catalog oids, after mixed ops), GC-participating delete, virgin-store no-ops, CAS re-put, stage rollback, ranged reads; proven by the exact-count sweep-outcome shape ADR-007's gate uses. - Performance — durable puts 60–65 MB/s at ≥1 MiB, within 5–10% of durable local fs at 16–128 MiB and beating it below 1 MiB on the POC box (fsync-dominated); cached gets 70–180 MB/s single-stream, ~700 MB/s aggregate at 16 readers — 20–50× behind page-cache fs single-stream, named honestly (see Consequences).
- Ops posture — the tx-scoped-descriptor question resolved in
pg's favor (descriptorless
lo_get(oid, off, len)windows: pool- friendly, stateless, ranged reads never pay the tx cost); LO catalog churn page-granular and reusing; LO creation transactional — crash orphans structurally zero.
Decision
The large tier ships pg-lo as its second engine — postgres Large
Objects — behind the same Backend trait, feature pg-lo (default-
off), constructor-selected per node. REQ-2's deployment choice set
(ADR-008) extends from {shared media, re-routing} to include the
consolidation option as a shipped engine.
- Engine shape (POC #7 C1 is the contract): LOs hold the bytes;
the companion table
lo_entries (key bytea PK, loid oid, size bigint, committed_at)is the contract authority —list(),size(),has(), CAS-existence are all table-level. Never trust or readpg_largeobjectfor addressing (the C6 sweep validates against the table only). - Put = one transaction: BEGIN →
lo_create→ chunkedlowrite(512 KiB statements — the measured knee; statement count, not LOBLKSIZE, is what costs) → companion-row insert → COMMIT. (Ownership corrected by ADR-010 §2: the store core holds this tx; the engine contributes LO writes + the companion-row insert into it.) Pin-before-publish (ADR-005/008) rides this shape: entry row + pin row commit together with the content — the pin-row insert is the store's contribution (hosted on the fleet's kv engine under fleet mode, ADR-012 §4 — same pg instance, one tx). Unknown-length puts stage into the LO and commit the row on completion; rollback discards both. - Get =
lo_get(oid, off, len)windows by default (POC #7 C4): the large-tier handle closes over the companion row (oid + length) — immutable, so the row is authoritative — not over a tx-scoped descriptor. Ranged reads are single stateless statements, pool-friendly (acquire p99 1–6 ms ≤ pool size). The held-descriptor tx posture exists as the fallback (measured: no advantage even for sequential full reads). Reads ridelo_*64-bit variants end-to-end (lo_lseek64/int8 offsets;lo_lseek's int4 cap is a C7 trap the engine must not inherit). - Delete = one tx (
lo_unlink+ row delete) under the ADR-008 delete-window arbitration (table-level; nothing fs-specific). - Posture deltas, as operator requirements (the honest cost of the
engine, POC #7 C5/C6):
- LO catalog space is reused but never returned — the engine's
inverted-fs property. The engine ADR requirement: rel-size
monitoring (
pg_total_relation_size('pg_largeobject')) named in ops docs;VACUUM FULL/pg_repackis the only shrink path — an operator runbook line, never automatic. - Autovacuum inherited; keep it on.
pg_largeobjectis an ordinary catalog heap (autovacuum applies, measured active); a deployment disabling autovacuum gets an operator-visible note. - Orphan-recovery sweep exists for the legacy/bypass class only
(committed LOs without companion rows — mixed/legacy stores). The
crash story is structurally clean: LO lifecycle is transactional
in modern postgres; kill mid-tx leaves zero pages, zero
half-commits (cleaner than
local's stage-file residue, which needs its own recovery sweep).
- LO catalog space is reused but never returned — the engine's
inverted-fs property. The engine ADR requirement: rel-size
monitoring (
- Performance deltas, as named capacity facts (not hidden, not
solved): single-stream cached get runs 20–50× behind page-cache
local fs at ≥1 MiB (server-side LO page-walk + copy ceiling,
~65–70 MB/s per in-flight statement); the fleet serving picture
aggregates ~700 MB/s at 16 readers (server-side parallelism) —
which is REQ-2's relevant number — and exceeds-threshold single
clone streams get ~120 MB/s. Deployments whose large-blob serving
is dominated by single-stream page-cache-bound reads stay better
served by
local-on-shared-media; pg-lo's case is the consolidated fleet (one engine to operate) whose serving burden is many-user aggregate. A third engine (S3-like) remains the ADR-003 door; nothing here touches it. - redb-class guardrail maintained: this is the second engine the admission door ships (ADR-007: kv; ADR-009: large) — both on the same measured-evidence shape; no third engine is opened by this one.
Consequences
Positive
- REQ-2's consolidation option is real and shipped: a fleet node runs postgres only — kv (bytea) tier + large (LO) tier from one engine, no shared-media requirement, no re-routing topology, no consumer-ops dependency for pool content.
- The engine adds ~400 lines of impl over an already-fixed contract — the additive shape both ADR-007 and ADR-008 projected, now evidenced.
- Crash hygiene is better than
local(transactional LO lifecycle vs. stage-file residue); the engine's only recovery sweep targets a class (bypassed-LO) that healthy deployments never produce.
Negative
- Single-stream cached get is 20–50× behind local fs — a real,
named cost. Deployments with page-cache-friendly media and single-
dominant readers should stay on
localengines. (Aggregate serving — the fleet case — scales to ~700 MB/s on the POC box; per-box numbers shift with cores.) - LO catalog space is never returned (monitor + occasional
operator-run
VACUUM FULL), the inverse oflocal's immediate-return. - Feature/maintenance surface: the large tier now carries two engines (like the kv tier), each with posture deltas to keep honest (ADR-007's negative, inherited).
Neutral
- The
Backendtrait is unchanged (ADR-008'ssizeprobe is the last trait change — pg-lo fits the contract as it stands). - Driver stack: unchanged (tokio-postgres + deadpool; SQL
lo_*functions,postgres_large_objectcrate stays avoided). - The
pg-lofeature defaults off; CI's--all-featuresgate adds it to the sweep-outcome suite (dockerized postgres:16, same shape as thepostgreskv engine's gate).
References
docs/research/poc-pglo-findings.md(POC #7, findings C1–C7 — the evidence base);poc-pglo-spec.md(the gate this ADR's legs mirror)- requirements.md — REQ-2 (the fleet topology this engine serves); REQ-1/3/4 unchanged postures
- ADR-003 (the trait contract + admission door exercised twice now), ADR-004 (tier ≠ engine, large tier edition), ADR-005/008 (the GC invariants this engine participates in — table-level, nothing fs-specific), ADR-007 (the engine-admission precedent and driver stack)
- backends-and-dispatch.md (the large-tier engine surface this amends); overview.md; store-api.md