The ADR-008 pg-lo admission POC ran in a standalone crate (/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait contract (including size), 10/10 exact-count sweep-outcome contract tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432, POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions, no new dependency. Gate verdict: passed, with named deltas. - Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and below 1 MiB beating — durable local fs on this fsync-slow disk; cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16 readers (20-50x behind page-cache fs — the honest named delta) - Contract: companion table is the list()/size()/CAS authority (never the catalogs); stage-then-commit; GC-participating lo_unlink delete - Handles: the tx-scoped descriptor is real but pool-compatible via descriptorless lo_get(oid, off, len) windows — window gets keep handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor is the fallback posture - Vacuum: pg_largeobject pages churn-reused, never returned; tracked by autovacuum; rel-size monitoring named as an ops requirement - Crash/orphan: LO creation is transactional — kill/terminate mid-write-tx leaves zero orphan pages; the only orphan class is a committed LO bypassing the companion table (planted, reaped by the ~7 ms/oid sweep; committed content survives byte-exact) - Harness lessons: lo_lseek is int4 — the 64 variants are the >2 GiB discipline; shared-table parallel tests are unsound (per-test CREATE DATABASE isolation) Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed; register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated (duplicate bullet removed) + backends-and-dispatch/open-questions cross-references. Verification: cargo test --release (10 passed), clippy -D warnings, fmt --check in /workspace/alkblobs-pglo-poc.
12 KiB
status, last_updated
| status | last_updated |
|---|---|
| draft | 2026-10-03 |
Backends and dispatch
What this is
The physical storage layer: the Backend trait contract every backend
implements, the two shipped backends (kv, fs), and the size-threshold
dispatch that routes between them. The trait is the crate's second
stable seam (with the key encoding, ADR-002): backends must survive
digest-layer evolution, so the boundary stays dumb.
What the Backend trait is (and is not)
The trait is a CAS-tier contract: it accepts, serves, and enumerates
immutable content-addressed entries (hash → bytes), and its list()/
delete() shapes exist to feed GC (ADR-005). That is its whole job, and
the contract assumes it: complete-list-equals-sweep-safety only makes
sense for a store where every entry is a standalone immutable blob.
It follows that the trait is not a general-purpose storage abstraction,
and a downstream must never implement it to host data that is not
content-addressed blobs — a manifest table, a queryable index, or any
mutable/structured state. Such storage is consumer-owned, beside the
pool (a downstream's own sqlite/whatever, ADR-004): it needs no
digest-byte addressing, no list()-completeness, and no sweep
participation, so the CAS-tier contract fits it worse than it fits any
purpose-built option. The wrong turn to rule out is "my data isn't
blobs, therefore I implement a Backend for it"; nothing a consumer needs
is ever reachable through that door.
The Backend trait contract (ADR-003)
- Opaque byte keys, opaque byte values. Backends never learn what a
digest is; typed
Keyconverts at the store boundary (ADR-002). This is also the persistence boundary — a kv row or an fs filename has no schema-migration story, so the byte layout must be self-describing (algorithm-tagged keys). - Methods:
has/get/put/delete/list/name/size(the length probe added by ADR-008). Streaming put/get shapes ride the store layer's seams (store-api.md); backends provide byte- or handle-level primitives underneath (geton the fs tier yields a file handle, not a loaded buffer).size(key)returns the entry's length orNoneif absent — metadata only, never a content fetch: kv engines probe the indexedsizecolumn both row shapes carry (POC #5 B5); thelocalfs engine stats the file. Store-levelstatand pool accounting ride this method. list()complete by contract. A malformed list is as deadly as an incomplete one: list correctness is only observable through GC (POC #1 finding 2 — the redb key-vs-value trap produced a sweep that deleted the wrong blobs, silently). Therefore:- every backend implementation must prove list correctness through sweep-outcome tests (the invariant's test gate lives in store-api.md);
- an implementation that cannot enumerate (the rudolfs S3 anti-lesson) must not ship — GC is a structural requirement, not an optional extra (ADR-005).
- Virgin-store reads are no-ops: read paths treat a missing table as absent/empty (POC #1 finding 2).
- Namespace-blind: backends never see namespaces, tenants, or
reference structure (ADR-005 — the rudolfs inversion; physical
storage is
hash → bytesflat).
Shipped backends
Two tiers, exactly — this counts trait implementations this crate ships as CAS tiers, the dispatch routes between them; it is the complete set the problem requires, and the "third backend" fear is a category error fixed in ADR-004. It says nothing about other storage existing in a deployment: a downstream runs whatever else it needs on its own media, beside the pool, above the store's seams (see "What the Backend trait is"). Tier count is not engine count: the kv tier carries two engines (sqlite, postgres) behind one trait — ADR-007; per-node engine selection is a constructor parameter.
kv backend (feature kv, default-on); engines: sqlite (default) + postgres (feature postgres, default-off)
Small blobs — kv is the tier name; the engine is a per-node
constructor choice between the shipped engines (ADR-007; ADR-003
carries sqlite's on-disk format trade: file compat is upstream
sqlite's guarantee). The shipped trait impls write plain
content-addressed rows into the chosen engine, nothing else. A
third engine remains the substitution door as amended: a new ADR
with its own sweep-safety proof (the door is no longer hypothetical —
ADR-007 is its first in-repo exercise); and a downstream's other
storage — schemas, manifests, whatever it likes including some other
sqlite file — never comes near this tier at all, per "What the
Backend trait is".
The collapse ADR-007 buys: the classical downstream stack (git server, vfs node) needed three storage systems — kv + fs + relational. Here the kv tier rides the relational engine itself, so a node provisions one relational engine + fs; consumer tables (refs, manifests, queues) sit beside the pool on the same engine or file by the deployment's own choice (co-tenancy note: same-file consumer writes share sqlite's single-writer ceiling — that sharing is the node's trade; separate files remain available).
Evidence (POC #3 finding A4, first-party measured): sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime — most git objects, workspace files, manifests), with the crossover at ~128-256 KiB where fs stops paying the B-tree row rewrite and wins.
sqlite engine (the default): WAL + synchronous=NORMAL as the
shipped durability tier (matching what the benchmark measured and what
iroh's store ships); bounded reads; has as an EXISTS probe; prepared
statements; read paths tolerate the no-tables-yet database (see
contract). Solo economics win the single-machine case by measurement;
the WAL single-writer ceiling (~1.2k objects/s under contention,
POC #5 B3) is an order above single-node push rates.
postgres engine (feature postgres, default-off; ADR-007): the
cross-machine-writers / many-client-node / already-running-pg engine.
Contract-identical row shape (key bytea PK / value bytea / size),
ON CONFLICT DO NOTHING CAS, complete list() as a stable cursor
over a vacuuming table, batch-delete via key = ANY($1). Impl
requirements per POC #5 finding B5: configurable synchronous_commit
posture (the parity knob with the sqlite shipped tier), pooled
connections with per-connection prepared-statement discipline, stated
autovacuum + fillfactor tuning, ~40× small-tier storage overhead
accepted (B6). Sweep-safety proof via the same exact-count
sweep-outcome CI gate, run against dockerized postgres under
--all-features. Concurrency evidence (POC #5 B3/B4): gets/puts scale
near-linearly to ~37k puts/s at 24 conns — the only engine whose
throughput increases under concurrency.
fs backend (feature fs, default-on); engines: local (default) + pg-lo (candidate)
Large blobs. Like the kv tier (ADR-007), the fs tier is one contract, multiple engines — the medium is a per-node constructor choice (ADR-008; REQ-2 in requirements.md makes the engine concept load-bearing for fleets):
localengine (the default; the ADR-003 backend as originally specced): flat sharded layout:{hex-prefix}/{hex-prefix}/{hash}sharding survives from iroh's conclusion (limits directory size on huge pools); stage-then-commit- rename for the two-pass unknown-length path; pread-based range reads (POC #3 finding A2 — local range serving is sound, e.g. for packfiles).pg-loengine (candidate, not shipped; ADR-008): postgres Large Objects as the fs tier's storage — the same engine-behind- one-trait move as ADR-007, one tier over. Admission is gated on POC #7's measured evidence (passed 2026-10-03:docs/research/poc-pglo-findings.md— durable put ≈ fs's durable put,lo_getwindow gets, companion-table authority, clean crash-orphan behavior) and its own engine ADR (seedocs/research/poc-pglo-spec.md).- Fleet locality contract (ADR-008): over one shared pool, the fs
tier is either shared media (every node's
localroot on the same fleet-shared media — soundness caveats and the required media guarantees in ADR-008) or re-routed (one storage node serves pool large-blob content via the ops surface; client nodes are kv-tier-only over pool content). Mixed per-node-local fs tiers over one pool is the partitioning failure this contract exists to prevent — a documented deployment invariant the constructor cannot fully prove but must be declared against (ADR-008 names the enforceable seam and the detection symptom).
mem backend (feature mem, default-off)
BTreeMap-shaped ephemeral backend for tests and in-process
ephemerality. A testing/utility tier, never a production story.
Implements the full trait contract including size (trivially) — it
exists partly as the reference impl every other backend's conformance
tests can mirror.
Size-threshold dispatch (ADR-003)
- Routing is a pure function of content length. Same content ⇒ same length ⇒ same backend; re-puts are deterministic. No content ever migrates between backends — the migration question existed only to patch the unknown-length asymmetry, which ADR-003's pre-threshold buffering eliminates by construction.
- Default threshold: 128 KiB (constructor-tunable). Midpoint of the measured flat zone (POC #3 A4). Re-tuning per deployment media is a constructor parameter, not an API change.
- Get fall-through: small-tier miss queries the large tier (deterministic, cheap — a stat probe).
- Per-namespace or per-tenant backend configuration: rejected (ADR-003 §Consequences — it would re-weld namespacing into the physical layer, the rudolfs anti-pattern ADR-005 inverts).
Where a new backend or engine could come from
The trait is open to future implementations (network stores, S3-like
tiers), but nothing in the current consumer set requires one, and the
contract is deliberately hostile to half-implementations (complete
list(), GC-participating delete). Any future backend — or a third
kv engine or second fs engine — is a new ADR carrying its own
sweep-safety story (ADR-007 is the precedent for a kv engine
addition; ADR-008 named pg-lo as the candidate second fs engine with
POC #7 as its admission evidence — the door is exercised twice). This
crate's roadmap is not blocked on one (see open-questions.md — alkfs
intake may name needs externally; OQ-08).
Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| 003 | Backend contract & dispatch | opaque keys, complete list, pure-function routing, no migration |
| 004 | Two backends, no third | scope boundary against manifest-layer absorption |
| 005 | Namespace-blindness | backends see hashes only |
| 007 | Two kv engines | sqlite (default) + postgres behind one trait; one relational engine + fs per node |
| 008 | size probe + fleet GC + fs engines |
trait length probe; DB-backed pins/advisory-locked sweeper for fleets; fs tier engine-selectable (local default, pg-lo candidate) |
Open Questions
- OQ-08: alkfs requirement intake may name storage requirements (e.g., durability tiers, sync-friendly layouts) that touch this layer — open, external owner (alkfs Phase 0).
References
docs/research/poc-trait-dispatch-findings.mdfindings 1/2/6docs/research/poc-largeblob-findings.mdfindings A2/A4 (+ the re-runnable benchmark harness)docs/research/poc-postgres-kv-findings.md— the pg arm's measured curves + engine-posture deltas (ADR-007's evidence base)docs/research/poc-redb-kv-findings.md— redb ruled out at a durability-tier mismatch (POC #6; recorded in ADR-007)docs/research/iroh-blobs-eval.md— the fs-layout conclusions borrowed (sharding, crash ordering, inline thresholds rejected as weld)- rudolfs notes —
list()anti-lesson, decorator alternative noted and not adopted (threshold dispatch chose the simpler policy; ADR-003 §Context) - ADR-003, ADR-004, ADR-005; store-api.md (the invariants backends must satisfy)