Files
alkblobs/docs/architecture/backends-and-dispatch.md
T
glm-5.3-flash cee97defba docs(research): POC #7 — postgres Large Objects as the fs-tier pg-lo engine, passed
The ADR-008 pg-lo admission POC ran in a standalone crate
(/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait
contract (including size), 10/10 exact-count sweep-outcome contract
tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432,
POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions,
no new dependency.

Gate verdict: passed, with named deltas.

- Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and
  below 1 MiB beating — durable local fs on this fsync-slow disk;
  cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16
  readers (20-50x behind page-cache fs — the honest named delta)
- Contract: companion table is the list()/size()/CAS authority (never
  the catalogs); stage-then-commit; GC-participating lo_unlink delete
- Handles: the tx-scoped descriptor is real but pool-compatible via
  descriptorless lo_get(oid, off, len) windows — window gets keep
  handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor
  is the fallback posture
- Vacuum: pg_largeobject pages churn-reused, never returned; tracked
  by autovacuum; rel-size monitoring named as an ops requirement
- Crash/orphan: LO creation is transactional — kill/terminate
  mid-write-tx leaves zero orphan pages; the only orphan class is a
  committed LO bypassing the companion table (planted, reaped by the
  ~7 ms/oid sweep; committed content survives byte-exact)
- Harness lessons: lo_lseek is int4 — the 64 variants are the
  >2 GiB discipline; shared-table parallel tests are unsound (per-test
  CREATE DATABASE isolation)

Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed;
register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated
(duplicate bullet removed) + backends-and-dispatch/open-questions
cross-references.

Verification: cargo test --release (10 passed), clippy -D warnings,
fmt --check in /workspace/alkblobs-pglo-poc.
2026-10-03 04:23:06 +00:00

12 KiB
Raw Blame History

status, last_updated
status last_updated
draft 2026-10-03

Backends and dispatch

What this is

The physical storage layer: the Backend trait contract every backend implements, the two shipped backends (kv, fs), and the size-threshold dispatch that routes between them. The trait is the crate's second stable seam (with the key encoding, ADR-002): backends must survive digest-layer evolution, so the boundary stays dumb.

What the Backend trait is (and is not)

The trait is a CAS-tier contract: it accepts, serves, and enumerates immutable content-addressed entries (hash → bytes), and its list()/ delete() shapes exist to feed GC (ADR-005). That is its whole job, and the contract assumes it: complete-list-equals-sweep-safety only makes sense for a store where every entry is a standalone immutable blob.

It follows that the trait is not a general-purpose storage abstraction, and a downstream must never implement it to host data that is not content-addressed blobs — a manifest table, a queryable index, or any mutable/structured state. Such storage is consumer-owned, beside the pool (a downstream's own sqlite/whatever, ADR-004): it needs no digest-byte addressing, no list()-completeness, and no sweep participation, so the CAS-tier contract fits it worse than it fits any purpose-built option. The wrong turn to rule out is "my data isn't blobs, therefore I implement a Backend for it"; nothing a consumer needs is ever reachable through that door.

The Backend trait contract (ADR-003)

  • Opaque byte keys, opaque byte values. Backends never learn what a digest is; typed Key converts at the store boundary (ADR-002). This is also the persistence boundary — a kv row or an fs filename has no schema-migration story, so the byte layout must be self-describing (algorithm-tagged keys).
  • Methods: has / get / put / delete / list / name / size (the length probe added by ADR-008). Streaming put/get shapes ride the store layer's seams (store-api.md); backends provide byte- or handle-level primitives underneath (get on the fs tier yields a file handle, not a loaded buffer). size(key) returns the entry's length or None if absent — metadata only, never a content fetch: kv engines probe the indexed size column both row shapes carry (POC #5 B5); the local fs engine stats the file. Store-level stat and pool accounting ride this method.
  • list() complete by contract. A malformed list is as deadly as an incomplete one: list correctness is only observable through GC (POC #1 finding 2 — the redb key-vs-value trap produced a sweep that deleted the wrong blobs, silently). Therefore:
    • every backend implementation must prove list correctness through sweep-outcome tests (the invariant's test gate lives in store-api.md);
    • an implementation that cannot enumerate (the rudolfs S3 anti-lesson) must not ship — GC is a structural requirement, not an optional extra (ADR-005).
  • Virgin-store reads are no-ops: read paths treat a missing table as absent/empty (POC #1 finding 2).
  • Namespace-blind: backends never see namespaces, tenants, or reference structure (ADR-005 — the rudolfs inversion; physical storage is hash → bytes flat).

Shipped backends

Two tiers, exactly — this counts trait implementations this crate ships as CAS tiers, the dispatch routes between them; it is the complete set the problem requires, and the "third backend" fear is a category error fixed in ADR-004. It says nothing about other storage existing in a deployment: a downstream runs whatever else it needs on its own media, beside the pool, above the store's seams (see "What the Backend trait is"). Tier count is not engine count: the kv tier carries two engines (sqlite, postgres) behind one trait — ADR-007; per-node engine selection is a constructor parameter.

kv backend (feature kv, default-on); engines: sqlite (default) + postgres (feature postgres, default-off)

Small blobs — kv is the tier name; the engine is a per-node constructor choice between the shipped engines (ADR-007; ADR-003 carries sqlite's on-disk format trade: file compat is upstream sqlite's guarantee). The shipped trait impls write plain content-addressed rows into the chosen engine, nothing else. A third engine remains the substitution door as amended: a new ADR with its own sweep-safety proof (the door is no longer hypothetical — ADR-007 is its first in-repo exercise); and a downstream's other storage — schemas, manifests, whatever it likes including some other sqlite file — never comes near this tier at all, per "What the Backend trait is".

The collapse ADR-007 buys: the classical downstream stack (git server, vfs node) needed three storage systems — kv + fs + relational. Here the kv tier rides the relational engine itself, so a node provisions one relational engine + fs; consumer tables (refs, manifests, queues) sit beside the pool on the same engine or file by the deployment's own choice (co-tenancy note: same-file consumer writes share sqlite's single-writer ceiling — that sharing is the node's trade; separate files remain available).

Evidence (POC #3 finding A4, first-party measured): sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime — most git objects, workspace files, manifests), with the crossover at ~128-256 KiB where fs stops paying the B-tree row rewrite and wins.

sqlite engine (the default): WAL + synchronous=NORMAL as the shipped durability tier (matching what the benchmark measured and what iroh's store ships); bounded reads; has as an EXISTS probe; prepared statements; read paths tolerate the no-tables-yet database (see contract). Solo economics win the single-machine case by measurement; the WAL single-writer ceiling (~1.2k objects/s under contention, POC #5 B3) is an order above single-node push rates.

postgres engine (feature postgres, default-off; ADR-007): the cross-machine-writers / many-client-node / already-running-pg engine. Contract-identical row shape (key bytea PK / value bytea / size), ON CONFLICT DO NOTHING CAS, complete list() as a stable cursor over a vacuuming table, batch-delete via key = ANY($1). Impl requirements per POC #5 finding B5: configurable synchronous_commit posture (the parity knob with the sqlite shipped tier), pooled connections with per-connection prepared-statement discipline, stated autovacuum + fillfactor tuning, ~40× small-tier storage overhead accepted (B6). Sweep-safety proof via the same exact-count sweep-outcome CI gate, run against dockerized postgres under --all-features. Concurrency evidence (POC #5 B3/B4): gets/puts scale near-linearly to ~37k puts/s at 24 conns — the only engine whose throughput increases under concurrency.

fs backend (feature fs, default-on); engines: local (default) + pg-lo (candidate)

Large blobs. Like the kv tier (ADR-007), the fs tier is one contract, multiple engines — the medium is a per-node constructor choice (ADR-008; REQ-2 in requirements.md makes the engine concept load-bearing for fleets):

  • local engine (the default; the ADR-003 backend as originally specced): flat sharded layout: {hex-prefix}/{hex-prefix}/{hash} sharding survives from iroh's conclusion (limits directory size on huge pools); stage-then-commit- rename for the two-pass unknown-length path; pread-based range reads (POC #3 finding A2 — local range serving is sound, e.g. for packfiles).
  • pg-lo engine (candidate, not shipped; ADR-008): postgres Large Objects as the fs tier's storage — the same engine-behind- one-trait move as ADR-007, one tier over. Admission is gated on POC #7's measured evidence (passed 2026-10-03: docs/research/poc-pglo-findings.md — durable put ≈ fs's durable put, lo_get window gets, companion-table authority, clean crash-orphan behavior) and its own engine ADR (see docs/research/poc-pglo-spec.md).
  • Fleet locality contract (ADR-008): over one shared pool, the fs tier is either shared media (every node's local root on the same fleet-shared media — soundness caveats and the required media guarantees in ADR-008) or re-routed (one storage node serves pool large-blob content via the ops surface; client nodes are kv-tier-only over pool content). Mixed per-node-local fs tiers over one pool is the partitioning failure this contract exists to prevent — a documented deployment invariant the constructor cannot fully prove but must be declared against (ADR-008 names the enforceable seam and the detection symptom).

mem backend (feature mem, default-off)

BTreeMap-shaped ephemeral backend for tests and in-process ephemerality. A testing/utility tier, never a production story. Implements the full trait contract including size (trivially) — it exists partly as the reference impl every other backend's conformance tests can mirror.

Size-threshold dispatch (ADR-003)

  • Routing is a pure function of content length. Same content ⇒ same length ⇒ same backend; re-puts are deterministic. No content ever migrates between backends — the migration question existed only to patch the unknown-length asymmetry, which ADR-003's pre-threshold buffering eliminates by construction.
  • Default threshold: 128 KiB (constructor-tunable). Midpoint of the measured flat zone (POC #3 A4). Re-tuning per deployment media is a constructor parameter, not an API change.
  • Get fall-through: small-tier miss queries the large tier (deterministic, cheap — a stat probe).
  • Per-namespace or per-tenant backend configuration: rejected (ADR-003 §Consequences — it would re-weld namespacing into the physical layer, the rudolfs anti-pattern ADR-005 inverts).

Where a new backend or engine could come from

The trait is open to future implementations (network stores, S3-like tiers), but nothing in the current consumer set requires one, and the contract is deliberately hostile to half-implementations (complete list(), GC-participating delete). Any future backend — or a third kv engine or second fs engine — is a new ADR carrying its own sweep-safety story (ADR-007 is the precedent for a kv engine addition; ADR-008 named pg-lo as the candidate second fs engine with POC #7 as its admission evidence — the door is exercised twice). This crate's roadmap is not blocked on one (see open-questions.md — alkfs intake may name needs externally; OQ-08).

Design Decisions

ADR Decision Summary
003 Backend contract & dispatch opaque keys, complete list, pure-function routing, no migration
004 Two backends, no third scope boundary against manifest-layer absorption
005 Namespace-blindness backends see hashes only
007 Two kv engines sqlite (default) + postgres behind one trait; one relational engine + fs per node
008 size probe + fleet GC + fs engines trait length probe; DB-backed pins/advisory-locked sweeper for fleets; fs tier engine-selectable (local default, pg-lo candidate)

Open Questions

  • OQ-08: alkfs requirement intake may name storage requirements (e.g., durability tiers, sync-friendly layouts) that touch this layer — open, external owner (alkfs Phase 0).

References

  • docs/research/poc-trait-dispatch-findings.md findings 1/2/6
  • docs/research/poc-largeblob-findings.md findings A2/A4 (+ the re-runnable benchmark harness)
  • docs/research/poc-postgres-kv-findings.md — the pg arm's measured curves + engine-posture deltas (ADR-007's evidence base)
  • docs/research/poc-redb-kv-findings.md — redb ruled out at a durability-tier mismatch (POC #6; recorded in ADR-007)
  • docs/research/iroh-blobs-eval.md — the fs-layout conclusions borrowed (sharding, crash ordering, inline thresholds rejected as weld)
  • rudolfs notes — list() anti-lesson, decorator alternative noted and not adopted (threshold dispatch chose the simpler policy; ADR-003 §Context)
  • ADR-003, ADR-004, ADR-005; store-api.md (the invariants backends must satisfy)