From 395d58cf3fc6d5101726b552aef02eef66c88444 Mon Sep 17 00:00:00 2001 From: "glm-5.3-flash" Date: Fri, 2 Oct 2026 19:29:42 +0000 Subject: [PATCH] =?UTF-8?q?docs(architecture):=20ADR-007=20=E2=80=94=20two?= =?UTF-8?q?=20kv=20engines=20(sqlite=20+=20postgres)=20behind=20one=20trai?= =?UTF-8?q?t?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Dissolve OQ-10's circular "standing offer" framing and record the engine decision the POC evidence already supported: - ADR-007 (new): kv tier ships sqlite (default) + postgres (feature `postgres`, default-off), constructor-selected per node; records the deployment economics (classical kv+fs+relational downstream stack collapses to one relational engine + fs), the pg impl contract (POC #5 B5 posture deltas), the CI sweep-safety gate under --all-features, redb ruled out (POC #6), and posture defaults as deferred cost in the ADR-006 pattern - ADR-003: amended status (engine pin widened by ADR-007); kv entry notes the widened engine set - ADR-004: reconciliation note — the "two backends" count is tiers, not engines - backends-and-dispatch.md: tier≠engine distinction, pg engine section, co-tenancy note (same-file consumer writes share sqlite's single-writer ceiling), substitution-door precedent - open-questions.md: OQ-10 resolved-by-ADR-007, with the why-this-is-not-Schrödinger's-code record (a deployment running alkblobs cannot pre-exist the pg engine — OQ-09 precedent); parked index narrowed to OQ-07/OQ-08 - overview.md: layer map + dependency posture reflect the `postgres` feature; README.md: ADR-007 in tables + corollary applications - phase-0.md / pg+redb findings: promotion + supersession notes Trigger framing survives only as deployment guidance (which engine a node chooses — topology, not throughput), not as a gate on the crate's own work. Verified: cargo test, clippy --all-targets -D warnings, fmt --check. Docs-only change (Phase 1 architecture); no implementation yet. --- docs/architecture/README.md | 16 +- docs/architecture/backends-and-dispatch.md | 89 ++++++--- .../003-backend-contract-and-dispatch.md | 7 +- .../decisions/004-two-backends-no-third.md | 9 + .../007-two-kv-engines-sqlite-and-postgres.md | 177 ++++++++++++++++++ docs/architecture/open-questions.md | 86 ++++----- docs/architecture/overview.md | 16 +- docs/research/phase-0.md | 12 ++ docs/research/poc-postgres-kv-findings.md | 45 +++-- docs/research/poc-redb-kv-findings.md | 12 +- 10 files changed, 369 insertions(+), 100 deletions(-) create mode 100644 docs/architecture/decisions/007-two-kv-engines-sqlite-and-postgres.md diff --git a/docs/architecture/README.md b/docs/architecture/README.md index 9b61aa6..5b293ce 100644 --- a/docs/architecture/README.md +++ b/docs/architecture/README.md @@ -27,6 +27,15 @@ there — not inherited as hedges: two backends (kv small / fs large — both required by measured scale economics); the path→hash mapping layer (the alkfs/alknet "vfs" shape) is a consumer-layer concern and must never become a third backend. +- **ADR-007** widens the kv tier's engine set rather than hedging it: + sqlite (default) + postgres (feature-gated) behind one trait. The + deployment economics this records: a downstream node (git server, vfs) + classically needed three storage systems — kv + fs + relational — and + this design collapses them to **one relational engine + fs**, with the + engine a per-node constructor choice. Resolved from evidence in hand + (POCs #3/#5/#6); the would-be "deployment trigger" framing was caught + at review as circular hedging (a deployment running alkblobs cannot + pre-exist the engine) and closed on the OQ-09 precedent. ## Architecture Documents @@ -50,6 +59,7 @@ there — not inherited as hedges: | [004](decisions/004-two-backends-no-third.md) | Exactly two backends; manifest layers stay above | Accepted | | [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Pooled CAS, liveness seams, mark-and-sweep, delete windows | Accepted | | [006](decisions/006-verification-posture-and-transfer-encoding.md) | Whole-blob verification; transfer encoding excluded by scoping | Accepted | +| [007](decisions/007-two-kv-engines-sqlite-and-postgres.md) | Two kv engines (sqlite + postgres) behind one trait; constructor-selected | Accepted | ## Open Questions @@ -57,7 +67,8 @@ Tracked in [open-questions.md](open-questions.md). The Phase 0 register (OQ-BL-01..06) is promoted there with its resolutions; two questions remain parked — OQ-07/OQ-08 externally-owned (alkgit seam mapping, alkfs intake; carried for visibility, gating nothing here). OQ-09 -(namespace-visibility default) resolved closed-by-default. **Deferral +(namespace-visibility default) resolved closed-by-default; OQ-10 +(second kv engine) resolved by ADR-007. **Deferral policy (the "Schrödinger's code" rule):** a *decision this crate needs before shipping* may not be deferred on a dependency that is itself waiting for this crate to exist. Two parking kinds remain legitimate: @@ -68,7 +79,8 @@ crate) that gate no decision here — full definitions in the header of because it recurs: a fact this crate must create (its own first deployment, its own first consumer) is never a deciding input — decisions stand on evidence in hand, and future needs reopen via named -requirements, not via waiting. +requirements, not via waiting. (Corollary applications so far: OQ-09 +and OQ-10 both resolved on it.) ## Lifecycle diff --git a/docs/architecture/backends-and-dispatch.md b/docs/architecture/backends-and-dispatch.md index e89d2ca..b74dd63 100644 --- a/docs/architecture/backends-and-dispatch.md +++ b/docs/architecture/backends-and-dispatch.md @@ -61,34 +61,66 @@ is ever reachable through that door. ## Shipped backends -Two, exactly — this counts *trait implementations this crate ships*, the -CAS tiers the dispatch routes between; it is the complete set the -problem requires, and the "third backend" fear is a category error fixed -in ADR-004. It says nothing about other storage existing in a deployment: -a downstream runs whatever else it needs on its own media, beside the -pool, above the store's seams (see "What the Backend trait is"). +Two tiers, exactly — this counts *trait implementations this crate +ships* as CAS tiers, the dispatch routes between them; it is the +complete set the problem requires, and the "third backend" fear is a +category error fixed in ADR-004. It says nothing about other storage +existing in a deployment: a downstream runs whatever else it needs on +its own media, beside the pool, above the store's seams (see "What the +Backend trait is"). **Tier count is not engine count**: the kv tier +carries two engines (sqlite, postgres) behind one trait — ADR-007; +per-node engine selection is a constructor parameter. -### kv backend (feature `kv`, default-on; shipped engine: sqlite) +### kv backend (feature `kv`, default-on); engines: sqlite (default) + postgres (feature `postgres`, default-off) -Small blobs — **`kv` is the tier name; sqlite is the pinned shipped -engine** (ADR-003's consequences carry that trade: sqlite's on-disk -format compat is upstream's guarantee). The shipped trait impl writes -plain content-addressed rows into it, nothing else. "Pinned engine, not -a kv abstraction" does not open a substitution seam inside this tier: -a downstream wanting a different engine there would be new ADRs (a new -impl, its own sweep-safety proof); and a downstream's *other* storage — -schemas, manifests, whatever it likes including some other sqlite file — -never comes near this tier at all, per "What the Backend trait is". +Small blobs — **`kv` is the tier name; the engine is a per-node +constructor choice** between the shipped engines (ADR-007; ADR-003 +carries sqlite's on-disk format trade: file compat is upstream +sqlite's guarantee). The shipped trait impls write plain +content-addressed rows into the chosen engine, nothing else. A +*third* engine remains the substitution door as amended: a new ADR +with its own sweep-safety proof (the door is no longer hypothetical — +ADR-007 is its first in-repo exercise); and a downstream's *other* +storage — schemas, manifests, whatever it likes including some other +sqlite file — never comes near this tier at all, per "What the +Backend trait is". + +The collapse ADR-007 buys: the classical downstream stack (git +server, vfs node) needed three storage systems — kv + fs + relational. +Here the kv tier rides the relational engine itself, so a node +provisions **one relational engine + fs**; consumer tables (refs, +manifests, queues) sit beside the pool on the same engine or file by +the deployment's own choice (co-tenancy note: same-file consumer +writes share sqlite's single-writer ceiling — that sharing is the +node's trade; separate files remain available). Evidence (POC #3 finding A4, first-party measured): sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime — most git objects, workspace files, manifests), with the crossover at ~128-256 KiB where fs stops paying the B-tree row rewrite and wins. -- Bounded reads; `has` as an EXISTS probe; prepared statements. -- WAL + `synchronous=NORMAL` as the shipped durability tier (matching - what the benchmark measured and what iroh's store ships). -- Read paths tolerate the no-tables-yet database (see contract). +**sqlite engine** (the default): WAL + `synchronous=NORMAL` as the +shipped durability tier (matching what the benchmark measured and what +iroh's store ships); bounded reads; `has` as an EXISTS probe; prepared +statements; read paths tolerate the no-tables-yet database (see +contract). Solo economics win the single-machine case by measurement; +the WAL single-writer ceiling (~1.2k objects/s under contention, +POC #5 B3) is an order above single-node push rates. + +**postgres engine** (feature `postgres`, default-off; ADR-007): the +cross-machine-writers / many-client-node / already-running-pg engine. +Contract-identical row shape (`key bytea PK / value bytea / size`), +`ON CONFLICT DO NOTHING` CAS, complete `list()` as a stable cursor +over a vacuuming table, batch-delete via `key = ANY($1)`. Impl +requirements per POC #5 finding B5: configurable `synchronous_commit` +posture (the parity knob with the sqlite shipped tier), pooled +connections with per-connection prepared-statement discipline, stated +autovacuum + fillfactor tuning, ~40× small-tier storage overhead +accepted (B6). Sweep-safety proof via the same exact-count +sweep-outcome CI gate, run against dockerized postgres under +`--all-features`. Concurrency evidence (POC #5 B3/B4): gets/puts scale +near-linearly to ~37k puts/s at 24 conns — the only engine whose +throughput *increases* under concurrency. ### fs backend (feature `fs`, default-on) @@ -119,15 +151,17 @@ ephemerality. A testing/utility tier, never a production story. (ADR-003 §Consequences — it would re-weld namespacing into the physical layer, the rudolfs anti-pattern ADR-005 inverts). -## Where a *new* backend could come from +## Where a *new* backend or engine could come from The trait is open to future implementations (network stores, S3-like tiers), but nothing in the current consumer set requires one, and the contract is deliberately hostile to half-implementations (complete -`list()`, GC-participating `delete`). Any future backend is a new ADR -carrying its own sweep-safety story. This crate's roadmap is not -blocked on one (see open-questions.md — alkfs intake may name needs -externally; OQ-08). +`list()`, GC-participating `delete`). Any future backend — or third kv +engine — is a new ADR carrying its own sweep-safety story (ADR-007 is +the precedent for an engine addition: measured evidence first, impl +contract naming the durability/ops posture deltas). This crate's +roadmap is not blocked on one (see open-questions.md — alkfs intake +may name needs externally; OQ-08). ## Design Decisions @@ -136,6 +170,7 @@ externally; OQ-08). | [003](decisions/003-backend-contract-and-dispatch.md) | Backend contract & dispatch | opaque keys, complete list, pure-function routing, no migration | | [004](decisions/004-two-backends-no-third.md) | Two backends, no third | scope boundary against manifest-layer absorption | | [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Namespace-blindness | backends see hashes only | +| [007](decisions/007-two-kv-engines-sqlite-and-postgres.md) | Two kv engines | sqlite (default) + postgres behind one trait; one relational engine + fs per node | ## Open Questions @@ -148,6 +183,10 @@ externally; OQ-08). - `docs/research/poc-trait-dispatch-findings.md` findings 1/2/6 - `docs/research/poc-largeblob-findings.md` findings A2/A4 (+ the re-runnable benchmark harness) +- `docs/research/poc-postgres-kv-findings.md` — the pg arm's measured + curves + engine-posture deltas (ADR-007's evidence base) +- `docs/research/poc-redb-kv-findings.md` — redb ruled out at a + durability-tier mismatch (POC #6; recorded in ADR-007) - `docs/research/iroh-blobs-eval.md` — the fs-layout conclusions borrowed (sharding, crash ordering, inline thresholds rejected as weld) - rudolfs notes — `list()` anti-lesson, decorator alternative noted and diff --git a/docs/architecture/decisions/003-backend-contract-and-dispatch.md b/docs/architecture/decisions/003-backend-contract-and-dispatch.md index 5e031d9..ac0733c 100644 --- a/docs/architecture/decisions/003-backend-contract-and-dispatch.md +++ b/docs/architecture/decisions/003-backend-contract-and-dispatch.md @@ -2,7 +2,9 @@ ## Status -Accepted +Accepted (engine pin amended by ADR-007 — the kv tier ships two +engines, sqlite and postgres, behind this contract; the dispatch layer +and trait are unchanged) ## Context @@ -53,7 +55,8 @@ Evidence in hand (no further POC or consumer-waiting needed): faster than fs at 1-16 KiB (POC #3 A4, first-party benchmark), the regime holding most git objects, workspace files, and manifests. WAL + `synchronous=NORMAL` as shipped durability; prepared - statements; no-tables-yet tolerance. + statements; no-tables-yet tolerance. (The tier's engine set was + widened to include postgres by ADR-007.) - **`fs` (default-on)** for large blobs — flat `{hex-prefix}/{hex-prefix}/{hash}` sharding (iroh's surviving conclusion); stage-then-commit-rename; pread range reads (sound diff --git a/docs/architecture/decisions/004-two-backends-no-third.md b/docs/architecture/decisions/004-two-backends-no-third.md index 30c2d5d..0f6a5e9 100644 --- a/docs/architecture/decisions/004-two-backends-no-third.md +++ b/docs/architecture/decisions/004-two-backends-no-third.md @@ -36,6 +36,15 @@ and by contract.** (A non-production ephemeral `mem` tier exists for tests — ADR-003; it is not a storage story and does not extend this boundary.) +**Scope note (post-ADR-007 reconciliation): the "two" here counts +*tiers*, not engines.** The kv tier is one backend; ADR-007 widened the +engines behind it to a set (sqlite default + postgres, feature-gated) +so a node's kv rides its relational engine of choice. The count that +matters to this ADR — trait impls *this crate ships* as CAS tiers — +remains two (kv + fs); a second kv engine does not make a third +backend, and nothing in the engine set touches the manifest-layer +boundary below. + 1. **Two production backends is the complete set.** kv/sqlite (small) + fs (large) cover the scale economics with measured evidence (ADR-003); no other physical tier is needed by any current or diff --git a/docs/architecture/decisions/007-two-kv-engines-sqlite-and-postgres.md b/docs/architecture/decisions/007-two-kv-engines-sqlite-and-postgres.md new file mode 100644 index 0000000..85374ba --- /dev/null +++ b/docs/architecture/decisions/007-two-kv-engines-sqlite-and-postgres.md @@ -0,0 +1,177 @@ +# ADR-007: Two kv engines — sqlite and postgres — one trait, constructor-selected + +## Status + +Accepted + +## Context + +ADR-003/004 settled the crate's physical shape: two backends (kv for +small blobs, fs for large), size-threshold dispatch between them, and +the "two backends" count as a *tier* count. Within the kv tier, sqlite +was pinned as the shipped engine, on single-connection economics (POC +#3 A4). That pin left a real downstream shape unserved: the classical +stack for the consumers this crate exists for — git servers, vfs/ +appfile nodes, replicators — is *three* storage systems (a kv for small +content, fs for large content, a relational database for refs, +manifests, ACLs, queues). Forcing such a node to run sqlite *beside* +its relational engine multiplies engines, WAL files, backup stories, +and ops postures for no benefit. + +POC #5 (`poc-postgres-kv-findings.md`, findings B1-B6) measured +postgres holding the kv tier's full contract and POC #6 +(`poc-redb-kv-findings.md`, findings C1-C6) ruled out redb: its +durability API cannot express the shipped tier (crash-consistent +without per-commit fsync). Together with POC #3 A4 this is four-way +triangulated engine evidence, complete. + +The deciding facts are all in hand and none is a fact only a future +deployment of this crate could create (the deferral rule — see the +Schrödinger's-code corollary in README.md, and the OQ-09 precedent): + +- **Topology, not throughput, is the trigger** — but topology is a + property of the node the operator builds *today*, not a future + event: a private single-machine node has one writer and sqlite wins + by measurement (44k solo puts/s at 1-16 KiB, POC #3 A4, reproduced + in POC #5); a multi-machine node writing one pool needs postgres + structurally, not merely faster (sqlite's WAL single-writer ceiling + ~1.2k objects/s degrades under contention and flatlines — latency + climbs without bound — while postgres scales near-linearly to ~37k + puts/s across workers, POC #5 B3/B4). +- **The engine cost delta is measured and bounded**: ~180-line impl on + the same table shape (B5), with named posture deltas (see Decision). +- **The sweep-safety gate is testable in CI**, not in production: list + correctness is observable only through GC (POC #1 finding 2), and + the exact-count sweep-outcome test shape runs against any engine the + contract covers — including dockerized postgres under `--all-features`. + +Recording this as an open question gated on "a replicator-shaped +deployment materializes" would be circular hedging: such a deployment +runs alkblobs, so it cannot cross the trigger until the pg engine +ships, and the engine would not ship until the deployment exists. +That deferral shape was already rejected once, on the same reasoning, +in OQ-09. + +## Decision + +**The kv tier ships two engines — sqlite (default) and postgres — +behind one `Backend` trait; engine selection is a constructor +parameter on the backend's configuration.** A node provisions one +relational engine plus the fs tier; the kv tier rides whichever +engine the operator chose. This collapses the classical three-engine +downstream stack (kv + fs + relational) to two, and makes which +relational engine a per-node choice rather than a crate fork. + +1. **sqlite stays the default engine**: zero-ops, no daemon, solo + economics (44k solo puts/s at 1-16 KiB), and consumer co-tenancy + rides the same file (honker's WAL-NORMAL defaults are literally the + shipped tier). Single-machine nodes are sqlite by measurement, and + remain so even at sustained rates well above observed push traffic + (~1.2k/s WAL ceiling vs. single-node push rates an order below it). +2. **postgres ships as an engine, feature-gated (`postgres`, + default-off)**: `tokio-postgres` + `deadpool-postgres`, the POC #5 + arm's stack. It exists for the topology cases measured in POC #5 + B3/B4: cross-machine write access to one pool, many-client + replicator/hub nodes, and consolidated-ops deployments already + running postgres. A postgres node's consumer tables (refs, + manifests, queues) co-tenant the same instance by choice — the + deployment's own consolidation, beside the pool per ADR-004. +3. **The `Backend` trait does not change.** Both engines implement the + identical contract: opaque byte keys/values, CAS semantics + (`INSERT ... ON CONFLICT DO NOTHING` maps sqlite's + `INSERT OR IGNORE`), complete `list()`, GC-participating `delete`, + virgin-store no-op reads, namespace-blindness. The dispatch layer + (size threshold, kv↔fs routing, fall-through) is engine-blind and + unchanged. +4. **The pg engine's impl contract carries POC #5 finding B5's deltas + as requirements** (they are the difference between a port and an + engine posture): + - `synchronous_commit` posture per pool/session, configurable — + the parity knob with sqlite's `synchronous=NORMAL`-WAL shipped + tier (latency moves ~2-3× across settings on the POC disk); + - pooled connections with per-connection prepared-statement + discipline (fresh sessions cost ~20 ms, B2; prepared statements + die with sessions — deadpool's statement cache encodes the + pattern); + - autovacuum tuning stated for the churning CAS table, plus + `fillfactor` (90 measured) — a real maintenance surface sqlite + does not have; + - ~40× storage overhead at the 1-4 KiB tier accepted as the + small-tier cost of a page-based server engine (B6) — documented, + not hidden. +5. **Sweep-safety proof per ADR-003's admission gate**: for pg, `list()` + is a stable `SELECT key` cursor over a vacuuming table (MVCC-ghost + stability across a long sweep), batch-delete via `key = ANY($1)`; + the exact-count sweep-outcome tests (store-api.md invariants 4/5) + run in CI against dockerized postgres under `--all-features`, same + gate shape as every backend. +6. **Non-production posture defaults are named, deferred costs** (the + ADR-006 pattern, not hedges): the B5 tuning numbers are POC-derived + from one dev-box disk; the first real pg deployment verifies them + against its media and records deltas. No decision here waits on + that verification. +7. **redb stays ruled out** (POC #6) — recorded here so the question + does not reopen implicitly: its durability API cannot express the + shipped crash-consistent-without-per-commit-fsync tier, and its + ~40/s write scale-out is below even sqlite's WAL ceiling. + +The ADR-003 substitution door ("a downstream wanting a different +engine there would be new ADRs") is now exercised in-repo, ahead of +any downstream needing it: the downstream's third-engine question +collapses to its own deployment choice among shipped engines. + +## Consequences + +**Positive** + +- The classical downstream stack lands at two engines per node (one + relational + fs), eliminating the "sqlite *beside* the real + database" duplication this crate's consumers would otherwise carry + forever. +- Per-node engine choice (constructor parameter, per ADR-003/004) + makes a heterogeneous fleet — personal nodes on sqlite, multi-tenant + hubs on postgres — a deployment topology, not a fork; everything + above the trait is written once and is engine-blind. +- The evidence for both engines is first-party measured (POC #3 A4, + POC #5 B1-B6, POC #6 C1-C6) and complete; this ADR stands on + evidence in hand, no deferred decision remains in the engine story. +- CI covers the pg backend's contract via the same sweep-outcome test + gate as sqlite, behind `--all-features`. + +**Negative** + +- The pg engine adds a dependency cluster (tokio-postgres, + deadpool-postgres) behind a default-off feature — the base crate + stays lean per convention 8, but the crate now carries a second + engine impl and its posture surface to maintain. +- Postgres storage overhead at the small tier (~40×) is real (B6) and + accepted; deployments dominated by tiny blobs on constrained disks + stay better served by sqlite. +- Two engines means two durability/ops postures to keep honest across + upstream changes (sqlite version/file-compat upstream guarantees; + pg autovacuum behavior is config-dependent). + +**Neutral** + +- Implementation sequencing: sqlite is Phase 1's mainline; the pg impl + is parallelizable against the fixed contract (POC #5 scaffold + exists). No consumer wait gates it; it lands as normal Phase 1 + implementation work. +- A third engine (S3-like, etc.) remains the ADR-003 path: a new ADR + with its own sweep-safety proof. Nothing in this decision closes or + opens that door further. + +## References + +- `docs/research/poc-postgres-kv-findings.md` (POC #5, findings B1-B6) +- `docs/research/poc-redb-kv-findings.md` (POC #6, findings C1-C6) +- `docs/research/poc-largeblob-findings.md` (POC #3, finding A4) +- `docs/research/poc-trait-dispatch-findings.md` (POC #1, finding 2) +- ADR-003 (the trait contract and dispatch this decision rides), + ADR-004 (tier count ≠ engine count; the beside-the-pool boundary), + ADR-006 (the deferred-cost pattern this decision's posture defaults + follow) +- [open-questions.md](../open-questions.md) — OQ-10, resolved by this + ADR; OQ-09 (the anti-circularity precedent) +- [backends-and-dispatch.md](../backends-and-dispatch.md) — the spec + surface this ADR amends \ No newline at end of file diff --git a/docs/architecture/open-questions.md b/docs/architecture/open-questions.md index 1a190d8..ee39081 100644 --- a/docs/architecture/open-questions.md +++ b/docs/architecture/open-questions.md @@ -31,10 +31,10 @@ make before shipping. Phase 0's residual list was swept under this rule: every residual either resolved into an ADR (the evidence was already in hand) or re-owned as below. -**Index of active OQs:** OQ-07, OQ-08, OQ-10 (externally-owned, carried -for visibility). OQ-09 is resolved (recorded below). Promoted Phase 0 -questions OQ-BL-01..06 are recorded here with their resolutions for -traceability. +**Index of active OQs:** OQ-07, OQ-08 (externally-owned, carried +for visibility). OQ-09 and OQ-10 are resolved (recorded below). +Promoted Phase 0 questions OQ-BL-01..06 are recorded here with their +resolutions for traceability. --- @@ -81,51 +81,33 @@ traceability. - **Origin**: `docs/research/poc-postgres-kv-findings.md`, `poc-redb-kv-findings.md`, [backends-and-dispatch.md](backends-and-dispatch.md) -- **Status**: externally-owned (the deployment trigger) — carried here - as a standing offer, not a blocker -- **Priority**: medium-high (the evidence base is done; only the - deployment trigger is outstanding) -- **Owner**: whichever consumer deployment crosses the trigger - (per OQ-07: alkgit's replicator is the named candidate) -- **Question**: when a replicator-shaped deployment materializes - (multi-tenant public node: many users pushing, gossip catch-up, bulk - ingest), open the postgres-kv ADR. The measured case for it: - sqlite's WAL single-writer ceiling is ~1.2k objects/s under - contention and *flatlines* (latency climbs without bound under - overload) while postgres scales out to ~28-37k puts/s across - workers with group-commit amortization (POC #5 findings B3/B4); - single-user/private/small-replicator nodes stay sqlite by - measurement (POC #5 A4/B1). Also settled by the same POCs: redb is - ruled out (POC #6 — its durability API cannot express the shipped - WAL-NORMAL-tier: crash-consistent without per-commit fsync), and - the trigger is **topology, not throughput** — cross-MACHINE write - access to one pool is the condition that makes postgres mandatory - rather than merely faster. -- **Shape of the work when triggered** (per-node choice is a - constructor parameter, per ADR-003/004; the impl is parallelizable - against the fixed contract — two engines, one trait): - 1. `PgKv` backend impl (~180 lines, POC #5 scaffold exists) - with the engine-posture deltas POC #5 finding B5 named: - per-pool `synchronous_commit` posture, pooled-connection + - prepared-statement discipline (prepared statements die with - sessions), autovacuum tuning for a churning CAS table, - fillfactor, ~40× storage overhead accepted as the small-tier - cost; - 2. the sweep-safety proof ADR-003's admission gate requires: - `list()` as a stable `SELECT key` cursor over a vacuuming table - (MVCC-ghost stability across a long sweep), batch-delete via - `key = ANY($1)`, exact-count sweep-outcome tests (the POC #1 - finding-2 gate applied to pg); - 3. consumer-layer note: honker's notify/queue/stream features ride - sqlite's file; the pg deployment's parity stack is - `LISTEN`/`NOTIFY` + pgboss-rs (honker's own README points - postgres users that direction) — a deployment concern, not a - store-crate one. -- **Resolution**: none — waiting on the deployment trigger. The ADR - opens the day a real replicator consumer exists to keep its - sweep-safety proof honest; building it earlier would leave an impl - with no first consumer. -- **Cross-references**: OQ-07, ADR-003, ADR-004, POC #5, POC #6 +- **Status**: resolved — by ADR-007 (2026-10-02) +- **Resolution**: **the kv tier ships two engines — sqlite (default) + and postgres (feature `postgres`, default-off) — behind one + `Backend` trait; engine selection is a per-node constructor + parameter** ([ADR-007](decisions/007-two-kv-engines-sqlite-and-postgres.md)). + The evidence base was already complete (POC #3 A4's sqlite solo + curves; POC #5 B1-B6's pg concurrency scale-out and posture deltas; + POC #6 C1-C6 ruling out redb at a durability-tier mismatch) — all + deciding facts were in hand, and the only thing the prior framing + left outstanding was this crate's own build sequence, which is + implementation sequencing, not an open architecture question. +- **Why this is not Schrödinger's code** (recorded because the + question was once framed as an externally-owned standing offer): the + proposed trigger — "a replicator-shaped deployment materializes" — + is a fact only this crate could create (a deployment *running + alkblobs* capable of pg cannot pre-exist the pg engine), so gating + on it was circular hedging, the exact corollary the README names and + the precedent OQ-09 resolved on. The topology-not-throughput + decision rule and the measured curves stand on evidence in hand; + the deployment is where the *costs* named by the ADR (POC-derived + posture defaults) get verified — a deferred cost in the ADR-006 + pattern, not a gating input. +- **Reopen condition** (not a parked question): a third engine is the + ADR-003/007 substitution door — a new ADR with its own measured + evidence and sweep-safety proof. "Some future deployment might want + it" does not open it. +- **Cross-references**: OQ-07, ADR-003, ADR-004, ADR-007, POC #3, POC #5, POC #6 --- @@ -236,13 +218,17 @@ traceability. |---|---|---| | OQ-07 | externally-owned | alkgit's architecture process (answerable on paper anytime; gates nothing here) | | OQ-08 | externally-owned | alkfs Phase 0 intake | -| OQ-10 | externally-owned (standing offer) | the first multi-tenant replicator deployment (POC #5 is the ready evidence base; the ADR opens on the trigger) | OQ-09 was `deferred(scope)` on the first embedded ops deployment — a deciding fact that could only exist once the ops module ships, i.e. a wait that never collapses. It resolved to closed-by-default (see its entry above); its tracker task is deleted. +OQ-10 was briefly recorded as an externally-owned "standing offer" on +a deployment trigger — the same invalid shape, caught at review: the +trigger was a fact only this crate could create. It resolved to +ADR-007 (two kv engines; see its entry above). + OQ-07/OQ-08 are owned by other repos' processes and are not alkblobs tracker tasks (their outcome arrives *through* their owners, not through any artifact this repo creates). \ No newline at end of file diff --git a/docs/architecture/overview.md b/docs/architecture/overview.md index 77f9c1c..5ad67dd 100644 --- a/docs/architecture/overview.md +++ b/docs/architecture/overview.md @@ -60,7 +60,10 @@ or three times. │ size-threshold dispatch (ADR-003) │ └───────┬───────────────────┬───────────────────────────────┘ ▼ ▼ - kv backend fs backend (ADR-003: exactly two) + kv backend fs backend (ADR-003: exactly two tiers; + │ ADR-007: kv = 2 engines) + ▼ + engine per node: sqlite (default) | postgres (feature "postgres") transport: none, anywhere. alkgit/alkfs or the embedder dials; an alkcall Connection is handed in at the ops boundary. @@ -71,13 +74,15 @@ or three times. | Layer | Dependencies | Notes | |---|---|---| | base crate (store core + backends) | `tokio`, `thiserror`, `parking_lot`, plus `rusqlite` behind the `kv` feature and nothing behind `fs`/`mem` | no alkcall, no serialization frameworks | +| `postgres` feature (ADR-007) | adds `tokio-postgres` + `deadpool-postgres` behind the kv tier's second engine, default-off | same `Backend` contract, own sweep-safety CI gate under `--all-features` | | `ops` feature | adds `alkcall` (feature-gated, default-off) | JSON op specs + binary channels + `AccessControl`; ADR-001 | | consumer crates | depend on alkblobs (and on alkcall directly when they speak ops) | alkgit, alkfs | Backends are feature-gated; both `kv` and `fs` ship default-on -(batteries-included, the alkgit feature-model pattern); the `ops` -feature is default-off. Wasm: the backends are not a wasm story; a wasm -client is a consumer of the `ops` surface (which rides alkcall, itself +(batteries-included, the alkgit feature-model pattern); the `kv` tier's +postgres engine (`postgres`, ADR-007) and the `ops` feature are +default-off. Wasm: the backends are not a wasm story; a wasm client is a +consumer of the `ops` surface (which rides alkcall, itself wasm-clean), never an in-process embedder (ADR-001 §Consequences). ## Non-goals (each enforced by an ADR) @@ -105,6 +110,7 @@ wasm-clean), never an in-process embedder (ADR-001 §Consequences). | [004](decisions/004-two-backends-no-third.md) | Two backends, no third | manifest layers are consumer-side | | [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Pooled CAS & GC | flat pool, registered liveness, delete windows | | [006](decisions/006-verification-posture-and-transfer-encoding.md) | Verification posture | whole-blob verification; chunk trees excluded by scoping | +| [007](decisions/007-two-kv-engines-sqlite-and-postgres.md) | Two kv engines | sqlite default + postgres behind one trait; one relational engine + fs per node | ## Open Questions @@ -117,6 +123,8 @@ crate level: - **OQ-09**: default namespace visibility in network ops — resolved closed-by-default (deny-unlisted; reopen only on a named open-read requirement) +- **OQ-10**: second kv engine — resolved by ADR-007 (sqlite + postgres + shipped behind one trait; per-node constructor choice) ## References diff --git a/docs/research/phase-0.md b/docs/research/phase-0.md index e4f3ced..9be0a11 100644 --- a/docs/research/phase-0.md +++ b/docs/research/phase-0.md @@ -617,6 +617,18 @@ self-contained POC runs as a standalone crate in the global workspace with findings written into `docs/research/` here. Findings always land in `docs/research/` regardless of where the code lives. +Post-convergence promotion (2026-10-02, Phase 1): #5's pg engine case +and #6's redb guardrail became an accepted decision — ADR-007 (two kv +engines: sqlite default + postgres feature-gated, behind one trait, +constructor-selected; "two backends" counts tiers, not engines). An +interim framing recorded in OQ-10 (postgres as an externally-owned +"standing offer" opening on a first multi-tenant replicator +deployment) was caught at review as circular hedging — the trigger +was a fact only this crate could create — and dissolved on the same +reasoning as the Schrödinger's-code rule's README corollary (the +OQ-09 precedent). The POC evidence needed no strengthening; only the +decision's recording did. + ## Convergence Phase 0's final step per `docs/sdd_process.md`: *converge on a diff --git a/docs/research/poc-postgres-kv-findings.md b/docs/research/poc-postgres-kv-findings.md index f5fd824..5d7fe18 100644 --- a/docs/research/poc-postgres-kv-findings.md +++ b/docs/research/poc-postgres-kv-findings.md @@ -238,12 +238,14 @@ a deployment running the *fs tier anyway* has none of either cost. differences (synchronous_commit posture, pooling discipline, autovacuum surface, ~40× storage floor) are the ADR content any postgres Backend impl would carry. -- **No ADR action recommended from a single POC.** The decision is - Phase 1's when a replicator-shaped consumer exists to measure against; - this POC's register entry is the evidence input, not a recommendation. - SQLite stays the pinned shipped engine (single-node economics, no - daemon to run, zero-ops); postgres is the *named future ADR* if the - multi-client replicator shape materializes. +- **Decision status (as originally recorded): no ADR action from a + single POC at the time of writing** — sqlite stayed the pinned + engine, postgres the named candidate. *Amended 2026-10-02:* the + decision has since been made on the accumulated evidence of this + POC + POC #6 (with POC #3's A4 sqlite baseline) — ADR-007 ships + postgres as the kv tier's second engine (feature-gated, + constructor-selected); this section's measured content above stands + as the ADR's evidence base. - **OQ-BL-02 residuals**: unchanged in substance. The one new residual: if a postgres Backend were adopted, its GC-sweep implementation is `TRUNCATE`-friendly (whole-table dead-set rebuilds) but the @@ -253,7 +255,8 @@ a deployment running the *fs tier anyway* has none of either cost. ## Engine deployment posture (recorded 2026-10-02 — the two-engine story, post-POC #4/#6) The three alkblobs-relevant git/vfs use cases, mapped onto the measured -curves; this is the conclusion OQ-10 carries: +curves; this was the conclusion OQ-10 carried before its resolution +in ADR-007): | Engine | Covers | Evidence | |---|---|---| @@ -278,11 +281,29 @@ that are engine-blind. Work-above-the-trait (store core, GC mechanism, ops, consumers) is written once; engine work parallelizes against the fixed contract (two engines, one trait, same sweep-safety proof shape). -Sequencing (recorded in OQ-10): sqlite stays Phase 1's mainline; the pg -ADR opens on the deployment trigger, not before — building it earlier -would leave an impl with no first consumer to keep the sweep-safety -proof honest. The evidence base (this doc + POC #6's redb guardrail) is -complete and waiting. +The deployment economics paragraph this table encodes, now stated +outright (recorded in ADR-007): the classical downstream stack — a git +server, a vfs node — needs three storage systems (kv + fs + relational); +this design collapses the kv into the relational engine, so a node +provisions **one relational engine + fs**, and which relational engine +is the operator's per-node choice among the shipped engines. + +Sequencing (amended 2026-10-02 by ADR-007): postgres **ships** as the +kv tier's second engine (feature `postgres`, default-off), +constructor-selected per node — sqlite stays the default and Phase 1's +mainline; the pg impl is normal Phase 1 implementation work, +parallelizable against the fixed contract (two engines, one trait, +same sweep-safety proof shape, CI-gated via `--all-features`). The +earlier framing here and in OQ-10 — a "standing offer" whose ADR +opens on a first multi-tenant replicator deployment — was caught at +review as circular hedging: such a deployment runs alkblobs, so it +cannot cross the trigger until the pg engine ships, which the trigger +then forbids shipping. Superseded by +[ADR-007](../architecture/decisions/007-two-kv-engines-sqlite-and-postgres.md), +which records the decision this POC's evidence supports; the +deployment-trigger rule survives in corrected form as deployment +guidance (which engine a *node* chooses — ADR-007's decision rule), +not as a gate on the crate's own work. ## POC quality notes diff --git a/docs/research/poc-redb-kv-findings.md b/docs/research/poc-redb-kv-findings.md index 1f6b183..e77ee4e 100644 --- a/docs/research/poc-redb-kv-findings.md +++ b/docs/research/poc-redb-kv-findings.md @@ -205,11 +205,13 @@ go up. - **Reads: no case for redb on merit** — page-cache reads are what any of these give the large tier for free; sqlite's 19 µs probe was already fine. -- **ADR action: none.** SQLite's pin now has four-way triangulated - evidence (POC #3 fs, POC #4 pg, POC #6 redb, plus the lineage - citation). If a future consumer wanted redb-shaped durability - trade-offs (ephemeral cache tier), that's a *different* tier than - the kv engine slot. +- **Anti-recommendation adopted (ADR-007, 2026-10-02): redb stays + ruled out** — recorded in the decision so the question does not + reopen implicitly. SQLite's engine pin now has four-way triangulated + evidence (POC #3 fs, POC #4/#5 pg, POC #6 redb, plus the lineage + citation), and ADR-007 ships the kv tier as sqlite + postgres. If a + future consumer wanted redb-shaped durability trade-offs (ephemeral + cache tier), that's a *different* tier than the kv engine slot. ## POC quality notes