From e403b22e2637c57b95fa6d16ee8c0fc6f93300a5 Mon Sep 17 00:00:00 2001 From: "glm-5.3-flash" Date: Fri, 2 Oct 2026 08:13:29 +0000 Subject: [PATCH] docs(architecture): clarify Backend trait scope and kv/sqlite framing MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The backends spec read as if the Backend trait were a general storage abstraction and 'two backends, exactly' banned other storage in a deployment — so a downstream wanting sqlite manifests appeared forced into implementing a third backend, the exact weld ADR-004 forbids. - backends-and-dispatch.md: state plainly that the trait is a CAS-tier contract (immutable digest-keyed entries, GC-shaped list/delete) and never a host for manifest/index storage; consumer storage is beside-the-pool. 'Two, exactly' counts trait impls shipped, not storage per deployment. kv: tier name vs pinned sqlite engine made explicit. - ADR-004: add the explicit corollary under the consumer-seams point. --- docs/architecture/backends-and-dispatch.md | 43 ++++++++++++++++--- .../decisions/004-two-backends-no-third.md | 11 +++++ 2 files changed, 49 insertions(+), 5 deletions(-) diff --git a/docs/architecture/backends-and-dispatch.md b/docs/architecture/backends-and-dispatch.md index b326328..e89d2ca 100644 --- a/docs/architecture/backends-and-dispatch.md +++ b/docs/architecture/backends-and-dispatch.md @@ -1,6 +1,6 @@ --- status: draft -last_updated: 2026-10-01 +last_updated: 2026-10-02 --- # Backends and dispatch @@ -13,6 +13,25 @@ dispatch that routes between them. The trait is the crate's second stable seam (with the key encoding, ADR-002): backends must survive digest-layer evolution, so the boundary stays dumb. +## What the Backend trait is (and is not) + +The trait is a **CAS-tier contract**: it accepts, serves, and enumerates +immutable content-addressed entries (`hash → bytes`), and its `list()`/ +`delete()` shapes exist to feed GC (ADR-005). That is its whole job, and +the contract assumes it: complete-list-equals-sweep-safety only makes +sense for a store where every entry is a standalone immutable blob. + +It follows that the trait is *not* a general-purpose storage abstraction, +and a downstream must never implement it to host data that is not +content-addressed blobs — a manifest table, a queryable index, or any +mutable/structured state. Such storage is **consumer-owned, beside the +pool** (a downstream's own sqlite/whatever, ADR-004): it needs no +digest-byte addressing, no `list()`-completeness, and no sweep +participation, so the CAS-tier contract fits it worse than it fits any +purpose-built option. The wrong turn to rule out is "my data isn't +blobs, therefore I implement a Backend for it"; nothing a consumer needs +is ever reachable through that door. + ## The Backend trait contract (ADR-003) - **Opaque byte keys, opaque byte values.** Backends never learn what a @@ -42,12 +61,26 @@ digest-layer evolution, so the boundary stays dumb. ## Shipped backends -Two, exactly — this is the complete set the problem requires; the -"third backend" fear is a category error fixed in ADR-004. +Two, exactly — this counts *trait implementations this crate ships*, the +CAS tiers the dispatch routes between; it is the complete set the +problem requires, and the "third backend" fear is a category error fixed +in ADR-004. It says nothing about other storage existing in a deployment: +a downstream runs whatever else it needs on its own media, beside the +pool, above the store's seams (see "What the Backend trait is"). -### kv backend (feature `kv`, default-on; sqlite) +### kv backend (feature `kv`, default-on; shipped engine: sqlite) -Small blobs. Evidence (POC #3 finding A4, first-party measured): +Small blobs — **`kv` is the tier name; sqlite is the pinned shipped +engine** (ADR-003's consequences carry that trade: sqlite's on-disk +format compat is upstream's guarantee). The shipped trait impl writes +plain content-addressed rows into it, nothing else. "Pinned engine, not +a kv abstraction" does not open a substitution seam inside this tier: +a downstream wanting a different engine there would be new ADRs (a new +impl, its own sweep-safety proof); and a downstream's *other* storage — +schemas, manifests, whatever it likes including some other sqlite file — +never comes near this tier at all, per "What the Backend trait is". + +Evidence (POC #3 finding A4, first-party measured): sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime — most git objects, workspace files, manifests), with the crossover at ~128-256 KiB where fs stops paying the B-tree row rewrite and wins. diff --git a/docs/architecture/decisions/004-two-backends-no-third.md b/docs/architecture/decisions/004-two-backends-no-third.md index 21d31c2..30c2d5d 100644 --- a/docs/architecture/decisions/004-two-backends-no-third.md +++ b/docs/architecture/decisions/004-two-backends-no-third.md @@ -57,6 +57,17 @@ boundary.) large-file ranges; - pinned batch puts — manifest-atomic multi-file writes are protected in flight (ADR-005). + + Explicit corollary (added after a downstream misread): "consumers + build it on their own storage" means a consumer's manifest storage — + sqlite or whatever else — lives *beside the pool*, never behind the + `Backend` trait. The trait is a CAS-tier contract (immutable + digest-keyed entries, complete list, GC-participating delete); it + cannot express manifest storage, so a consumer implementing a + Backend to host a manifest was never a live option this boundary + forecloses — the boundary forecloses the crate *absorbing* such + layers, and the "no third backend" count is trait impls this crate + ships, not storage that may exist in a deployment. 4. **Namespacing at the store level is logical only** (ADR-005): reference tables above, flat bytes below. If a consumer wants physical isolation per tenant/repo, that is two pool deployments (a