ADR-010 (new): the Backend trait's I/O seams pinned before any code — 'get' returns a crate-internal read cursor (kv engines materialize; 'local' pread loop; pg-lo lo_get windows), 'put' has two named forms (whole-value kv / staged-put large); GC state is store-core-owned via a crate-internal contract-tested engine-state seam on the SQL-backed engines (backends never learn liveness — ADR-005 verbatim); the store core holds the joint entry+pin transaction (corrects ADR-009 §2's ownership statement); kv-tier fleet-validity rule added (mirror of ADR-008 §3); ops pin-token TTL race resolved (the token IS the pin; renewal rides have/re-put; pause-past-TTL falls to delete-then-recover); feature graph pinned (pg-lo does not imply postgres). ADR-011 (new): vocabulary pinned — tier (contract, exactly two) vs engine (concrete impl, exactly five); store instance / node / fleet split (fleet = pool-sharing, not node count — fixes requirements.md's self-contradiction with REQ-2); mem demoted from 'backend'/'testing tier' to the kv tier's third engine (contract-reference engine); fs tier renamed 'large' including the feature name (last free moment before code exists); constructor modes pinned (dual-tier default, kv-only, mem-only) plus the composite fleet-validity predicate. Consistency round across all specs and ADR-003/004/007/008/009: stale pg-lo passages resolved per ADR-009's supersession; ADR-008 owner 'node id' → store-instance id; threshold wording corrected (conservative edge, not midpoint, of the 128-256 KiB crossover zone); SweepReport shape specced; mem classification unified; 008/009 ADR files renamed to match the tier rename; README deferral-policy recap aligned (third category = decided-but-sequenced work, not a parking kind); POC crate list completed. Verification: two independent architecture review rounds (the first found 4 criticals — unpinned trait I/O shapes, missing fleet-state seam, node/fleet self-contradiction, pin-token/TTL conflict — all resolved; final round: zero criticals); all markdown links resolve; ADR tables complete (11 ADRs).
6.3 KiB
ADR-004: Exactly two backends; manifest layers stay above the crate
Status
Accepted (the count's vocabulary is pinned by ADR-011: the "two" here
are tiers — kv + large, the latter renamed from fs; "backend"
names the same thing at a different altitude. The mem item below is
the kv tier's third engine, per ADR-011 §2, not an extra tier —
the pre-ADR-011 "testing tier" carve-out is superseded. The I/O
seam-level refinements of the trait contract are ADR-010's)
Context
An early Phase 0 framing — "using two backends is ugly" — shaped several rounds before being edited away, but the correction it was reaching for arrived only in the Phase 1 review, and needs recording so no future round reabsorbs the concern incorrectly:
- There are always two physical tiers in this problem: storage for small content (fast, kv/sqlite) and storage for large content (the large tier / filesystem). This is not ugly, it is the measured shape of the problem (POC #3 A4: crossover ~128-256 KiB; SQLite ~9-10× faster below it).
- The "two backends is ugly" sentiment was actually about a third layer in the old alknet-filesystem research — the virtual filesystem / appfile stack: path trees, branches, tombstones, filename↔hash maps living alongside the blob store (in that research, literally a co-equal SQLite store beside iroh-blobs).
If left implicit, this misframing has two failure modes: (a) this crate absorbs a path-tree/manifest layer ("three backends"), growing tables, schemas, and branching semantics that belong to consumers; (b) this crate shapes itself such that a future vfs builder must hack around it. alkfs is a pending consumer precisely of that manifest layer — this ADR is the boundary it will build against.
Decision
This crate ships exactly two production tiers and nothing more;
every manifest/namespace/path-tree layer is a consumer, by construction
and by contract. (A non-production ephemeral mem engine exists for
tests — ADR-003, resolved to a kv-tier engine by ADR-011 §2; it is not
a storage story and does not extend this boundary.)
Scope note (post-ADR-007 reconciliation, vocabulary finalized by
ADR-011): the "two" here counts tiers, not engines. The kv tier is
one contract; ADR-007 widened the engines behind it to a set (sqlite
default + postgres, feature-gated; mem the ephemeral reference engine,
ADR-011) so a node's kv rides its relational engine of choice. The
count that matters to this ADR — trait impls this crate ships as CAS
tiers — remains two (kv + large, the tier ADR-011 renamed from fs);
a second kv engine does not make a third tier, and nothing in the
engine set touches the manifest-layer boundary below.
-
Two production tiers is the complete set. kv/sqlite (small)
- large (the tier formerly named
fs) cover the scale economics with measured evidence (ADR-003); no other physical tier is needed by any current or planned consumer's blob needs.
- large (the tier formerly named
-
The "third backend" is not this crate's layer. The vfs/appfile shape — path→hash mapping, branches/tombstones/deltas per workspace, the alknet-filesystem lineage — is storage above the pool: reference sets over hashes (ADR-005's namespace machinery), not a Backend implementation. Consumers build it on their own storage (alkfs's path tree, alkgit's refs).
-
The store provides the seams the manifest layer needs, and nothing that forces it to hack:
- hash-addressed put/get with a verified, git-interop-compatible addressing scheme (ADR-002) — manifests reference entries by the same canonical addresses git uses;
- liveness registration so manifest-holds maps to GC roots without the store learning manifest formats (ADR-005);
stat/read_range— enough for serving path-resolved content and large-file ranges;- pinned batch puts — manifest-atomic multi-file writes are protected in flight (ADR-005).
Explicit corollary (added after a downstream misread): "consumers build it on their own storage" means a consumer's manifest storage — sqlite or whatever else — lives beside the pool, never behind the
Backendtrait. The trait is a CAS-tier contract (immutable digest-keyed entries, complete list, GC-participating delete); it cannot express manifest storage, so a consumer implementing a Backend to host a manifest was never a live option this boundary forecloses — the boundary forecloses the crate absorbing such layers, and the "no third backend" count is trait impls this crate ships, not storage that may exist in a deployment. -
Namespacing at the store level is logical only (ADR-005): reference tables above, flat bytes below. If a consumer wants physical isolation per tenant/repo, that is two pool deployments (a topology choice), never a namespace flag inside one.
Consequences
Positive
- The crate's scope stays testable and small: two tiers, one pool, one hash domain. No schema/versioning surface for manifest formats — consumers iterate on theirs freely.
- alkgit and alkfs build their (very different) mapping layers on one shared pool without either bending the store; their manifests interop at the address level by construction.
- The alknet-filesystem lineage is honored as shape guidance for consumers (branches, write sessions, chain walks), with zero of its storage-layer mechanics (SQLite path tables, honker wiring, CRDT sync) leaking in here.
Negative
- Consumers each re-build the small amount of manifest plumbing they need (a reference table and a root registration). This is deliberate: the manifest needs of alkgit (git refs) and alkfs (path trees) are different enough that a shared one would fit neither.
Neutral
- If a future consumer did want a first-party manifest layer, it
would be a sibling crate over this one (e.g.
alkfsitself), not an extension of this crate.
References
docs/research/phase-0.md(settled approach: pooled CAS; OQ-BL-05)/workspace/@alkdev/alknet/docs/research/alknet-filesystem/poc-summary.md— the historical lineage this ADR explicitly bounds (not a design input)- ADR-003 (the two backends), ADR-005 (namespaces as reference sets)
- overview.md (consumer map: alkgit, alkfs); backends-and-dispatch.md