Files
alkblobs/docs/architecture/decisions/011-vocabulary-tiers-engines-and-fleet-state.md
glm-5.3-flash 7b9d904b8a docs(architecture): ADR-012 — pre-decomposition consistency rulings; tree verified single-valued
Full-tree review before task decomposition found five composition
defects (mechanisms specced correctly in isolation, composition
unruled) and a set of caller-facing gaps. ADR-012 rules each:

- Sweeps are never errors: aborts are GcAbortCause report data in
  Ok(SweepReport) (ProtectFailed / SweeperLock / NoLivenessSources);
  GcAborted retired from the error enum; direct-delete refusal is
  the GcRefuse error covering the full protection set
- Pin token gains its wire shape: blobs/put response {token, digest};
  blobs/have's token-renewal form (one digest + token); token
  validity domain = the minting serving node, process-lifetime
  mapping
- Fleet mode is an explicit constructor declaration (fleet: true),
  never inferred from engine choice
- All fleet GC state hosts on the fleet's kv engine (postgres) — one
  arbitration domain; large=pg-lo fleet nodes required onto the same
  pg instance (composite predicate's new clause); large=local fleet
  puts pin-row-first, publish-second
- Engine-state seam reduced: sqlite pin/sweep-lock bodies dropped
  (dead machinery); non-SQL engines stage delete-window candidates
  in-process; one-window-host rule per instance
- Facade clarifications: fall-through for all key-addressed ops,
  kv-only put-time rejection, mem+local dual-tier valid, error-model
  member/return-shape ruling (trait/facade family split), has ->
  bool, PinState variants, fleet liveness-table registration form,
  window executor = the next sweep

Alignment edits across all specs and ADR-005/008/009/010/011
(bracketed corrections per the established pattern); OQ-11 (pg-only-kv
feature graph) added to the parked index for auditability.

Verification: two independent review passes; all findings resolved;
verdict READY for task decomposition.
2026-10-03 07:14:45 +00:00

12 KiB

ADR-011: Vocabulary — tiers, engines, store instances, fleets; mem joins the kv tier; the fs tier is renamed large

Status

Accepted (fleet-mode activation, the GC-state host, the sweep-abort shape, and kv-only's put-time rejection point are ruled by ADR-012 — this ADR's constructor table is closed under ADR-012 §4's host rule; §4's "constructor-time error" wording reads as ADR-012 §6.2's put-time rejection)

Context

The engine story (ADR-007/008/009) grew the physical layer faster than its vocabulary grew, and the pre-decomposition review found the drift is now self-contradictory in the anchor document itself:

  1. Node vs store instance. requirements.md defines Node as "one running store instance" — then REQ-2 states "a replicator node may run multiple store instances (a fleet)" and Fleet as "two or more nodes over one shared pool". Under the written definition, one process hosting two instances is simultaneously one node (by definition) and two nodes (by REQ-2), and the GC arbitration domain ("node is the GC arbitration domain") is undefined for the exact topology ADR-008 exists for. ADR-010 fixes pin ownership; the pin's owner id and the node/instance distinction it needs are still unpinned vocabulary.
  2. The "fs" tier name is now wrong. The tier was named when its only engine was a filesystem (ADR-003's local layout). After ADR-009, half its engines are postgres Large Objects — the name describes one engine's medium, not the tier. Worse, "fs" collides with the dispatch language: "large content goes to fs" never said why — the tier's actual routing criterion is content length (ADR-003's size-threshold dispatch). The docs already speak the true name informally: store-api.md says "small tier"/"large tier".
  3. mem's classification drifts. backends-and-dispatch.md calls it a "backend" in one place and a "testing/utility tier" in another; ADR-003 lists it under "shipped backends … plus a testing tier"; ADR-004 calls it an ephemeral tier and carves it out of the two-backend count. In truth it is neither a tier nor a separate backend: it implements exactly the kv tier's contract (including size, trivially), it is the contract-reference engine every other engine's conformance tests mirror, and "engine" names what it is (ADR-007 established that an engine is a concrete implementation of one tier's contract — requirements.md already defines the word this way).
  4. Tier and Backend were never formally defined even though ADR-004's scope note ("the 'two' here counts tiers, not engines") leans on both words. Also undecided: whether the feature names follow the tier rename now or after code exists — the last free moment is now (zero code; after implementation the rename breaks crates.io/feature ergonomics for every consumer).

Deciding facts are in hand: REQ-2's fleet fact, the engine set's final shape (ADR-007/009), the dispatch mechanism (ADR-003), and ADR-010's GC-state-home table. This is a vocabulary-and-boundary decision, not an architecture change: ADR-004's tier≠engine note already established the shape; this ADR gives it final, consistent names.

Candidates weighed for the tier rename:

  • lo/los — names the postgres engine (Large Objects), so "the lo tier's local engine" would be backwards. Rejected.
  • lfs — workable, echoes git-lfs specifically; but the tier also serves packfiles and any oversized workspace artifact, and the echo binds the crate's naming to one consumer's protocol shape. Rejected.
  • large — the dispatch-true name (size-threshold routing), the word the specs already use informally ("large tier"), medium-agnostic. Chosen. For symmetry in prose, the kv tier stays kv (its routing criterion — "small" — lives in the prose as "small blobs" and needs no rename: kv names the mechanism class, and small as a feature/tier name would read as an adjective; the pair's asymmetry is cosmetic, the load-bearing correction is that neither tier is named after a medium).

Feature-name mirroring: fs → large; pg-lo keeps its name (it names the engine's mechanism, accurate); postgres keeps its name (it names the kv engine); mem keeps its name; kv and ops keep their names.

Decision

1. Vocabulary (the pinned definitions; requirements.md is the canonical copy)

  • Tier — one contract (a Backend trait implementation boundary) serving one routing class of the pool. The crate ships exactly two tiers: kv (small content) and large (over the size-threshold content). Tier count (2) is ADR-004's count.
  • Engine — a concrete storage implementation of one tier's contract (the existing definition, unchanged). Engine count (5) is this ADR's count.
  • Backend — the contract (the Backend trait) plus its shipped engine set, i.e. the tier as shipped. "Backend" and "tier" name the same thing at different altitudes; the crate says "tier" for counting, "the Backend trait" for the contract. "Two backends" reads forever as "two tiers".
  • Store instance — one running instantiation of the store facade (one engine set — one or two engines per the constructor mode, §4 — plus dispatch and GC state). The GC arbitration domain for single-instance topologies; the pin-row owner id under fleet engines (ADR-010's owner; this corrects ADR-008's "node id" wording, which predates the instance/node split — a store instance is what owns pins).
  • Node — one running process (an embedder's binary). A node hosts one or more store instances. Node is the process/ops unit: alkcall connections, ACL grants, sweep cadence, and deployment identity attach to nodes; GC arbitration and pin ownership attach to store instances.
  • Fleet — two or more store instances over one shared pool (via a shared SQL engine). Deliberately not defined by node count: two instances in one process over one postgres pool are a fleet under this definition (and REQ-2's phrasing "a replicator node may run multiple store instances (a fleet)" reads true as written); two processes on two sqlite files are two pools, never a fleet.
  • Deployment — unchanged (one operator's whole system).

2. mem joins the kv tier as its third engine

mem is the kv tier's ephemeral engine (feature mem, default-off): a BTreeMap-shaped implementation of the kv contract, including size (trivially) and list; never fleet-valid (ADR-010 §3's rule and this ADR §4's constructor table); the contract-reference engine — every other engine's conformance tests mirror its suite. It is not a tier, not a third backend, and the ADR-004 count ("two") survives untouched because tier and engine were never the same count — this ADR is that distinction's final form.

3. The fs tier is renamed large

In every document, status line, feature name, and constructor parameter: tier fs → tier large; feature fs → feature large (default-on, as fs was); dispatch prose says "large tier"/"small tier" as store-api.md already does; the local engine's name is kept (its meaning — this node's own filesystem — stays the constructor's meaning even under the shared-media fleet posture, where it is explicitly not only-node-local; the constructor's explicit shared-media assertion, ADR-008 §3, is what disambiguates).

4. The constructor table (the load-bearing copy of the fleet-validity summary; ADR-010 §3 points here)

Tier Engine Feature Default Fleet-valid
kv sqlite — (tier default) yes (tier default-on) no
kv postgres postgres no yes
kv mem mem no no
large local — (tier default) yes (tier default-on) shared media only, explicit assertion
large pg-lo pg-lo no yes

Constructor modes (dispatch shapes) — pinned here (implied by ADR-008 §3's re-routing bullet and the mem tests, but never named):

  • dual-tier (the default): kv + large, size-threshold dispatch between them (ADR-003). Engine pairs beyond the default (kv = mem + large = local) are valid dual-tier constructions — ADR-012 §6.3.
  • kv-only (ADR-008's "kv-tier-only dispatch", now named): a constructor mode with no large tier; over-threshold puts reject at put time — known-length immediately, unknown-length at threshold overflow (the original "constructor-time error" wording is corrected by ADR-012 §6.2: lengths are not known at construction). The client-node re-routing posture handles over-threshold content via the ops surface instead — its composition, not a store mode.
  • mem-only (test/embedder-ephemeral mode): the mem engine alone, no dispatch threshold in effect (everything is local and small); a testing posture, explicitly not a production story.
  • Single-tier SQL modes (postgres-only, pg-lo-only) are not constructor modes this crate ships: the two-tier shape is ADR-004's requirement (both tiers are load-bearing by measured economics); "postgres-only" as in one SQL instance serving both tiers already exists (kv=postgres + large=pg-lo — the ADR-009 consolidation), which is the dual-tier mode over one engine — under fleet mode that one instance is a requirement, not a coincidence (ADR-012 §4).

5. Deferral-policy wording (README alignment)

The deferral policy's categories are: deferred(scope), externally-owned questions, and decided-but-sequenced work — three names, defined in open-questions.md's header. README.md's recap says "two parking kinds" and is corrected to name all three (the third is not a parking kind — it is work, not a parked question — hence this ADR's wording here rather than a silent edit).

Consequences

Positive

  • The anchor document's self-contradiction (node/fleet) becomes a three-level vocabulary (store instance / node / deployment) whose levels map one-to-one onto the mechanisms (GC domain / process-ops unit / operator whole).
  • The fs→large rename removes the last medium-named tier; feature names, tier names, and dispatch prose now agree, before any code exists to break.
  • mem's classification is stated once, as the contract-reference kv engine; the "plus a testing tier" carve-outs dissolve.
  • The constructor table gives task decomposition a closed set: 5 engine implementations, 3 constructor modes, 1 dispatch policy.

Negative

  • Vocabulary churn in every doc that says "fs" (the rename is mechanical but touches all specs and ADR-003/004/007/008/009 status lines and reference sections). Zero code exists — this is the last free moment; after Phase 1 starts, feature renames break consumers.
  • "Backend"/"tier" as interchangeable words risks reader confusion with ADR-003's "Backend trait" — mitigated by the rule that the trait is always named with "trait" ("the Backend trait"), never bare "the backend".

Neutral

  • No engine's contract, feature flag semantics, or constructor parameter changes; mem gains the word "engine" and loses the word "backend".
  • REQ-2's meaning is unchanged (the fleet definition is re-derived from the same fact, not re-fact'd).

References

  • ADR-003 (the trait contract; its "shipped backends" list is now read through this ADR's vocabulary), ADR-004 (tier≠engine — the scope note this ADR gives final names), ADR-007 (kv engines; constructor selection), ADR-008 (the second tier's engine candidate and the fleet topology; the vocabulary it assumed is now pinned), ADR-009 (pg-lo admission; the consolidation option whose constructor mode this ADR names dual-tier), ADR-010 (the I/O seams and GC-state home this vocabulary must fit), ADR-012 (fleet-mode activation + GC-state host + kv-only put-time rejection — the constructor table's final closure)
  • requirements.md (the canonical vocabulary copy, REQ-1..4); backends-and-dispatch.md (the tier/engine table's spec home)
  • docs/research/poc-trait-dispatch-findings.md finding 1 (the reference-impl role mem already played in the POCs); poc-pglo-findings.md (the gate the engine suite mirrors)