- last_updated annotations containing ': ' ('family composition: ...',
'feature graph re-pinned: ...', 'REQ-5 added: ...') broke Gitea's
frontmatter parser in docs/architecture/README.md,
backends-and-dispatch.md, and requirements.md
- last_updated is now a bare date; the annotation lives in a quoted
notes field across all annotated architecture docs
- convention documented in .opencode/agents/architect.md: bare
last_updated date, quoted notes, quotes required for ': ' scalars
Verification: all docs/**/*.md frontmatter parsed with yaml.safe_load
(16/16 OK); cargo test, clippy --all-targets -D warnings, fmt --check pass
9.8 KiB
status, last_updated, notes
| status | last_updated | notes |
|---|---|---|
| draft | 2026-10-11 | ADR-013 — sweeper-lock semantics pinned as the named exclusive sweep lease; provider unchanged; composition boundary referenced |
Pooling, namespaces, and GC
What this is
The layer that makes the store pooled: one content-addressed pool per node, dedup across every consumer by construction, with garbage collection that is safe against concurrent writes. The pool property is the reason both demanding consumers exist (overview.md); GC is its price, and this document specifies the mechanism the POCs validated (ADR-005).
The pooled CAS
- One pool per deployment; a store instance is one GC arbitration
domain (the engine/vocabulary definitions live in requirements.md
/ ADR-011: a node — one process — may host several instances; the
instance is where pins and arbitration live for single-instance
topologies). Repos, workspaces, and appfile stores are all
sets of hash references over the same CAS — never walled-off
per-repo stores. Git's
alternates/object-pool mechanism (GitLab object pools) is the same idea done awkwardly; the pool generalizes it to unrelated repos and to p2p replication. A fleet (REQ-2) — two or more store instances sharing one pool — is the one topology where the GC mechanism below extends beyond one instance: ADR-008 pins, sweeper locking, and delete arbitration move into the shared engine; the invariants are unchanged — and the sweeper lock's semantics are the named exclusive sweep lease (one sweeper per pool; atomic test-and-set; holdership ends at release/owner death/engine-session loss; contention is theSweeperLockreport abort — ADR-013 §2, shipped provider: the pg advisory lock). Where this crate's GC state (fleet rows and locks) physically lives, and who holds the joint entry+pin transaction, is ruled by ADR-010 §2: the store core owns all GC state; SQL-backed engines host it through a crate-internal, contract-tested seam; backends never learn liveness. - One address space. Under the canonical hash (ADR-002), the same file content has the same address whether it entered via a git oid or a workspace manifest — cross-consumer dedup is free by construction, not a feature.
- Physically flat, logically namespaced (the rudolfs inversion).
The byte layer is
hash → bytes, backends namespace-blind; the namespace is a reference set above the crate:
Namespaces (logical, above the pool)
- A namespace is
namespace → set of root hashesheld by the consumer (git refs are the names; an LFS pointer file in a tree is the reference; an alkfs path tree is the manifest). The store never walks consumer-side manifests — it stays structure-blind. - The store-side counterpart is the liveness seam: consumers
register live-shared liveness sources (ADR-005, §Decision —
register_liveness_source; the POC's copy-semantics failure finding 7 is the pinning reason). Namespace tables may also exist purely in-memory for ephemeral consumers; nothing in the store persists namespaces — where roots live when a deployment has no SQL engine (no kv engine with rows) is a non-question by design: roots live with the consumer, whose references are as durable as they need to be. - ACL mapping for network ops: the namespace is the resource
(
resource_id_pathselects it) — see ops-surface.md.
Mark-and-sweep GC (ADR-005)
Validated mechanism (POC #1 findings 3/7/8 — exact sweep counts, abort semantics, delete-then-recover as byte-identical re-put):
- Liveness sources (roots):
- consumer-registered liveness sources (live-shared, above),
- put-path pins (RAII; batch scope available for multi-put writes — manifest writes are exactly this),
- the protect callback consulted before each sweep: it may add
externally-known hashes or abort the run (report abort —
GcAbortCause::ProtectFailed, nothing deleted; ADR-012 §1: sweeps are never errors; a flaky protection source skips the sweep rather than risk deletion — iroh'sProtectOutcome::Abortconclusion, adopted).
- Sweep: enumerate the whole pool (backends' complete
list(), ADR-003), compute the live set, batch-delete the dead (batch-sized, ~100/batch, iroh's proven shape). Delete-then-recover is a re-put — byte-identical under the same key (validated). Overdue delete-window candidates are re-arbitrated and executed by the sweep that finds them — the executor is the next sweep (ADR-012 §6.7); SQL-backed engines stage candidates durably (engine rows),local/memstage in-process and execute the window inline in the same sweep (ADR-012 §5). - Traversal ownership: lean. Liveness computation beyond "these roots exist" is the consumer's job, handed over via the seam. The store never learns manifest formats. If a real consumer's live-set computation proves too heavy, store-native traversal is a new ADR (documented constraint, not a parked hedge — the un-pause condition is an external fact: a measured consumer cost).
Delete windows (the sweep-vs-put race)
The race (POC #1 finding 3): a pin added between sweep-"list" and
sweep-"delete" is a lost pin — a blob deleted while in flight. The
architecture resolves it as a delete window with three phases —
mark (enumerate + live-set), per-key arbitration under the
pin/liveness lock at delete time, batched commit — resting on two
ordering invariants: put commits pin before publish, and deletes
arbitrate under the pin lock at delete time (full protocol + proof:
ADR-005 §Sweep; iroh delete_set.rs's ProtectHandle/protect-cancel
shape and its serialized-actor pattern are the re-borrowed prior art —
their serialized actor is exactly "arbitrate at delete time"; the
window batches the arbitration).
Test-asserted property: a visible pool entry is never deleted while
liveness (pin or registered source) protects it, and an in-flight put
is never deleted by a sweep started before it committed. Direct
delete(key) refuses protected keys (the typed GcRefuse error; the
full protection set per ADR-005's invariant — ADR-012 §6.4).
has()/get() during a window observe either the old or the new
state; there is no torn observation (the pool's entries are immutable;
existence flips atomically per key).
GC scheduling and scope (ADR-005)
- Explicit
sweep(); embedder owns cadence. No background timers in the crate. An embedder that wants interval sweeps adds them above (one call). Consumer-driven sweeps (after a manifest prune, say) are the same API. - Sweep scoping: whole-pool is the base. Namespace-scoped sweeps run by computing a narrower live set from that namespace's liveness source; the pool remains shared (scoping is a liveness computation choice, not a physical partition).
- Accounting (per-namespace sizes): computed from reference sets by
consumers; the store reports pool-level totals only (
list+stat). Physical-layer accounting would re-introduce namespace physics — rejected with the per-namespace engine config (ADR-003 §Consequences).
Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| 005 | Pooled CAS & GC | one pool, registered liveness, mark-and-sweep, delete windows, no ambient scheduling |
| 003 | list() contract | complete lists are the GC substrate |
| 004 | Structure-blindness | manifests/refs never sink into the store |
| 008 | Fleet GC extension | DB-backed pins, advisory-locked single sweeper, SQL delete arbitration — for shared-pool (fleet) topologies only; the in-process protocol is unchanged elsewhere |
| 010 | GC-state home | the store core owns all GC state; SQL engines host it via the contract-tested engine-state seam; pin-tx ownership ruled |
| 012 | Abort-as-data + window executor | sweep aborts are GcAbortCause report data, never errors; the next sweep executes overdue window rows; non-SQL engines stage in-process; GC state hosts on the fleet's kv engine |
| 013 | Sweep lease + composition boundary | the fleet sweeper lock's semantics = named exclusive sweep lease (provider unchanged: pg advisory lock); reactivity/cadence composition lives above the facade (alkstore the sibling substrate), never in-core |
Open Questions
None owned by this document. Shared-pool GC coordination within a deployment is now in-crate (fleet mechanism: ADR-008). What remains a replicator-policy-layer property (above the crate) is cross-pool coordination — gossip sync between separately-pooled replicator nodes deciding what each keeps — governed by the same seam; it raises a question here only when a consumer names a requirement (OQ-07/OQ-08 intake), which is re-entry via a new decision at a new ADR, not a parked question: no decision here waits on a consumer, and a future requirement does not create one retroactively.
References
docs/research/phase-0.mdOQ-BL-05 (settled decisions + residuals)docs/research/poc-trait-dispatch-findings.mdfindings 3/7/8docs/research/iroh-blobs-eval.md— gc.rs / delete_set.rs reads (the re-borrowed conclusions)- rudolfs — the physical-namespace anti-pattern (inverted)
- gix-odb — alternates/pool prior art for the p2p case
- ADR-005; ADR-012 (the abort-as-data ruling, the window executor, the GC-state host); ADR-013 (§2 the sweep lease; §1 the composition boundary); store-api.md (Pin semantics, sweep API); ops-surface.md (namespace-as-resource ACL)