docs(research): consolidate GC/pooling round into phase-0

- iroh-blobs GC mechanics verified against current checkout:
  mark-and-sweep (named tags + TempTag RAII + ProtectCb with abort),
  format-aware mark traversal, DeleteSet write-safety, no namespaces
  (flat tags) — our namespace concept maps onto their machinery
- rudolfs prior art section: StorageKey(ns, oid) composite key +
  decorator storage stack; THE ANTI-PATTERN inverted (physically
  namespaced storage defeats cross-tenant dedup); list() is
  load-bearing; verify-as-decorator; streaming LFSObject + fanout
- OQ-BL-03: oid<->hash mapping DISSOLVED — git's oid derivation is
  itself a hash algorithm (domain-separated preamble); per-algorithm
  namespaces; mapping only reappears for transfer-vs-storage hashes
- OQ-BL-05 consolidated: physically flat/logically namespaced; tags
  are the root table; references stay above the crate; mark-and-sweep
  with protect-callback lean (a) + TempTag pinning; traversal
  ownership (a/b/c) is the Phase 1 ADR
- POC register updated: #1 adds list() completeness + pinning
  contract, #2 reframed as validating oid-dissolution, #4 as the GC
  worked example; 'what each POC decides' section added
This commit is contained in:
glm-5.3-flash committed 2026-10-01 06:23:32 +00:00
1 parent 5e0a594cfd
commit 55450d6bf0
1 file changed
+230 -43
+230 -43
View File
@@ -1,8 +1,9 @@
--- ---
status: draft status: draft
last_updated: 2026-10-01 (dedup/p2p discussion folded in: the consumers last_updated: 2026-10-01 (GC/pooling round: iroh-blobs mark/sweep verified
section re-worked around the pooled-CAS insight, new OQ-BL-05, OQ-BL-01 against the current checkout, rudolfs prior art added, oid↔hash mapping
and OQ-BL-04 refined; un-numbered sections otherwise as the setup draft) dissolved into OQ-BL-03 as per-algorithm namespaces, OQ-BL-05 consolidated
around tags-as-roots; earlier rounds: dedup/p2p consumer round, setup draft)
--- ---
# alkblobs — Phase 0 (Exploration) # alkblobs — Phase 0 (Exploration)
@@ -98,9 +99,12 @@ Guiding principles, inherited from the alk* family:
reassembled anywhere particular. Backends behind an injected seam reassembled anywhere particular. Backends behind an injected seam
(alktty `TtyBackend` / alktunnels pump-halves precedent). (alktty `TtyBackend` / alktunnels pump-halves precedent).
2. **Hashes are data, not transport identity.** Multiple hash 2. **Hashes are data, not transport identity.** Multiple hash
algorithms must coexist (git SHA-1/SHA-256, BLAKE3, maybe bao-shaped algorithms must coexist (git SHA-1/SHA-256, BLAKE3), each with its
chunk trees as one encoding among several). How this is abstracted — own domain-separated derivation ("git-blob-sha256" hashes
trait, enum, per-backend config — is an open question (OQ-BL-03). `blob <len>\0 + content`, not bare content — OQ-BL-03's resolution);
consumers address the pool by the algorithm their format demands,
no mapping layer. The abstraction shape (trait, enum, per-backend
config) is an open question (OQ-BL-03).
3. **Borrow conclusions, not wire surface.** iroh-blobs' tickets, 3. **Borrow conclusions, not wire surface.** iroh-blobs' tickets,
postcard serialization, and provider protocol are design-welded postcard serialization, and provider protocol are design-welded
choices we do not inherit. Its *store-shape* lessons (kv + flat choices we do not inherit. Its *store-shape* lessons (kv + flat
@@ -113,7 +117,8 @@ Guiding principles, inherited from the alk* family:
transfer (bao outboard encoding). The alkgit case already has transfer (bao outboard encoding). The alkgit case already has
verified content (git objects are hash-addressed by git itself). verified content (git objects are hash-addressed by git itself).
Which verification story the crate owns — bao trees, per-blob Which verification story the crate owns — bao trees, per-blob
digests, backend-native — is open (OQ-BL-04). digests, backend-native, or the rudolfs verify-as-decorator shape —
is open (OQ-BL-04).
6. **Pooling is the point.** The dedup wins both demanding consumers 6. **Pooling is the point.** The dedup wins both demanding consumers
want (Forknet-style multi-repo OSS content, agent swarms) require want (Forknet-style multi-repo OSS content, agent swarms) require
that repos/workspaces are *sets of hash references* over one pooled that repos/workspaces are *sets of hash references* over one pooled
@@ -149,6 +154,11 @@ setup:
content-addressed store. This is the property both demanding content-addressed store. This is the property both demanding
consumers actually need; everything else (GC, manifests, git consumers actually need; everything else (GC, manifests, git
semantics) lives around it. Details open (OQ-BL-05). semantics) lives around it. Details open (OQ-BL-05).
- **Physically flat, logically namespaced** (2026-10-01 GC/pooling
round; the rudolfs inversion): the byte layer is a flat
dedup-by-construction CAS; namespaces are reference tables above it
(GC roots, sweep scoping, ACL boundary). Backends stay
namespace-blind.
- **Whole-file CAS is the default granularity**; chunking is scoped - **Whole-file CAS is the default granularity**; chunking is scoped
to the large-binary-blob case (principle 7), not the default path. to the large-binary-blob case (principle 7), not the default path.
@@ -165,6 +175,79 @@ diverge from. **The store eval should be written against the current
checkout** — alknet's older research refers to an older iroh-blobs and checkout** — alknet's older research refers to an older iroh-blobs and
its conclusions must be re-verified rather than inherited. its conclusions must be re-verified rather than inherited.
**GC mechanics — verified against the current checkout (2026-10-01,
`src/store/gc.rs`, `src/util/temp_tag.rs`, `src/store/fs/delete_set.rs`).
This is the load-bearing prior art for OQ-BL-05:**
- **Mark-and-sweep, not refcounting, for persistent liveness.**
`gc_mark_task` collects roots from three sources: persistent *named
tags* (a flat `tags-0` table, name → `HashAndFormat`), in-memory *
`TempTag`s* (RAII-refcounted, `#[must_use]`, `.leak()` for
pin-until-exit-of-process), and an injectable `ProtectCb` — a
callback GC consults before each run that can add externally-known
hashes or **abort the run** (`ProtectOutcome::Abort`; a flaky
protection source skips the sweep rather than risking deletion).
- **Format-aware mark traversal** — a non-raw root (HashSeq/collection)
contributes all reachable children to the live set via the bao hash
stream (gc.rs:63-78). Protecting a collection protects everything
reachable from it.
- **Sweep** lists the whole store and batch-deletes (~100/batch)
anything not live. `list()` being complete is therefore load-bearing
for GC — relevant to the backend trait contract (OQ-BL-02).
- **Write-safety against a concurrent sweep** — the fs backend carries
a separate `DeleteSet` transaction layer (`ProtectHandle` /
mark-for-delete / protect-cancel / commit); `TempTag` holds a refcount
so a blob put but not yet referenced cannot vanish mid-write. Temp
tags are scoped (batch scope / process scope) with per-scope counters.
- **No namespaces in their model** — flat named tags over a flat
pool. Our namespace concept (OQ-BL-05) lands on their machinery as
policy over what counts as a root: `namespace → root manifest →
(mark-walk) → {oids}` is tags plus traversal, not new storage
concepts.
### rudolfs — git-lfs server; the composite-key + decorator-storage prior art
`/workspace/rudolfs` (v0.3.8, MIT, read-only reference; the alknet
research doc is
`/workspace/@alkdev/alknet/docs/research/references/gitlfs/
rudolfs-reference.md`). A git-lfs server whose **storage layer**
(`src/storage/`) is the most direct ancestor of our composite-key +
dispatch shape:
- **`StorageKey = (Namespace, Oid)`** — a composite key over a flat CAS
(`Namespace = (org, project)` strings, `Oid = SHA-256`). All backend
operations take the composite key.
- **`Storage` trait** — `get`/`put`/`size`/`delete`/`list`/`
total_size`/`max_size`/`public_url`/`upload_url`, with
`LFSObject = (len: u64, ByteStream)` — the streaming-first put/get
shape (pinned boxed async byte stream), the natural large-blob API
and likely the channels ops-surface shape.
- **Decorator composition** — `Verify ↔ Encrypted ↔ Cached(LRU →
permanent) ↔ Retrying → s3/disk`. The `Cached` decorator is the
appfile shape in LRU form: "fast store in front of fallback" as a
composable wrapper, an alternative dispatch policy to size
thresholds. `fanout()` duplicates one stream into two lock-step
copies (serve + persist) — precisely "serve the fetch while
persisting on receipt" for networked gets.
- **Verify-as-decorator (OQ-BL-04 input)** — streaming SHA-256 check on
both put and get paths with auto-purge of a corrupted tier.
- **Footnote if at-rest encryption is ever considered** — nonce derived
from the oid: deterministic and correct, but a key rotation
invalidates every object and breaks dedup across keys (the known
encryption-vs-dedup tension).
- **THE ANTI-PATTERN (inverted, not inherited)** — namespaces are
*physical*: `s3://{org}/{project}/{sha256}` — identical content in
two orgs is stored twice. It buys tenant isolation by paying the
cross-tenant dedup that is our whole reason for pooled CAS. The
reconciliation recorded in OQ-BL-05: keep the composite-key API
shape, put the namespace half in metadata (reference tables) above
the byte layer, and keep physical storage flat (`oid → bytes`,
backends namespace-blind).
- **`list()` is load-bearing** — the S3 backend punts on `list()`
(returns an empty stream) and its `delete()` is a no-op, so it can
*never* GC. A lesson for the backend trait contract: list/sweep
support is a real requirement, not an optional extra.
### gix-odb — git's own object database (alkgit's baseline) ### gix-odb — git's own object database (alkgit's baseline)
`/workspace/git-oxide/gix-odb` (read-only reference). The backend `/workspace/git-oxide/gix-odb` (read-only reference). The backend
@@ -245,13 +328,40 @@ committing.
### OQ-BL-03: Hash abstraction — trait, enum, or per-backend config? ### OQ-BL-03: Hash abstraction — trait, enum, or per-backend config?
The hard requirement: git SHA-1/SHA-256 and BLAKE3 must coexist (alkgit The hard requirement: git SHA-1/SHA-256 and BLAKE3 must coexist (alkgit
conflict). Sub-questions: is the hash algorithm a *store* parameter conflict).
**Major resolution (2026-10-01, from the GC/pooling round): the
oid↔hash *mapping* problem dissolves rather than gets solved.** The
assumed shape was "two identifiers per blob (consumer oid, store hash)
+ a mapping table." That table is only necessary if the store insists
on one canonical hash. Since multi-hash is already committed, the
resolution is: **"git's oid derivation" is itself one of the store's
hash algorithms.** A store entry keyed by `git-blob-sha256(content)` *is*
the oid; the consumer (alkgit's odb) addresses the pool directly, zero
indirection. Two consequences:
- **Algorithms must allow domain separation in their input** — git's
oid hashes `blob <len>\0 + content` (a preamble, and uncompressed
content), so the hash-algorithm abstraction is not a bare `fn(content)
→ digest`; an algorithm may define its own preamble/preprocessing.
Per-algorithm namespace tags (`git-sha1`, `git-sha256`, `blake3`)
then coexist in one flat pool collision-free.
- **The mapping problem only reappears for transfer vs storage** — if
content arrives over the network keyed by one hash (e.g. BLAKE3 for
bao verification) and is stored under another (a git oid). The git
case doesn't need this (git objects are self-verifying, so no
transfer-verification layer); whether any consumer ever needs a
cross-hash registration (verify-under-A, store-under-B) is deferred
until a consumer asks.
Remaining sub-questions: is the hash algorithm a *store* parameter
(one algorithm per store instance, chosen by the consumer) or (one algorithm per store instance, chosen by the consumer) or
*per-blob* data (multi-algorithm within one store)? Does verification *per-blob* data (multi-algorithm within one store — the resolution
(bao trees) get per-algorithm treatment or does bao stay BLAKE3-bound above implies per-blob, or rather per-algorithm-namespace within one
(git objects don't need bao verification anyway)? This interacts with store)? Does verification (bao trees) get per-algorithm treatment or
OQ-BL-01 and the wire-format question — a wire ADR must follow whatever does bao stay BLAKE3-bound (git objects don't need bao verification
this settles. anyway)? This interacts with OQ-BL-01 and the wire-format question — a
wire ADR must follow whatever this settles.
### OQ-BL-04: Verification and chunking story ### OQ-BL-04: Verification and chunking story
@@ -277,46 +387,98 @@ git-lfs-shaped large files vs how agent workspaces transfer large
artifacts, and whether range reads are a store API or a artifacts, and whether range reads are a store API or a
reassemble-above concern. reassemble-above concern.
### OQ-BL-05: Pooling and GC — namespaces, lifetime, and the pack tension ### OQ-BL-05: Pooling and GC — namespaces as reference sets over a flat CAS
**The decision from the discussion (2026-10-01, principle 6):** one **The decisions from the discussions (2026-10-01, principles 6; the
pooled CAS per node; repos/workspaces are sets of hash references GC/pooling round; the rudolfs collision):**
(manifests/trees) rather than isolated stores — cross-repo dedup by
construction. The open questions are the mechanics:
- **Namespace shape** — does the store know namespaces at all 1. **One pooled CAS per node**; repos/workspaces are sets of hash
(per-repo/partition prefixes), or is the pool flat with structure references (manifests/trees) rather than isolated stores —
living entirely in the manifests above? Flat is simpler and dedups cross-repo dedup by construction.
unconditionally; namespaces buy cheap garbage identification and 2. **Physically flat, logically namespaced** (the rudolfs inversion).
per-consumer isolation (multi-tenant nodes may need it). Physical layer: `hash → bytes`, flat, backends never see the
- **GC / lifetime** — a pooled store never deletes by default, so namespace (`{hex-prefix}/{hex-prefix}/{hash}` sharding survives —
something must decide when a blob dies. Options: refcounting fan-out, not partitioning). Logical layer: `namespace → {hashes}`
maintained by put/manifest-write, mark-from-roots over registered reference tables, which simultaneously give GC roots, per-namespace
root manifests (replicators know their contract-derived heads sweeps, accounting, and the ACL boundary (alkcall `AccessControl`
locally — tractable), or lease/expiry schemes. This is a gates the namespace for network ops). rudolfs' composite
correctness surface (a GC bug is silent data loss), so it likely `StorageKey(ns, oid)` API shape is kept; only its physical
earns a POC or at least a worked example before Phase 1 pins it. namespacing is inverted.
- **The pack tension (alkgit-shaped, constrains the store surface)** — 3. **Tags are the root table; references stay above the crate.**
iroh-blobs' named `Tag` (flat name → hash+format table) is the
"an external consumer cares about this hash" concept — namespaces
map onto it as `namespace → root manifest tag → mark-walk → {oids}`.
Git-flavored: refs (`refs/heads/*`) *are* the names; the consumer
registers them as tags. LFS pointer files are the same shape in
another hat — the pointer (oid + size) in the tree *is* the
reference; no extra tag needed because the consumer's own
structure holds it. **Store stays structure-blind**: it does not
walk consumer-side manifests.
**The GC picture, assembled from verified prior art:**
- **Mechanism: mark-and-sweep from registered roots** (iroh-blobs'
model, verified — see §Prior art). Roots = registered namespace/root
tags + put-path RAII pins (TempTag-shaped) + a protect callback with
abort semantics (`ProtectOutcome` — protection-source errors skip the
sweep rather than risk deletion). The store does sweep; liveness
beyond "these roots exist" is the consumer's job (it walks its own
manifest format) — the callback seam.
- **Traversal ownership is THE open sub-question.** Iroh can walk
children in-mark because HashSeq is store-native; our
manifests/trees live above the crate, so either **(a) generalized
protect callback** — the layer above computes liveness and hands
hash sets to GC (store stays structure-blind; simplest; liveness
computation is the consumer's), **(b) store-native manifest format**
(a hashseq-like encoding so mark can traverse — cheap GC, but drags
structure into the store layer and creates a wire-ish format), or
**(c) reference-tracker seam** — roots registered with a
children-of(hash) callback; the store does mark/sweep through it.
Tentative lean: **(a) + temp-tag pinning**, revisiting (b) only if
traversal-per-sweep proves expensive (a large multi-tenant node
could pass large hash sets each run; incremental sweeps are an
in-(a) option since the callback is injectable). The three options
are a Phase 1 ADR, decided with the first consumer in sight.
- **Temp-tag pinning is needed regardless** of which option wins: a
blob put but not yet referenced by any namespace must survive
mid-write before any sweep sees it (iroh's `TempTag`/batch-scope
pattern).
- **Per-namespace sweeps** become possible with the reference tables
(delete only within a reclamation scope) — the multi-tenant
refinement the flat iroh model can't express; not needed at minimum.
- **Delete-then-recover semantics** — CAS append-only by nature; a
swept-but-somehow-referenced blob is a re-put, not corruption. The
dangerous direction is only deleting liveness (data loss), which the
abort-callback + pinning layers exist to prevent.
**Still open (the residual mechanics):** namespace registry shape
(namespaced root tags vs separate reference tables), GC scheduling
(interval + injectable protect vs consumer-driven sweeps), and the pack
tension below.
**The pack tension (alkgit-shaped, constrains the store surface)** —
git packfiles are pack-efficient but blob one object-store-per-pack git packfiles are pack-efficient but blob one object-store-per-pack
(bad for cross-repo dedup); loose-per-object in the pooled store (bad for cross-repo dedup); loose-per-object in the pooled store dedups
dedups but is fs-inefficient at git's object volumes. Small git but is fs-inefficient at git's object volumes. Small git objects in the
objects in the kv backend largely dissolve this for the common kv backend largely dissolve this for the common case; the residual
case; the residual question is what the store must support for question is what the store must support for large/packed content —
large/packed content — per-object granularity, range reads into per-object granularity, range reads into packed blobs, or both
packed blobs, or both (interacts with OQ-BL-04's range-read (interacts with OQ-BL-04's range-read question — likely read them
question — likely read them together). together).
### OQ-BL-06: POC register (draft) ### OQ-BL-06: POC register (draft)
Numbered POCs, opened as research reaches them (findings land in Numbered POCs, opened as research reaches them (findings land in
`docs/research/`; worktree placement per the SDD process): `docs/research/`; worktree placement per the SDD process). The GC round
reshaped the register — see the discussion at the end for what each POC
is actually deciding:
| # | What | Status | Where | | # | What | Status | Where |
|---|------|--------|-------| |---|------|--------|-------|
| 1 | Backend-trait + dual-dispatch shape (kv small / fs large) | **Pending** — likely first POC | findings file TBD | | 1 | Backend-trait + dual-dispatch shape (kv small / fs large); trait must include `list()` complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path | **Pending** — first POC | findings file TBD |
| 2 | Multi-hash store (SHA-256 + BLAKE3 coexisting) | **Pending** — rides #1's data | findings file TBD | | 2 | Multi-hash store (git-sha256 + BLAKE3 coexisting in one flat pool, per-algorithm namespaces; oid = a store entry keyed by git's derivation) | **Pending** — rides #1's data | findings file TBD |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads) | **Pending** | findings file TBD | | 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) | **Pending** — reading-and-design POC | findings file TBD |
| 4 | Pooled CAS + GC (manifest-referenced blobs, delete-then-recover semantics, refcount vs mark-from-roots) | **Pending** — after #1 (needs a backend to pool over) | findings file TBD | | 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Pending** — after #1/#2 (needs a backend to pool over) | findings file TBD |
Sequencing note: #1 and #2 are probably one worktree (the dispatch POC Sequencing note: #1 and #2 are probably one worktree (the dispatch POC
naturally exercises two hash algorithms); #3 is a reading-and-design naturally exercises two hash algorithms); #3 is a reading-and-design
@@ -325,6 +487,31 @@ re-verification; #4's GC half matters most — it is a correctness
surface and should be a worked example at minimum. The pack tension surface and should be a worked example at minimum. The pack tension
(OQ-BL-05) rides #3's reading rather than earning its own POC yet. (OQ-BL-05) rides #3's reading rather than earning its own POC yet.
**What each POC is deciding (2026-10-01, post-GC round):**
- **POC #1 decides the backend trait's minimum contract** — and the GC
round already sharpened it: `list()` must be complete (rudolfs'
S3 backend punting on list is *why* it can never GC), and the put
path must support not-yet-referenced pinning (TempTag-shaped RAII,
iroh-blobs' batch-scope pattern). The trait is a one-way door; this
POC is where the shape gets evidence rather than instinct.
- **POC #2 validates "oid = a store entry keyed by git's own
derivation"** — the multi-hash resolution in OQ-BL-03 (git-blob
prefix as an algorithm's domain-separated input, per-algorithm
namespaces in one flat pool). If it works, the alkgit odb binds
directly with no mapping layer; if it fights back (e.g. the
preamble/preprocessing abstraction gets ugly), we learn that before
the trait is written, not after.
- **POC #3 is the reading-and-design leg** — against the *current*
iroh-blobs checkout (the store eval) + re-verifying alknet's probe.
Also where the pack tension is analyzed (no independent POC yet).
- **POC #4 is the GC worked example** — the correctness surface of
OQ-BL-05: namespace reference tables + mark-and-sweep with
protect-callback + TempTag pinning + delete-then-recover, exercised
until the failure modes (pin missed, sweep racing a write, live-set
miscomputed) are named rather than theoretical. Expected to be the
POC that actually constrains the trait in return.
POC placement conventions (inherited from alksocks/alktunnels): a POC POC placement conventions (inherited from alksocks/alktunnels): a POC
that needs code from this repo runs in a worktree/branch that needs code from this repo runs in a worktree/branch
(`.worktrees/research/<task-id>/` per the SDD process); a (`.worktrees/research/<task-id>/` per the SDD process); a