docs(research): consolidate GC/pooling round into phase-0

- iroh-blobs GC mechanics verified against current checkout:
  mark-and-sweep (named tags + TempTag RAII + ProtectCb with abort),
  format-aware mark traversal, DeleteSet write-safety, no namespaces
  (flat tags) — our namespace concept maps onto their machinery
- rudolfs prior art section: StorageKey(ns, oid) composite key +
  decorator storage stack; THE ANTI-PATTERN inverted (physically
  namespaced storage defeats cross-tenant dedup); list() is
  load-bearing; verify-as-decorator; streaming LFSObject + fanout
- OQ-BL-03: oid<->hash mapping DISSOLVED — git's oid derivation is
  itself a hash algorithm (domain-separated preamble); per-algorithm
  namespaces; mapping only reappears for transfer-vs-storage hashes
- OQ-BL-05 consolidated: physically flat/logically namespaced; tags
  are the root table; references stay above the crate; mark-and-sweep
  with protect-callback lean (a) + TempTag pinning; traversal
  ownership (a/b/c) is the Phase 1 ADR
- POC register updated: #1 adds list() completeness + pinning
  contract, #2 reframed as validating oid-dissolution, #4 as the GC
  worked example; 'what each POC decides' section added
This commit is contained in:
glm-5.3-flash committed 2026-10-01 06:23:32 +00:00
1 parent 5e0a594cfd
commit 55450d6bf0
1 file changed
+231 -44
+231 -44
View File
@@ -1,8 +1,9 @@
---
status: draft
last_updated: 2026-10-01 (dedup/p2p discussion folded in: the consumers
section re-worked around the pooled-CAS insight, new OQ-BL-05, OQ-BL-01
and OQ-BL-04 refined; un-numbered sections otherwise as the setup draft)
last_updated: 2026-10-01 (GC/pooling round: iroh-blobs mark/sweep verified
against the current checkout, rudolfs prior art added, oid↔hash mapping
dissolved into OQ-BL-03 as per-algorithm namespaces, OQ-BL-05 consolidated
around tags-as-roots; earlier rounds: dedup/p2p consumer round, setup draft)
---
# alkblobs — Phase 0 (Exploration)
@@ -98,9 +99,12 @@ Guiding principles, inherited from the alk* family:
reassembled anywhere particular. Backends behind an injected seam
(alktty `TtyBackend` / alktunnels pump-halves precedent).
2. **Hashes are data, not transport identity.** Multiple hash
algorithms must coexist (git SHA-1/SHA-256, BLAKE3, maybe bao-shaped
chunk trees as one encoding among several). How this is abstracted —
trait, enum, per-backend config — is an open question (OQ-BL-03).
algorithms must coexist (git SHA-1/SHA-256, BLAKE3), each with its
own domain-separated derivation ("git-blob-sha256" hashes
`blob <len>\0 + content`, not bare content — OQ-BL-03's resolution);
consumers address the pool by the algorithm their format demands,
no mapping layer. The abstraction shape (trait, enum, per-backend
config) is an open question (OQ-BL-03).
3. **Borrow conclusions, not wire surface.** iroh-blobs' tickets,
postcard serialization, and provider protocol are design-welded
choices we do not inherit. Its *store-shape* lessons (kv + flat
@@ -113,7 +117,8 @@ Guiding principles, inherited from the alk* family:
transfer (bao outboard encoding). The alkgit case already has
verified content (git objects are hash-addressed by git itself).
Which verification story the crate owns — bao trees, per-blob
digests, backend-native — is open (OQ-BL-04).
digests, backend-native, or the rudolfs verify-as-decorator shape —
is open (OQ-BL-04).
6. **Pooling is the point.** The dedup wins both demanding consumers
want (Forknet-style multi-repo OSS content, agent swarms) require
that repos/workspaces are *sets of hash references* over one pooled
@@ -149,6 +154,11 @@ setup:
content-addressed store. This is the property both demanding
consumers actually need; everything else (GC, manifests, git
semantics) lives around it. Details open (OQ-BL-05).
- **Physically flat, logically namespaced** (2026-10-01 GC/pooling
round; the rudolfs inversion): the byte layer is a flat
dedup-by-construction CAS; namespaces are reference tables above it
(GC roots, sweep scoping, ACL boundary). Backends stay
namespace-blind.
- **Whole-file CAS is the default granularity**; chunking is scoped
to the large-binary-blob case (principle 7), not the default path.
@@ -165,6 +175,79 @@ diverge from. **The store eval should be written against the current
checkout** — alknet's older research refers to an older iroh-blobs and
its conclusions must be re-verified rather than inherited.
**GC mechanics — verified against the current checkout (2026-10-01,
`src/store/gc.rs`, `src/util/temp_tag.rs`, `src/store/fs/delete_set.rs`).
This is the load-bearing prior art for OQ-BL-05:**
- **Mark-and-sweep, not refcounting, for persistent liveness.**
`gc_mark_task` collects roots from three sources: persistent *named
tags* (a flat `tags-0` table, name → `HashAndFormat`), in-memory *
`TempTag`s* (RAII-refcounted, `#[must_use]`, `.leak()` for
pin-until-exit-of-process), and an injectable `ProtectCb` — a
callback GC consults before each run that can add externally-known
hashes or **abort the run** (`ProtectOutcome::Abort`; a flaky
protection source skips the sweep rather than risking deletion).
- **Format-aware mark traversal** — a non-raw root (HashSeq/collection)
contributes all reachable children to the live set via the bao hash
stream (gc.rs:63-78). Protecting a collection protects everything
reachable from it.
- **Sweep** lists the whole store and batch-deletes (~100/batch)
anything not live. `list()` being complete is therefore load-bearing
for GC — relevant to the backend trait contract (OQ-BL-02).
- **Write-safety against a concurrent sweep** — the fs backend carries
a separate `DeleteSet` transaction layer (`ProtectHandle` /
mark-for-delete / protect-cancel / commit); `TempTag` holds a refcount
so a blob put but not yet referenced cannot vanish mid-write. Temp
tags are scoped (batch scope / process scope) with per-scope counters.
- **No namespaces in their model** — flat named tags over a flat
pool. Our namespace concept (OQ-BL-05) lands on their machinery as
policy over what counts as a root: `namespace → root manifest →
(mark-walk) → {oids}` is tags plus traversal, not new storage
concepts.
### rudolfs — git-lfs server; the composite-key + decorator-storage prior art
`/workspace/rudolfs` (v0.3.8, MIT, read-only reference; the alknet
research doc is
`/workspace/@alkdev/alknet/docs/research/references/gitlfs/
rudolfs-reference.md`). A git-lfs server whose **storage layer**
(`src/storage/`) is the most direct ancestor of our composite-key +
dispatch shape:
- **`StorageKey = (Namespace, Oid)`** — a composite key over a flat CAS
(`Namespace = (org, project)` strings, `Oid = SHA-256`). All backend
operations take the composite key.
- **`Storage` trait** — `get`/`put`/`size`/`delete`/`list`/`
total_size`/`max_size`/`public_url`/`upload_url`, with
`LFSObject = (len: u64, ByteStream)` — the streaming-first put/get
shape (pinned boxed async byte stream), the natural large-blob API
and likely the channels ops-surface shape.
- **Decorator composition** — `Verify ↔ Encrypted ↔ Cached(LRU →
permanent) ↔ Retrying → s3/disk`. The `Cached` decorator is the
appfile shape in LRU form: "fast store in front of fallback" as a
composable wrapper, an alternative dispatch policy to size
thresholds. `fanout()` duplicates one stream into two lock-step
copies (serve + persist) — precisely "serve the fetch while
persisting on receipt" for networked gets.
- **Verify-as-decorator (OQ-BL-04 input)** — streaming SHA-256 check on
both put and get paths with auto-purge of a corrupted tier.
- **Footnote if at-rest encryption is ever considered** — nonce derived
from the oid: deterministic and correct, but a key rotation
invalidates every object and breaks dedup across keys (the known
encryption-vs-dedup tension).
- **THE ANTI-PATTERN (inverted, not inherited)** — namespaces are
*physical*: `s3://{org}/{project}/{sha256}` — identical content in
two orgs is stored twice. It buys tenant isolation by paying the
cross-tenant dedup that is our whole reason for pooled CAS. The
reconciliation recorded in OQ-BL-05: keep the composite-key API
shape, put the namespace half in metadata (reference tables) above
the byte layer, and keep physical storage flat (`oid → bytes`,
backends namespace-blind).
- **`list()` is load-bearing** — the S3 backend punts on `list()`
(returns an empty stream) and its `delete()` is a no-op, so it can
*never* GC. A lesson for the backend trait contract: list/sweep
support is a real requirement, not an optional extra.
### gix-odb — git's own object database (alkgit's baseline)
`/workspace/git-oxide/gix-odb` (read-only reference). The backend
@@ -245,13 +328,40 @@ committing.
### OQ-BL-03: Hash abstraction — trait, enum, or per-backend config?
The hard requirement: git SHA-1/SHA-256 and BLAKE3 must coexist (alkgit
conflict). Sub-questions: is the hash algorithm a *store* parameter
conflict).
**Major resolution (2026-10-01, from the GC/pooling round): the
oid↔hash *mapping* problem dissolves rather than gets solved.** The
assumed shape was "two identifiers per blob (consumer oid, store hash)
+ a mapping table." That table is only necessary if the store insists
on one canonical hash. Since multi-hash is already committed, the
resolution is: **"git's oid derivation" is itself one of the store's
hash algorithms.** A store entry keyed by `git-blob-sha256(content)` *is*
the oid; the consumer (alkgit's odb) addresses the pool directly, zero
indirection. Two consequences:
- **Algorithms must allow domain separation in their input** — git's
oid hashes `blob <len>\0 + content` (a preamble, and uncompressed
content), so the hash-algorithm abstraction is not a bare `fn(content)
→ digest`; an algorithm may define its own preamble/preprocessing.
Per-algorithm namespace tags (`git-sha1`, `git-sha256`, `blake3`)
then coexist in one flat pool collision-free.
- **The mapping problem only reappears for transfer vs storage** — if
content arrives over the network keyed by one hash (e.g. BLAKE3 for
bao verification) and is stored under another (a git oid). The git
case doesn't need this (git objects are self-verifying, so no
transfer-verification layer); whether any consumer ever needs a
cross-hash registration (verify-under-A, store-under-B) is deferred
until a consumer asks.
Remaining sub-questions: is the hash algorithm a *store* parameter
(one algorithm per store instance, chosen by the consumer) or
*per-blob* data (multi-algorithm within one store)? Does verification
(bao trees) get per-algorithm treatment or does bao stay BLAKE3-bound
(git objects don't need bao verification anyway)? This interacts with
OQ-BL-01 and the wire-format question — a wire ADR must follow whatever
this settles.
*per-blob* data (multi-algorithm within one store — the resolution
above implies per-blob, or rather per-algorithm-namespace within one
store)? Does verification (bao trees) get per-algorithm treatment or
does bao stay BLAKE3-bound (git objects don't need bao verification
anyway)? This interacts with OQ-BL-01 and the wire-format question — a
wire ADR must follow whatever this settles.
### OQ-BL-04: Verification and chunking story
@@ -277,46 +387,98 @@ git-lfs-shaped large files vs how agent workspaces transfer large
artifacts, and whether range reads are a store API or a
reassemble-above concern.
### OQ-BL-05: Pooling and GC — namespaces, lifetime, and the pack tension
### OQ-BL-05: Pooling and GC — namespaces as reference sets over a flat CAS
**The decision from the discussion (2026-10-01, principle 6):** one
pooled CAS per node; repos/workspaces are sets of hash references
(manifests/trees) rather than isolated stores — cross-repo dedup by
construction. The open questions are the mechanics:
**The decisions from the discussions (2026-10-01, principles 6; the
GC/pooling round; the rudolfs collision):**
- **Namespace shape** — does the store know namespaces at all
(per-repo/partition prefixes), or is the pool flat with structure
living entirely in the manifests above? Flat is simpler and dedups
unconditionally; namespaces buy cheap garbage identification and
per-consumer isolation (multi-tenant nodes may need it).
- **GC / lifetime** — a pooled store never deletes by default, so
something must decide when a blob dies. Options: refcounting
maintained by put/manifest-write, mark-from-roots over registered
root manifests (replicators know their contract-derived heads
locally — tractable), or lease/expiry schemes. This is a
correctness surface (a GC bug is silent data loss), so it likely
earns a POC or at least a worked example before Phase 1 pins it.
- **The pack tension (alkgit-shaped, constrains the store surface)** —
git packfiles are pack-efficient but blob one object-store-per-pack
(bad for cross-repo dedup); loose-per-object in the pooled store
dedups but is fs-inefficient at git's object volumes. Small git
objects in the kv backend largely dissolve this for the common
case; the residual question is what the store must support for
large/packed content — per-object granularity, range reads into
packed blobs, or both (interacts with OQ-BL-04's range-read
question — likely read them together).
1. **One pooled CAS per node**; repos/workspaces are sets of hash
references (manifests/trees) rather than isolated stores —
cross-repo dedup by construction.
2. **Physically flat, logically namespaced** (the rudolfs inversion).
Physical layer: `hash → bytes`, flat, backends never see the
namespace (`{hex-prefix}/{hex-prefix}/{hash}` sharding survives —
fan-out, not partitioning). Logical layer: `namespace → {hashes}`
reference tables, which simultaneously give GC roots, per-namespace
sweeps, accounting, and the ACL boundary (alkcall `AccessControl`
gates the namespace for network ops). rudolfs' composite
`StorageKey(ns, oid)` API shape is kept; only its physical
namespacing is inverted.
3. **Tags are the root table; references stay above the crate.**
iroh-blobs' named `Tag` (flat name → hash+format table) is the
"an external consumer cares about this hash" concept — namespaces
map onto it as `namespace → root manifest tag → mark-walk → {oids}`.
Git-flavored: refs (`refs/heads/*`) *are* the names; the consumer
registers them as tags. LFS pointer files are the same shape in
another hat — the pointer (oid + size) in the tree *is* the
reference; no extra tag needed because the consumer's own
structure holds it. **Store stays structure-blind**: it does not
walk consumer-side manifests.
**The GC picture, assembled from verified prior art:**
- **Mechanism: mark-and-sweep from registered roots** (iroh-blobs'
model, verified — see §Prior art). Roots = registered namespace/root
tags + put-path RAII pins (TempTag-shaped) + a protect callback with
abort semantics (`ProtectOutcome` — protection-source errors skip the
sweep rather than risk deletion). The store does sweep; liveness
beyond "these roots exist" is the consumer's job (it walks its own
manifest format) — the callback seam.
- **Traversal ownership is THE open sub-question.** Iroh can walk
children in-mark because HashSeq is store-native; our
manifests/trees live above the crate, so either **(a) generalized
protect callback** — the layer above computes liveness and hands
hash sets to GC (store stays structure-blind; simplest; liveness
computation is the consumer's), **(b) store-native manifest format**
(a hashseq-like encoding so mark can traverse — cheap GC, but drags
structure into the store layer and creates a wire-ish format), or
**(c) reference-tracker seam** — roots registered with a
children-of(hash) callback; the store does mark/sweep through it.
Tentative lean: **(a) + temp-tag pinning**, revisiting (b) only if
traversal-per-sweep proves expensive (a large multi-tenant node
could pass large hash sets each run; incremental sweeps are an
in-(a) option since the callback is injectable). The three options
are a Phase 1 ADR, decided with the first consumer in sight.
- **Temp-tag pinning is needed regardless** of which option wins: a
blob put but not yet referenced by any namespace must survive
mid-write before any sweep sees it (iroh's `TempTag`/batch-scope
pattern).
- **Per-namespace sweeps** become possible with the reference tables
(delete only within a reclamation scope) — the multi-tenant
refinement the flat iroh model can't express; not needed at minimum.
- **Delete-then-recover semantics** — CAS append-only by nature; a
swept-but-somehow-referenced blob is a re-put, not corruption. The
dangerous direction is only deleting liveness (data loss), which the
abort-callback + pinning layers exist to prevent.
**Still open (the residual mechanics):** namespace registry shape
(namespaced root tags vs separate reference tables), GC scheduling
(interval + injectable protect vs consumer-driven sweeps), and the pack
tension below.
**The pack tension (alkgit-shaped, constrains the store surface)** —
git packfiles are pack-efficient but blob one object-store-per-pack
(bad for cross-repo dedup); loose-per-object in the pooled store dedups
but is fs-inefficient at git's object volumes. Small git objects in the
kv backend largely dissolve this for the common case; the residual
question is what the store must support for large/packed content —
per-object granularity, range reads into packed blobs, or both
(interacts with OQ-BL-04's range-read question — likely read them
together).
### OQ-BL-06: POC register (draft)
Numbered POCs, opened as research reaches them (findings land in
`docs/research/`; worktree placement per the SDD process):
`docs/research/`; worktree placement per the SDD process). The GC round
reshaped the register — see the discussion at the end for what each POC
is actually deciding:
| # | What | Status | Where |
|---|------|--------|-------|
| 1 | Backend-trait + dual-dispatch shape (kv small / fs large) | **Pending** — likely first POC | findings file TBD |
| 2 | Multi-hash store (SHA-256 + BLAKE3 coexisting) | **Pending** — rides #1's data | findings file TBD |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads) | **Pending** | findings file TBD |
| 4 | Pooled CAS + GC (manifest-referenced blobs, delete-then-recover semantics, refcount vs mark-from-roots) | **Pending** — after #1 (needs a backend to pool over) | findings file TBD |
| 1 | Backend-trait + dual-dispatch shape (kv small / fs large); trait must include `list()` complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path | **Pending** — first POC | findings file TBD |
| 2 | Multi-hash store (git-sha256 + BLAKE3 coexisting in one flat pool, per-algorithm namespaces; oid = a store entry keyed by git's derivation) | **Pending** — rides #1's data | findings file TBD |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) | **Pending** — reading-and-design POC | findings file TBD |
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Pending** — after #1/#2 (needs a backend to pool over) | findings file TBD |
Sequencing note: #1 and #2 are probably one worktree (the dispatch POC
naturally exercises two hash algorithms); #3 is a reading-and-design
@@ -325,6 +487,31 @@ re-verification; #4's GC half matters most — it is a correctness
surface and should be a worked example at minimum. The pack tension
(OQ-BL-05) rides #3's reading rather than earning its own POC yet.
**What each POC is deciding (2026-10-01, post-GC round):**
- **POC #1 decides the backend trait's minimum contract** — and the GC
round already sharpened it: `list()` must be complete (rudolfs'
S3 backend punting on list is *why* it can never GC), and the put
path must support not-yet-referenced pinning (TempTag-shaped RAII,
iroh-blobs' batch-scope pattern). The trait is a one-way door; this
POC is where the shape gets evidence rather than instinct.
- **POC #2 validates "oid = a store entry keyed by git's own
derivation"** — the multi-hash resolution in OQ-BL-03 (git-blob
prefix as an algorithm's domain-separated input, per-algorithm
namespaces in one flat pool). If it works, the alkgit odb binds
directly with no mapping layer; if it fights back (e.g. the
preamble/preprocessing abstraction gets ugly), we learn that before
the trait is written, not after.
- **POC #3 is the reading-and-design leg** — against the *current*
iroh-blobs checkout (the store eval) + re-verifying alknet's probe.
Also where the pack tension is analyzed (no independent POC yet).
- **POC #4 is the GC worked example** — the correctness surface of
OQ-BL-05: namespace reference tables + mark-and-sweep with
protect-callback + TempTag pinning + delete-then-recover, exercised
until the failure modes (pin missed, sweep racing a write, live-set
miscomputed) are named rather than theoretical. Expected to be the
POC that actually constrains the trait in return.
POC placement conventions (inherited from alksocks/alktunnels): a POC
that needs code from this repo runs in a worktree/branch
(`.worktrees/research/<task-id>/` per the SDD process); a