diff --git a/docs/research/phase-0.md b/docs/research/phase-0.md index c811daf..648fbda 100644 --- a/docs/research/phase-0.md +++ b/docs/research/phase-0.md @@ -1,8 +1,9 @@ --- status: draft -last_updated: 2026-10-01 (dedup/p2p discussion folded in: the consumers -section re-worked around the pooled-CAS insight, new OQ-BL-05, OQ-BL-01 -and OQ-BL-04 refined; un-numbered sections otherwise as the setup draft) +last_updated: 2026-10-01 (GC/pooling round: iroh-blobs mark/sweep verified +against the current checkout, rudolfs prior art added, oid↔hash mapping +dissolved into OQ-BL-03 as per-algorithm namespaces, OQ-BL-05 consolidated +around tags-as-roots; earlier rounds: dedup/p2p consumer round, setup draft) --- # alkblobs — Phase 0 (Exploration) @@ -98,9 +99,12 @@ Guiding principles, inherited from the alk* family: reassembled anywhere particular. Backends behind an injected seam (alktty `TtyBackend` / alktunnels pump-halves precedent). 2. **Hashes are data, not transport identity.** Multiple hash - algorithms must coexist (git SHA-1/SHA-256, BLAKE3, maybe bao-shaped - chunk trees as one encoding among several). How this is abstracted — - trait, enum, per-backend config — is an open question (OQ-BL-03). + algorithms must coexist (git SHA-1/SHA-256, BLAKE3), each with its + own domain-separated derivation ("git-blob-sha256" hashes + `blob \0 + content`, not bare content — OQ-BL-03's resolution); + consumers address the pool by the algorithm their format demands, + no mapping layer. The abstraction shape (trait, enum, per-backend + config) is an open question (OQ-BL-03). 3. **Borrow conclusions, not wire surface.** iroh-blobs' tickets, postcard serialization, and provider protocol are design-welded choices we do not inherit. Its *store-shape* lessons (kv + flat @@ -113,7 +117,8 @@ Guiding principles, inherited from the alk* family: transfer (bao outboard encoding). The alkgit case already has verified content (git objects are hash-addressed by git itself). Which verification story the crate owns — bao trees, per-blob - digests, backend-native — is open (OQ-BL-04). + digests, backend-native, or the rudolfs verify-as-decorator shape — + is open (OQ-BL-04). 6. **Pooling is the point.** The dedup wins both demanding consumers want (Forknet-style multi-repo OSS content, agent swarms) require that repos/workspaces are *sets of hash references* over one pooled @@ -149,6 +154,11 @@ setup: content-addressed store. This is the property both demanding consumers actually need; everything else (GC, manifests, git semantics) lives around it. Details open (OQ-BL-05). +- **Physically flat, logically namespaced** (2026-10-01 GC/pooling + round; the rudolfs inversion): the byte layer is a flat + dedup-by-construction CAS; namespaces are reference tables above it + (GC roots, sweep scoping, ACL boundary). Backends stay + namespace-blind. - **Whole-file CAS is the default granularity**; chunking is scoped to the large-binary-blob case (principle 7), not the default path. @@ -165,6 +175,79 @@ diverge from. **The store eval should be written against the current checkout** — alknet's older research refers to an older iroh-blobs and its conclusions must be re-verified rather than inherited. +**GC mechanics — verified against the current checkout (2026-10-01, +`src/store/gc.rs`, `src/util/temp_tag.rs`, `src/store/fs/delete_set.rs`). +This is the load-bearing prior art for OQ-BL-05:** + +- **Mark-and-sweep, not refcounting, for persistent liveness.** + `gc_mark_task` collects roots from three sources: persistent *named + tags* (a flat `tags-0` table, name → `HashAndFormat`), in-memory * + `TempTag`s* (RAII-refcounted, `#[must_use]`, `.leak()` for + pin-until-exit-of-process), and an injectable `ProtectCb` — a + callback GC consults before each run that can add externally-known + hashes or **abort the run** (`ProtectOutcome::Abort`; a flaky + protection source skips the sweep rather than risking deletion). +- **Format-aware mark traversal** — a non-raw root (HashSeq/collection) + contributes all reachable children to the live set via the bao hash + stream (gc.rs:63-78). Protecting a collection protects everything + reachable from it. +- **Sweep** lists the whole store and batch-deletes (~100/batch) + anything not live. `list()` being complete is therefore load-bearing + for GC — relevant to the backend trait contract (OQ-BL-02). +- **Write-safety against a concurrent sweep** — the fs backend carries + a separate `DeleteSet` transaction layer (`ProtectHandle` / + mark-for-delete / protect-cancel / commit); `TempTag` holds a refcount + so a blob put but not yet referenced cannot vanish mid-write. Temp + tags are scoped (batch scope / process scope) with per-scope counters. +- **No namespaces in their model** — flat named tags over a flat + pool. Our namespace concept (OQ-BL-05) lands on their machinery as + policy over what counts as a root: `namespace → root manifest → + (mark-walk) → {oids}` is tags plus traversal, not new storage + concepts. + +### rudolfs — git-lfs server; the composite-key + decorator-storage prior art + +`/workspace/rudolfs` (v0.3.8, MIT, read-only reference; the alknet +research doc is +`/workspace/@alkdev/alknet/docs/research/references/gitlfs/ +rudolfs-reference.md`). A git-lfs server whose **storage layer** +(`src/storage/`) is the most direct ancestor of our composite-key + +dispatch shape: + +- **`StorageKey = (Namespace, Oid)`** — a composite key over a flat CAS + (`Namespace = (org, project)` strings, `Oid = SHA-256`). All backend + operations take the composite key. +- **`Storage` trait** — `get`/`put`/`size`/`delete`/`list`/` + total_size`/`max_size`/`public_url`/`upload_url`, with + `LFSObject = (len: u64, ByteStream)` — the streaming-first put/get + shape (pinned boxed async byte stream), the natural large-blob API + and likely the channels ops-surface shape. +- **Decorator composition** — `Verify ↔ Encrypted ↔ Cached(LRU → + permanent) ↔ Retrying → s3/disk`. The `Cached` decorator is the + appfile shape in LRU form: "fast store in front of fallback" as a + composable wrapper, an alternative dispatch policy to size + thresholds. `fanout()` duplicates one stream into two lock-step + copies (serve + persist) — precisely "serve the fetch while + persisting on receipt" for networked gets. +- **Verify-as-decorator (OQ-BL-04 input)** — streaming SHA-256 check on + both put and get paths with auto-purge of a corrupted tier. +- **Footnote if at-rest encryption is ever considered** — nonce derived + from the oid: deterministic and correct, but a key rotation + invalidates every object and breaks dedup across keys (the known + encryption-vs-dedup tension). +- **THE ANTI-PATTERN (inverted, not inherited)** — namespaces are + *physical*: `s3://{org}/{project}/{sha256}` — identical content in + two orgs is stored twice. It buys tenant isolation by paying the + cross-tenant dedup that is our whole reason for pooled CAS. The + reconciliation recorded in OQ-BL-05: keep the composite-key API + shape, put the namespace half in metadata (reference tables) above + the byte layer, and keep physical storage flat (`oid → bytes`, + backends namespace-blind). +- **`list()` is load-bearing** — the S3 backend punts on `list()` + (returns an empty stream) and its `delete()` is a no-op, so it can + *never* GC. A lesson for the backend trait contract: list/sweep + support is a real requirement, not an optional extra. + ### gix-odb — git's own object database (alkgit's baseline) `/workspace/git-oxide/gix-odb` (read-only reference). The backend @@ -245,13 +328,40 @@ committing. ### OQ-BL-03: Hash abstraction — trait, enum, or per-backend config? The hard requirement: git SHA-1/SHA-256 and BLAKE3 must coexist (alkgit -conflict). Sub-questions: is the hash algorithm a *store* parameter +conflict). + +**Major resolution (2026-10-01, from the GC/pooling round): the +oid↔hash *mapping* problem dissolves rather than gets solved.** The +assumed shape was "two identifiers per blob (consumer oid, store hash) ++ a mapping table." That table is only necessary if the store insists +on one canonical hash. Since multi-hash is already committed, the +resolution is: **"git's oid derivation" is itself one of the store's +hash algorithms.** A store entry keyed by `git-blob-sha256(content)` *is* +the oid; the consumer (alkgit's odb) addresses the pool directly, zero +indirection. Two consequences: + +- **Algorithms must allow domain separation in their input** — git's + oid hashes `blob \0 + content` (a preamble, and uncompressed + content), so the hash-algorithm abstraction is not a bare `fn(content) + → digest`; an algorithm may define its own preamble/preprocessing. + Per-algorithm namespace tags (`git-sha1`, `git-sha256`, `blake3`) + then coexist in one flat pool collision-free. +- **The mapping problem only reappears for transfer vs storage** — if + content arrives over the network keyed by one hash (e.g. BLAKE3 for + bao verification) and is stored under another (a git oid). The git + case doesn't need this (git objects are self-verifying, so no + transfer-verification layer); whether any consumer ever needs a + cross-hash registration (verify-under-A, store-under-B) is deferred + until a consumer asks. + +Remaining sub-questions: is the hash algorithm a *store* parameter (one algorithm per store instance, chosen by the consumer) or -*per-blob* data (multi-algorithm within one store)? Does verification -(bao trees) get per-algorithm treatment or does bao stay BLAKE3-bound -(git objects don't need bao verification anyway)? This interacts with -OQ-BL-01 and the wire-format question — a wire ADR must follow whatever -this settles. +*per-blob* data (multi-algorithm within one store — the resolution +above implies per-blob, or rather per-algorithm-namespace within one +store)? Does verification (bao trees) get per-algorithm treatment or +does bao stay BLAKE3-bound (git objects don't need bao verification +anyway)? This interacts with OQ-BL-01 and the wire-format question — a +wire ADR must follow whatever this settles. ### OQ-BL-04: Verification and chunking story @@ -277,46 +387,98 @@ git-lfs-shaped large files vs how agent workspaces transfer large artifacts, and whether range reads are a store API or a reassemble-above concern. -### OQ-BL-05: Pooling and GC — namespaces, lifetime, and the pack tension +### OQ-BL-05: Pooling and GC — namespaces as reference sets over a flat CAS -**The decision from the discussion (2026-10-01, principle 6):** one -pooled CAS per node; repos/workspaces are sets of hash references -(manifests/trees) rather than isolated stores — cross-repo dedup by -construction. The open questions are the mechanics: +**The decisions from the discussions (2026-10-01, principles 6; the +GC/pooling round; the rudolfs collision):** -- **Namespace shape** — does the store know namespaces at all - (per-repo/partition prefixes), or is the pool flat with structure - living entirely in the manifests above? Flat is simpler and dedups - unconditionally; namespaces buy cheap garbage identification and - per-consumer isolation (multi-tenant nodes may need it). -- **GC / lifetime** — a pooled store never deletes by default, so - something must decide when a blob dies. Options: refcounting - maintained by put/manifest-write, mark-from-roots over registered - root manifests (replicators know their contract-derived heads - locally — tractable), or lease/expiry schemes. This is a - correctness surface (a GC bug is silent data loss), so it likely - earns a POC or at least a worked example before Phase 1 pins it. -- **The pack tension (alkgit-shaped, constrains the store surface)** — - git packfiles are pack-efficient but blob one object-store-per-pack - (bad for cross-repo dedup); loose-per-object in the pooled store - dedups but is fs-inefficient at git's object volumes. Small git - objects in the kv backend largely dissolve this for the common - case; the residual question is what the store must support for - large/packed content — per-object granularity, range reads into - packed blobs, or both (interacts with OQ-BL-04's range-read - question — likely read them together). +1. **One pooled CAS per node**; repos/workspaces are sets of hash + references (manifests/trees) rather than isolated stores — + cross-repo dedup by construction. +2. **Physically flat, logically namespaced** (the rudolfs inversion). + Physical layer: `hash → bytes`, flat, backends never see the + namespace (`{hex-prefix}/{hex-prefix}/{hash}` sharding survives — + fan-out, not partitioning). Logical layer: `namespace → {hashes}` + reference tables, which simultaneously give GC roots, per-namespace + sweeps, accounting, and the ACL boundary (alkcall `AccessControl` + gates the namespace for network ops). rudolfs' composite + `StorageKey(ns, oid)` API shape is kept; only its physical + namespacing is inverted. +3. **Tags are the root table; references stay above the crate.** + iroh-blobs' named `Tag` (flat name → hash+format table) is the + "an external consumer cares about this hash" concept — namespaces + map onto it as `namespace → root manifest tag → mark-walk → {oids}`. + Git-flavored: refs (`refs/heads/*`) *are* the names; the consumer + registers them as tags. LFS pointer files are the same shape in + another hat — the pointer (oid + size) in the tree *is* the + reference; no extra tag needed because the consumer's own + structure holds it. **Store stays structure-blind**: it does not + walk consumer-side manifests. + +**The GC picture, assembled from verified prior art:** + +- **Mechanism: mark-and-sweep from registered roots** (iroh-blobs' + model, verified — see §Prior art). Roots = registered namespace/root + tags + put-path RAII pins (TempTag-shaped) + a protect callback with + abort semantics (`ProtectOutcome` — protection-source errors skip the + sweep rather than risk deletion). The store does sweep; liveness + beyond "these roots exist" is the consumer's job (it walks its own + manifest format) — the callback seam. +- **Traversal ownership is THE open sub-question.** Iroh can walk + children in-mark because HashSeq is store-native; our + manifests/trees live above the crate, so either **(a) generalized + protect callback** — the layer above computes liveness and hands + hash sets to GC (store stays structure-blind; simplest; liveness + computation is the consumer's), **(b) store-native manifest format** + (a hashseq-like encoding so mark can traverse — cheap GC, but drags + structure into the store layer and creates a wire-ish format), or + **(c) reference-tracker seam** — roots registered with a + children-of(hash) callback; the store does mark/sweep through it. + Tentative lean: **(a) + temp-tag pinning**, revisiting (b) only if + traversal-per-sweep proves expensive (a large multi-tenant node + could pass large hash sets each run; incremental sweeps are an + in-(a) option since the callback is injectable). The three options + are a Phase 1 ADR, decided with the first consumer in sight. +- **Temp-tag pinning is needed regardless** of which option wins: a + blob put but not yet referenced by any namespace must survive + mid-write before any sweep sees it (iroh's `TempTag`/batch-scope + pattern). +- **Per-namespace sweeps** become possible with the reference tables + (delete only within a reclamation scope) — the multi-tenant + refinement the flat iroh model can't express; not needed at minimum. +- **Delete-then-recover semantics** — CAS append-only by nature; a + swept-but-somehow-referenced blob is a re-put, not corruption. The + dangerous direction is only deleting liveness (data loss), which the + abort-callback + pinning layers exist to prevent. + +**Still open (the residual mechanics):** namespace registry shape +(namespaced root tags vs separate reference tables), GC scheduling +(interval + injectable protect vs consumer-driven sweeps), and the pack +tension below. + +**The pack tension (alkgit-shaped, constrains the store surface)** — +git packfiles are pack-efficient but blob one object-store-per-pack +(bad for cross-repo dedup); loose-per-object in the pooled store dedups +but is fs-inefficient at git's object volumes. Small git objects in the +kv backend largely dissolve this for the common case; the residual +question is what the store must support for large/packed content — +per-object granularity, range reads into packed blobs, or both +(interacts with OQ-BL-04's range-read question — likely read them +together). ### OQ-BL-06: POC register (draft) Numbered POCs, opened as research reaches them (findings land in -`docs/research/`; worktree placement per the SDD process): +`docs/research/`; worktree placement per the SDD process). The GC round +reshaped the register — see the discussion at the end for what each POC +is actually deciding: | # | What | Status | Where | |---|------|--------|-------| -| 1 | Backend-trait + dual-dispatch shape (kv small / fs large) | **Pending** — likely first POC | findings file TBD | -| 2 | Multi-hash store (SHA-256 + BLAKE3 coexisting) | **Pending** — rides #1's data | findings file TBD | -| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads) | **Pending** | findings file TBD | -| 4 | Pooled CAS + GC (manifest-referenced blobs, delete-then-recover semantics, refcount vs mark-from-roots) | **Pending** — after #1 (needs a backend to pool over) | findings file TBD | +| 1 | Backend-trait + dual-dispatch shape (kv small / fs large); trait must include `list()` complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path | **Pending** — first POC | findings file TBD | +| 2 | Multi-hash store (git-sha256 + BLAKE3 coexisting in one flat pool, per-algorithm namespaces; oid = a store entry keyed by git's derivation) | **Pending** — rides #1's data | findings file TBD | +| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) | **Pending** — reading-and-design POC | findings file TBD | +| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Pending** — after #1/#2 (needs a backend to pool over) | findings file TBD | Sequencing note: #1 and #2 are probably one worktree (the dispatch POC naturally exercises two hash algorithms); #3 is a reading-and-design @@ -325,6 +487,31 @@ re-verification; #4's GC half matters most — it is a correctness surface and should be a worked example at minimum. The pack tension (OQ-BL-05) rides #3's reading rather than earning its own POC yet. +**What each POC is deciding (2026-10-01, post-GC round):** + +- **POC #1 decides the backend trait's minimum contract** — and the GC + round already sharpened it: `list()` must be complete (rudolfs' + S3 backend punting on list is *why* it can never GC), and the put + path must support not-yet-referenced pinning (TempTag-shaped RAII, + iroh-blobs' batch-scope pattern). The trait is a one-way door; this + POC is where the shape gets evidence rather than instinct. +- **POC #2 validates "oid = a store entry keyed by git's own + derivation"** — the multi-hash resolution in OQ-BL-03 (git-blob + prefix as an algorithm's domain-separated input, per-algorithm + namespaces in one flat pool). If it works, the alkgit odb binds + directly with no mapping layer; if it fights back (e.g. the + preamble/preprocessing abstraction gets ugly), we learn that before + the trait is written, not after. +- **POC #3 is the reading-and-design leg** — against the *current* + iroh-blobs checkout (the store eval) + re-verifying alknet's probe. + Also where the pack tension is analyzed (no independent POC yet). +- **POC #4 is the GC worked example** — the correctness surface of + OQ-BL-05: namespace reference tables + mark-and-sweep with + protect-callback + TempTag pinning + delete-then-recover, exercised + until the failure modes (pin missed, sweep racing a write, live-set + miscomputed) are named rather than theoretical. Expected to be the + POC that actually constrains the trait in return. + POC placement conventions (inherited from alksocks/alktunnels): a POC that needs code from this repo runs in a worktree/branch (`.worktrees/research//` per the SDD process); a