diff --git a/docs/research/phase-0.md b/docs/research/phase-0.md index 7538cc1..c811daf 100644 --- a/docs/research/phase-0.md +++ b/docs/research/phase-0.md @@ -1,8 +1,8 @@ --- status: draft -last_updated: 2026-09-30 (initial setup draft: vision, prior art, open -questions from the setup discussion; un-numbered — renumber -OQ-BL-01..NN as the register solidifies) +last_updated: 2026-10-01 (dedup/p2p discussion folded in: the consumers +section re-worked around the pooled-CAS insight, new OQ-BL-05, OQ-BL-01 +and OQ-BL-04 refined; un-numbered sections otherwise as the setup draft) --- # alkblobs — Phase 0 (Exploration) @@ -27,7 +27,7 @@ from any transport/protocol layer and tolerant of hash-algorithm conflicts (the alkgit case), with auth-gated network operations riding the alkcall seam when they exist. -**Why this crate exists (two converging consumers):** +**Why this crate exists (three converging consumers):** 1. **alkgit** (planning phase) needs an object backend that is *not* its proposed default file-based backend ("kind of gross") — blobs @@ -46,6 +46,36 @@ the alkcall seam when they exist. mapping. alknet's own research hit the core awkwardness of dispatching across more than one backend at a time — which is exactly a store-layer problem this crate should own. +3. **Agent workspaces** — many agents working simultaneously in + workspaces that are mostly identical to each other's, each making + small edits in specific areas. Per-agent stores multiply the near- + identical content; a shared content-addressed store dedups it by + construction, and a path→hash manifest per workspace makes the + per-agent delta just its changed hashes (the alknet-filesystem + probe shows the manifest mapping is straightforward). + +**The p2p shape and what it pins (added 2026-10-01, from the dedup +discussion).** Two alkgit use cases exist: self-hosted git (where +cross-repo dedup matters little — you don't fork your own repos) and +**p2p git for OSS projects**, the demanding one. The p2p sketch: +ownership/ACL live in a smart contract on a low-fee network; +replicators are push/clone endpoints that watch the contract and cache +its state; repos sync via iroh-style gossip (hashes, not content) and +pull/sync on demand. Neither alkgit approach described so far takes +advantage of dedup — which is exactly a store-layer property. The +implications for this crate: + +- Both hard consumers (p2p replicators, agent swarms) reduce to the + same shape: **a pooled CAS per node**, where repos / workspaces are + *sets of hash references* (trees/manifests), not separate object + stores. Cross-repo dedup is then by construction; git's own + `alternates`/object-pool mechanism is the same idea done awkwardly. +- Because everything is hash-addressed, the p2p wire story collapses + to "announce hashes, diff hash sets, fetch missing blobs verified" — + which is precisely the ops surface OQ-BL-01 left provisional. The + consumer exists now: if p2p git is a confirmed alkgit direction, + the ops surface upgrades from "leaning store-only" to a confirmed + have/need + fetch op family, auth-gated via the alkcall seam. **The scope line (current posture, revisit as evidence arrives):** this crate is the *store*, not the *transport*. iroh-blobs welds store @@ -53,8 +83,13 @@ this crate is the *store*, not the *transport*. iroh-blobs welds store posture here is the opposite split — put/get/verify against a hash is its own layer, and any provider/protocol/ops surface lives above it or in feature-gated modules (the alk* inversion-point pattern). The ops -surface (a channels-native "give me the blob with this hash" op family, -auth-gated via `AccessControl`) is a Phase 0/1 question, not assumed. +surface (a channels-native "have/need + fetch blob with this hash" op +family, auth-gated via `AccessControl`) is a Phase 0/1 question — the +p2p-git consumer (§The p2p shape) upgrades it from speculative to +likely, but the crate-boundary decision itself is still open +(OQ-BL-01). What stays *above* the crate regardless: repos/workspaces +as reference sets (trees/manifests), path→hash mapping, git semantics, +the contract/gossip/replicator policy layer. Guiding principles, inherited from the alk* family: @@ -79,6 +114,20 @@ Guiding principles, inherited from the alk* family: verified content (git objects are hash-addressed by git itself). Which verification story the crate owns — bao trees, per-blob digests, backend-native — is open (OQ-BL-04). +6. **Pooling is the point.** The dedup wins both demanding consumers + want (Forknet-style multi-repo OSS content, agent swarms) require + that repos/workspaces are *sets of hash references* over one pooled + CAS per node, not walled-off per-repo stores. Pooling forces a GC + story (OQ-BL-05) and a multi-backend dispatch story (OQ-BL-02). +7. **Whole-file CAS as the default; chunking earns its keep only + where it must.** For source-code-scale content, whole-file hashing + captures the agent/edit delta perfectly and needs no chunk tree; + fixed-size chunk trees are actually *hostile* to insertion/deletion + edits (byte shifts cascade). Content-defined chunking (CDC — + fastcdc/restic-style, shift-resistant) only pays on large binary + files receiving small edits — the git-lfs weakness. iroh-blobs/bao + use fixed-size trees; do not inherit that as the default. Where + chunking lives (store encoding vs manifest layer) is OQ-BL-04. ## What is already known (settled, thin) @@ -94,6 +143,14 @@ setup: blobs, filesystem fallback for large — the appfile shape — but *how* multi-backend dispatch works is exactly what alknet's research ran into, so it earns research and probably POCs (OQ-BL-02). +- **Pooled CAS, not per-repo stores** (agreed in the 2026-10-01 + dedup discussion — the load-bearing scope decision so far): repos + and agent workspaces are sets of hash references over one shared + content-addressed store. This is the property both demanding + consumers actually need; everything else (GC, manifests, git + semantics) lives around it. Details open (OQ-BL-05). +- **Whole-file CAS is the default granularity**; chunking is scoped + to the large-binary-blob case (principle 7), not the default path. ## Prior art @@ -118,7 +175,11 @@ whether this crate can sit *under or beside* gix-odb semantics — git objects are already content-addressed and verified by git's own model, so the crate's value there is the large-blob story (git-lfs-shaped) and a better small-object backend than the proposed default, without -fighting git's hash model. +fighting git's hash model. Note for the p2p case: git's `alternates` +/ object-pool mechanism (e.g. GitLab object pools for fork networks) +is the prior art for cross-repo dedup among related repos — the +pooled-CAS posture (OQ-BL-05) generalizes it to unrelated repos and +to p2p replication. ### alknet's appfile external-store probe — the direct ancestor @@ -150,14 +211,23 @@ final set into Phase 1's `docs/architecture/open-questions.md`. ### OQ-BL-01: Crate scope — store-only, or store + ops surface? The store (put/get/verify) is clearly in. What about the network ops -layer — a channels-native "fetch/provide blob" op family on the -alkcall substrate, auth-gated via `AccessControl`? iroh-blobs has it -(provider protocol); our posture is to separate it. Options: (a) -store-only crate, ops in a sibling crate later; (b) store + optional -feature-gated ops module here; (c) undecided pending the first -consumer's shape. Leaning (a)/(b) per the inversion-point pattern, but -the alkgit + alknet consumer needs should decide — neither has confirmed -a *networked* transfer requirement yet. +layer — a channels-native op family on the alkcall substrate, +auth-gated via `AccessControl`? iroh-blobs has it (provider protocol); +our posture is to separate it. Options: (a) store-only crate, ops in a +sibling crate later; (b) store + optional feature-gated ops module +here; (c) undecided pending the first consumer's shape. + +Original lean was (a)/(b) on the inversion-point pattern. **Updated +2026-10-01:** the p2p-git consumer (§The p2p shape) makes the ops +surface likely rather than speculative — because everything is +hash-addressed, p2p sync *is* "announce hashes, diff hash sets, fetch +missing blobs verified," which is an op family on this store. Whether +that family lives here (feature-gated) or in the sibling that owns +replicator/gossip policy is the remaining question; the deciding input +is what alkgit's replicator actually needs and whether alknet's +appfile case ever transfers over the network. Even so, the *shape* of +the ops (have/need + verified fetch, ACL-gated) is now firm enough to +plan against. ### OQ-BL-02: Multi-backend dispatch — the appfile problem @@ -185,17 +255,58 @@ this settles. ### OQ-BL-04: Verification and chunking story -iroh-blobs' verification is bao outboard encoding over BLAKE3 chunk -trees. Git objects are self-verifying under git's model. Raw large -blobs (git-lfs-shaped, appfile large files) need *some* verification -story from us. Options: bao (borrow the *conclusion*, dependency -posture per convention 7), our own digest scheme, or pluggable -verification per blob-kind. Also: does chunking exist at the store -layer at all (git objects are unit blobs; appfile items may be large -files wanting range reads)? Range-read support is an open API-shape -question. +iroh-blobs' verification is bao outboard encoding over fixed-size +BLAKE3 chunk trees; git objects are self-verifying under git's own +model; raw large blobs (git-lfs-shaped, appfile large files) need +*some* verification story from us. Options: bao (borrow the +*conclusion*, dependency posture per convention 7), our own digest +scheme, or pluggable verification per blob-kind. -### OQ-BL-05: POC register (draft) +**Chunking is scoped (2026-10-01, principle 7), not open-ended:** the +default is whole-file CAS — content-defined chunking (fastcdc/ +restic-style; shift-resistant, unlike bao's fixed trees) is an +encoding consideration *only* for large binary blobs receiving small +edits, and git objects are unit blobs that never want chunking. The +remaining question is *where* variable chunking lives if adopted: in +the store's encoding layer (a blob is stored as chunk tree + the +store serves ranges) or in the manifest layer above (chunks are +themselves small blobs; the store stays whole-file-only — the dedup- +friendlier option, since chunks are then pooled and cross-file shared +like any other content). Decision input: how alkgit handles +git-lfs-shaped large files vs how agent workspaces transfer large +artifacts, and whether range reads are a store API or a +reassemble-above concern. + +### OQ-BL-05: Pooling and GC — namespaces, lifetime, and the pack tension + +**The decision from the discussion (2026-10-01, principle 6):** one +pooled CAS per node; repos/workspaces are sets of hash references +(manifests/trees) rather than isolated stores — cross-repo dedup by +construction. The open questions are the mechanics: + +- **Namespace shape** — does the store know namespaces at all + (per-repo/partition prefixes), or is the pool flat with structure + living entirely in the manifests above? Flat is simpler and dedups + unconditionally; namespaces buy cheap garbage identification and + per-consumer isolation (multi-tenant nodes may need it). +- **GC / lifetime** — a pooled store never deletes by default, so + something must decide when a blob dies. Options: refcounting + maintained by put/manifest-write, mark-from-roots over registered + root manifests (replicators know their contract-derived heads + locally — tractable), or lease/expiry schemes. This is a + correctness surface (a GC bug is silent data loss), so it likely + earns a POC or at least a worked example before Phase 1 pins it. +- **The pack tension (alkgit-shaped, constrains the store surface)** — + git packfiles are pack-efficient but blob one object-store-per-pack + (bad for cross-repo dedup); loose-per-object in the pooled store + dedups but is fs-inefficient at git's object volumes. Small git + objects in the kv backend largely dissolve this for the common + case; the residual question is what the store must support for + large/packed content — per-object granularity, range reads into + packed blobs, or both (interacts with OQ-BL-04's range-read + question — likely read them together). + +### OQ-BL-06: POC register (draft) Numbered POCs, opened as research reaches them (findings land in `docs/research/`; worktree placement per the SDD process): @@ -205,11 +316,14 @@ Numbered POCs, opened as research reaches them (findings land in | 1 | Backend-trait + dual-dispatch shape (kv small / fs large) | **Pending** — likely first POC | findings file TBD | | 2 | Multi-hash store (SHA-256 + BLAKE3 coexisting) | **Pending** — rides #1's data | findings file TBD | | 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads) | **Pending** | findings file TBD | +| 4 | Pooled CAS + GC (manifest-referenced blobs, delete-then-recover semantics, refcount vs mark-from-roots) | **Pending** — after #1 (needs a backend to pool over) | findings file TBD | Sequencing note: #1 and #2 are probably one worktree (the dispatch POC naturally exercises two hash algorithms); #3 is a reading-and-design POC against the current iroh-blobs checkout plus the alknet probe -re-verification. +re-verification; #4's GC half matters most — it is a correctness +surface and should be a worked example at minimum. The pack tension +(OQ-BL-05) rides #3's reading rather than earning its own POC yet. POC placement conventions (inherited from alksocks/alktunnels): a POC that needs code from this repo runs in a worktree/branch @@ -229,5 +343,5 @@ in `docs/research/` regardless of where the code lives. 3. **Read gix-odb** (`/workspace/git-oxide/gix-odb`) for the alkgit baseline: its storage layout, object model, and where a blob-store crate under/beside it earns its keep. -4. **Open POCs** per OQ-BL-05, in the order above. +4. **Open POCs** per OQ-BL-06, in the order above. 5. Converge: recommended approach + final OQ register → Phase 1. \ No newline at end of file