docs(research): fold p2p-git/dedup discussion into phase-0

- Consumers: add agent workspaces (third consumer); add the p2p shape
  section (contract/gossip/replicator sketch) and pin what it implies
  here: pooled CAS per node, repos/workspaces as hash-reference sets
- Principles: add 6 (pooling is the point) and 7 (whole-file CAS
  default; CDC scoped to large-binary edits; fixed-size trees not
  inherited)
- OQ-BL-01: ops surface upgraded from speculative to likely (have/need
  + verified fetch), crate-boundary question remains
- OQ-BL-04: chunking scoped; new sub-question — chunking in store
  encoding vs manifest layer
- New OQ-BL-05: pooling/GC — namespace shape, lifetime strategies,
  the pack tension (constrains store surface)
- POC register: add #4 (pooled CAS + GC); note pack tension rides #3
- gix-odb: note alternates/object-pools as cross-repo dedup prior art
This commit is contained in:
glm-5.3-flash committed 2026-10-01 04:30:54 +00:00
1 parent 7f7f70b220
commit 5e0a594cfd
1 file changed
+141 -27
+141 -27
View File
@@ -1,8 +1,8 @@
---
status: draft
last_updated: 2026-09-30 (initial setup draft: vision, prior art, open
questions from the setup discussion; un-numbered — renumber
OQ-BL-01..NN as the register solidifies)
last_updated: 2026-10-01 (dedup/p2p discussion folded in: the consumers
section re-worked around the pooled-CAS insight, new OQ-BL-05, OQ-BL-01
and OQ-BL-04 refined; un-numbered sections otherwise as the setup draft)
---
# alkblobs — Phase 0 (Exploration)
@@ -27,7 +27,7 @@ from any transport/protocol layer and tolerant of hash-algorithm
conflicts (the alkgit case), with auth-gated network operations riding
the alkcall seam when they exist.
**Why this crate exists (two converging consumers):**
**Why this crate exists (three converging consumers):**
1. **alkgit** (planning phase) needs an object backend that is *not*
its proposed default file-based backend ("kind of gross") — blobs
@@ -46,6 +46,36 @@ the alkcall seam when they exist.
mapping. alknet's own research hit the core awkwardness of
dispatching across more than one backend at a time — which is
exactly a store-layer problem this crate should own.
3. **Agent workspaces** — many agents working simultaneously in
workspaces that are mostly identical to each other's, each making
small edits in specific areas. Per-agent stores multiply the near-
identical content; a shared content-addressed store dedups it by
construction, and a path→hash manifest per workspace makes the
per-agent delta just its changed hashes (the alknet-filesystem
probe shows the manifest mapping is straightforward).
**The p2p shape and what it pins (added 2026-10-01, from the dedup
discussion).** Two alkgit use cases exist: self-hosted git (where
cross-repo dedup matters little — you don't fork your own repos) and
**p2p git for OSS projects**, the demanding one. The p2p sketch:
ownership/ACL live in a smart contract on a low-fee network;
replicators are push/clone endpoints that watch the contract and cache
its state; repos sync via iroh-style gossip (hashes, not content) and
pull/sync on demand. Neither alkgit approach described so far takes
advantage of dedup — which is exactly a store-layer property. The
implications for this crate:
- Both hard consumers (p2p replicators, agent swarms) reduce to the
same shape: **a pooled CAS per node**, where repos / workspaces are
*sets of hash references* (trees/manifests), not separate object
stores. Cross-repo dedup is then by construction; git's own
`alternates`/object-pool mechanism is the same idea done awkwardly.
- Because everything is hash-addressed, the p2p wire story collapses
to "announce hashes, diff hash sets, fetch missing blobs verified" —
which is precisely the ops surface OQ-BL-01 left provisional. The
consumer exists now: if p2p git is a confirmed alkgit direction,
the ops surface upgrades from "leaning store-only" to a confirmed
have/need + fetch op family, auth-gated via the alkcall seam.
**The scope line (current posture, revisit as evidence arrives):**
this crate is the *store*, not the *transport*. iroh-blobs welds store
@@ -53,8 +83,13 @@ this crate is the *store*, not the *transport*. iroh-blobs welds store
posture here is the opposite split — put/get/verify against a hash is
its own layer, and any provider/protocol/ops surface lives above it or
in feature-gated modules (the alk* inversion-point pattern). The ops
surface (a channels-native "give me the blob with this hash" op family,
auth-gated via `AccessControl`) is a Phase 0/1 question, not assumed.
surface (a channels-native "have/need + fetch blob with this hash" op
family, auth-gated via `AccessControl`) is a Phase 0/1 question — the
p2p-git consumer (§The p2p shape) upgrades it from speculative to
likely, but the crate-boundary decision itself is still open
(OQ-BL-01). What stays *above* the crate regardless: repos/workspaces
as reference sets (trees/manifests), path→hash mapping, git semantics,
the contract/gossip/replicator policy layer.
Guiding principles, inherited from the alk* family:
@@ -79,6 +114,20 @@ Guiding principles, inherited from the alk* family:
verified content (git objects are hash-addressed by git itself).
Which verification story the crate owns — bao trees, per-blob
digests, backend-native — is open (OQ-BL-04).
6. **Pooling is the point.** The dedup wins both demanding consumers
want (Forknet-style multi-repo OSS content, agent swarms) require
that repos/workspaces are *sets of hash references* over one pooled
CAS per node, not walled-off per-repo stores. Pooling forces a GC
story (OQ-BL-05) and a multi-backend dispatch story (OQ-BL-02).
7. **Whole-file CAS as the default; chunking earns its keep only
where it must.** For source-code-scale content, whole-file hashing
captures the agent/edit delta perfectly and needs no chunk tree;
fixed-size chunk trees are actually *hostile* to insertion/deletion
edits (byte shifts cascade). Content-defined chunking (CDC —
fastcdc/restic-style, shift-resistant) only pays on large binary
files receiving small edits — the git-lfs weakness. iroh-blobs/bao
use fixed-size trees; do not inherit that as the default. Where
chunking lives (store encoding vs manifest layer) is OQ-BL-04.
## What is already known (settled, thin)
@@ -94,6 +143,14 @@ setup:
blobs, filesystem fallback for large — the appfile shape — but
*how* multi-backend dispatch works is exactly what alknet's research
ran into, so it earns research and probably POCs (OQ-BL-02).
- **Pooled CAS, not per-repo stores** (agreed in the 2026-10-01
dedup discussion — the load-bearing scope decision so far): repos
and agent workspaces are sets of hash references over one shared
content-addressed store. This is the property both demanding
consumers actually need; everything else (GC, manifests, git
semantics) lives around it. Details open (OQ-BL-05).
- **Whole-file CAS is the default granularity**; chunking is scoped
to the large-binary-blob case (principle 7), not the default path.
## Prior art
@@ -118,7 +175,11 @@ whether this crate can sit *under or beside* gix-odb semantics — git
objects are already content-addressed and verified by git's own model,
so the crate's value there is the large-blob story (git-lfs-shaped)
and a better small-object backend than the proposed default, without
fighting git's hash model.
fighting git's hash model. Note for the p2p case: git's `alternates`
/ object-pool mechanism (e.g. GitLab object pools for fork networks)
is the prior art for cross-repo dedup among related repos — the
pooled-CAS posture (OQ-BL-05) generalizes it to unrelated repos and
to p2p replication.
### alknet's appfile external-store probe — the direct ancestor
@@ -150,14 +211,23 @@ final set into Phase 1's `docs/architecture/open-questions.md`.
### OQ-BL-01: Crate scope — store-only, or store + ops surface?
The store (put/get/verify) is clearly in. What about the network ops
layer — a channels-native "fetch/provide blob" op family on the
alkcall substrate, auth-gated via `AccessControl`? iroh-blobs has it
(provider protocol); our posture is to separate it. Options: (a)
store-only crate, ops in a sibling crate later; (b) store + optional
feature-gated ops module here; (c) undecided pending the first
consumer's shape. Leaning (a)/(b) per the inversion-point pattern, but
the alkgit + alknet consumer needs should decide — neither has confirmed
a *networked* transfer requirement yet.
layer — a channels-native op family on the alkcall substrate,
auth-gated via `AccessControl`? iroh-blobs has it (provider protocol);
our posture is to separate it. Options: (a) store-only crate, ops in a
sibling crate later; (b) store + optional feature-gated ops module
here; (c) undecided pending the first consumer's shape.
Original lean was (a)/(b) on the inversion-point pattern. **Updated
2026-10-01:** the p2p-git consumer (§The p2p shape) makes the ops
surface likely rather than speculative — because everything is
hash-addressed, p2p sync *is* "announce hashes, diff hash sets, fetch
missing blobs verified," which is an op family on this store. Whether
that family lives here (feature-gated) or in the sibling that owns
replicator/gossip policy is the remaining question; the deciding input
is what alkgit's replicator actually needs and whether alknet's
appfile case ever transfers over the network. Even so, the *shape* of
the ops (have/need + verified fetch, ACL-gated) is now firm enough to
plan against.
### OQ-BL-02: Multi-backend dispatch — the appfile problem
@@ -185,17 +255,58 @@ this settles.
### OQ-BL-04: Verification and chunking story
iroh-blobs' verification is bao outboard encoding over BLAKE3 chunk
trees. Git objects are self-verifying under git's model. Raw large
blobs (git-lfs-shaped, appfile large files) need *some* verification
story from us. Options: bao (borrow the *conclusion*, dependency
posture per convention 7), our own digest scheme, or pluggable
verification per blob-kind. Also: does chunking exist at the store
layer at all (git objects are unit blobs; appfile items may be large
files wanting range reads)? Range-read support is an open API-shape
question.
iroh-blobs' verification is bao outboard encoding over fixed-size
BLAKE3 chunk trees; git objects are self-verifying under git's own
model; raw large blobs (git-lfs-shaped, appfile large files) need
*some* verification story from us. Options: bao (borrow the
*conclusion*, dependency posture per convention 7), our own digest
scheme, or pluggable verification per blob-kind.
### OQ-BL-05: POC register (draft)
**Chunking is scoped (2026-10-01, principle 7), not open-ended:** the
default is whole-file CAS — content-defined chunking (fastcdc/
restic-style; shift-resistant, unlike bao's fixed trees) is an
encoding consideration *only* for large binary blobs receiving small
edits, and git objects are unit blobs that never want chunking. The
remaining question is *where* variable chunking lives if adopted: in
the store's encoding layer (a blob is stored as chunk tree + the
store serves ranges) or in the manifest layer above (chunks are
themselves small blobs; the store stays whole-file-only — the dedup-
friendlier option, since chunks are then pooled and cross-file shared
like any other content). Decision input: how alkgit handles
git-lfs-shaped large files vs how agent workspaces transfer large
artifacts, and whether range reads are a store API or a
reassemble-above concern.
### OQ-BL-05: Pooling and GC — namespaces, lifetime, and the pack tension
**The decision from the discussion (2026-10-01, principle 6):** one
pooled CAS per node; repos/workspaces are sets of hash references
(manifests/trees) rather than isolated stores — cross-repo dedup by
construction. The open questions are the mechanics:
- **Namespace shape** — does the store know namespaces at all
(per-repo/partition prefixes), or is the pool flat with structure
living entirely in the manifests above? Flat is simpler and dedups
unconditionally; namespaces buy cheap garbage identification and
per-consumer isolation (multi-tenant nodes may need it).
- **GC / lifetime** — a pooled store never deletes by default, so
something must decide when a blob dies. Options: refcounting
maintained by put/manifest-write, mark-from-roots over registered
root manifests (replicators know their contract-derived heads
locally — tractable), or lease/expiry schemes. This is a
correctness surface (a GC bug is silent data loss), so it likely
earns a POC or at least a worked example before Phase 1 pins it.
- **The pack tension (alkgit-shaped, constrains the store surface)** —
git packfiles are pack-efficient but blob one object-store-per-pack
(bad for cross-repo dedup); loose-per-object in the pooled store
dedups but is fs-inefficient at git's object volumes. Small git
objects in the kv backend largely dissolve this for the common
case; the residual question is what the store must support for
large/packed content — per-object granularity, range reads into
packed blobs, or both (interacts with OQ-BL-04's range-read
question — likely read them together).
### OQ-BL-06: POC register (draft)
Numbered POCs, opened as research reaches them (findings land in
`docs/research/`; worktree placement per the SDD process):
@@ -205,11 +316,14 @@ Numbered POCs, opened as research reaches them (findings land in
| 1 | Backend-trait + dual-dispatch shape (kv small / fs large) | **Pending** — likely first POC | findings file TBD |
| 2 | Multi-hash store (SHA-256 + BLAKE3 coexisting) | **Pending** — rides #1's data | findings file TBD |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads) | **Pending** | findings file TBD |
| 4 | Pooled CAS + GC (manifest-referenced blobs, delete-then-recover semantics, refcount vs mark-from-roots) | **Pending** — after #1 (needs a backend to pool over) | findings file TBD |
Sequencing note: #1 and #2 are probably one worktree (the dispatch POC
naturally exercises two hash algorithms); #3 is a reading-and-design
POC against the current iroh-blobs checkout plus the alknet probe
re-verification.
re-verification; #4's GC half matters most — it is a correctness
surface and should be a worked example at minimum. The pack tension
(OQ-BL-05) rides #3's reading rather than earning its own POC yet.
POC placement conventions (inherited from alksocks/alktunnels): a POC
that needs code from this repo runs in a worktree/branch
@@ -229,5 +343,5 @@ in `docs/research/` regardless of where the code lives.
3. **Read gix-odb** (`/workspace/git-oxide/gix-odb`) for the alkgit
baseline: its storage layout, object model, and where a blob-store
crate under/beside it earns its keep.
4. **Open POCs** per OQ-BL-05, in the order above.
4. **Open POCs** per OQ-BL-06, in the order above.
5. Converge: recommended approach + final OQ register → Phase 1.