- AGENTS.md: repo conventions adapted from alkcall/alksocks (store posture, multi-hash requirement, no-iroh-wire rule; no fuzz gate yet) - Cargo.toml + src/lib.rs: bare lean lib skeleton (tokio async, thiserror) - docs/research/phase-0.md: Phase 0 draft — vision, prior art (iroh-blobs store, gix-odb, alknet appfile probe, alkcall), OQ register, POC plan - .gitignore, LICENSE-APACHE/MIT - Re-pointed stale alknet/alkcall references in .opencode/agents/ and docs/sdd_process.md Verification: cargo build/test, cargo clippy --all-targets -- -D warnings, cargo fmt --check, cargo doc --no-deps — all pass
12 KiB
status: draft last_updated: 2026-09-30 (initial setup draft: vision, prior art, open questions from the setup discussion; un-numbered — renumber OQ-BL-01..NN as the register solidifies)
alkblobs — Phase 0 (Exploration)
This document captures Phase 0 (Exploration) for the alkblobs crate:
vision, guiding principles, prior art, and open questions (OQ-BL-01..NN).
Phase 0's objective per docs/sdd_process.md: capture vision and guiding
principles; research options; validate approaches; converge on a
recommended approach. It is the input to Phase 1 (Architecture), where
the Architect will produce docs/architecture/ specs, ADRs, and the
open-questions tracker.
Drafted 2026-09-30, emerging from the initial setup discussion. Nothing is converged yet — this is the working sketch, not a specification.
Vision and guiding principles
One sentence: content-addressed blob storage in the alk* family — a multi-hash put/get/verify store with pluggable backends (small blobs in a kv/sqlite-ish store, large blobs on a filesystem fallback), split from any transport/protocol layer and tolerant of hash-algorithm conflicts (the alkgit case), with auth-gated network operations riding the alkcall seam when they exist.
Why this crate exists (two converging consumers):
- alkgit (planning phase) needs an object backend that is not its proposed default file-based backend ("kind of gross") — blobs stored under git's own object hashes, but with iroh-blobs-like handling for large blobs (where git typically delegates to git-lfs; we could instead follow an iroh-blobs-shaped large-blob path). The conflict: iroh-blobs is BLAKE3-only and content-hashes with bao verification, while git's object database is SHA-1/SHA-256 with git's own object format. A store that insists on one hash algorithm cannot serve git objects directly.
- The alknet rewrite (the original alk* project, being decomposed and improved) needs the "appfile" external-store shape: small blobs in a kv/sqlite store (faster than the filesystem for small items — true beyond sqlite: it applies to the kv store iroh-blobs uses too), large blobs on a filesystem fallback, with filename↔hash mapping. alknet's own research hit the core awkwardness of dispatching across more than one backend at a time — which is exactly a store-layer problem this crate should own.
The scope line (current posture, revisit as evidence arrives): this crate is the store, not the transport. iroh-blobs welds store
- provider protocol + tickets + postcard into one crate; the defining
posture here is the opposite split — put/get/verify against a hash is
its own layer, and any provider/protocol/ops surface lives above it or
in feature-gated modules (the alk* inversion-point pattern). The ops
surface (a channels-native "give me the blob with this hash" op family,
auth-gated via
AccessControl) is a Phase 0/1 question, not assumed.
Guiding principles, inherited from the alk* family:
- Substrate-agnostic by construction. The store must not know
whether bytes arrive over the network, from a local writer, or get
reassembled anywhere particular. Backends behind an injected seam
(alktty
TtyBackend/ alktunnels pump-halves precedent). - Hashes are data, not transport identity. Multiple hash algorithms must coexist (git SHA-1/SHA-256, BLAKE3, maybe bao-shaped chunk trees as one encoding among several). How this is abstracted — trait, enum, per-backend config — is an open question (OQ-BL-03).
- Borrow conclusions, not wire surface. iroh-blobs' tickets, postcard serialization, and provider protocol are design-welded choices we do not inherit. Its store-shape lessons (kv + flat backends, verification flow, chunking) are fair game per prior-art reading.
- Producer/consumer vocabulary for any network-facing surface;
authorization via alkcall's
AccessControl/identity seam, never an in-band invented scheme (convention 6). - Verify where it matters. iroh-blobs' core virtue is verified transfer (bao outboard encoding). The alkgit case already has verified content (git objects are hash-addressed by git itself). Which verification story the crate owns — bao trees, per-blob digests, backend-native — is open (OQ-BL-04).
What is already known (settled, thin)
Almost nothing is pinned — deliberately. The only postures agreed at setup:
- Not a fork of iroh-blobs. A downstream store crate inspired by
its
src/storework, diverging deliberately on hashing and wire surface (seeiroh-blobs-evalbelow when written). - Multi-hash from day one (hard requirement from alkgit — do not hardcode BLAKE3).
- Backend pluralism assumed, not designed. kv/sqlite for small blobs, filesystem fallback for large — the appfile shape — but how multi-backend dispatch works is exactly what alknet's research ran into, so it earns research and probably POCs (OQ-BL-02).
Prior art
iroh-blobs — the shape inspiration (evaluated, not the base)
/workspace/iroh-blobs (fresh upstream checkout; read-only reference).
Specifically src/store: the kv backend (small blobs, in-process) and
flat file backend (large blobs) split; bao outboard encoding and
verification flow; chunking. Its BLAKE3-only hashing, tickets, postcard
serialization, and provider protocol are the design-welded choices we
diverge from. The store eval should be written against the current
checkout — alknet's older research refers to an older iroh-blobs and
its conclusions must be re-verified rather than inherited.
gix-odb — git's own object database (alkgit's baseline)
/workspace/git-oxide/gix-odb (read-only reference). The backend
alkgit currently plans against; the hashing-algorithm baseline git
actually uses (SHA-1/SHA-256), git's loose-object and packfile layout,
and git's already-hash-addressed object model. The alkgit question is
whether this crate can sit under or beside gix-odb semantics — git
objects are already content-addressed and verified by git's own model,
so the crate's value there is the large-blob story (git-lfs-shaped)
and a better small-object backend than the proposed default, without
fighting git's hash model.
alknet's appfile external-store probe — the direct ancestor
/workspace/@alkdev/alknet/docs/research/alknet-filesystem/
(alknet-blobs-external-store-probe.md, poc-summary.md) — written
against an older iroh-blobs; conclusions re-verify in this Phase 0.
The load-bearing bits: the appfile shape (small blobs in kv/sqlite,
large on fs fallback), the filename↔hash mapping problem, and the
multi-backend dispatch pain the probe hit (which is this crate's
reason to own that dispatch). Old-data warning: any API or behavior
claims there describe an older upstream; re-check against
/workspace/iroh-blobs as checked out today.
alkcall — the substrate
/workspace/@alkdev/alkcall (the call + channels RPC crate). The
authorization seam for network-facing blob ops (AccessControl,
producer/consumer vocabulary, alkcall ADR-022/037) and, if blob
transfer rides channels, the established data-path patterns
(BiStream, two-pump pump_bidi ADR-050). Whether transport even
belongs in this crate is itself open (OQ-BL-01); alkcall is the
substrate it composes with when it does.
Open Questions
Numbering is provisional until the register solidifies; promote the
final set into Phase 1's docs/architecture/open-questions.md.
OQ-BL-01: Crate scope — store-only, or store + ops surface?
The store (put/get/verify) is clearly in. What about the network ops
layer — a channels-native "fetch/provide blob" op family on the
alkcall substrate, auth-gated via AccessControl? iroh-blobs has it
(provider protocol); our posture is to separate it. Options: (a)
store-only crate, ops in a sibling crate later; (b) store + optional
feature-gated ops module here; (c) undecided pending the first
consumer's shape. Leaning (a)/(b) per the inversion-point pattern, but
the alkgit + alknet consumer needs should decide — neither has confirmed
a networked transfer requirement yet.
OQ-BL-02: Multi-backend dispatch — the appfile problem
Small blobs in kv/sqlite, large on fs fallback — the shape is agreed, the mechanics are not: how does put/get pick a backend (size thresholds? per-algorithm routing? per-namespace config?), how does it work when a blob should migrate between backends, and what happens on a get when the "wrong" backend was probed? alknet's probe hit exactly this ("dealing with more than one backend at a time"); its findings need re-verification against the current iroh-blobs before relying on them. Expected shape: a backend trait + a dispatch layer, but the trait shape is a one-way door once written — deserves a POC before committing.
OQ-BL-03: Hash abstraction — trait, enum, or per-backend config?
The hard requirement: git SHA-1/SHA-256 and BLAKE3 must coexist (alkgit conflict). Sub-questions: is the hash algorithm a store parameter (one algorithm per store instance, chosen by the consumer) or per-blob data (multi-algorithm within one store)? Does verification (bao trees) get per-algorithm treatment or does bao stay BLAKE3-bound (git objects don't need bao verification anyway)? This interacts with OQ-BL-01 and the wire-format question — a wire ADR must follow whatever this settles.
OQ-BL-04: Verification and chunking story
iroh-blobs' verification is bao outboard encoding over BLAKE3 chunk trees. Git objects are self-verifying under git's model. Raw large blobs (git-lfs-shaped, appfile large files) need some verification story from us. Options: bao (borrow the conclusion, dependency posture per convention 7), our own digest scheme, or pluggable verification per blob-kind. Also: does chunking exist at the store layer at all (git objects are unit blobs; appfile items may be large files wanting range reads)? Range-read support is an open API-shape question.
OQ-BL-05: POC register (draft)
Numbered POCs, opened as research reaches them (findings land in
docs/research/; worktree placement per the SDD process):
| # | What | Status | Where |
|---|---|---|---|
| 1 | Backend-trait + dual-dispatch shape (kv small / fs large) | Pending — likely first POC | findings file TBD |
| 2 | Multi-hash store (SHA-256 + BLAKE3 coexisting) | Pending — rides #1's data | findings file TBD |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads) | Pending | findings file TBD |
Sequencing note: #1 and #2 are probably one worktree (the dispatch POC naturally exercises two hash algorithms); #3 is a reading-and-design POC against the current iroh-blobs checkout plus the alknet probe re-verification.
POC placement conventions (inherited from alksocks/alktunnels): a POC
that needs code from this repo runs in a worktree/branch
(.worktrees/research/<task-id>/ per the SDD process); a
self-contained POC runs as a standalone crate in the global workspace
with findings written into docs/research/ here. Findings always land
in docs/research/ regardless of where the code lives.
Phase 0 plan (next steps)
- Write
iroh-blobs-eval.mdagainst the current checkout, focused onsrc/store(kv + flat backends, bao, chunking) — the conclusions inventory we may borrow, and the weld points we won't. - Re-verify alknet's probe (
alknet-blobs-external-store-probe.md,poc-summary.md) against the current upstream; mark what carried over. - Read gix-odb (
/workspace/git-oxide/gix-odb) for the alkgit baseline: its storage layout, object model, and where a blob-store crate under/beside it earns its keep. - Open POCs per OQ-BL-05, in the order above.
- Converge: recommended approach + final OQ register → Phase 1.