docs(research): POC #3 passed — large-blob streaming + sqlite-vs-fs benchmark

- POC #3 (register OQ-BL-06) executed and passed: streaming LFSObject-shaped
  put/get at the store layer; fanout seam (serve-while-persisting) validated
  as broadcast above the backends; fs range reads with the verification gap
  documented (preamble makes slice-vs-whole verification impossible —
  chunk-tree encoding sharpened to a transfer-layer conditional)
- sqlite-vs-fs micro-benchmark: first-party anchor for the small-blob
  belief (~9-10x sqlite at 1-16 KiB, crossover ~128-256 KiB, 128 KiB
  default lean); harness re-runnable on real media
- Stage-file hygiene tests caught a real leak (one temp file per failed
  put) — fixed; all failure paths converge on discard
- Pack tension resolved in shape from the gix-odb read: packs as large
  blobs served by range reads; object-level dedup lives in the loose tier
- iroh-blobs-eval.md written against the current checkout (e82cbdc /
  v0.103.0): actor/message-protocol pattern, fs backend layout,
  verification, do-not-inherit + keep lists
- alknet probe re-check demoted to historical cross-check only: alknet's
  code is not a design input (decomposed; substrate is alkcall); its
  small-blob perf claim was asserted-never-measured, superseded by A4
- phase-0 register complete: no open POCs; next step is convergence ->
  Phase 1

Verification: POC crate (standalone /workspace/alkblobs-largeblob-poc)
14 tests pass, clippy -D warnings clean, fmt clean; digests byte-exact
vs the real git hash-object CLI on both streaming put paths
This commit is contained in:
glm-5.3-flash committed 2026-10-01 08:14:51 +00:00
1 parent 493516424a
commit 6b278fb490
3 files changed
+653 -67

No files matched your search

+198
View File
@@ -0,0 +1,198 @@
---
status: draft
title: "iroh-blobs store eval — against the current checkout"
last_updated: 2026-10-02 (POC #3 reading; checkout e82cbdc, v0.103.0)
---
# iroh-blobs store eval (current checkout)
Evaluated against `/workspace/iroh-blobs` @ `e82cbdc` ("Release 0.103.0",
2026-06-15 — the checkout *is* the 0.103.0 release, so alknet's older
probe claims about "published 0.103" verify directly against it).
Scope: the store layer (`src/store/**`, `src/api/**`, `src/util/**`).
The wire surface (tickets, postcard, provider protocol) is out of scope
per convention 7 — we are not adopting it. This eval feeds OQ-BL-02
(backend shape), OQ-BL-04 (verification/chunking), and the POC #3
large-blob design. GC mechanics were already verified separately
(see phase-0 §Prior art; the gc.rs read predates POC #1).
## The headline: there is no Store trait — and that is instructive
iroh-blobs 0.103 has **no `Store` trait at all**. The store API is a
message protocol: an enum of RPC requests (`src/api/proto.rs:90-139`,
`#[rpc_requests]`-derived `Command`), and each backend (MemStore,
FsStore, ReadonlyMemStore) is an *actor* consuming that command set
over `irpc` channels. The user-facing `Store` struct is just a client
handle; `Store::connect/listen` make the same message enum work over
the network. The abstraction is "same messages locally and remote."
Lessons for alkblobs:
- **The command-enum-as-contract is a working pattern** (it is what
makes iroh-blobs' remote/local symmetry possible), but it carries the
whole irpc machinery into the store layer and *is* the reason their
store is welded to their transport. POC #1's lean `Backend` trait is
the smaller hammer for our split: the trait stays dumb (opaque bytes),
and any ops surface above it (alkcall) speaks its own protocol.
- **Re-borrowed conclusion, not mechanism: per-hash serialized actors.**
FsStore runs one serialized state machine *per hash* (entity manager,
`src/store/fs/util/entity_manager.rs`) with parallelism across hashes;
per-hash state is a `watch::Sender<BaoFileStorage>` (cheap snapshots,
observable progress, idle-recycled actor pool). This is a genuinely
good concurrency shape for Phase 1's sweep-vs-put story: "one writer
per content address, readers snapshot-observe" matches CAS semantics
naturally. iroh's DeleteSet was already our named prior art; the
per-hash-actor adds a second reusable conclusion.
## What the fs backend actually looks like on disk (0.103)
```
blobs.db redb — hash → EntryState (+tags, + inline data/outboard)
data/<hash>.data complete + partial canonical data
data/<hash>.obao4 outboard (bao pre-order, 64-byte pairs)
data/<hash>.sizes4 per-chunk-group size slots
data/<hash>.bitfield blake3-checksummed postcard bitfield (coverage)
temp/<uuid>.temp staging (must be same-device as data/)
```
- **Entry state machine per blob** (`src/store/fs/entry_state.rs`):
`Complete { data_location: Inline|Owned|External, outboard_location:
Inline|Owned|NotNeeded }` / `Partial { size: Option<u64> }`. The
inline/owned split is the small-blob/large-blob dispatch *inside one
backend*: ≤16 KiB data lives entirely in the redb `inline-data` row;
≥~1 MiB is data file + outboard file; between, data file + inline
outboard. Thresholds are `InlineOptions { max_data_inlined: 16 KiB,
max_outboard_inlined: 16 KiB }` (`fs/options.rs:71-78`).
- **16 KiB is exactly one bao chunk group** (16 × 1024 B chunks), so
every blob ≤ 16 KiB has an outboard of size *zero* — a "small blob"
is one redb row keyed by plain `blake3(content)`, no tree material at
all (`fs/import.rs:160-221`, `DESIGN.md:71`). iroh couples "hash
overhead" and "storage tier" at the same boundary; under our git-oid
derivation there is no outboard to elide (whole-blob digest only) and
the tier boundary is a pure dispatch choice — POC #3's benchmark says
~64-128 KiB, not 16 KiB (see `poc-largeblob-findings.md`, B-finding).
- **Crash-consistency story, borrowed as a conclusion** (`DESIGN.md:
126-142`, "files are hard"): persist data+outboard+sizes *first*, the
checksummed bitfield *last*, with truncate-to-zero-on-read marking a
bitfile dirty under crash. If the bitfield is missing/corrupt, the
entry re-validates by full local hash check. Lazy persistence (on
actor idle/shutdown only), never per-chunk fsync. For alkblobs: our
put path is *stage-temp → hash → one rename*, which needs no coverage
bitfield at all — we only ever need this machinery if we adopt
partial/streaming-download entries (range-serve before complete).
Recorded as a conditional, like bao.
- **Temp-file flow** (<`fs/import.rs`>): byte-stream puts stage to a
temp file (after a 16 KiB in-memory peek), path-import reflinks-or-
copies (`reflink-copy` crate) to temp, outboard encodes in a second
pass, then `fs::rename` atomizes both files before the redb row lands
*awaiting* the transaction commit. `ExportMode::TryReference` is the
mirror: rename the store's own file into the caller's tree and mark
the entry `External`. (TryReference is a neat dedup trick for the
appfile/workspace consumer — "move into the store without copying" —
Phase 2 note, not Phase 1.)
- **Metadata batching**: the redb actor batches read commands (10k msgs
/ 1 s window) and write commands (1000 / 500 ms) into single
transactions (`meta.rs:785-856`) — "the secret to fast download
speeds is to not touch the metadata database at all" (`DESIGN.md:61`).
sqlite's WAL gives comparable batching for free at the smaller scale
our store keys live at; the *conclusion* (batch metadata writes,
never per-blob-fsync) transfers.
## Verification & chunking (OQ-BL-04 input)
- Fixed 1024 B bao chunks, grouped 16-wide (iroh `IROH_BLOCK_SIZE`,
`store/mod.rs:17-18`); no content-defined chunking exists in-tree.
- Outboard encoding is a **custom single-pass encoder over blake3
hazmat APIs** (`src/util.rs:224-336`) — post-order chunk iteration,
`set_input_offset`/`finalize_non_root` merges, buffered 1 MiB reads.
iroh wrote their own rather than using bao-tree's `CreateOutboard`;
the *shape* of "encode outboard while streaming through chunks once"
is the pattern any chunk-tree encoder follows, bao or CDC.
- **Get verification is subtree-local**: `export_bao(hash, ranges)` runs
`traverse_ranges_validated` and lazily reads only the parent pairs the
requested chunks need. Byte-range `export_ranges` is deliberately
*unvalidated* (bitfield-superset check only; `DataReader` is documented
"a reader for the unvalidated data file", `fs.rs:337-343`).
- **For our canonical digest this exposes a hard fact** (POC #3 finding
A2, reproduced independently): the git-blob derivation verifies *the
whole content* (preamble includes the length). There is no sub-range
relationship between `blob <len>\0+content` and any slice of it, so
range reads under our hash cannot be verified against the blob's own
digest, period. iroh gets range-verification *from the chunk tree*, not
from the hash algorithm. If ranged transfer ever needs per-range
verification, the options stay as OQ-BL-04 lists them (caller-carried
slice digests = transport-integrity only; conditional chunk-tree
encoding; or backend-native). POC #3's verdict: the local-store case
doesn't need it (pages fail via EIO, git blobs self-verify on read);
the *networked large-blob* case is the only real consumer of a chunk
tree — which matches the "conditional, confined to the encoding layer"
posture already recorded.
## Small-blob handling & streaming (POC #3 design input)
- **Streaming put exists in three shapes**; none hashes on the fly:
`add_stream` (bidirectional `ImportByteStream` channel; stages to
memory then temp file, computes outboard in a **second pass**,
`fs/import.rs:290-352, 372-419`), `import_bao` (pre-verified items,
no re-check), and `import_bao_reader` (network decode-verify-push,
`blobs.rs:446-488` — the provider's push handler). There is **no
put-from-`AsyncRead` API**; callers adapt to `Stream<io::Result<Bytes>>`.
- POC #3 found the same constraint from our side *before* reading this
in detail: the git preamble needs the length at hash time, so an
unknown-length stream *must* stage-then-hash. iroh's architecture
arrives at the same stage-then-encode shape (for bao-tree reasons
instead). **The one-pass alternative exists only for known-length
puts** — which is exactly the git/alkgit dominant case (a staged file
has a stat; a protocol offer carries a size). Unknown-length remains
the two-pass shape for both crates; POC #3's finding A1 records why
(and that the second pass can read from the *staged file*, so it's
page-cache-cheap, not network-repetitive).
- **Serve-while-persisting is not a bolt-on in iroh — it is the
download path**: `get_blob_ranges_impl` opens the import *before*
reading from the network and `tokio::try_join!`s the network decode
loop with the store import (`src/api/remote.rs:884-944`). POC #3's
fanout seam (finding A3) rebuilt this at our layer with a broadcast
channel; the mechanism differs (single reader broadcasting vs their
per-item verified forwarding) but the conclusion matches: the fetch
handler should be "decode-verify → broadcast → (store arm + subscriber
arms)", all driven from one receive loop.
- **`BlobReader`** implements `AsyncRead + AsyncSeek` built on
`export_ranges` per read call (`api/blobs/reader.rs`); seek-from-end
is an upstream TODO. Our fs range-read surface covers the same need.
## Performance posture (no in-tree benches; code choices as evidence)
No `benches/`, no criterion in the checkout. The perf story is in
DESIGN.md + code micro-choices: redb write/read batching; partial
entries live entirely outside redb until completion; zero-copy
`Bytes` plumbing end to end; reflink-first import/export; 1 MiB copy
buffers with `yield_now()` fairness; per-hash channel caps bounding
backpressure. The two transferable *conclusions* for us: batch
metadata transactions (never per-blob fsync), and zero-copy/`Bytes`
at the streaming seams. Our own sqlite-vs-fs numbers are in the POC #3
findings (first-party, measured).
## Do-not-inherit list (re-confirmed against 0.103)
- The `Command`-actor + irpc store abstraction (welds store to their
transport stack).
- redb-specific entry-state modeling if we go sqlite (their tables and
`EntryState` postcard encoding are theirs; our sqlite schema is
`key → (value, size)` and our dispatch is store-layer policy).
- Fixed 1024 B chunk trees as the default granularity (principle 7
already rejects; nothing new in the eval changes it).
- Tickets/postcard/provider — out of scope as always.
## What we keep (conclusions only)
1. Per-hash serialized state machines + snapshot-observation for
concurrent access (entity-manager pattern).
2. Stage-temp → atomic rename as the universal put commit primitive.
3. Metadata write batching (amortize sync cost).
4. Zero-copy `Bytes` at streaming seams.
5. Serve-while-persisting as the fetch-handler shape (POC #3 validated
at our layer).
6. Crash-consistency ordering (data first, coverage/index last, with
dirty-marking) — only if partial entries ever exist.
7. The inline threshold concept, re-sized by our own benchmark
(~64-128 KiB crossover, git-blob workload) instead of their 16 KiB.
+132 -67
View File
@@ -1,9 +1,10 @@
---
status: draft
last_updated: 2026-10-02 (POC #1 executed and passed — see
poc-trait-dispatch-findings.md; POC #2 absorbed earlier; register
updated; earlier rounds: hash simplification, GC/pooling, dedup/p2p,
setup draft)
last_updated: 2026-10-02 (POC #3 executed and passed — see
poc-largeblob-findings.md plus iroh-blobs-eval.md; POC #1 executed and
passed — see poc-trait-dispatch-findings.md; POC #2 absorbed earlier;
register complete; earlier rounds: hash simplification, GC/pooling,
dedup/p2p, setup draft)
---
# alkblobs — Phase 0 (Exploration)
@@ -28,7 +29,7 @@ git-family derivation; legacy SHA-1-tolerant), with pluggable backends
fallback), split from any transport/protocol layer, with auth-gated
network operations riding the alkcall seam when they exist.
**Why this crate exists (three converging consumers):**
**Why this crate exists (the converging consumers):**
1. **alkgit** (planning phase) needs an object backend that is *not*
its proposed default file-based backend ("kind of gross") — blobs
@@ -41,21 +42,23 @@ network operations riding the alkcall seam when they exist.
first-class entries with zero indirection — the conflict was an
artifact of inheriting iroh-blobs' BLAKE3-only choice, and
once that inheritance was rejected the conflict went with it.
2. **The alknet rewrite** (the original alk* project, being decomposed
and improved) needs the "appfile" external-store shape: small blobs
in a kv/sqlite store (faster than the filesystem for small items —
true beyond sqlite: it applies to the kv store iroh-blobs uses
too), large blobs on a filesystem fallback, with filename↔hash
mapping. alknet's own research hit the core awkwardness of
dispatching across more than one backend at a time — which is
exactly a store-layer problem this crate should own.
2. **Workspace/file-shaped storage in the alk* family** (the "appfile"
shape, originally surfaced by old alknet research — historical
lineage only; the consumers below are current): small blobs need a
store faster than the filesystem (now first-party measured, POC #3
finding A4: sqlite ~9-10× faster at 1-16 KiB), large blobs want a
filesystem fallback, with a filename↔hash mapping above. Dispatching
across backends at the same time is a store-layer problem this crate
should own (OQ-BL-02; POC #1 validated the seam, POC #3 the
streaming + threshold evidence).
3. **Agent workspaces** — many agents working simultaneously in
workspaces that are mostly identical to each other's, each making
small edits in specific areas. Per-agent stores multiply the near-
identical content; a shared content-addressed store dedups it by
construction, and a path→hash manifest per workspace makes the
per-agent delta just its changed hashes (the alknet-filesystem
probe shows the manifest mapping is straightforward).
per-agent delta just its changed hashes (a path→hash manifest is
straightforward by construction; the old alknet probe's manifest
mapping sketch showed the same).
**The p2p shape and what it pins (added 2026-10-01, from the dedup
discussion).** Two alkgit use cases exist: self-hosted git (where
@@ -158,10 +161,10 @@ Almost nothing is pinned — deliberately. The postures agreed so far:
as primary, legacy SHA-1 tolerated, BLAKE3 demoted to a conditional
large-blob encoding consideration. Details and the reasoning trail
in OQ-BL-03.
- **Backend pluralism assumed, not designed.** kv/sqlite for small
blobs, filesystem fallback for large — the appfile shape — but
*how* multi-backend dispatch works is exactly what alknet's research
ran into, so it earns research and probably POCs (OQ-BL-02).
- **Backend pluralism assumed, then designed.** kv/sqlite for small
blobs, filesystem fallback for large — the appfile shape — with
*how* multi-backend dispatch works validated by POC #1 (the lean
trait seam) and POC #3 (streaming + threshold evidence) (OQ-BL-02).
- **Pooled CAS, not per-repo stores** (agreed in the 2026-10-01
dedup discussion — the load-bearing scope decision so far): repos
and agent workspaces are sets of hash references over one shared
@@ -181,13 +184,13 @@ Almost nothing is pinned — deliberately. The postures agreed so far:
### iroh-blobs — the shape inspiration (evaluated, not the base)
`/workspace/iroh-blobs` (fresh upstream checkout; read-only reference).
Specifically `src/store`: the kv backend (small blobs, in-process) and
flat file backend (large blobs) split; bao outboard encoding and
verification flow; chunking. Its BLAKE3-only hashing, tickets, postcard
serialization, and provider protocol are the design-welded choices we
diverge from. **The store eval should be written against the current
checkout** — alknet's older research refers to an older iroh-blobs and
its conclusions must be re-verified rather than inherited.
Specifically `src/store`: **the full store eval is now written —
`docs/research/iroh-blobs-eval.md` (2026-10-02, against checkout
e82cbdc / v0.103.0)**: no Store trait (message-protocol actor pattern),
fs backend on-disk layout (redb entry-state + inline thresholds + crash
ordering), verification (subtree-local, chunk-tree-driven), streaming
shapes, and the do-not-inherit / keep (conclusions-only) lists.
Previously-read GC mechanics (kept here for the detail):
**GC mechanics — verified against the current checkout (2026-10-01,
`src/store/gc.rs`, `src/util/temp_tag.rs`, `src/store/fs/delete_set.rs`).
@@ -278,17 +281,22 @@ is the prior art for cross-repo dedup among related repos — the
pooled-CAS posture (OQ-BL-05) generalizes it to unrelated repos and
to p2p replication.
### alknet's appfile external-store probe — the direct ancestor
### alknet's appfile external-store probe — historical reference only
`/workspace/@alkdev/alknet/docs/research/alknet-filesystem/`
(`alknet-blobs-external-store-probe.md`, `poc-summary.md`) — written
against an older iroh-blobs; conclusions re-verify in this Phase 0.
The load-bearing bits: the appfile shape (small blobs in kv/sqlite,
large on fs fallback), the filename↔hash mapping problem, and the
multi-backend dispatch pain the probe hit (which is this crate's
reason to own that dispatch). Old-data warning: any API or behavior
claims there describe an older upstream; re-check against
`/workspace/iroh-blobs` as checked out today.
against an older iroh-blobs. **Demoted 2026-10-02: alknet's code is
not a design input for this crate** (it is old, was poorly planned,
and has been decomposed into the alk* family; the substrate is
alkcall). What remains relevant from it is only the *shape* lineage —
the appfile pattern (small blobs in kv/sqlite, large on fs fallback,
filename↔hash mapping) and the multi-backend dispatch pain it
recorded, both of which this crate now owns by its own validated
design (POC #1, POC #3). Its re-verification ran 2026-10-02 as a
historical cross-check (`poc-largeblob-findings.md`, final section):
architectural claims about iroh-blobs held up, its performance claims
were asserted-never-measured and are superseded by POC #3's
first-party benchmark.
### alkcall — the substrate
@@ -321,8 +329,9 @@ hash-addressed, p2p sync *is* "announce hashes, diff hash sets, fetch
missing blobs verified," which is an op family on this store. Whether
that family lives here (feature-gated) or in the sibling that owns
replicator/gossip policy is the remaining question; the deciding input
is what alkgit's replicator actually needs and whether alknet's
appfile case ever transfers over the network. Even so, the *shape* of
is what alkgit's replicator actually needs and whether the
appfile-shaped consumers ever transfer blobs over the network. Even
so, the *shape* of
the ops (have/need + verified fetch, ACL-gated) is now firm enough to
plan against.
@@ -344,6 +353,18 @@ anchors the small-blob belief the dispatch leans on; streaming put/
get shape rides POC #3's findings (`Vec<u8>` was the *contract*
question, not the performance question).
**Updated 2026-10-02 (POC #3, `poc-largeblob-findings.md`):** the
small-blob belief is now first-party measured — sqlite beats fs ~9-10×
at 1-16 KiB, crossover ~128-256 KiB (finding A4; harness included in
the POC crate, re-runnable on real media). Default Phase 1 threshold
lean: 128 KiB (tunable constructor param per POC #1's contract).
Streaming put/get shape also landed (findings A1/A2/A3): the put seam
is `(Option<len>, stream)` with known-length as the encouraged one-pass
path and unknown-length staging-then-hashing (forced by the git
preamble); the ops layer's fetch handler should be the broadcast
fanout shape. An unknown-length-vs-known-length dispatch asymmetry
(finding A6) is a named Phase 1 decision, options recorded.
### OQ-BL-03: Hash abstraction — RESOLVED: canonical git-family hash
**RESOLVED 2026-10-01 (the hash simplification round).** The original
@@ -457,6 +478,24 @@ its verified whole-contents as git-blob-sha-256 pool entries (a
`git-blob-sha-256` digest over the reassembled content), so chunked
and whole-file paths dedup against each other in the one pool.
**Sharpened by POC #3 (2026-10-02, `poc-largeblob-findings.md`
finding A2):** per-range verification is *impossible* under the
canonical git-blob digest — the preamble hashes the length, so no
slice has any hash relationship to the whole (iroh gets range
verification from the bao *chunk tree*, not from BLAKE3 the
algorithm). Consequences carried: (i) local range reads need no
verification beyond transport-integrity slice digests the caller
carries out-of-band (the store's `read_range` returns one); (ii) the
only genuine consumer of a chunk-tree encoding is *networked ranged
fetch of large blobs* — making the bao-like encoding a transfer-layer
conditional, not a storage design; (iii) git blobs self-verify under
git's own model when read whole, and small-tier range reads slice
whole values in memory. The whole-file CAS posture is strengthened,
not changed. Also added to Phase 1 surface notes: a cheap
`stat(digest)` (length/type probe; gix's header-only read is git's
cheapest primitive and the POC's `LfsObject.len` already covers the
get side).
### OQ-BL-05: Pooling and GC — namespaces as reference sets over a flat CAS
**The decisions from the discussions (2026-10-01, principles 6; the
@@ -546,6 +585,24 @@ large/packed content — per-object granularity, range reads into
packed blobs, or both (interacts with OQ-BL-04's range-read
question — likely read them together).
**Resolved in shape by POC #3 (2026-10-02, gix-odb analysis in
`poc-largeblob-findings.md`):** gix-odb's own model confirms the two
halves — alternates/multi-pack-index give shared-read across packs
with no dedup/GC/large-blob/streaming value-add (gix-lfs is an empty
placeholder), and gitoxide's whole pack-ID-stability machinery exists
only because packs are immutable-with-stable-IDs; a per-object
appending pool has no such IDs and avoids that machinery entirely.
The evidence-supported resolution: **git objects enter the pool as
loose-equivalent kv entries (small tier, the common case); if packfile
serving is ever wanted, packs are stored as large blobs and served by
the store's range-read surface** (git's own .idx does per-object
offset lookup — the store is a flat byte server for the pack; finding
A2 says local range serving is fine). Dedup at pack granularity is
weak, but pack-level dedup across unrelated repos was never the goal;
object-level dedup lives in the loose tier. Phase 1 ADR should record
this as the default resolution, revisit only with a real consumer
demand.
### OQ-BL-06: POC register (draft)
Numbered POCs, opened as research reaches them (findings land in
@@ -557,19 +614,15 @@ is actually deciding:
|---|------|--------|-------|
| 1 | Backend-trait + dispatch shape (kv small / fs large); trait must include `list()` complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path; the hash-enum abstraction with git-blob-sha-256's domain-separated preamble + git-sha-1 tolerance case | **Passed 2026-10-02** — `poc-trait-dispatch-findings.md` (8 findings; code: `/workspace/alkblobs-trait-poc`) | findings landed |
| 2 | ~~Multi-hash store~~ **Absorbed into #1** (2026-10-01 hash round): the canonical-hash resolution removed the "two families coexisting" question; what remains (the preamble abstraction + SHA-1 tolerance) is POC #1's trait work | **Absorbed** (and validated by #1) | — |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) | **Pending** — reading-and-design POC + benchmark | findings file TBD |
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) | **Passed 2026-10-02** — `poc-largeblob-findings.md` (6 findings A1-A6 + pack-tension analysis + historical alknet probe re-check; code: `/workspace/alkblobs-largeblob-poc`) + `iroh-blobs-eval.md` | findings landed |
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Covered in miniature by #1** (2026-10-02): namespace tables + sweep + recover all validated single-threaded; the concurrency half (sweep-vs-put race, batch-scope pins) is Phase 1 implementation work, not a POC gate | `poc-trait-dispatch-findings.md` findings 3/7/8 |
Sequencing note: #1 is done (single opening worktree; multi-hash
subsumed; git-preamble hashing validated byte-exact vs the git CLI);
#3 is a reading-and-design POC against the current iroh-blobs checkout
plus the alknet probe re-verification, with its micro-benchmark
pull-out; #4's single-threaded core is covered by #1 (see the
register row) and the residual is implementation work. The pack
tension (OQ-BL-05) rides #3's reading rather than earning its own POC
yet.
Sequencing note: #1 and #3 are done; #2 absorbed/validated under #1;
#4's single-threaded core is covered by #1 (see the register row) and
the residual is implementation work. **The POC register is complete —
Phase 0's remaining step is convergence → Phase 1.**
**What each POC is deciding (updated 2026-10-02, post-POC-#1):**
**What each POC is deciding (updated 2026-10-02, post-POC-#3):**
- **POC #1 — DONE.** Decided the backend trait's minimum contract
(opaque byte keys, `list()` complete-by-contract, put-path
@@ -585,13 +638,21 @@ yet.
evaporated when BLAKE3's only reason-for-being (iroh-blobs
inheritance) was rejected, and its residual content validated under
#1.
- **POC #3 is the next (and last) POC** — reading-and-design against
the *current* iroh-blobs checkout (the store eval) + re-verifying
alknet's probe + the **micro-benchmark pull-out** (the "kv/sqlite
is faster for small blobs" belief underlies the dual-backend design
and traces to *old* alknet research — anchor it: small objects,
1KB-64KB, both backends, random + sequential). The pack tension
(OQ-BL-05) is analyzed here without its own POC.
- **POC #3 is DONE** — the last POC. Streaming large-blob path
validated (`LFSObject` put/get; known-length one-pass vs unknown-
length stage-then-hash, forced by the git preamble — finding A1);
the fanout seam (serve-while-persisting) works as a broadcast above
the backends (A3); range reads work with the verification gap
documented honestly (A2 — impossible under the canonical digest,
sharpening the chunk-tree question to "transfer-layer conditional");
the sqlite-vs-fs benchmark anchored the small-blob belief first-party
(~9-10× at 1-16 KiB, crossover ~128-256 KiB — A4); stage-file hygiene
tests caught a real leak (A5). The pack tension analyzed from
gix-odb and resolved in shape (packs-as-large-blobs served by range
reads). The iroh-blobs store eval landed as `iroh-blobs-eval.md`;
alknet's probe re-check is recorded as historical cross-check only
(per the 2026-10-02 clarification, alknet's code is not a design
input — the substrate is alkcall).
- **POC #4's single-threaded core is covered by #1** (namespace
tables + mark-and-sweep + protect-callback abort + delete-then-
recover all validated; `poc-trait-dispatch-findings.md` findings
@@ -609,17 +670,21 @@ in `docs/research/` regardless of where the code lives.
## Phase 0 plan (next steps)
1. ~~Write `iroh-blobs-eval.md`~~ — still open, rides POC #3's
reading against the *current* checkout (kv + flat backends, bao,
chunking; the GC mechanics eval is already done — see
`poc-trait-dispatch-findings.md`'s context and phase-0's verified
GC section).
2. **Re-verify alknet's probe** (`alknet-blobs-external-store-probe.md`,
`poc-summary.md`) against the current upstream; mark what carried
over — rides POC #3 too.
3. **Read gix-odb** (`/workspace/git-oxide/gix-odb`) for the alkgit
baseline: its storage layout, object model, and where a blob-store
crate under/beside it earns its keep.
4. **POC #3** (the last open POC): reading-and-design + the
sqlite-vs-fs micro-benchmark per OQ-BL-06.
5. Converge: recommended approach + final OQ register → Phase 1.
1. ~~Write `iroh-blobs-eval.md`~~ — **done 2026-10-02** (POC #3's
reading against the current checkout; `docs/research/
iroh-blobs-eval.md` — store traits/API, fs backend layout, bao/-
verification conclusions, small-blob handling, do-not-inherit and
keep-lists).
2. ~~Re-verify alknet's probe~~ — **done 2026-10-02**, recorded as a
historical cross-check only in `poc-largeblob-findings.md`
(alknet's code is not a design input; substrate is alkcall). The
one load-bearing output: the probe's small-blob performance claim
was asserted-never-measured; POC #3's benchmark is the anchor.
3. ~~Read gix-odb for the alkgit baseline~~ — **done 2026-10-02**
(POC #3's pack-tension analysis; conclusions folded into OQ-BL-05
and the findings doc; gix-lfs is a placeholder crate, the
large-blob lane is open as assumed).
4. ~~POC #3~~ — **done 2026-10-02**, `poc-largeblob-findings.md`.
5. **Converge**: recommended approach + final OQ register → Phase 1.
All gates passed; the register rows above carry the empirical
inputs (thresholds, seams, lessons) into the SDD's Phase 1.
+323
View File
@@ -0,0 +1,323 @@
---
status: passed
title: "POC #3 — large-blob streaming path + sqlite-vs-fs micro-benchmark"
last_updated: 2026-10-02
---
# POC: large-blob path (streaming, fanout, range reads) + small-object micro-benchmark — findings
> **POC register #3**, per `docs/research/phase-0.md` OQ-BL-06. Code:
> standalone crate `/workspace/alkblobs-largeblob-poc` (findings land
> here regardless, per the established convention). Companions from the
> same POC's reading work: `iroh-blobs-eval.md` (fresh-store eval
> against checkout e82cbdc / v0.103.0) and the alknet probe re-check
> (below — recorded as **historical cross-check only**).
> Date: 2026-10-02. Status: **passed**, 14 tests, clippy `-D warnings`
> clean, fmt clean. Digests cross-checked against the real `git
> hash-object` CLI (git 2.43), extending POC #1's validation to the
> streaming paths.
## What this POC set out to decide
From the phase-0 register:
1. **The large-blob path** — does an `LFSObject`-shaped `(len, stream)`
put/get work at the store layer, and what does the git-blob preamble
do to streaming hashing?
2. **The fanout seam** — "serve the fetch while persisting on receipt"
(rudolfs' `fanout()`; iroh's `get_blob_ranges_impl`) — does it work
at our layer, above the backends?
3. **Range reads on the fs backend** — cheap partial reads with *some*
verification story (OQ-BL-04).
4. **The small-object sqlite-vs-fs micro-benchmark** — anchor the
"kv is faster than fs for small blobs" belief first-party
(OQ-BL-02's "dual-belief anchor").
5. **The pack tension** (OQ-BL-05) — analyzed from the gix-odb read,
no POC of its own.
Note on standing: **alknet's code is not a design input** — it is old,
poorly planned, decomposed into the alk* family; its probe docs are
treated as references to old research only. The substrate this crate
rides is alkcall (like alkhttp/alksocks/alktty/alktunnels).
## Result summary
| Question | Verdict |
|----------|---------|
| Streaming put (`LFSObject` shape) | **Yes** — two paths: known-length one-pass; unknown-length stage-then-hash (forced by the preamble, finding A1) |
| Streaming get without whole-blob load | **Yes** — fs arm streams a file handle; kv arm one bounded read |
| Fanout (serve while persisting) | **Yes** — one reader task, broadcast channel, store arm + N subscriber arms (finding A3) |
| Range reads on fs | **Yes** — pread-window only; **verification has a real gap** (finding A2) |
| sqlite faster than fs for small blobs | **Confirmed and measured** — ~10× below 16 KiB; crossover ~128-256 KiB (B-finding) |
| Digest = git oid on streaming paths | **Byte-exact vs `git hash-object`** for both put paths |
| Stage-file hygiene | **Verified** — failures (stream error, overflow, early-EOF) leave zero stage files (finding A5) |
## Findings
### Finding A1: the git-blob preamble splits streaming into exactly two paths — and the split is forced, not chosen
The canonical derivation is `H("blob <len>\0" + content)`. The length
is *inside the hashed input*, so:
- **Known length** ⇒ write the preamble into the hasher at construction,
feed content chunks incrementally, finalize at stream end: **one pass
over the content**, no staging needed for hashing (staging to a temp
file still happens for the fs tier, but the hash is done by the time
the rename commits). This covers the dominant git and protocol cases:
a pre-staged file has its length from `stat`, a network offer carries
a size, an HTTP PUT can carry `Content-Length`.
- **Unknown length** ⇒ the preamble can't be written until the stream
ends, so content must be staged (file, not memory) and hashed in a
second pass over the staged bytes. **Two passes are forced by the
derivation**, not by our choice of hash.
iroh-blobs independently arrives at the same stage-then-encode shape
(their `add_stream` buffers then computes outboard in a second pass,
`fs/import.rs:290-419`), for bao-tree reasons instead of
preamble reasons. Their `import_bao_reader` sidesteps it by knowing
both hash and size upfront (a protocol-level offer) — the same
"known-length is the normal case" assumption, enforced by their wire
format.
**Verdict carried to Phase 1:** the store's put seam should be
`(Option<len>, stream)` exactly as built. Known-length puts must be the
encouraged path (documented as such on the API); unknown-length is
supported but slower by one page-cache read. Do *not* invent a
length-prefix-on-the-wire scheme to "fix" unknown-length puts — the
wire carries sizes naturally (it must, for the CAS routing).
### Finding A2: range reads cannot verify against the canonical digest — the preamble includes the length, so slices have no hash relationship to the whole
Confirmed at implementation level (and independently by the
iroh-blobs eval): `blob <len>\0 + content` gives *no* verifiable
relation between `H(whole)` and any `content[a..b]`. iroh gets
per-range verification from the **bao chunk tree**, not from BLAKE3 the
algorithm — the tree is the mechanism, the hash is just its leaf
function.
POC #3's `read_range` therefore returns the slice plus a *slice digest*
(`git-blob-sha-256` over the slice — a transport-integrity check the
caller carries out-of-band, e.g. from the offering side). Consequences:
- **Local-store range reads need no more** — the OS pages honest media;
EIO surfaces read errors; git objects (the small tier) self-verify
when git reads them whole.
- **The only genuine consumer of a chunk tree is networked large-blob
transfer where the caller wants per-chunk integrity.** This sharpens
OQ-BL-04's "conditional": the bao-like encoding is not about
*storage* at all, it is a *transfer* encoding. If/when p2p git needs
verified ranged fetch of LFS-shaped artifacts, the encoding layer
(BLAKE3-confined per OQ-BL-03 tier 3) registers its whole-content
root back into the pool as an ordinary git-blob-sha-256 entry.
- Practical corollary: for git-blob-shaped content, byte-range *serving*
is fine (packfiles will want it; see pack tension below) but per-range
*verification* is out of scope for the store layer today.
### Finding A3: the fanout seam is a broadcast channel above the store, and late-joiners degrade to get-after-commit
Built `put_stream_fanout(stream, n)`: one reader task consumes the
source stream exactly once and broadcasts each chunk to (a) the store
put arm and (b) N subscriber feeds (`mpsc<io::Result<Bytes>>`);
subscribers read while the put is still landing. Tests show 3
subscribers each receive the full 256 KiB byte-identical while the put
completes, and the put receipt matches the git oid.
Design notes validated:
- **One read, many consumers** — the source stream must never be
polled twice; anything that wants a second copy subscribes. This
matches iroh's `get_blob_ranges_impl` shape (network decode loop
joined with the store import) and rudolfs' `fanout()` (two lock-step
copies from one stream). All three arrive at "broadcast from a single
received stream" — treat that as the settled seam shape.
- **Backpressure discipline:** bounded channels (8 items in the POC)
make a slow subscriber stall the broadcast — which is *correct* for a
put (the put must land) but means subscriber count is bounded and
slow-subscriber handling is a policy question for the ops layer
(drop-and-late-join is the escape hatch; see next note).
- **Late joiners:** a consumer arriving after the live stream has
passed simply waits for the commit then does a normal verified
`get_stream` — the CAS guarantees the bytes are identical. The POC
didn't build this (subscriber set is fixed at start) but the degrade
path is exactly the normal get; worth documenting in the ops layer.
### Finding A4 (benchmark): sqlite beats fs decisively for small blobs; crossover ~128-256 KiB; fsync discipline dominates both arms
Release build, per-arm timing loops, 2k distinct keys/values random
working set, `synchronous=NORMAL` WAL sqlite vs stage+rename fs puts /
open+read fs gets. Representative run (3 s/arm):
```
backend size puts/s gets/s put p50µs put p99µs get p50µs get p99µs
sqlite 1024 48319 49508 20.0 37.0 19.0 34.0
fs 1024 5306 3357 205.0 302.0 288.0 384.0
sqlite 16384 14129 14166 70.0 125.0 68.0 107.0
fs 16384 4203 2399 257.0 395.0 433.0 599.0
sqlite 65536 4028 4048 252.0 364.0 243.0 344.0
fs 65536 3212 1852 339.0 507.0 540.0 688.0
sqlite 131072 2067 2042 494.0 698.0 482.0 699.0
fs 131072 1892 1652 593.0 879.0 619.0 809.0
sqlite 262144 981 1003 1033.0 1792.0 994.0 1176.0
fs 262144 1244 1576 822.0 1181.0 741.0 964.0
```
Readings:
- **~9-10× sqlite advantage at 1-16 KiB** (the git small-blob regime —
iroh's DESIGN.md claim and the old citations hold for our workload).
- **Crossover at ~128-256 KiB**: at 256 KiB fs wins puts (no B-tree
rewrite of a 256 KiB row) and gets. **The dispatch threshold is a
tunable**, placed anywhere in 64-256 KiB; POC #1's contract already
makes it a constructor parameter. Phase 1 default: 128 KiB (midpoint
of the flat zone) — revisit with the storage-medium reality of real
consumers (tmpfs vs disk; the POC SSD regime is one data point).
- **fs puts cost ~200 µs floor even at 1 KiB** — that is the
create+write+rename syscall chain; batching stage files (a shared
preallocated pool) or O_TMPFILE could shave it, but the sqlite arm
makes the effort moot below the crossover.
- Caveats recorded honestly: single-node tmpdir on this dev box
(backend-specific tuning — e.g. iroh's redb batch windows — could
shift the fs curve), no fsync-in-loop on the fs arm (matching NORMAL
WAL durability, sqlite arm likewise not full-fsync; both arms are
"crash-tolerant-ish", so the comparison is fair *as measured*: both
are the durability tier a store would ship).
- The old alknet claim "<100 KB faster in sqlite" is **confirmed in
direction, refined in boundary**: the honest crossover is higher
(~128-256 KiB) and sqlite is never *pathological* below it.
### Finding A5: stage-file hygiene is testable and the tests earned their keep (again)
The `stream_failure_cleans_stage_files` test asserted zero files left in
the data dir after a mid-stream connection failure — and caught a real
bug: the original `chunk?` error path exited *before* discarding the
stage file, leaking one temp file per failed put. Fixed by restructuring
the loops to route all failure paths (stream error, overflow, early-EOF,
write error) through one discard. Two lessons compound with POC #1's:
- **Resource-hygiene invariants need dedicated tests**; code review
missed what the exact-count assertion caught (echoing POC #1's exact
deletion counts).
- The store's put path now has the shape Phase 1 should codify: **every
failure path converges on stage-discard; every success path converges
on commit-rename** — no early returns that bypass cleanup.
### Finding A6: re-put dedup across the two put paths holds; unknown-length re-put of a small blob lands on fs (a dispatch asymmetry worth one Phase 1 decision)
The same content put via known-length (small) and unknown-length
(staged) paths produces the identical digest and coexists as one entry.
But routing differs: an unknown-length put *always* stages through the
fs arm in the POC (routing must wait for the realized size), so a small
blob arriving via an unknown-length stream occupies an fs file where the
same digest's known-length put would occupy a kv row. If both happen,
the pool holds the content twice (different physical tiers).
Options for Phase 1, decided there, not here (recorded honestly — the
POC as built has the asymmetry): (a) post-commit tier-correction sweep
(move small fs entries into kv during idle GC), (b) require `len` for
the small tier (reject unknown-length small puts — simple, slightly
hostile to streaming producers), (c) accept the duplication (it's
temporary if (a) runs) — or (d) route *unknown-length* puts through a
memory buffer up to the threshold before staging (bounded-memory
reconciliation: small unknown puts never touch fs). Given the
benchmark (small = cheap everywhere), (d) looks natural but should be
weighed against batch-scope pinning from POC #1 finding 3.
## The pack tension (OQ-BL-05 input, from the gix-odb read — analysis only)
gix-odb (v0.84) audit conclusions, in alkblobs terms:
- **Git's own pooling is administratively shared packs**: alternates/
multi-pack-index give one flat OID→location lookup across many packs
(slot-map + `ArcSwap` snapshots for lock-free reads), but there is no
cross-store dedup, no refcounting, no GC, no large-blob handling, no
streaming reads — every object decompresses whole into memory (writes
stream; reads do not). gix-lfs is an empty placeholder crate. The
value-adds alkblobs planned are confirmed unclaimed.
- **Packs are pack-efficient because they delta-chain within
themselves; a per-object CAS pool fights that in two ways.** First,
delta chains want objects stored *near their bases*; a flat per-oid
pool scatters them, so serving a delta-heavy pack read pattern off the
pool re-inflates from scattered loose entries (loose objects are
git's *dedup unit and fs bottleneck simultaneously* — gix's own
discovery docs flag loose-object proliferation as a server-scale
performance problem). Second, packs are immutable-with-stable-IDs,
and gitoxide's whole generation/consolidation machinery exists to
rebind pack IDs when packs churn — a blob pool that *appends per
object* avoids that machinery entirely (there are no IDs to rebind;
the address is the content).
- **Resolution shape that the evidence supports** (Phase 1 ADR input):
git objects enter the pool as **loose-equivalent kv entries** on the
small tier (the common case — git's median blob is small), and the
pack question is *deferred to serving*, not storage: if/when serving
git fetches wants packfile byte-range reads (OQ-BL-05's pack
tension), the options are (i) store packs as large blobs and serve
ranges *out of the pack* via the store's range-read (the pack itself
is the container; per-object addressing rides git's own .idx — the
store is just a flat byte server for it, which finding A2 says is
fine locally), or (ii) keep loose-per-object as the only form and
let consumers synthesize packs above. (i) collapses the tension
without new store concepts: **the pool stores whole files — including
packs — and ranged serving is its range-read surface.** Dedup at pack
granularity is weak, but pack-level dedup across unrelated repos was
never the goal; object-level dedup lives in the loose tier.
- **Concrete consequence for the store API:** a `metadata(oid)`-shaped
cheap size/type probe (gix's header-only read is git's cheapest
primitive) should exist in Phase 1's surface — the POC's
`LfsObject.len` covers the get side already; make it an explicit
`stat(digest)` op.
## Alknet probe re-check (historical cross-check only — not design input)
Run for completeness, per the register line that flagged it; per the
session's clarification, **alknet's old code/probe conclusions carry no
decision weight for this crate** (decomposed substrate; alkcall is the
substrate). What the re-check contributes, stripped to the factual:
- The probe's architectural claims about iroh-blobs (no Store trait;
`Command` enum seam; three-actor pattern; the four pub(crate)
blockers; the four redb tables; the 16 KiB hybrid) all verified
against the current checkout — the eval above re-states them
independently.
- The probe's performance claims ("faster than the filesystem for
small items") were **asserted-with-citation, never measured** — the
POC #3 benchmark is the first first-party anchor (finding A4).
- Stale: the version-gap open question (alknet-core is on iroh 1.0
now); `run_gc` is *not* downstream-importable (private module) —
irrelevant here, we are not reusing their GC.
## Consequences for the phase-0 OQ register
- **OQ-BL-02 (dispatch):** the small-blob belief is now first-party
measured (A4). Threshold ~128 KiB default; benchmark harness exists
(`cargo run --release --bin bench_small`) for re-running on real
media. Remaining above-trait concerns (migration policy, per-
namespace config) unchanged — no POC gate.
- **OQ-BL-04 (verification/chunking):** sharpened — range verification
is impossible under the canonical digest (A2); the chunk-tree
encoding is a *transfer-layer* conditional whose only consumer is
networked ranged fetch. Whole-file CAS posture unchanged and
strengthened. `stat(digest)` surfaces in the Phase 1 API notes.
- **OQ-BL-05 (pooling/GC):** pack tension resolved in shape — packs
(if stored) are just large blobs served by range reads; object-level
dedup lives in the loose tier. No new store concepts needed.
- **POC register:** #3 is done; no POCs remain open. Phase 0's next
step is convergence → Phase 1.
## POC quality notes
- 14 tests (12 integration `large_blob.rs`, 2 CLI cross-check
`git_cli.rs`); digest validation against real `git hash-object`
(`--stdin`, repo initialized `--object-format=sha256` — the POC #1
gotcha holds: the flag is not a hash-object flag).
- clippy `-D warnings` clean (0 warnings), `cargo fmt` clean.
- Deliberate scope limits: single-put-at-a-time (no batch concurrency —
POC #1's finding 3 covers the pinning story), no persistence-across-
restart tests (same rationale as POC #1), the fanout seam is
subscriber-count-fixed (no late-joiner implementation — degrade path
documented instead), benchmark is one dev box + one working-set
profile (harness is included for re-running).
- Known micro-detail: `has()` on the kv arm uses a prepared
`SELECT 1 ... EXISTS`; `get_stream` dispatch does two kv round-trips
(size, then get) — a schema returning the value directly would avoid
one; left as-is to keep the size probe identical for both tiers in
the benchmark comparison.