docs(research): POC #3 passed — large-blob streaming + sqlite-vs-fs benchmark
- POC #3 (register OQ-BL-06) executed and passed: streaming LFSObject-shaped put/get at the store layer; fanout seam (serve-while-persisting) validated as broadcast above the backends; fs range reads with the verification gap documented (preamble makes slice-vs-whole verification impossible — chunk-tree encoding sharpened to a transfer-layer conditional) - sqlite-vs-fs micro-benchmark: first-party anchor for the small-blob belief (~9-10x sqlite at 1-16 KiB, crossover ~128-256 KiB, 128 KiB default lean); harness re-runnable on real media - Stage-file hygiene tests caught a real leak (one temp file per failed put) — fixed; all failure paths converge on discard - Pack tension resolved in shape from the gix-odb read: packs as large blobs served by range reads; object-level dedup lives in the loose tier - iroh-blobs-eval.md written against the current checkout (e82cbdc / v0.103.0): actor/message-protocol pattern, fs backend layout, verification, do-not-inherit + keep lists - alknet probe re-check demoted to historical cross-check only: alknet's code is not a design input (decomposed; substrate is alkcall); its small-blob perf claim was asserted-never-measured, superseded by A4 - phase-0 register complete: no open POCs; next step is convergence -> Phase 1 Verification: POC crate (standalone /workspace/alkblobs-largeblob-poc) 14 tests pass, clippy -D warnings clean, fmt clean; digests byte-exact vs the real git hash-object CLI on both streaming put paths
This commit is contained in:
1 parent
493516424a
commit
6b278fb490
3 files changed
+653
-67
No files matched your search
@@ -0,0 +1,198 @@
|
||||
---
|
||||
status: draft
|
||||
title: "iroh-blobs store eval — against the current checkout"
|
||||
last_updated: 2026-10-02 (POC #3 reading; checkout e82cbdc, v0.103.0)
|
||||
---
|
||||
|
||||
# iroh-blobs store eval (current checkout)
|
||||
|
||||
Evaluated against `/workspace/iroh-blobs` @ `e82cbdc` ("Release 0.103.0",
|
||||
2026-06-15 — the checkout *is* the 0.103.0 release, so alknet's older
|
||||
probe claims about "published 0.103" verify directly against it).
|
||||
Scope: the store layer (`src/store/**`, `src/api/**`, `src/util/**`).
|
||||
The wire surface (tickets, postcard, provider protocol) is out of scope
|
||||
per convention 7 — we are not adopting it. This eval feeds OQ-BL-02
|
||||
(backend shape), OQ-BL-04 (verification/chunking), and the POC #3
|
||||
large-blob design. GC mechanics were already verified separately
|
||||
(see phase-0 §Prior art; the gc.rs read predates POC #1).
|
||||
|
||||
## The headline: there is no Store trait — and that is instructive
|
||||
|
||||
iroh-blobs 0.103 has **no `Store` trait at all**. The store API is a
|
||||
message protocol: an enum of RPC requests (`src/api/proto.rs:90-139`,
|
||||
`#[rpc_requests]`-derived `Command`), and each backend (MemStore,
|
||||
FsStore, ReadonlyMemStore) is an *actor* consuming that command set
|
||||
over `irpc` channels. The user-facing `Store` struct is just a client
|
||||
handle; `Store::connect/listen` make the same message enum work over
|
||||
the network. The abstraction is "same messages locally and remote."
|
||||
|
||||
Lessons for alkblobs:
|
||||
|
||||
- **The command-enum-as-contract is a working pattern** (it is what
|
||||
makes iroh-blobs' remote/local symmetry possible), but it carries the
|
||||
whole irpc machinery into the store layer and *is* the reason their
|
||||
store is welded to their transport. POC #1's lean `Backend` trait is
|
||||
the smaller hammer for our split: the trait stays dumb (opaque bytes),
|
||||
and any ops surface above it (alkcall) speaks its own protocol.
|
||||
- **Re-borrowed conclusion, not mechanism: per-hash serialized actors.**
|
||||
FsStore runs one serialized state machine *per hash* (entity manager,
|
||||
`src/store/fs/util/entity_manager.rs`) with parallelism across hashes;
|
||||
per-hash state is a `watch::Sender<BaoFileStorage>` (cheap snapshots,
|
||||
observable progress, idle-recycled actor pool). This is a genuinely
|
||||
good concurrency shape for Phase 1's sweep-vs-put story: "one writer
|
||||
per content address, readers snapshot-observe" matches CAS semantics
|
||||
naturally. iroh's DeleteSet was already our named prior art; the
|
||||
per-hash-actor adds a second reusable conclusion.
|
||||
|
||||
## What the fs backend actually looks like on disk (0.103)
|
||||
|
||||
```
|
||||
blobs.db redb — hash → EntryState (+tags, + inline data/outboard)
|
||||
data/<hash>.data complete + partial canonical data
|
||||
data/<hash>.obao4 outboard (bao pre-order, 64-byte pairs)
|
||||
data/<hash>.sizes4 per-chunk-group size slots
|
||||
data/<hash>.bitfield blake3-checksummed postcard bitfield (coverage)
|
||||
temp/<uuid>.temp staging (must be same-device as data/)
|
||||
```
|
||||
|
||||
- **Entry state machine per blob** (`src/store/fs/entry_state.rs`):
|
||||
`Complete { data_location: Inline|Owned|External, outboard_location:
|
||||
Inline|Owned|NotNeeded }` / `Partial { size: Option<u64> }`. The
|
||||
inline/owned split is the small-blob/large-blob dispatch *inside one
|
||||
backend*: ≤16 KiB data lives entirely in the redb `inline-data` row;
|
||||
≥~1 MiB is data file + outboard file; between, data file + inline
|
||||
outboard. Thresholds are `InlineOptions { max_data_inlined: 16 KiB,
|
||||
max_outboard_inlined: 16 KiB }` (`fs/options.rs:71-78`).
|
||||
- **16 KiB is exactly one bao chunk group** (16 × 1024 B chunks), so
|
||||
every blob ≤ 16 KiB has an outboard of size *zero* — a "small blob"
|
||||
is one redb row keyed by plain `blake3(content)`, no tree material at
|
||||
all (`fs/import.rs:160-221`, `DESIGN.md:71`). iroh couples "hash
|
||||
overhead" and "storage tier" at the same boundary; under our git-oid
|
||||
derivation there is no outboard to elide (whole-blob digest only) and
|
||||
the tier boundary is a pure dispatch choice — POC #3's benchmark says
|
||||
~64-128 KiB, not 16 KiB (see `poc-largeblob-findings.md`, B-finding).
|
||||
- **Crash-consistency story, borrowed as a conclusion** (`DESIGN.md:
|
||||
126-142`, "files are hard"): persist data+outboard+sizes *first*, the
|
||||
checksummed bitfield *last*, with truncate-to-zero-on-read marking a
|
||||
bitfile dirty under crash. If the bitfield is missing/corrupt, the
|
||||
entry re-validates by full local hash check. Lazy persistence (on
|
||||
actor idle/shutdown only), never per-chunk fsync. For alkblobs: our
|
||||
put path is *stage-temp → hash → one rename*, which needs no coverage
|
||||
bitfield at all — we only ever need this machinery if we adopt
|
||||
partial/streaming-download entries (range-serve before complete).
|
||||
Recorded as a conditional, like bao.
|
||||
- **Temp-file flow** (<`fs/import.rs`>): byte-stream puts stage to a
|
||||
temp file (after a 16 KiB in-memory peek), path-import reflinks-or-
|
||||
copies (`reflink-copy` crate) to temp, outboard encodes in a second
|
||||
pass, then `fs::rename` atomizes both files before the redb row lands
|
||||
*awaiting* the transaction commit. `ExportMode::TryReference` is the
|
||||
mirror: rename the store's own file into the caller's tree and mark
|
||||
the entry `External`. (TryReference is a neat dedup trick for the
|
||||
appfile/workspace consumer — "move into the store without copying" —
|
||||
Phase 2 note, not Phase 1.)
|
||||
- **Metadata batching**: the redb actor batches read commands (10k msgs
|
||||
/ 1 s window) and write commands (1000 / 500 ms) into single
|
||||
transactions (`meta.rs:785-856`) — "the secret to fast download
|
||||
speeds is to not touch the metadata database at all" (`DESIGN.md:61`).
|
||||
sqlite's WAL gives comparable batching for free at the smaller scale
|
||||
our store keys live at; the *conclusion* (batch metadata writes,
|
||||
never per-blob-fsync) transfers.
|
||||
|
||||
## Verification & chunking (OQ-BL-04 input)
|
||||
|
||||
- Fixed 1024 B bao chunks, grouped 16-wide (iroh `IROH_BLOCK_SIZE`,
|
||||
`store/mod.rs:17-18`); no content-defined chunking exists in-tree.
|
||||
- Outboard encoding is a **custom single-pass encoder over blake3
|
||||
hazmat APIs** (`src/util.rs:224-336`) — post-order chunk iteration,
|
||||
`set_input_offset`/`finalize_non_root` merges, buffered 1 MiB reads.
|
||||
iroh wrote their own rather than using bao-tree's `CreateOutboard`;
|
||||
the *shape* of "encode outboard while streaming through chunks once"
|
||||
is the pattern any chunk-tree encoder follows, bao or CDC.
|
||||
- **Get verification is subtree-local**: `export_bao(hash, ranges)` runs
|
||||
`traverse_ranges_validated` and lazily reads only the parent pairs the
|
||||
requested chunks need. Byte-range `export_ranges` is deliberately
|
||||
*unvalidated* (bitfield-superset check only; `DataReader` is documented
|
||||
"a reader for the unvalidated data file", `fs.rs:337-343`).
|
||||
- **For our canonical digest this exposes a hard fact** (POC #3 finding
|
||||
A2, reproduced independently): the git-blob derivation verifies *the
|
||||
whole content* (preamble includes the length). There is no sub-range
|
||||
relationship between `blob <len>\0+content` and any slice of it, so
|
||||
range reads under our hash cannot be verified against the blob's own
|
||||
digest, period. iroh gets range-verification *from the chunk tree*, not
|
||||
from the hash algorithm. If ranged transfer ever needs per-range
|
||||
verification, the options stay as OQ-BL-04 lists them (caller-carried
|
||||
slice digests = transport-integrity only; conditional chunk-tree
|
||||
encoding; or backend-native). POC #3's verdict: the local-store case
|
||||
doesn't need it (pages fail via EIO, git blobs self-verify on read);
|
||||
the *networked large-blob* case is the only real consumer of a chunk
|
||||
tree — which matches the "conditional, confined to the encoding layer"
|
||||
posture already recorded.
|
||||
|
||||
## Small-blob handling & streaming (POC #3 design input)
|
||||
|
||||
- **Streaming put exists in three shapes**; none hashes on the fly:
|
||||
`add_stream` (bidirectional `ImportByteStream` channel; stages to
|
||||
memory then temp file, computes outboard in a **second pass**,
|
||||
`fs/import.rs:290-352, 372-419`), `import_bao` (pre-verified items,
|
||||
no re-check), and `import_bao_reader` (network decode-verify-push,
|
||||
`blobs.rs:446-488` — the provider's push handler). There is **no
|
||||
put-from-`AsyncRead` API**; callers adapt to `Stream<io::Result<Bytes>>`.
|
||||
- POC #3 found the same constraint from our side *before* reading this
|
||||
in detail: the git preamble needs the length at hash time, so an
|
||||
unknown-length stream *must* stage-then-hash. iroh's architecture
|
||||
arrives at the same stage-then-encode shape (for bao-tree reasons
|
||||
instead). **The one-pass alternative exists only for known-length
|
||||
puts** — which is exactly the git/alkgit dominant case (a staged file
|
||||
has a stat; a protocol offer carries a size). Unknown-length remains
|
||||
the two-pass shape for both crates; POC #3's finding A1 records why
|
||||
(and that the second pass can read from the *staged file*, so it's
|
||||
page-cache-cheap, not network-repetitive).
|
||||
- **Serve-while-persisting is not a bolt-on in iroh — it is the
|
||||
download path**: `get_blob_ranges_impl` opens the import *before*
|
||||
reading from the network and `tokio::try_join!`s the network decode
|
||||
loop with the store import (`src/api/remote.rs:884-944`). POC #3's
|
||||
fanout seam (finding A3) rebuilt this at our layer with a broadcast
|
||||
channel; the mechanism differs (single reader broadcasting vs their
|
||||
per-item verified forwarding) but the conclusion matches: the fetch
|
||||
handler should be "decode-verify → broadcast → (store arm + subscriber
|
||||
arms)", all driven from one receive loop.
|
||||
- **`BlobReader`** implements `AsyncRead + AsyncSeek` built on
|
||||
`export_ranges` per read call (`api/blobs/reader.rs`); seek-from-end
|
||||
is an upstream TODO. Our fs range-read surface covers the same need.
|
||||
|
||||
## Performance posture (no in-tree benches; code choices as evidence)
|
||||
|
||||
No `benches/`, no criterion in the checkout. The perf story is in
|
||||
DESIGN.md + code micro-choices: redb write/read batching; partial
|
||||
entries live entirely outside redb until completion; zero-copy
|
||||
`Bytes` plumbing end to end; reflink-first import/export; 1 MiB copy
|
||||
buffers with `yield_now()` fairness; per-hash channel caps bounding
|
||||
backpressure. The two transferable *conclusions* for us: batch
|
||||
metadata transactions (never per-blob fsync), and zero-copy/`Bytes`
|
||||
at the streaming seams. Our own sqlite-vs-fs numbers are in the POC #3
|
||||
findings (first-party, measured).
|
||||
|
||||
## Do-not-inherit list (re-confirmed against 0.103)
|
||||
|
||||
- The `Command`-actor + irpc store abstraction (welds store to their
|
||||
transport stack).
|
||||
- redb-specific entry-state modeling if we go sqlite (their tables and
|
||||
`EntryState` postcard encoding are theirs; our sqlite schema is
|
||||
`key → (value, size)` and our dispatch is store-layer policy).
|
||||
- Fixed 1024 B chunk trees as the default granularity (principle 7
|
||||
already rejects; nothing new in the eval changes it).
|
||||
- Tickets/postcard/provider — out of scope as always.
|
||||
|
||||
## What we keep (conclusions only)
|
||||
|
||||
1. Per-hash serialized state machines + snapshot-observation for
|
||||
concurrent access (entity-manager pattern).
|
||||
2. Stage-temp → atomic rename as the universal put commit primitive.
|
||||
3. Metadata write batching (amortize sync cost).
|
||||
4. Zero-copy `Bytes` at streaming seams.
|
||||
5. Serve-while-persisting as the fetch-handler shape (POC #3 validated
|
||||
at our layer).
|
||||
6. Crash-consistency ordering (data first, coverage/index last, with
|
||||
dirty-marking) — only if partial entries ever exist.
|
||||
7. The inline threshold concept, re-sized by our own benchmark
|
||||
(~64-128 KiB crossover, git-blob workload) instead of their 16 KiB.
|
||||
+132
-67
@@ -1,9 +1,10 @@
|
||||
---
|
||||
status: draft
|
||||
last_updated: 2026-10-02 (POC #1 executed and passed — see
|
||||
poc-trait-dispatch-findings.md; POC #2 absorbed earlier; register
|
||||
updated; earlier rounds: hash simplification, GC/pooling, dedup/p2p,
|
||||
setup draft)
|
||||
last_updated: 2026-10-02 (POC #3 executed and passed — see
|
||||
poc-largeblob-findings.md plus iroh-blobs-eval.md; POC #1 executed and
|
||||
passed — see poc-trait-dispatch-findings.md; POC #2 absorbed earlier;
|
||||
register complete; earlier rounds: hash simplification, GC/pooling,
|
||||
dedup/p2p, setup draft)
|
||||
---
|
||||
|
||||
# alkblobs — Phase 0 (Exploration)
|
||||
@@ -28,7 +29,7 @@ git-family derivation; legacy SHA-1-tolerant), with pluggable backends
|
||||
fallback), split from any transport/protocol layer, with auth-gated
|
||||
network operations riding the alkcall seam when they exist.
|
||||
|
||||
**Why this crate exists (three converging consumers):**
|
||||
**Why this crate exists (the converging consumers):**
|
||||
|
||||
1. **alkgit** (planning phase) needs an object backend that is *not*
|
||||
its proposed default file-based backend ("kind of gross") — blobs
|
||||
@@ -41,21 +42,23 @@ network operations riding the alkcall seam when they exist.
|
||||
first-class entries with zero indirection — the conflict was an
|
||||
artifact of inheriting iroh-blobs' BLAKE3-only choice, and
|
||||
once that inheritance was rejected the conflict went with it.
|
||||
2. **The alknet rewrite** (the original alk* project, being decomposed
|
||||
and improved) needs the "appfile" external-store shape: small blobs
|
||||
in a kv/sqlite store (faster than the filesystem for small items —
|
||||
true beyond sqlite: it applies to the kv store iroh-blobs uses
|
||||
too), large blobs on a filesystem fallback, with filename↔hash
|
||||
mapping. alknet's own research hit the core awkwardness of
|
||||
dispatching across more than one backend at a time — which is
|
||||
exactly a store-layer problem this crate should own.
|
||||
2. **Workspace/file-shaped storage in the alk* family** (the "appfile"
|
||||
shape, originally surfaced by old alknet research — historical
|
||||
lineage only; the consumers below are current): small blobs need a
|
||||
store faster than the filesystem (now first-party measured, POC #3
|
||||
finding A4: sqlite ~9-10× faster at 1-16 KiB), large blobs want a
|
||||
filesystem fallback, with a filename↔hash mapping above. Dispatching
|
||||
across backends at the same time is a store-layer problem this crate
|
||||
should own (OQ-BL-02; POC #1 validated the seam, POC #3 the
|
||||
streaming + threshold evidence).
|
||||
3. **Agent workspaces** — many agents working simultaneously in
|
||||
workspaces that are mostly identical to each other's, each making
|
||||
small edits in specific areas. Per-agent stores multiply the near-
|
||||
identical content; a shared content-addressed store dedups it by
|
||||
construction, and a path→hash manifest per workspace makes the
|
||||
per-agent delta just its changed hashes (the alknet-filesystem
|
||||
probe shows the manifest mapping is straightforward).
|
||||
per-agent delta just its changed hashes (a path→hash manifest is
|
||||
straightforward by construction; the old alknet probe's manifest
|
||||
mapping sketch showed the same).
|
||||
|
||||
**The p2p shape and what it pins (added 2026-10-01, from the dedup
|
||||
discussion).** Two alkgit use cases exist: self-hosted git (where
|
||||
@@ -158,10 +161,10 @@ Almost nothing is pinned — deliberately. The postures agreed so far:
|
||||
as primary, legacy SHA-1 tolerated, BLAKE3 demoted to a conditional
|
||||
large-blob encoding consideration. Details and the reasoning trail
|
||||
in OQ-BL-03.
|
||||
- **Backend pluralism assumed, not designed.** kv/sqlite for small
|
||||
blobs, filesystem fallback for large — the appfile shape — but
|
||||
*how* multi-backend dispatch works is exactly what alknet's research
|
||||
ran into, so it earns research and probably POCs (OQ-BL-02).
|
||||
- **Backend pluralism assumed, then designed.** kv/sqlite for small
|
||||
blobs, filesystem fallback for large — the appfile shape — with
|
||||
*how* multi-backend dispatch works validated by POC #1 (the lean
|
||||
trait seam) and POC #3 (streaming + threshold evidence) (OQ-BL-02).
|
||||
- **Pooled CAS, not per-repo stores** (agreed in the 2026-10-01
|
||||
dedup discussion — the load-bearing scope decision so far): repos
|
||||
and agent workspaces are sets of hash references over one shared
|
||||
@@ -181,13 +184,13 @@ Almost nothing is pinned — deliberately. The postures agreed so far:
|
||||
### iroh-blobs — the shape inspiration (evaluated, not the base)
|
||||
|
||||
`/workspace/iroh-blobs` (fresh upstream checkout; read-only reference).
|
||||
Specifically `src/store`: the kv backend (small blobs, in-process) and
|
||||
flat file backend (large blobs) split; bao outboard encoding and
|
||||
verification flow; chunking. Its BLAKE3-only hashing, tickets, postcard
|
||||
serialization, and provider protocol are the design-welded choices we
|
||||
diverge from. **The store eval should be written against the current
|
||||
checkout** — alknet's older research refers to an older iroh-blobs and
|
||||
its conclusions must be re-verified rather than inherited.
|
||||
Specifically `src/store`: **the full store eval is now written —
|
||||
`docs/research/iroh-blobs-eval.md` (2026-10-02, against checkout
|
||||
e82cbdc / v0.103.0)**: no Store trait (message-protocol actor pattern),
|
||||
fs backend on-disk layout (redb entry-state + inline thresholds + crash
|
||||
ordering), verification (subtree-local, chunk-tree-driven), streaming
|
||||
shapes, and the do-not-inherit / keep (conclusions-only) lists.
|
||||
Previously-read GC mechanics (kept here for the detail):
|
||||
|
||||
**GC mechanics — verified against the current checkout (2026-10-01,
|
||||
`src/store/gc.rs`, `src/util/temp_tag.rs`, `src/store/fs/delete_set.rs`).
|
||||
@@ -278,17 +281,22 @@ is the prior art for cross-repo dedup among related repos — the
|
||||
pooled-CAS posture (OQ-BL-05) generalizes it to unrelated repos and
|
||||
to p2p replication.
|
||||
|
||||
### alknet's appfile external-store probe — the direct ancestor
|
||||
### alknet's appfile external-store probe — historical reference only
|
||||
|
||||
`/workspace/@alkdev/alknet/docs/research/alknet-filesystem/`
|
||||
(`alknet-blobs-external-store-probe.md`, `poc-summary.md`) — written
|
||||
against an older iroh-blobs; conclusions re-verify in this Phase 0.
|
||||
The load-bearing bits: the appfile shape (small blobs in kv/sqlite,
|
||||
large on fs fallback), the filename↔hash mapping problem, and the
|
||||
multi-backend dispatch pain the probe hit (which is this crate's
|
||||
reason to own that dispatch). Old-data warning: any API or behavior
|
||||
claims there describe an older upstream; re-check against
|
||||
`/workspace/iroh-blobs` as checked out today.
|
||||
against an older iroh-blobs. **Demoted 2026-10-02: alknet's code is
|
||||
not a design input for this crate** (it is old, was poorly planned,
|
||||
and has been decomposed into the alk* family; the substrate is
|
||||
alkcall). What remains relevant from it is only the *shape* lineage —
|
||||
the appfile pattern (small blobs in kv/sqlite, large on fs fallback,
|
||||
filename↔hash mapping) and the multi-backend dispatch pain it
|
||||
recorded, both of which this crate now owns by its own validated
|
||||
design (POC #1, POC #3). Its re-verification ran 2026-10-02 as a
|
||||
historical cross-check (`poc-largeblob-findings.md`, final section):
|
||||
architectural claims about iroh-blobs held up, its performance claims
|
||||
were asserted-never-measured and are superseded by POC #3's
|
||||
first-party benchmark.
|
||||
|
||||
### alkcall — the substrate
|
||||
|
||||
@@ -321,8 +329,9 @@ hash-addressed, p2p sync *is* "announce hashes, diff hash sets, fetch
|
||||
missing blobs verified," which is an op family on this store. Whether
|
||||
that family lives here (feature-gated) or in the sibling that owns
|
||||
replicator/gossip policy is the remaining question; the deciding input
|
||||
is what alkgit's replicator actually needs and whether alknet's
|
||||
appfile case ever transfers over the network. Even so, the *shape* of
|
||||
is what alkgit's replicator actually needs and whether the
|
||||
appfile-shaped consumers ever transfer blobs over the network. Even
|
||||
so, the *shape* of
|
||||
the ops (have/need + verified fetch, ACL-gated) is now firm enough to
|
||||
plan against.
|
||||
|
||||
@@ -344,6 +353,18 @@ anchors the small-blob belief the dispatch leans on; streaming put/
|
||||
get shape rides POC #3's findings (`Vec<u8>` was the *contract*
|
||||
question, not the performance question).
|
||||
|
||||
**Updated 2026-10-02 (POC #3, `poc-largeblob-findings.md`):** the
|
||||
small-blob belief is now first-party measured — sqlite beats fs ~9-10×
|
||||
at 1-16 KiB, crossover ~128-256 KiB (finding A4; harness included in
|
||||
the POC crate, re-runnable on real media). Default Phase 1 threshold
|
||||
lean: 128 KiB (tunable constructor param per POC #1's contract).
|
||||
Streaming put/get shape also landed (findings A1/A2/A3): the put seam
|
||||
is `(Option<len>, stream)` with known-length as the encouraged one-pass
|
||||
path and unknown-length staging-then-hashing (forced by the git
|
||||
preamble); the ops layer's fetch handler should be the broadcast
|
||||
fanout shape. An unknown-length-vs-known-length dispatch asymmetry
|
||||
(finding A6) is a named Phase 1 decision, options recorded.
|
||||
|
||||
### OQ-BL-03: Hash abstraction — RESOLVED: canonical git-family hash
|
||||
|
||||
**RESOLVED 2026-10-01 (the hash simplification round).** The original
|
||||
@@ -457,6 +478,24 @@ its verified whole-contents as git-blob-sha-256 pool entries (a
|
||||
`git-blob-sha-256` digest over the reassembled content), so chunked
|
||||
and whole-file paths dedup against each other in the one pool.
|
||||
|
||||
**Sharpened by POC #3 (2026-10-02, `poc-largeblob-findings.md`
|
||||
finding A2):** per-range verification is *impossible* under the
|
||||
canonical git-blob digest — the preamble hashes the length, so no
|
||||
slice has any hash relationship to the whole (iroh gets range
|
||||
verification from the bao *chunk tree*, not from BLAKE3 the
|
||||
algorithm). Consequences carried: (i) local range reads need no
|
||||
verification beyond transport-integrity slice digests the caller
|
||||
carries out-of-band (the store's `read_range` returns one); (ii) the
|
||||
only genuine consumer of a chunk-tree encoding is *networked ranged
|
||||
fetch of large blobs* — making the bao-like encoding a transfer-layer
|
||||
conditional, not a storage design; (iii) git blobs self-verify under
|
||||
git's own model when read whole, and small-tier range reads slice
|
||||
whole values in memory. The whole-file CAS posture is strengthened,
|
||||
not changed. Also added to Phase 1 surface notes: a cheap
|
||||
`stat(digest)` (length/type probe; gix's header-only read is git's
|
||||
cheapest primitive and the POC's `LfsObject.len` already covers the
|
||||
get side).
|
||||
|
||||
### OQ-BL-05: Pooling and GC — namespaces as reference sets over a flat CAS
|
||||
|
||||
**The decisions from the discussions (2026-10-01, principles 6; the
|
||||
@@ -546,6 +585,24 @@ large/packed content — per-object granularity, range reads into
|
||||
packed blobs, or both (interacts with OQ-BL-04's range-read
|
||||
question — likely read them together).
|
||||
|
||||
**Resolved in shape by POC #3 (2026-10-02, gix-odb analysis in
|
||||
`poc-largeblob-findings.md`):** gix-odb's own model confirms the two
|
||||
halves — alternates/multi-pack-index give shared-read across packs
|
||||
with no dedup/GC/large-blob/streaming value-add (gix-lfs is an empty
|
||||
placeholder), and gitoxide's whole pack-ID-stability machinery exists
|
||||
only because packs are immutable-with-stable-IDs; a per-object
|
||||
appending pool has no such IDs and avoids that machinery entirely.
|
||||
The evidence-supported resolution: **git objects enter the pool as
|
||||
loose-equivalent kv entries (small tier, the common case); if packfile
|
||||
serving is ever wanted, packs are stored as large blobs and served by
|
||||
the store's range-read surface** (git's own .idx does per-object
|
||||
offset lookup — the store is a flat byte server for the pack; finding
|
||||
A2 says local range serving is fine). Dedup at pack granularity is
|
||||
weak, but pack-level dedup across unrelated repos was never the goal;
|
||||
object-level dedup lives in the loose tier. Phase 1 ADR should record
|
||||
this as the default resolution, revisit only with a real consumer
|
||||
demand.
|
||||
|
||||
### OQ-BL-06: POC register (draft)
|
||||
|
||||
Numbered POCs, opened as research reaches them (findings land in
|
||||
@@ -557,19 +614,15 @@ is actually deciding:
|
||||
|---|------|--------|-------|
|
||||
| 1 | Backend-trait + dispatch shape (kv small / fs large); trait must include `list()` complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path; the hash-enum abstraction with git-blob-sha-256's domain-separated preamble + git-sha-1 tolerance case | **Passed 2026-10-02** — `poc-trait-dispatch-findings.md` (8 findings; code: `/workspace/alkblobs-trait-poc`) | findings landed |
|
||||
| 2 | ~~Multi-hash store~~ **Absorbed into #1** (2026-10-01 hash round): the canonical-hash resolution removed the "two families coexisting" question; what remains (the preamble abstraction + SHA-1 tolerance) is POC #1's trait work | **Absorbed** (and validated by #1) | — |
|
||||
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) | **Pending** — reading-and-design POC + benchmark | findings file TBD |
|
||||
| 3 | Large-blob path (iroh-blobs store read under current checkout; fs fallback + range reads; streaming `LFSObject`-shaped put/get + `fanout` seam) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) | **Passed 2026-10-02** — `poc-largeblob-findings.md` (6 findings A1-A6 + pack-tension analysis + historical alknet probe re-check; code: `/workspace/alkblobs-largeblob-poc`) + `iroh-blobs-eval.md` | findings landed |
|
||||
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | **Covered in miniature by #1** (2026-10-02): namespace tables + sweep + recover all validated single-threaded; the concurrency half (sweep-vs-put race, batch-scope pins) is Phase 1 implementation work, not a POC gate | `poc-trait-dispatch-findings.md` findings 3/7/8 |
|
||||
|
||||
Sequencing note: #1 is done (single opening worktree; multi-hash
|
||||
subsumed; git-preamble hashing validated byte-exact vs the git CLI);
|
||||
#3 is a reading-and-design POC against the current iroh-blobs checkout
|
||||
plus the alknet probe re-verification, with its micro-benchmark
|
||||
pull-out; #4's single-threaded core is covered by #1 (see the
|
||||
register row) and the residual is implementation work. The pack
|
||||
tension (OQ-BL-05) rides #3's reading rather than earning its own POC
|
||||
yet.
|
||||
Sequencing note: #1 and #3 are done; #2 absorbed/validated under #1;
|
||||
#4's single-threaded core is covered by #1 (see the register row) and
|
||||
the residual is implementation work. **The POC register is complete —
|
||||
Phase 0's remaining step is convergence → Phase 1.**
|
||||
|
||||
**What each POC is deciding (updated 2026-10-02, post-POC-#1):**
|
||||
**What each POC is deciding (updated 2026-10-02, post-POC-#3):**
|
||||
|
||||
- **POC #1 — DONE.** Decided the backend trait's minimum contract
|
||||
(opaque byte keys, `list()` complete-by-contract, put-path
|
||||
@@ -585,13 +638,21 @@ yet.
|
||||
evaporated when BLAKE3's only reason-for-being (iroh-blobs
|
||||
inheritance) was rejected, and its residual content validated under
|
||||
#1.
|
||||
- **POC #3 is the next (and last) POC** — reading-and-design against
|
||||
the *current* iroh-blobs checkout (the store eval) + re-verifying
|
||||
alknet's probe + the **micro-benchmark pull-out** (the "kv/sqlite
|
||||
is faster for small blobs" belief underlies the dual-backend design
|
||||
and traces to *old* alknet research — anchor it: small objects,
|
||||
1KB-64KB, both backends, random + sequential). The pack tension
|
||||
(OQ-BL-05) is analyzed here without its own POC.
|
||||
- **POC #3 is DONE** — the last POC. Streaming large-blob path
|
||||
validated (`LFSObject` put/get; known-length one-pass vs unknown-
|
||||
length stage-then-hash, forced by the git preamble — finding A1);
|
||||
the fanout seam (serve-while-persisting) works as a broadcast above
|
||||
the backends (A3); range reads work with the verification gap
|
||||
documented honestly (A2 — impossible under the canonical digest,
|
||||
sharpening the chunk-tree question to "transfer-layer conditional");
|
||||
the sqlite-vs-fs benchmark anchored the small-blob belief first-party
|
||||
(~9-10× at 1-16 KiB, crossover ~128-256 KiB — A4); stage-file hygiene
|
||||
tests caught a real leak (A5). The pack tension analyzed from
|
||||
gix-odb and resolved in shape (packs-as-large-blobs served by range
|
||||
reads). The iroh-blobs store eval landed as `iroh-blobs-eval.md`;
|
||||
alknet's probe re-check is recorded as historical cross-check only
|
||||
(per the 2026-10-02 clarification, alknet's code is not a design
|
||||
input — the substrate is alkcall).
|
||||
- **POC #4's single-threaded core is covered by #1** (namespace
|
||||
tables + mark-and-sweep + protect-callback abort + delete-then-
|
||||
recover all validated; `poc-trait-dispatch-findings.md` findings
|
||||
@@ -609,17 +670,21 @@ in `docs/research/` regardless of where the code lives.
|
||||
|
||||
## Phase 0 plan (next steps)
|
||||
|
||||
1. ~~Write `iroh-blobs-eval.md`~~ — still open, rides POC #3's
|
||||
reading against the *current* checkout (kv + flat backends, bao,
|
||||
chunking; the GC mechanics eval is already done — see
|
||||
`poc-trait-dispatch-findings.md`'s context and phase-0's verified
|
||||
GC section).
|
||||
2. **Re-verify alknet's probe** (`alknet-blobs-external-store-probe.md`,
|
||||
`poc-summary.md`) against the current upstream; mark what carried
|
||||
over — rides POC #3 too.
|
||||
3. **Read gix-odb** (`/workspace/git-oxide/gix-odb`) for the alkgit
|
||||
baseline: its storage layout, object model, and where a blob-store
|
||||
crate under/beside it earns its keep.
|
||||
4. **POC #3** (the last open POC): reading-and-design + the
|
||||
sqlite-vs-fs micro-benchmark per OQ-BL-06.
|
||||
5. Converge: recommended approach + final OQ register → Phase 1.
|
||||
1. ~~Write `iroh-blobs-eval.md`~~ — **done 2026-10-02** (POC #3's
|
||||
reading against the current checkout; `docs/research/
|
||||
iroh-blobs-eval.md` — store traits/API, fs backend layout, bao/-
|
||||
verification conclusions, small-blob handling, do-not-inherit and
|
||||
keep-lists).
|
||||
2. ~~Re-verify alknet's probe~~ — **done 2026-10-02**, recorded as a
|
||||
historical cross-check only in `poc-largeblob-findings.md`
|
||||
(alknet's code is not a design input; substrate is alkcall). The
|
||||
one load-bearing output: the probe's small-blob performance claim
|
||||
was asserted-never-measured; POC #3's benchmark is the anchor.
|
||||
3. ~~Read gix-odb for the alkgit baseline~~ — **done 2026-10-02**
|
||||
(POC #3's pack-tension analysis; conclusions folded into OQ-BL-05
|
||||
and the findings doc; gix-lfs is a placeholder crate, the
|
||||
large-blob lane is open as assumed).
|
||||
4. ~~POC #3~~ — **done 2026-10-02**, `poc-largeblob-findings.md`.
|
||||
5. **Converge**: recommended approach + final OQ register → Phase 1.
|
||||
All gates passed; the register rows above carry the empirical
|
||||
inputs (thresholds, seams, lessons) into the SDD's Phase 1.
|
||||
@@ -0,0 +1,323 @@
|
||||
---
|
||||
status: passed
|
||||
title: "POC #3 — large-blob streaming path + sqlite-vs-fs micro-benchmark"
|
||||
last_updated: 2026-10-02
|
||||
---
|
||||
|
||||
# POC: large-blob path (streaming, fanout, range reads) + small-object micro-benchmark — findings
|
||||
|
||||
> **POC register #3**, per `docs/research/phase-0.md` OQ-BL-06. Code:
|
||||
> standalone crate `/workspace/alkblobs-largeblob-poc` (findings land
|
||||
> here regardless, per the established convention). Companions from the
|
||||
> same POC's reading work: `iroh-blobs-eval.md` (fresh-store eval
|
||||
> against checkout e82cbdc / v0.103.0) and the alknet probe re-check
|
||||
> (below — recorded as **historical cross-check only**).
|
||||
> Date: 2026-10-02. Status: **passed**, 14 tests, clippy `-D warnings`
|
||||
> clean, fmt clean. Digests cross-checked against the real `git
|
||||
> hash-object` CLI (git 2.43), extending POC #1's validation to the
|
||||
> streaming paths.
|
||||
|
||||
## What this POC set out to decide
|
||||
|
||||
From the phase-0 register:
|
||||
|
||||
1. **The large-blob path** — does an `LFSObject`-shaped `(len, stream)`
|
||||
put/get work at the store layer, and what does the git-blob preamble
|
||||
do to streaming hashing?
|
||||
2. **The fanout seam** — "serve the fetch while persisting on receipt"
|
||||
(rudolfs' `fanout()`; iroh's `get_blob_ranges_impl`) — does it work
|
||||
at our layer, above the backends?
|
||||
3. **Range reads on the fs backend** — cheap partial reads with *some*
|
||||
verification story (OQ-BL-04).
|
||||
4. **The small-object sqlite-vs-fs micro-benchmark** — anchor the
|
||||
"kv is faster than fs for small blobs" belief first-party
|
||||
(OQ-BL-02's "dual-belief anchor").
|
||||
5. **The pack tension** (OQ-BL-05) — analyzed from the gix-odb read,
|
||||
no POC of its own.
|
||||
|
||||
Note on standing: **alknet's code is not a design input** — it is old,
|
||||
poorly planned, decomposed into the alk* family; its probe docs are
|
||||
treated as references to old research only. The substrate this crate
|
||||
rides is alkcall (like alkhttp/alksocks/alktty/alktunnels).
|
||||
|
||||
## Result summary
|
||||
|
||||
| Question | Verdict |
|
||||
|----------|---------|
|
||||
| Streaming put (`LFSObject` shape) | **Yes** — two paths: known-length one-pass; unknown-length stage-then-hash (forced by the preamble, finding A1) |
|
||||
| Streaming get without whole-blob load | **Yes** — fs arm streams a file handle; kv arm one bounded read |
|
||||
| Fanout (serve while persisting) | **Yes** — one reader task, broadcast channel, store arm + N subscriber arms (finding A3) |
|
||||
| Range reads on fs | **Yes** — pread-window only; **verification has a real gap** (finding A2) |
|
||||
| sqlite faster than fs for small blobs | **Confirmed and measured** — ~10× below 16 KiB; crossover ~128-256 KiB (B-finding) |
|
||||
| Digest = git oid on streaming paths | **Byte-exact vs `git hash-object`** for both put paths |
|
||||
| Stage-file hygiene | **Verified** — failures (stream error, overflow, early-EOF) leave zero stage files (finding A5) |
|
||||
|
||||
## Findings
|
||||
|
||||
### Finding A1: the git-blob preamble splits streaming into exactly two paths — and the split is forced, not chosen
|
||||
|
||||
The canonical derivation is `H("blob <len>\0" + content)`. The length
|
||||
is *inside the hashed input*, so:
|
||||
|
||||
- **Known length** ⇒ write the preamble into the hasher at construction,
|
||||
feed content chunks incrementally, finalize at stream end: **one pass
|
||||
over the content**, no staging needed for hashing (staging to a temp
|
||||
file still happens for the fs tier, but the hash is done by the time
|
||||
the rename commits). This covers the dominant git and protocol cases:
|
||||
a pre-staged file has its length from `stat`, a network offer carries
|
||||
a size, an HTTP PUT can carry `Content-Length`.
|
||||
- **Unknown length** ⇒ the preamble can't be written until the stream
|
||||
ends, so content must be staged (file, not memory) and hashed in a
|
||||
second pass over the staged bytes. **Two passes are forced by the
|
||||
derivation**, not by our choice of hash.
|
||||
|
||||
iroh-blobs independently arrives at the same stage-then-encode shape
|
||||
(their `add_stream` buffers then computes outboard in a second pass,
|
||||
`fs/import.rs:290-419`), for bao-tree reasons instead of
|
||||
preamble reasons. Their `import_bao_reader` sidesteps it by knowing
|
||||
both hash and size upfront (a protocol-level offer) — the same
|
||||
"known-length is the normal case" assumption, enforced by their wire
|
||||
format.
|
||||
|
||||
**Verdict carried to Phase 1:** the store's put seam should be
|
||||
`(Option<len>, stream)` exactly as built. Known-length puts must be the
|
||||
encouraged path (documented as such on the API); unknown-length is
|
||||
supported but slower by one page-cache read. Do *not* invent a
|
||||
length-prefix-on-the-wire scheme to "fix" unknown-length puts — the
|
||||
wire carries sizes naturally (it must, for the CAS routing).
|
||||
|
||||
### Finding A2: range reads cannot verify against the canonical digest — the preamble includes the length, so slices have no hash relationship to the whole
|
||||
|
||||
Confirmed at implementation level (and independently by the
|
||||
iroh-blobs eval): `blob <len>\0 + content` gives *no* verifiable
|
||||
relation between `H(whole)` and any `content[a..b]`. iroh gets
|
||||
per-range verification from the **bao chunk tree**, not from BLAKE3 the
|
||||
algorithm — the tree is the mechanism, the hash is just its leaf
|
||||
function.
|
||||
|
||||
POC #3's `read_range` therefore returns the slice plus a *slice digest*
|
||||
(`git-blob-sha-256` over the slice — a transport-integrity check the
|
||||
caller carries out-of-band, e.g. from the offering side). Consequences:
|
||||
|
||||
- **Local-store range reads need no more** — the OS pages honest media;
|
||||
EIO surfaces read errors; git objects (the small tier) self-verify
|
||||
when git reads them whole.
|
||||
- **The only genuine consumer of a chunk tree is networked large-blob
|
||||
transfer where the caller wants per-chunk integrity.** This sharpens
|
||||
OQ-BL-04's "conditional": the bao-like encoding is not about
|
||||
*storage* at all, it is a *transfer* encoding. If/when p2p git needs
|
||||
verified ranged fetch of LFS-shaped artifacts, the encoding layer
|
||||
(BLAKE3-confined per OQ-BL-03 tier 3) registers its whole-content
|
||||
root back into the pool as an ordinary git-blob-sha-256 entry.
|
||||
- Practical corollary: for git-blob-shaped content, byte-range *serving*
|
||||
is fine (packfiles will want it; see pack tension below) but per-range
|
||||
*verification* is out of scope for the store layer today.
|
||||
|
||||
### Finding A3: the fanout seam is a broadcast channel above the store, and late-joiners degrade to get-after-commit
|
||||
|
||||
Built `put_stream_fanout(stream, n)`: one reader task consumes the
|
||||
source stream exactly once and broadcasts each chunk to (a) the store
|
||||
put arm and (b) N subscriber feeds (`mpsc<io::Result<Bytes>>`);
|
||||
subscribers read while the put is still landing. Tests show 3
|
||||
subscribers each receive the full 256 KiB byte-identical while the put
|
||||
completes, and the put receipt matches the git oid.
|
||||
|
||||
Design notes validated:
|
||||
|
||||
- **One read, many consumers** — the source stream must never be
|
||||
polled twice; anything that wants a second copy subscribes. This
|
||||
matches iroh's `get_blob_ranges_impl` shape (network decode loop
|
||||
joined with the store import) and rudolfs' `fanout()` (two lock-step
|
||||
copies from one stream). All three arrive at "broadcast from a single
|
||||
received stream" — treat that as the settled seam shape.
|
||||
- **Backpressure discipline:** bounded channels (8 items in the POC)
|
||||
make a slow subscriber stall the broadcast — which is *correct* for a
|
||||
put (the put must land) but means subscriber count is bounded and
|
||||
slow-subscriber handling is a policy question for the ops layer
|
||||
(drop-and-late-join is the escape hatch; see next note).
|
||||
- **Late joiners:** a consumer arriving after the live stream has
|
||||
passed simply waits for the commit then does a normal verified
|
||||
`get_stream` — the CAS guarantees the bytes are identical. The POC
|
||||
didn't build this (subscriber set is fixed at start) but the degrade
|
||||
path is exactly the normal get; worth documenting in the ops layer.
|
||||
|
||||
### Finding A4 (benchmark): sqlite beats fs decisively for small blobs; crossover ~128-256 KiB; fsync discipline dominates both arms
|
||||
|
||||
Release build, per-arm timing loops, 2k distinct keys/values random
|
||||
working set, `synchronous=NORMAL` WAL sqlite vs stage+rename fs puts /
|
||||
open+read fs gets. Representative run (3 s/arm):
|
||||
|
||||
```
|
||||
backend size puts/s gets/s put p50µs put p99µs get p50µs get p99µs
|
||||
sqlite 1024 48319 49508 20.0 37.0 19.0 34.0
|
||||
fs 1024 5306 3357 205.0 302.0 288.0 384.0
|
||||
sqlite 16384 14129 14166 70.0 125.0 68.0 107.0
|
||||
fs 16384 4203 2399 257.0 395.0 433.0 599.0
|
||||
sqlite 65536 4028 4048 252.0 364.0 243.0 344.0
|
||||
fs 65536 3212 1852 339.0 507.0 540.0 688.0
|
||||
sqlite 131072 2067 2042 494.0 698.0 482.0 699.0
|
||||
fs 131072 1892 1652 593.0 879.0 619.0 809.0
|
||||
sqlite 262144 981 1003 1033.0 1792.0 994.0 1176.0
|
||||
fs 262144 1244 1576 822.0 1181.0 741.0 964.0
|
||||
```
|
||||
|
||||
Readings:
|
||||
|
||||
- **~9-10× sqlite advantage at 1-16 KiB** (the git small-blob regime —
|
||||
iroh's DESIGN.md claim and the old citations hold for our workload).
|
||||
- **Crossover at ~128-256 KiB**: at 256 KiB fs wins puts (no B-tree
|
||||
rewrite of a 256 KiB row) and gets. **The dispatch threshold is a
|
||||
tunable**, placed anywhere in 64-256 KiB; POC #1's contract already
|
||||
makes it a constructor parameter. Phase 1 default: 128 KiB (midpoint
|
||||
of the flat zone) — revisit with the storage-medium reality of real
|
||||
consumers (tmpfs vs disk; the POC SSD regime is one data point).
|
||||
- **fs puts cost ~200 µs floor even at 1 KiB** — that is the
|
||||
create+write+rename syscall chain; batching stage files (a shared
|
||||
preallocated pool) or O_TMPFILE could shave it, but the sqlite arm
|
||||
makes the effort moot below the crossover.
|
||||
- Caveats recorded honestly: single-node tmpdir on this dev box
|
||||
(backend-specific tuning — e.g. iroh's redb batch windows — could
|
||||
shift the fs curve), no fsync-in-loop on the fs arm (matching NORMAL
|
||||
WAL durability, sqlite arm likewise not full-fsync; both arms are
|
||||
"crash-tolerant-ish", so the comparison is fair *as measured*: both
|
||||
are the durability tier a store would ship).
|
||||
- The old alknet claim "<100 KB faster in sqlite" is **confirmed in
|
||||
direction, refined in boundary**: the honest crossover is higher
|
||||
(~128-256 KiB) and sqlite is never *pathological* below it.
|
||||
|
||||
### Finding A5: stage-file hygiene is testable and the tests earned their keep (again)
|
||||
|
||||
The `stream_failure_cleans_stage_files` test asserted zero files left in
|
||||
the data dir after a mid-stream connection failure — and caught a real
|
||||
bug: the original `chunk?` error path exited *before* discarding the
|
||||
stage file, leaking one temp file per failed put. Fixed by restructuring
|
||||
the loops to route all failure paths (stream error, overflow, early-EOF,
|
||||
write error) through one discard. Two lessons compound with POC #1's:
|
||||
|
||||
- **Resource-hygiene invariants need dedicated tests**; code review
|
||||
missed what the exact-count assertion caught (echoing POC #1's exact
|
||||
deletion counts).
|
||||
- The store's put path now has the shape Phase 1 should codify: **every
|
||||
failure path converges on stage-discard; every success path converges
|
||||
on commit-rename** — no early returns that bypass cleanup.
|
||||
|
||||
### Finding A6: re-put dedup across the two put paths holds; unknown-length re-put of a small blob lands on fs (a dispatch asymmetry worth one Phase 1 decision)
|
||||
|
||||
The same content put via known-length (small) and unknown-length
|
||||
(staged) paths produces the identical digest and coexists as one entry.
|
||||
But routing differs: an unknown-length put *always* stages through the
|
||||
fs arm in the POC (routing must wait for the realized size), so a small
|
||||
blob arriving via an unknown-length stream occupies an fs file where the
|
||||
same digest's known-length put would occupy a kv row. If both happen,
|
||||
the pool holds the content twice (different physical tiers).
|
||||
|
||||
Options for Phase 1, decided there, not here (recorded honestly — the
|
||||
POC as built has the asymmetry): (a) post-commit tier-correction sweep
|
||||
(move small fs entries into kv during idle GC), (b) require `len` for
|
||||
the small tier (reject unknown-length small puts — simple, slightly
|
||||
hostile to streaming producers), (c) accept the duplication (it's
|
||||
temporary if (a) runs) — or (d) route *unknown-length* puts through a
|
||||
memory buffer up to the threshold before staging (bounded-memory
|
||||
reconciliation: small unknown puts never touch fs). Given the
|
||||
benchmark (small = cheap everywhere), (d) looks natural but should be
|
||||
weighed against batch-scope pinning from POC #1 finding 3.
|
||||
|
||||
## The pack tension (OQ-BL-05 input, from the gix-odb read — analysis only)
|
||||
|
||||
gix-odb (v0.84) audit conclusions, in alkblobs terms:
|
||||
|
||||
- **Git's own pooling is administratively shared packs**: alternates/
|
||||
multi-pack-index give one flat OID→location lookup across many packs
|
||||
(slot-map + `ArcSwap` snapshots for lock-free reads), but there is no
|
||||
cross-store dedup, no refcounting, no GC, no large-blob handling, no
|
||||
streaming reads — every object decompresses whole into memory (writes
|
||||
stream; reads do not). gix-lfs is an empty placeholder crate. The
|
||||
value-adds alkblobs planned are confirmed unclaimed.
|
||||
- **Packs are pack-efficient because they delta-chain within
|
||||
themselves; a per-object CAS pool fights that in two ways.** First,
|
||||
delta chains want objects stored *near their bases*; a flat per-oid
|
||||
pool scatters them, so serving a delta-heavy pack read pattern off the
|
||||
pool re-inflates from scattered loose entries (loose objects are
|
||||
git's *dedup unit and fs bottleneck simultaneously* — gix's own
|
||||
discovery docs flag loose-object proliferation as a server-scale
|
||||
performance problem). Second, packs are immutable-with-stable-IDs,
|
||||
and gitoxide's whole generation/consolidation machinery exists to
|
||||
rebind pack IDs when packs churn — a blob pool that *appends per
|
||||
object* avoids that machinery entirely (there are no IDs to rebind;
|
||||
the address is the content).
|
||||
- **Resolution shape that the evidence supports** (Phase 1 ADR input):
|
||||
git objects enter the pool as **loose-equivalent kv entries** on the
|
||||
small tier (the common case — git's median blob is small), and the
|
||||
pack question is *deferred to serving*, not storage: if/when serving
|
||||
git fetches wants packfile byte-range reads (OQ-BL-05's pack
|
||||
tension), the options are (i) store packs as large blobs and serve
|
||||
ranges *out of the pack* via the store's range-read (the pack itself
|
||||
is the container; per-object addressing rides git's own .idx — the
|
||||
store is just a flat byte server for it, which finding A2 says is
|
||||
fine locally), or (ii) keep loose-per-object as the only form and
|
||||
let consumers synthesize packs above. (i) collapses the tension
|
||||
without new store concepts: **the pool stores whole files — including
|
||||
packs — and ranged serving is its range-read surface.** Dedup at pack
|
||||
granularity is weak, but pack-level dedup across unrelated repos was
|
||||
never the goal; object-level dedup lives in the loose tier.
|
||||
- **Concrete consequence for the store API:** a `metadata(oid)`-shaped
|
||||
cheap size/type probe (gix's header-only read is git's cheapest
|
||||
primitive) should exist in Phase 1's surface — the POC's
|
||||
`LfsObject.len` covers the get side already; make it an explicit
|
||||
`stat(digest)` op.
|
||||
|
||||
## Alknet probe re-check (historical cross-check only — not design input)
|
||||
|
||||
Run for completeness, per the register line that flagged it; per the
|
||||
session's clarification, **alknet's old code/probe conclusions carry no
|
||||
decision weight for this crate** (decomposed substrate; alkcall is the
|
||||
substrate). What the re-check contributes, stripped to the factual:
|
||||
|
||||
- The probe's architectural claims about iroh-blobs (no Store trait;
|
||||
`Command` enum seam; three-actor pattern; the four pub(crate)
|
||||
blockers; the four redb tables; the 16 KiB hybrid) all verified
|
||||
against the current checkout — the eval above re-states them
|
||||
independently.
|
||||
- The probe's performance claims ("faster than the filesystem for
|
||||
small items") were **asserted-with-citation, never measured** — the
|
||||
POC #3 benchmark is the first first-party anchor (finding A4).
|
||||
- Stale: the version-gap open question (alknet-core is on iroh 1.0
|
||||
now); `run_gc` is *not* downstream-importable (private module) —
|
||||
irrelevant here, we are not reusing their GC.
|
||||
|
||||
## Consequences for the phase-0 OQ register
|
||||
|
||||
- **OQ-BL-02 (dispatch):** the small-blob belief is now first-party
|
||||
measured (A4). Threshold ~128 KiB default; benchmark harness exists
|
||||
(`cargo run --release --bin bench_small`) for re-running on real
|
||||
media. Remaining above-trait concerns (migration policy, per-
|
||||
namespace config) unchanged — no POC gate.
|
||||
- **OQ-BL-04 (verification/chunking):** sharpened — range verification
|
||||
is impossible under the canonical digest (A2); the chunk-tree
|
||||
encoding is a *transfer-layer* conditional whose only consumer is
|
||||
networked ranged fetch. Whole-file CAS posture unchanged and
|
||||
strengthened. `stat(digest)` surfaces in the Phase 1 API notes.
|
||||
- **OQ-BL-05 (pooling/GC):** pack tension resolved in shape — packs
|
||||
(if stored) are just large blobs served by range reads; object-level
|
||||
dedup lives in the loose tier. No new store concepts needed.
|
||||
- **POC register:** #3 is done; no POCs remain open. Phase 0's next
|
||||
step is convergence → Phase 1.
|
||||
|
||||
## POC quality notes
|
||||
|
||||
- 14 tests (12 integration `large_blob.rs`, 2 CLI cross-check
|
||||
`git_cli.rs`); digest validation against real `git hash-object`
|
||||
(`--stdin`, repo initialized `--object-format=sha256` — the POC #1
|
||||
gotcha holds: the flag is not a hash-object flag).
|
||||
- clippy `-D warnings` clean (0 warnings), `cargo fmt` clean.
|
||||
- Deliberate scope limits: single-put-at-a-time (no batch concurrency —
|
||||
POC #1's finding 3 covers the pinning story), no persistence-across-
|
||||
restart tests (same rationale as POC #1), the fanout seam is
|
||||
subscriber-count-fixed (no late-joiner implementation — degrade path
|
||||
documented instead), benchmark is one dev box + one working-set
|
||||
profile (harness is included for re-running).
|
||||
- Known micro-detail: `has()` on the kv arm uses a prepared
|
||||
`SELECT 1 ... EXISTS`; `get_stream` dispatch does two kv round-trips
|
||||
(size, then get) — a schema returning the value directly would avoid
|
||||
one; left as-is to keep the size probe identical for both tiers in
|
||||
the benchmark comparison.
|
||||
Reference in new issue
Block a user