The ADR-008 pg-lo admission POC ran in a standalone crate (/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait contract (including size), 10/10 exact-count sweep-outcome contract tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432, POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions, no new dependency. Gate verdict: passed, with named deltas. - Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and below 1 MiB beating — durable local fs on this fsync-slow disk; cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16 readers (20-50x behind page-cache fs — the honest named delta) - Contract: companion table is the list()/size()/CAS authority (never the catalogs); stage-then-commit; GC-participating lo_unlink delete - Handles: the tx-scoped descriptor is real but pool-compatible via descriptorless lo_get(oid, off, len) windows — window gets keep handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor is the fallback posture - Vacuum: pg_largeobject pages churn-reused, never returned; tracked by autovacuum; rel-size monitoring named as an ops requirement - Crash/orphan: LO creation is transactional — kill/terminate mid-write-tx leaves zero orphan pages; the only orphan class is a committed LO bypassing the companion table (planted, reaped by the ~7 ms/oid sweep; committed content survives byte-exact) - Harness lessons: lo_lseek is int4 — the 64 variants are the >2 GiB discipline; shared-table parallel tests are unsound (per-test CREATE DATABASE isolation) Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed; register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated (duplicate bullet removed) + backends-and-dispatch/open-questions cross-references. Verification: cargo test --release (10 passed), clippy -D warnings, fmt --check in /workspace/alkblobs-pglo-poc.
43 KiB
status, last_updated
| status | last_updated |
|---|---|
| converged | 2026-10-03 (POC |
alkblobs — Phase 0 (Exploration)
This document captures Phase 0 (Exploration) for the alkblobs crate:
vision, guiding principles, prior art, the validated approach, and the
open-question register (OQ-BL-01..06). Phase 0's objective per
docs/sdd_process.md: capture vision and guiding principles; research
options; validate approaches; converge on a recommended approach. That
objective is now met — all POC register rows passed or were absorbed,
and this document carries the converged recommendation into Phase 1
(Architecture), where the Architect will produce docs/architecture/
specs, ADRs, and the open-questions tracker.
Drafted 2026-09-30; converged 2026-10-02 after the POC rounds. The
chronological accretion from the research rounds has been folded into a
single settled posture; the evidence trail lives in the findings
documents (poc-trait-dispatch-findings.md, poc-largeblob-findings.md,
iroh-blobs-eval.md).
Vision and guiding principles
One sentence: content-addressed blob storage in the alk* family — a put/get/verify store keyed under one canonical content hash (the git-family derivation; legacy SHA-1-tolerant), with pluggable backends (small blobs in a kv/sqlite-ish store, large blobs on a filesystem fallback), split from any wire/ops-protocol layer (framing, op family, ACL ride the alkcall substrate when they exist — OQ-BL-01), which is itself transport-free (alkcall is a pure protocol crate).
Why this crate exists (the converging consumers):
- alkgit (planning phase) needs an object backend that is not its proposed default file-based backend ("kind of gross") — blobs stored under git's own object hashes, with an iroh-blobs-shaped large-blob path standing in where git typically delegates to git-lfs. The once-called "hash conflict" does not exist: git's oid derivation is the store's canonical hash (OQ-BL-03), so git objects land in the pool as first-class entries with zero indirection.
- Workspace/file-shaped storage in the alk family* (the "appfile" shape; old alknet research is historical lineage only): small blobs need a store faster than the filesystem (first-party measured: sqlite ~9-10× faster at 1-16 KiB, POC #3 finding A4), large blobs want a filesystem fallback, with a filename↔hash mapping above. Dispatching across backends is a store-layer problem this crate owns (OQ-BL-02; validated by POC #1's trait seam and POC #3's streaming + threshold evidence).
- Agent workspaces — many agents working simultaneously in workspaces that are mostly identical to each other's, each making small edits in specific areas. Per-agent stores multiply the near-identical content; a shared content-addressed store dedups it by construction, and a path→hash manifest per workspace makes each agent's delta just its changed hashes.
The p2p shape (from the dedup discussion). Two alkgit use cases exist: self-hosted git (where cross-repo dedup matters little — you don't fork your own repos) and p2p git for OSS projects, the demanding one. The p2p sketch: ownership/ACL live in a smart contract on a low-fee network; replicators are push/clone endpoints that watch the contract and cache its state; repos sync via iroh-style gossip (hashes, not content) and pull/sync on demand. Neither alkgit approach takes advantage of dedup — which is exactly a store-layer property. The implications for this crate:
- Both demanding consumers (p2p replicators, agent swarms) reduce to
the same shape: a pooled CAS per node, where repos / workspaces
are sets of hash references (trees/manifests), not separate object
stores. Cross-repo dedup is then by construction; git's own
alternates/object-pool mechanism is the same idea done awkwardly. With the canonical-hash resolution (OQ-BL-03) the dedup strengthens further: workspaces and git repos collide into one address domain — the same file content has the same address whether it entered via a git oid or a workspace manifest. - Because everything is hash-addressed, the p2p wire story collapses to "announce hashes, diff hash sets, fetch missing blobs verified" — exactly the ops surface OQ-BL-01 settled. The consumer exists; the ops surface is a confirmed have/need + fetch op family, riding the alkcall substrate (OQ-BL-01), whether it lives in this crate or a sibling.
The scope line: this crate is the store, not the wire/ops protocol. iroh-blobs welds store + provider protocol + tickets + postcard into one crate; the defining posture here is the opposite split — put/get/verify against a hash is its own layer, and any ops/protocol surface lives above it or in feature-gated modules (the alk* inversion-point pattern). "Above it" does not mean "without alkcall": the ops surface is alkcall-shaped (OQ-BL-01) — what the store layer itself must never grow is wire framing, op dispatch, ACL, or transport (alkcall carries none of the last either; the connection is handed to it). What stays above the crate regardless: repos/ workspaces as reference sets (trees/manifests), path→hash mapping, git semantics, the contract/gossip/replicator policy layer.
Guiding principles, inherited from the alk* family:
- Substrate-agnostic by construction. The store must not know
whether bytes arrive over the network, from a local writer, or get
reassembled anywhere particular. Backends behind an injected seam
(alktty
TtyBackend/ alktunnels pump-halves precedent). - One canonical hash; the git-family derivation. The store keys
entries under a single canonical algorithm — the git-blob
derivation (
blob <len>\0domain-separated input, SHA-256 for new content; OQ-BL-03, implementation-validated by POC #1). Legacy SHA-1 (existing git repos) stays tolerated via the same prefix; the abstraction is a small enum/capability set, and the key encoding is a one-way door. - Borrow conclusions, not wire surface. iroh-blobs' tickets,
postcard serialization, provider protocol, and Command-actor store
abstraction are design-welded choices we do not inherit. Its
store-shape conclusions (kv + flat backends, verification flow,
chunking limits, put-commit primitives) are fair game — see
iroh-blobs-eval.md. - Producer/consumer vocabulary for any network-facing surface;
authorization via alkcall's
AccessControl/identity seam, never an in-band invented scheme (convention 6). - Verify where it matters. Git objects self-verify under git's own model — the git case needs nothing from us. Per-range verification is impossible under the canonical digest (the preamble hashes the length; POC #3 finding A2), so verification is whole-blob for stored content; the only genuine consumer of a chunk-tree encoding is networked ranged fetch (a transfer-layer conditional — OQ-BL-04).
- Pooling is the point. The dedup wins both demanding consumers want (Forknet-style multi-repo OSS content, agent swarms) require that repos/workspaces are sets of hash references over one pooled CAS per node, not walled-off per-repo stores. Pooling forces a GC story (OQ-BL-05) and a multi-backend dispatch story (OQ-BL-02).
- Whole-file CAS is the default; chunking earns its keep only where it must. For source-code-scale content, whole-file hashing captures the agent/edit delta perfectly and needs no chunk tree; fixed-size chunk trees are actually hostile to insertion/deletion edits (byte shifts cascade). Content-defined chunking (CDC — fastcdc/restic-style, shift-resistant) only pays on large binary files receiving small edits — the git-lfs weakness. iroh-blobs/bao use fixed-size trees; that is not inherited as the default. Where chunking would live is scoped in OQ-BL-04.
The settled approach
These are the converged postures. None is a Phase 1 ADR yet; each becomes one (or is already recorded as resolved here) — the ADR backlog is in §Convergence.
- Not a fork of iroh-blobs. A downstream store crate that borrows
its store-shape conclusions (per
iroh-blobs-eval.md) while diverging deliberately on hashing and wire surface. - Canonical hash = the git-family derivation (OQ-BL-03,
RESOLVED):
git-blob-sha-256primary, legacygit-blob-sha-1tolerated, BLAKE3 demoted to a conditional large-blob transfer encoding. Implementation-validated against the real git CLI, including byte-exact external-oid interop (POC #1 finding 4/5). - Backend pluralism, dispatched by size threshold (OQ-BL-02,
base settled): the lean
Backendtrait contract (opaque byte keys,has/get/put/delete/list/name,list()complete by contract) with deterministic size-threshold routing — kv/sqlite for small blobs, fs fallback for large. Validated by POC #1 (trait seam) and POC #3 (streaming shapes, sqlite-vs-fs benchmark: ~9-10× at 1-16 KiB, crossover ~128-256 KiB). Phase 1 default threshold lean: 128 KiB (tunable constructor parameter). - Streaming put seam =
(Option<len>, stream)(OQ-BL-02 / finding A1): known-length one-pass hashing as the encouraged path; unknown-length staging-then-hash, forced by the git preamble, not a design choice. Get-side: small tier returns values, large tier streams; range reads return slice + out-of-band slice digest. - Pooled CAS, not per-repo stores (OQ-BL-05): repos and agent workspaces are sets of hash references over one shared content-addressed store. This is the property both demanding consumers need; everything else lives around it.
- Physically flat, logically namespaced (OQ-BL-05; the rudolfs
inversion): the byte layer is a flat dedup-by-construction CAS
(
hash → bytes, backends namespace-blind,{prefix}/{prefix}/{hash}sharding survives); namespaces are reference tables above it — GC roots, sweep scoping, accounting, ACL boundary. - Mark-and-sweep GC with protect-callback + RAII pinning (OQ-BL-05,
mechanism validated by POC #1 findings 3/7/8): pins protect
not-yet-referenced puts, liveness sources register into a seam,
protection-source failure aborts the sweep (iroh's
ProtectOutcome::Abortconclusion), delete-then-recover is a re-put. - Pack tension resolved in shape (OQ-BL-05): git objects enter the
pool as loose-equivalent kv entries (small tier, the common case);
if packfile serving is ever wanted, packs are stored as large blobs
and served through the store's range-read surface (git's own
.idxdoes per-object offset lookup). Object-level dedup lives in the loose tier; pack-level dedup across unrelated repos was never the goal. - Whole-file CAS is the default granularity; chunking is scoped to the networked-large-blob transfer case (principle 7 + OQ-BL-04).
Prior art
iroh-blobs — the shape inspiration (evaluated, not the base)
/workspace/iroh-blobs (read-only reference), evaluated against
checkout e82cbdc / v0.103.0 in docs/research/iroh-blobs-eval.md: no
Store trait (message-protocol actor pattern), fs backend on-disk
layout (redb entry-state + inline thresholds + crash ordering),
verification (subtree-local, chunk-tree-driven), streaming shapes, and
the do-not-inherit / keep (conclusions-only) lists.
GC mechanics — verified against that checkout (src/store/gc.rs,
src/util/temp_tag.rs, src/store/fs/delete_set.rs). This is the
load-bearing prior art for OQ-BL-05:
- Mark-and-sweep, not refcounting, for persistent liveness.
gc_mark_taskcollects roots from three sources: persistent named tags (a flattags-0table, name →HashAndFormat), in-memoryTempTags (RAII-refcounted,#[must_use],.leak()for pin-until-exit-of-process), and an injectableProtectCb— a callback GC consults before each run that can add externally-known hashes or abort the run (ProtectOutcome::Abort; a flaky protection source skips the sweep rather than risking deletion). - Format-aware mark traversal — a non-raw root (HashSeq/ collection) contributes all reachable children to the live set via the bao hash stream. Protecting a collection protects everything reachable from it.
- Sweep lists the whole store and batch-deletes (~100/batch)
anything not live.
list()being complete is therefore load-bearing for GC — carried into the backend trait contract (OQ-BL-02). - Write-safety against a concurrent sweep — the fs backend carries
a separate
DeleteSettransaction layer (ProtectHandle/ mark-for-delete / protect-cancel / commit);TempTagholds a refcount so a blob put but not yet referenced cannot vanish mid-write. Temp tags are scoped (batch scope / process scope). - No namespaces in their model — flat named tags over a flat pool.
Our namespace concept (OQ-BL-05) lands on their machinery as policy
over what counts as a root:
namespace → root manifest → (mark-walk) → {oids}is tags plus traversal, not new storage concepts.
rudolfs — git-lfs server; the composite-key + decorator-storage prior art
/workspace/rudolfs (v0.3.8, MIT, read-only reference; detailed notes:
/workspace/@alkdev/alknet/docs/research/references/gitlfs/rudolfs-reference.md).
Its storage layer (src/storage/) is the most direct ancestor of
the composite-key + dispatch shape:
StorageKey = (Namespace, Oid)— a composite key over a flat CAS (Namespace = (org, project)strings,Oid = SHA-256). All backend operations take the composite key.Storagetrait —get/put/size/delete/list/total_size/max_size/public_url/upload_url, withLFSObject = (len: u64, ByteStream)— the streaming-first put/get shape, the natural large-blob API and the template for the ops surface's fetch op.- Decorator composition —
Verify ↔ Encrypted ↔ Cached(LRU → permanent) ↔ Retrying → s3/disk. TheCacheddecorator is the appfile shape in LRU form: "fast store in front of fallback" as a composable wrapper, an alternative dispatch policy to size thresholds.fanout()duplicates one stream into two lock-step copies (serve + persist) — rebuilt at our layer and validated as a broadcast seam (POC #3 finding A3). - Verify-as-decorator (OQ-BL-04 input) — streaming SHA-256 check on both put and get paths with auto-purge of a corrupted tier.
- Footnote if at-rest encryption is ever considered — nonce derived from the oid: deterministic and correct, but key rotation invalidates every object and breaks dedup across keys.
- THE ANTI-PATTERN (inverted, not inherited) — namespaces are
physical:
s3://{org}/{project}/{sha256}— identical content in two orgs is stored twice. It buys tenant isolation by paying the cross-tenant dedup that is our whole reason for pooled CAS. The reconciliation (OQ-BL-05): keep the composite-key API shape, put the namespace half in reference tables above the byte layer, keep physical storage flat (oid → bytes, backends namespace-blind). list()is load-bearing — the S3 backend punts onlist()(empty stream) and itsdelete()is a no-op, so it can never GC. A backend-trait-contract lesson: list/sweep support is a real requirement, not an optional extra (POC #1 finding 2 sharpens it: a malformed list is as deadly as an incomplete one).
gix-odb — git's own object database (alkgit's baseline)
/workspace/git-oxide/gix-odb (v0.84; read-only reference). The
backend alkgit currently plans against; the hashing-algorithm baseline
git actually uses (SHA-1/SHA-256), git's loose-object and packfile
layout, and git's already-hash-addressed object model. This crate can
sit under or beside gix-odb semantics because git objects are already
content-addressed and self-verified by git's own model; the crate's
value there is the large-blob story (git-lfs-shaped) and a better
small-object backend than the proposed default, without fighting git's
hash model. Audit conclusions carried into OQ-BL-04/OQ-BL-05:
alternates/multi-pack-index give shared-read across packs with no
dedup/GC/large-blob/streaming value-add (gix-lfs is an empty
placeholder); gitoxide's pack-ID-stability machinery exists only
because packs are immutable-with-stable-IDs, which a per-object
appending pool avoids entirely. Note for the p2p case: git's
alternates/object-pool mechanism (e.g. GitLab object pools for fork
networks) is the prior art for cross-repo dedup among related repos —
the pooled-CAS posture generalizes it to unrelated repos and to p2p
replication.
alknet's appfile external-store probe — historical cross-check only
/workspace/@alkdev/alknet/docs/research/alknet-filesystem/ — written
against an older iroh-blobs. Not a design input: alknet's code is
old, was poorly planned, and has been decomposed into the alk* family;
the substrate is alkcall. The re-verification ran 2026-10-02 as part of
POC #3 (poc-largeblob-findings.md, final section): its architectural
claims about iroh-blobs held up, its performance claims were
asserted-never-measured and are superseded by POC #3's first-party
benchmark. What remains relevant is only the shape lineage — the
appfile pattern this crate now owns by its own validated design.
alkcall — the substrate (and the ops surface's foundation)
/workspace/@alkdev/alkcall (the call + channels RPC crate). It
contains no transport code — no networking dependencies (README:
"a pure protocol crate — no networking, no transport dependencies");
a Connection wraps an already-established byte stream that the
caller provides (dialing/listening happens above it, per ADR-012's
dial/call decoupling). What it owns is the wire/ops protocol: JSON
op family with JSON Schema validation and typed error schemas
(ADR-016), binary streaming via the channel_open marker (ADR-047)
in both directions (Sub ADR-021 for verified fetch, Pub/
HandlerKind::Sink ADR-046 for put), ACL via AccessControl +
ownership checks (ADR-011) under the ADR-017 privilege model, service
discovery, and N-channel multiplexing; established data-path patterns
(BiStream, two-pump pump_bidi ADR-050). The store layer itself
stays alkcall-free (principle 1: substrate-agnostic); alkcall enters
only at the ops module/sibling boundary — which is where transport,
if p2p git ever needs a specific one, enters above that.
Open Questions
Numbering stable (OQ-BL-01..06); the final set is promoted into Phase
1's docs/architecture/open-questions.md with these states.
OQ-BL-01: Crate scope — store-only, or store + ops surface? — RESOLVED (shape + substrate)
Two separate questions were welded together in the original framing and are now settled separately.
The substrate question — settled: the ops surface rides alkcall. (This was previously carried as an open lean; it is now a justified decision rather than an inherited assumption.) The original draft's "undecided pending the first consumer's shape" conflated two things: iroh-blobs' actual weld (no Store trait — its message-actor store is inseparable from the irpc protocol machinery, the thing convention 7 rejects) and the presence of a protocol substrate under the ops layer, which is a different layer entirely and is family-standard. The inverted-point posture constrains the store layer only: the store grows no wire framing, op dispatch, ACL, or transport internals. It does not argue for an ops layer that reinvents them.
Why alkcall specifically, rather than an in-band op scheme:
- Both op shapes exist there already — with validation. JSON ops
carry
OperationSpec+ JSON Schema validation, typed error schemas (ADR-016), External/Internal visibility, andfrom_calldiscovery. Binary streaming ops are thechannel_openmarker pattern (ADR-047): an op declares "my stream is binary" and the channels layer allocates a data channel — the two directions areSub(server→client stream, ADR-021: the verified-fetch case) andPub/Sink(client→server stream, ADR-046: the put case). Blob ops are exactly this mixed family — a validated JSON control plane (have/need announcements,stat, offer/carry lengths) plus binary data channels (bulk bytes, broadcast fanout per POC #3 A3). - ACL is solved there — reviewed machinery we should not parallel.
AccessControl(AND/OR scopes,resource_type/resource_id_pathownership checks per ADR-011) maps onto namespace gating directly (namespace = resource; put/fetch/delete sweeps map onto resource actions); ADR-017 fixes the privilege-escalation subtleties (internal calls switch authority context, never skip ACL) — a security model an in-band scheme would have to re-derive, unreviewed. Convention 6 already mandates this seam; an invented alternative would be a second authorization story in the family. - The store stays substrate-agnostic regardless. alkcall enters
only at the ops module/sibling layer, behind the store's public
surface — the same relationship alkhttp's handlers have to their
storage. If the ops surface were ever re-homed, the store layer
loses nothing (that independence is principle 1); the point is that
the ops surface, when it exists, does not re-implement what the
substrate provides — and neither layer grows transport, which has
no home in either crate (alkcall's
Connectionis handed in, not dialed).
The crate-boundary question — reduced to placement only: given an
alkcall-shaped op family, the remaining choice is where the
registration/adapter lives: (i) feature-gated ops module in this crate
(alkblobs/ops behind a feature flag; the alkcall dependency is
feature-gated so the base crate stays lean), or (ii) a sibling crate.
Lean: (i), matching the alk* feature-gate pattern (convention 8) — the
op family is store-shaped (hashes in, streams out) rather than
policy-shaped, so a sibling buys nothing until a second store-adjacent
op family appears. Deciding input stays what alkgit's replicator
wires first; the Phase 1 ADR records which placement shipped, not
whether alkcall is used.
Residual shape note (rides Phase 1 regardless): the have/need +
verified-fetch family is settled as the confirmed consumer; the
streaming fetch handler shape is the broadcast-fanout seam POC #3
validated (finding A3), and the put/get shapes ride POC #3's
LFSObject-shaped seams.
OQ-BL-02: Multi-backend dispatch — SETTLED (base); residuals named
Settled (POC #1 findings 1/2/6, POC #3 findings A1-A5):
- The lean
Backendtrait contract works: opaque byte keys (typedKeyconverts at the store-layer boundary only),has/get/put/delete/list/name, andlist()complete by contract — a malformed list is as deadly as an incomplete one (the redb key-vs-value trap; list correctness is only observable through GC). - Virgin-store read paths must treat a missing table as absent/empty.
- Size-threshold dual dispatch (kv small / fs large) is a thin layer with deterministic per-digest routing (same content ⇒ same length ⇒ same backend). Default Phase 1 threshold: 128 KiB (tunable constructor parameter; measured crossover ~128-256 KiB, sqlite ~9-10× faster at 1-16 KiB — benchmark harness included in the POC crate, re-runnable on real media).
- Streaming put seam:
(Option<len>, stream); known-length one-pass hashing as the encouraged path, unknown-length stage-then-hash forced by the git preamble (finding A1). Stage-file hygiene: every failure path converges on stage-discard, every success path on commit-rename (finding A5). - Ops-layer fetch handler: broadcast fanout above the backends (finding A3), late joiners degrade to get-after-commit.
Residuals (Phase 1 ADR inputs, no POC gate): migration-between-
backends policy; whether per-namespace backend config is ever needed;
the unknown-length-vs-known-length dispatch asymmetry (finding A6: an
unknown-length re-put of a small blob lands on fs — options (a)
tier-correction sweep, (b) require len for small tier, (c) accept
duplication, (d) bounded-memory pre-threshold buffering; (d) looks
natural, weigh against batch-scope pinning).
OQ-BL-03: Hash abstraction — RESOLVED: canonical git-family hash
Resolved 2026-10-01; implementation-validated by POC #1
(byte-exact vs git hash-object, 2026-10-02). The original
question — "trait, enum, or per-backend config for tolerating multiple
hash algorithms?" — was based on an inherited premise: iroh-blobs uses
BLAKE3, so we must. With that inheritance rejected (convention 7), the
multi-hash framing dissolved; what the consumers force replaced it:
- alkgit forces git's derivation — by protocol definition. Unavoidable.
- The vfs/appfile/workspace case forces nothing — any cryptographically sound content hash works; consumers see addresses, not algorithms.
- Nothing forces BLAKE3. Its virtues (merkle-native chunk trees, raw throughput) are irrelevant at our scales under the canonical digest.
The resolution, three tiers:
- Canonical:
git-blob-sha-256— git's oid derivation ("blob <len>\0" + content, SHA-256, uncompressed content) is the store's canonical algorithm. Git objects land in the pool as first-class entries (no indirection); one algorithm domain = one address space, so non-git consumers collide into the same pool by construction — cross-consumer dedup for free. - Tolerated legacy:
git-blob-sha-1— existing git repos are overwhelmingly SHA-1; the store accepts them (protocol necessity; git's own hardened-SHA-1 threat model inherited verbatim). - Conditional: BLAKE3 — demoted to an encoding consideration: if a bao-like chunk-tree transfer encoding is ever adopted (OQ-BL-04), BLAKE3 lives there, confined to the encoding layer; its verified whole-contents register into the pool as ordinary git-blob-sha-256 entries.
Security footnote: one algorithm across consumers is safe here — SHA-256 has no practical cross-protocol ambiguity at these input shapes, and the git-blob prefix domain-separates the derivation. The SHA-1 caveat is already git's accepted collision-hardened position.
Key-encoding mechanics (from POC #1 finding 5) — Phase 1 ADR: the
tagged key encoding ([tag][algorithm byte][digest]) validated;
algorithm identity must remain part of the key itself (a per-store
single-algorithm config would not have survived the external-git-oid
interop test); the single tag byte is probably needless (one key kind
today); fixed-length enum-tagged keys sort cleanly for the sweep's
live-set diffs. This encoding is the one-way door — byte layouts,
length prefixes, wire representation are the ADR's content, and the
wire-format question (OQ-BL-01) inherits whatever it settles.
OQ-BL-04: Verification and chunking — SETTLED (posture); encoding conditional
Settled posture:
- Whole-file CAS is the default granularity; git objects (and any small-tier content) self-verify by reading whole and checking the canonical digest — put/get paths hash-check.
- Per-range verification is impossible under the canonical digest
(POC #3 finding A2, independently reproduced in
iroh-blobs-eval.md): the preamble hashes the length, so no slice has any hash relationship to the whole. iroh gets range verification from the bao chunk tree, not from BLAKE3 the algorithm. Consequences: local range reads need no verification beyond a transport-integrity slice digest the caller carries out-of-band (the store'sread_rangereturns one); git objects self-verify under git's model when read whole; small-tier range reads slice whole values in memory. - The chunk-tree encoding is a transfer-layer conditional, not a storage design. Its only genuine consumer is networked ranged fetch of large blobs (p2p git + LFS-shaped artifacts). If adopted (BLAKE3-confined per OQ-BL-03 tier 3), its verified whole-contents register as git-blob-sha-256 pool entries, so chunked and whole-file paths dedup in the one pool. Content-defined chunking (fastcdc/ restic-style) would be the shift-resistant choice for small-edit large binaries; fixed-size trees (iroh/bao) are not inherited as the default. Git objects never want chunking — they are unit blobs.
- If chunks were ever materialized as themselves small blobs (the manifest-layer option), the store stays whole-file-only and chunks pool like any other content. Not designed; parked behind the same transfer-layer condition.
Phase 1 surface addendum: a cheap stat(digest) (length/type
probe) — gix's header-only read is git's cheapest primitive, and the
POC's LfsObject.len covers only the get side.
stat(digest) + read_range + stream form the large-blob serving
surface (also what packfile serving rides, OQ-BL-05).
OQ-BL-05: Pooling and GC — settled in shape; residual mechanics named
Settled decisions (2026-10-01 rounds; POC #1 findings 3/7/8; POC #3 pack analysis):
- One pooled CAS per node; repos/workspaces are sets of hash references (manifests/trees) rather than isolated stores — cross-repo dedup by construction.
- Physically flat, logically namespaced (the rudolfs inversion).
Physical layer:
hash → bytes, flat,{hex-prefix}/{hex-prefix}/{hash}sharding survives, backends never see the namespace. Logical layer:namespace → {hashes}reference tables giving GC roots, per- namespace sweeps, accounting, and the ACL boundary (for network ops, the namespace is the alkcallAccessControlresource — per OQ-BL-01's substrate settlement, withresource_id_pathselecting the namespace). - Tags are the root table; references stay above the crate.
iroh-blobs' named
Tagis the "an external consumer cares about this hash" concept — namespaces map onto it asnamespace → root manifest tag → {oids}. Git refs are the names; LFS pointer files are the same shape (the pointer in the tree is the reference). Store stays structure-blind: it does not walk consumer-side manifests. - GC mechanism: mark-and-sweep from registered roots — validated
by POC #1 (exact sweep counts, abort semantics, delete-then-recover
as re-put byte-identical under the same oid). Roots = registered
namespace/root tags + put-path RAII pins + a protect callback with
abort semantics (
ProtectOutcome::Abort— protection-source errors skip the sweep rather than risk deletion). Liveness computation beyond "these roots exist" is the consumer's job, via the callback seam. - Traversal ownership: lean (a) generalized protect callback + temp-tag pinning — the layer above computes liveness and hands hash sets to GC; the store stays structure-blind. Option (b) (store-native manifest traversal) stays parked unless a real consumer's live-set computation proves too heavy; option (c) (children-of callback walkthrough) superseded by the lean.
- Pack tension resolved in shape — git objects enter the pool as
loose-equivalent kv entries (small tier, the common case); if
packfile serving is ever wanted, packs are stored as large blobs
and served via the store's range-read surface (git's own
.idxdoes per-object offset lookup; local range serving is fine per finding A2). Dedup at pack granularity is weak, but pack-level dedup across unrelated repos was never the goal — object-level dedup lives in the loose tier. A per-object appending pool also avoids gitoxide's pack-ID-stability machinery entirely (there are no IDs to rebind; the address is the content).
Residuals (Phase 1 ADR inputs):
- Namespace registry shape (namespaced root tags vs separate reference
tables — POC #1 built in-memory
NamespaceTables; table persistence is mechanical Phase 1 work). - GC scheduling (interval + injectable protection vs consumer-driven sweeps; incremental sweeps are an in-(a) option since the callback is injectable).
- Concurrency — the sweep-vs-put race is a named Phase 1
requirement (POC #1 finding 3: the pin list read under a lock
leaves a real race window between "list" and "delete"). Prior art
named: iroh's DeleteSet/ProtectHandle transactions; the per-hash
serialized-actor + snapshot-observation pattern
(
iroh-blobs-eval.md"re-borrowed conclusion"). Batch-scope pins (iroh'sBatch::temp_tag) for multi-put writes (manifest writes are exactly this). - API-shape lesson to carry: the liveness seam must be live-shared —
name it
register_liveness_source(or have the source hold and register itself), neverinstall_into(POC #1 finding 7: the "install" verb implied copy semantics and broke liveness).
OQ-BL-06: POC register — COMPLETE
| # | What | Status | Where |
|---|---|---|---|
| 1 | Backend-trait + dispatch shape (kv small / fs large); trait must include list() complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path; the hash-enum abstraction with git-blob-sha-256's domain-separated preamble + git-sha-1 tolerance case |
Passed 2026-10-02 (8 findings) | poc-trait-dispatch-findings.md; code: /workspace/alkblobs-trait-poc |
| 2 | Absorbed, validated by #1 | — | |
| 3 | Large-blob path (streaming LFSObject-shaped put/get, fanout seam, fs range reads) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) |
Passed 2026-10-02 (6 findings A1-A6 + pack-tension analysis + alknet probe re-check) | poc-largeblob-findings.md + iroh-blobs-eval.md; code: /workspace/alkblobs-largeblob-poc |
| 4 | Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics | Covered in miniature by #1 — namespace tables + sweep + recover validated single-threaded; the concurrency half is Phase 1 implementation work, not a POC gate | poc-trait-dispatch-findings.md findings 3/7/8 |
| 5 | Postgres as the kv engine (added post-convergence): inherited sqlite/fs/pg benchmark arms + write/read concurrency scale-out probes | Passed 2026-10-02 (6 findings B1-B6: single-conn pg floor ~1 ms fsync-dominated, ~40× storage overhead; pg PUT scale-out ~12× at 12 conns / ~37k ops/s vs sqlite's ~1.2k WAL-serialized ceiling) | poc-postgres-kv-findings.md; code: /workspace/alkblobs-postgres-poc |
| 6 | redb as the kv engine (added post-convergence, same standing as #5): inherited sqlite/fs arms + redb durability decomposition + scale-out probe | Passed 2026-10-02 (6 findings C1-C6: the "2-7× over sqlite" claim inverted — sqlite ~430× over redb at crash-consistent puts; redb Immediate = 1 fdatasync/commit, 8-30 ms on this disk; reads ~530k/s but irrelevant; write scale-out flat ~42/s; ruled out at a durability-tier mismatch, not a benchmark quibble) | poc-redb-kv-findings.md; code: /workspace/alkblobs-redb-poc |
| 7 | Postgres Large Objects as the fs tier's pg-lo engine (added 2026-10-03, per ADR-008 naming it the candidate fs-tier engine for REQ-2 fleets): LO write/read curves at the packfile regime, tx-scoped handle cost under pooling, pg_largeobject/vacuum posture under churn, crash-orphan behavior |
Passed 2026-10-03 (findings C1-C7: companion-table engine ~400 lines, contract gate passed 10/10 exact-count tests; durable put 60-65 MB/s ≥1 MiB ≈ fs's durable put and beating it below 1 MiB on this fsync-slow disk; cached gets ~70-180 MB/s single-stream / ~0.7 GB/s aggregate over 16 readers — 20-50× behind page-cache fs, the honest named delta; tx-scoped handles pool-compatible via descriptorless lo_get windows which become the shipped get shape; LO creation transactional — crash orphans structurally zero, sweep only for legacy bypassing the companion table; LO catalog pages churn-reused/never returned, autovacuum applies; lo_lseek64 discipline) |
poc-pglo-findings.md (spec: poc-pglo-spec.md); code: /workspace/alkblobs-pglo-poc |
Sequencing outcome: #1 and #3 passed; #2 absorbed/validated under #1;
#4's single-threaded core covered by #1 (the mechanism choice it
gated — mark-and-sweep-with-protect-callback — is settled empirically;
races are construction details with iroh's DeleteSet as named prior
art). The POC register is complete — Phase 0 ended here; Phase 1
begins with the ADR backlog below. (#5 and #6 were added
post-convergence 2026-10-02 as ADR-003 substitution-seam inputs, not
Phase 0 gates: #5 — postgres can hold the kv contract but as a
durability/ops posture change, the multi-client-replicator case where it
scales and sqlite serializes; #6 — redb ruled out, its durability API
cannot express the shipped crash-consistent-without-per-commit-fsync
tier, see poc-redb-kv-findings.md. Together: sqlite's pin now has
four-way triangulated evidence.)
POC placement conventions (inherited from alksocks/alktunnels): a POC
that needs code from this repo runs in a worktree/branch
(.worktrees/research/<task-id>/ per the SDD process); a
self-contained POC runs as a standalone crate in the global workspace
with findings written into docs/research/ here. Findings always land
in docs/research/ regardless of where the code lives.
Post-convergence promotion (2026-10-02, Phase 1): #5's pg engine case and #6's redb guardrail became an accepted decision — ADR-007 (two kv engines: sqlite default + postgres feature-gated, behind one trait, constructor-selected; "two backends" counts tiers, not engines). An interim framing recorded in OQ-10 (postgres as an externally-owned "standing offer" opening on a first multi-tenant replicator deployment) was caught at review as circular hedging — the trigger was a fact only this crate could create — and dissolved on the same reasoning as the Schrödinger's-code rule's README corollary (the OQ-09 precedent). The POC evidence needed no strengthening; only the decision's recording did.
Convergence
Phase 0's final step per docs/sdd_process.md: converge on a
recommended approach. This section is that convergence — the
recommended approach, the evidence it stands on, and the ADR backlog
Phase 1 opens with.
Recommended approach (one paragraph)
Build a content-addressed blob store keyed by the
git-blob-sha-256 derivation (git-blob-sha-1 tolerated, BLAKE3
conditionally available only as a transfer-encoding leaf function):
a lean Backend trait (opaque byte keys, complete-by-contract
list()) with size-threshold dual dispatch — sqlite-like kv for
small blobs, flat sharded fs for large (default threshold 128 KiB,
constructor-tunable) — fronted by a store layer that owns typed keys,
streaming put seam (Option<len>, stream) with stage-then-rename
commit discipline, get/range-read surface (get, stat, read_range
with out-of-band slice digests), RAII put-path pinning, and
mark-and-sweep GC whose liveness sources register via a live-shared
seam (protect callback with abort semantics) over a physically flat,
logically namespaced pool. Everything structure-aware — manifests,
refs, path→hash mapping, git semantics, replicator/gossip policy —
lives above the crate; the network ops surface — have/need + verified
fetch with broadcast fanout, JSON-shaped control ops plus binary data
channels, ACL-gated — rides the alkcall substrate (OQ-BL-01), placed
as a feature-gated ops module here or a sibling crate, recorded in
Phase 1's boundary ADR.
Decision table (what was decided, on what evidence, what remains)
| Area | Decision | Evidence |
|---|---|---|
| Canonical hash | git-blob-sha-256 canonical; git-blob-sha-1 tolerated; BLAKE3 conditional (encoding-layer only) | OQ-BL-03 resolution; POC #1 findings 4/5 (byte-exact vs git CLI, tagged-key coexistence) |
| Key encoding direction | Algorithm identity inside the key; fixed-length enum-tagged keys; tag byte likely dropped | POC #1 finding 5 (interop test survives per-store config) |
| Backend contract | Lean Backend trait: opaque byte keys, has/get/put/delete/list/name, list() complete by contract, virgin-store reads are no-ops |
POC #1 findings 1/2 (redb traps); rudolfs S3 anti-lesson |
| Dispatch | Size-threshold dual routing, deterministic per digest; default 128 KiB | POC #1 finding 6; POC #3 A4 (sqlite ~9-10× at 1-16 KiB, crossover ~128-256 KiB) |
| Streaming | Put seam (Option<len>, stream); known-length one-pass encouraged, unknown-length stage-then-hash; stage-hygiene invariants tested |
POC #3 A1/A5; iroh-blobs-eval (same shape, different reasons) |
| Ops fetch shape | Broadcast fanout (one reader, store arm + subscriber arms), late-join = get-after-commit | POC #3 A3; iroh get_blob_ranges_impl; rudolfs fanout() |
| Verification | Whole-blob via canonical digest; range reads carry out-of-band slice digests; chunk-tree encoding is a transfer-layer conditional only | POC #3 A2; iroh-blobs-eval §Verification |
| Pooling | One pooled CAS per node; physically flat, logically namespaced (reference tables above the byte layer) | Dedup discussion; rudolfs inversion; POC #1 namespace tests |
| GC | Mark-and-sweep; roots = tags + RAII pins + protect callback (abort semantics); traversal ownership lean (a), consumer computes liveness | iroh gc.rs verified read; POC #1 findings 3/7/8 (exact counts, abort, recover) |
| Git packs | Loose-equivalent kv entries are the default storage; packs, if stored, are large blobs served by range reads | POC #3 gix-odb analysis; finding A2 confirms local range serving is sound |
| Naming/API lessons | register_liveness_source (live-shared), not install_into; stat(digest) in the surface |
POC #1 finding 7; POC #3 gix analysis |
Phase 1 ADR backlog (open with these)
| ADR candidate | Scope |
|---|---|
| Key encoding (OQ-BL-03 residual) | Byte layouts, length prefixes, algorithm discriminant, wire representation — the one-way door; must precede any wire consumer (OQ-BL-01) |
| Crate boundary & ops surface (OQ-BL-01 residual) | Ops surface rides alkcall (settled — validation, ADR-016 error schemas, ADR-047 binary channel-open ops, ADR-011/017 ACL); remaining: feature-gated ops module (lean) vs sibling crate, recorded in the Phase 1 ADR |
| GC concurrency & scheduling (OQ-BL-05 residual) | Sweep-vs-put race solution (DeleteSet-shaped transactions / per-hash actors), batch-scope pins, scheduling, namespace-table persistence |
| Dispatch policy residuals (OQ-BL-02 residual) | Migration-between-backends; per-namespace backend config; unknown-length put asymmetry (finding A6 options) |
| Transfer encoding (OQ-BL-04 conditional) | Whether/when to adopt a chunk-tree (BLAKE3-confined; CDC if chunking for edit-resistance) encoding for networked ranged fetch — deferred until that consumer exists |
| Phase 1 open-questions register | Promote OQ-BL-01..06 (with the states above: OQ-BL-03 fully resolved; OQ-BL-01/02/04/05 partially resolved with named residuals; OQ-BL-06 complete) into docs/architecture/open-questions.md |
Phase 0 plan (retrospective — complete)
Write— done 2026-10-02 (POC #3's reading against checkout e82cbdc / v0.103.0).iroh-blobs-eval.mdRe-verify alknet's probe— done 2026-10-02, historical cross-check only (poc-largeblob-findings.mdfinal section; alknet's code is not a design input — the substrate is alkcall).Read gix-odb for the alkgit baseline— done 2026-10-02 (pack tension folded into OQ-BL-05; gix-lfs is a placeholder crate).POC #3— done 2026-10-02,poc-largeblob-findings.md.Converge— done 2026-10-02: see §Convergence above. Inputs to Phase 1: the settled-approach list, the decision table, and the ADR backlog. Phase 1 opens withdocs/architecture/and the promoted open-questions register.