Files
alkblobs/docs/research/phase-0.md
T
glm-5.3-flash cee97defba docs(research): POC #7 — postgres Large Objects as the fs-tier pg-lo engine, passed
The ADR-008 pg-lo admission POC ran in a standalone crate
(/workspace/alkblobs-pglo-poc): PgLoBackend over the ADR-003/008 trait
contract (including size), 10/10 exact-count sweep-outcome contract
tests, clippy/fmt clean; dockerized postgres:16-alpine on :15432,
POC #5 driver stack (tokio-postgres + deadpool) via SQL lo_* functions,
no new dependency.

Gate verdict: passed, with named deltas.

- Performance: durable put 60-65 MB/s at >=1 MiB, within 1.5x of — and
  below 1 MiB beating — durable local fs on this fsync-slow disk;
  cached gets 70-180 MB/s single-stream, ~0.7 GB/s aggregate over 16
  readers (20-50x behind page-cache fs — the honest named delta)
- Contract: companion table is the list()/size()/CAS authority (never
  the catalogs); stage-then-commit; GC-participating lo_unlink delete
- Handles: the tx-scoped descriptor is real but pool-compatible via
  descriptorless lo_get(oid, off, len) windows — window gets keep
  handle-acquire p99 at 1-6 ms under readers <= pool; held descriptor
  is the fallback posture
- Vacuum: pg_largeobject pages churn-reused, never returned; tracked
  by autovacuum; rel-size monitoring named as an ops requirement
- Crash/orphan: LO creation is transactional — kill/terminate
  mid-write-tx leaves zero orphan pages; the only orphan class is a
  committed LO bypassing the companion table (planted, reaped by the
  ~7 ms/oid sweep; committed content survives byte-exact)
- Harness lessons: lo_lseek is int4 — the 64 variants are the
  >2 GiB discipline; shared-table parallel tests are unsound (per-test
  CREATE DATABASE isolation)

Docs: new poc-pglo-findings.md; poc-pglo-spec.md status passed;
register OQ-BL-06 #7 marked passed; ADR-008 pg-lo bullet updated
(duplicate bullet removed) + backends-and-dispatch/open-questions
cross-references.

Verification: cargo test --release (10 passed), clippy -D warnings,
fmt --check in /workspace/alkblobs-pglo-poc.
2026-10-03 04:23:06 +00:00

43 KiB
Raw Blame History

status, last_updated
status last_updated
converged 2026-10-03 (POC

alkblobs — Phase 0 (Exploration)

This document captures Phase 0 (Exploration) for the alkblobs crate: vision, guiding principles, prior art, the validated approach, and the open-question register (OQ-BL-01..06). Phase 0's objective per docs/sdd_process.md: capture vision and guiding principles; research options; validate approaches; converge on a recommended approach. That objective is now met — all POC register rows passed or were absorbed, and this document carries the converged recommendation into Phase 1 (Architecture), where the Architect will produce docs/architecture/ specs, ADRs, and the open-questions tracker.

Drafted 2026-09-30; converged 2026-10-02 after the POC rounds. The chronological accretion from the research rounds has been folded into a single settled posture; the evidence trail lives in the findings documents (poc-trait-dispatch-findings.md, poc-largeblob-findings.md, iroh-blobs-eval.md).

Vision and guiding principles

One sentence: content-addressed blob storage in the alk* family — a put/get/verify store keyed under one canonical content hash (the git-family derivation; legacy SHA-1-tolerant), with pluggable backends (small blobs in a kv/sqlite-ish store, large blobs on a filesystem fallback), split from any wire/ops-protocol layer (framing, op family, ACL ride the alkcall substrate when they exist — OQ-BL-01), which is itself transport-free (alkcall is a pure protocol crate).

Why this crate exists (the converging consumers):

  1. alkgit (planning phase) needs an object backend that is not its proposed default file-based backend ("kind of gross") — blobs stored under git's own object hashes, with an iroh-blobs-shaped large-blob path standing in where git typically delegates to git-lfs. The once-called "hash conflict" does not exist: git's oid derivation is the store's canonical hash (OQ-BL-03), so git objects land in the pool as first-class entries with zero indirection.
  2. Workspace/file-shaped storage in the alk family* (the "appfile" shape; old alknet research is historical lineage only): small blobs need a store faster than the filesystem (first-party measured: sqlite ~9-10× faster at 1-16 KiB, POC #3 finding A4), large blobs want a filesystem fallback, with a filename↔hash mapping above. Dispatching across backends is a store-layer problem this crate owns (OQ-BL-02; validated by POC #1's trait seam and POC #3's streaming + threshold evidence).
  3. Agent workspaces — many agents working simultaneously in workspaces that are mostly identical to each other's, each making small edits in specific areas. Per-agent stores multiply the near-identical content; a shared content-addressed store dedups it by construction, and a path→hash manifest per workspace makes each agent's delta just its changed hashes.

The p2p shape (from the dedup discussion). Two alkgit use cases exist: self-hosted git (where cross-repo dedup matters little — you don't fork your own repos) and p2p git for OSS projects, the demanding one. The p2p sketch: ownership/ACL live in a smart contract on a low-fee network; replicators are push/clone endpoints that watch the contract and cache its state; repos sync via iroh-style gossip (hashes, not content) and pull/sync on demand. Neither alkgit approach takes advantage of dedup — which is exactly a store-layer property. The implications for this crate:

  • Both demanding consumers (p2p replicators, agent swarms) reduce to the same shape: a pooled CAS per node, where repos / workspaces are sets of hash references (trees/manifests), not separate object stores. Cross-repo dedup is then by construction; git's own alternates/object-pool mechanism is the same idea done awkwardly. With the canonical-hash resolution (OQ-BL-03) the dedup strengthens further: workspaces and git repos collide into one address domain — the same file content has the same address whether it entered via a git oid or a workspace manifest.
  • Because everything is hash-addressed, the p2p wire story collapses to "announce hashes, diff hash sets, fetch missing blobs verified" — exactly the ops surface OQ-BL-01 settled. The consumer exists; the ops surface is a confirmed have/need + fetch op family, riding the alkcall substrate (OQ-BL-01), whether it lives in this crate or a sibling.

The scope line: this crate is the store, not the wire/ops protocol. iroh-blobs welds store + provider protocol + tickets + postcard into one crate; the defining posture here is the opposite split — put/get/verify against a hash is its own layer, and any ops/protocol surface lives above it or in feature-gated modules (the alk* inversion-point pattern). "Above it" does not mean "without alkcall": the ops surface is alkcall-shaped (OQ-BL-01) — what the store layer itself must never grow is wire framing, op dispatch, ACL, or transport (alkcall carries none of the last either; the connection is handed to it). What stays above the crate regardless: repos/ workspaces as reference sets (trees/manifests), path→hash mapping, git semantics, the contract/gossip/replicator policy layer.

Guiding principles, inherited from the alk* family:

  1. Substrate-agnostic by construction. The store must not know whether bytes arrive over the network, from a local writer, or get reassembled anywhere particular. Backends behind an injected seam (alktty TtyBackend / alktunnels pump-halves precedent).
  2. One canonical hash; the git-family derivation. The store keys entries under a single canonical algorithm — the git-blob derivation (blob <len>\0 domain-separated input, SHA-256 for new content; OQ-BL-03, implementation-validated by POC #1). Legacy SHA-1 (existing git repos) stays tolerated via the same prefix; the abstraction is a small enum/capability set, and the key encoding is a one-way door.
  3. Borrow conclusions, not wire surface. iroh-blobs' tickets, postcard serialization, provider protocol, and Command-actor store abstraction are design-welded choices we do not inherit. Its store-shape conclusions (kv + flat backends, verification flow, chunking limits, put-commit primitives) are fair game — see iroh-blobs-eval.md.
  4. Producer/consumer vocabulary for any network-facing surface; authorization via alkcall's AccessControl/identity seam, never an in-band invented scheme (convention 6).
  5. Verify where it matters. Git objects self-verify under git's own model — the git case needs nothing from us. Per-range verification is impossible under the canonical digest (the preamble hashes the length; POC #3 finding A2), so verification is whole-blob for stored content; the only genuine consumer of a chunk-tree encoding is networked ranged fetch (a transfer-layer conditional — OQ-BL-04).
  6. Pooling is the point. The dedup wins both demanding consumers want (Forknet-style multi-repo OSS content, agent swarms) require that repos/workspaces are sets of hash references over one pooled CAS per node, not walled-off per-repo stores. Pooling forces a GC story (OQ-BL-05) and a multi-backend dispatch story (OQ-BL-02).
  7. Whole-file CAS is the default; chunking earns its keep only where it must. For source-code-scale content, whole-file hashing captures the agent/edit delta perfectly and needs no chunk tree; fixed-size chunk trees are actually hostile to insertion/deletion edits (byte shifts cascade). Content-defined chunking (CDC — fastcdc/restic-style, shift-resistant) only pays on large binary files receiving small edits — the git-lfs weakness. iroh-blobs/bao use fixed-size trees; that is not inherited as the default. Where chunking would live is scoped in OQ-BL-04.

The settled approach

These are the converged postures. None is a Phase 1 ADR yet; each becomes one (or is already recorded as resolved here) — the ADR backlog is in §Convergence.

  • Not a fork of iroh-blobs. A downstream store crate that borrows its store-shape conclusions (per iroh-blobs-eval.md) while diverging deliberately on hashing and wire surface.
  • Canonical hash = the git-family derivation (OQ-BL-03, RESOLVED): git-blob-sha-256 primary, legacy git-blob-sha-1 tolerated, BLAKE3 demoted to a conditional large-blob transfer encoding. Implementation-validated against the real git CLI, including byte-exact external-oid interop (POC #1 finding 4/5).
  • Backend pluralism, dispatched by size threshold (OQ-BL-02, base settled): the lean Backend trait contract (opaque byte keys, has/get/put/delete/list/name, list() complete by contract) with deterministic size-threshold routing — kv/sqlite for small blobs, fs fallback for large. Validated by POC #1 (trait seam) and POC #3 (streaming shapes, sqlite-vs-fs benchmark: ~9-10× at 1-16 KiB, crossover ~128-256 KiB). Phase 1 default threshold lean: 128 KiB (tunable constructor parameter).
  • Streaming put seam = (Option<len>, stream) (OQ-BL-02 / finding A1): known-length one-pass hashing as the encouraged path; unknown-length staging-then-hash, forced by the git preamble, not a design choice. Get-side: small tier returns values, large tier streams; range reads return slice + out-of-band slice digest.
  • Pooled CAS, not per-repo stores (OQ-BL-05): repos and agent workspaces are sets of hash references over one shared content-addressed store. This is the property both demanding consumers need; everything else lives around it.
  • Physically flat, logically namespaced (OQ-BL-05; the rudolfs inversion): the byte layer is a flat dedup-by-construction CAS (hash → bytes, backends namespace-blind, {prefix}/{prefix}/{hash} sharding survives); namespaces are reference tables above it — GC roots, sweep scoping, accounting, ACL boundary.
  • Mark-and-sweep GC with protect-callback + RAII pinning (OQ-BL-05, mechanism validated by POC #1 findings 3/7/8): pins protect not-yet-referenced puts, liveness sources register into a seam, protection-source failure aborts the sweep (iroh's ProtectOutcome::Abort conclusion), delete-then-recover is a re-put.
  • Pack tension resolved in shape (OQ-BL-05): git objects enter the pool as loose-equivalent kv entries (small tier, the common case); if packfile serving is ever wanted, packs are stored as large blobs and served through the store's range-read surface (git's own .idx does per-object offset lookup). Object-level dedup lives in the loose tier; pack-level dedup across unrelated repos was never the goal.
  • Whole-file CAS is the default granularity; chunking is scoped to the networked-large-blob transfer case (principle 7 + OQ-BL-04).

Prior art

iroh-blobs — the shape inspiration (evaluated, not the base)

/workspace/iroh-blobs (read-only reference), evaluated against checkout e82cbdc / v0.103.0 in docs/research/iroh-blobs-eval.md: no Store trait (message-protocol actor pattern), fs backend on-disk layout (redb entry-state + inline thresholds + crash ordering), verification (subtree-local, chunk-tree-driven), streaming shapes, and the do-not-inherit / keep (conclusions-only) lists.

GC mechanics — verified against that checkout (src/store/gc.rs, src/util/temp_tag.rs, src/store/fs/delete_set.rs). This is the load-bearing prior art for OQ-BL-05:

  • Mark-and-sweep, not refcounting, for persistent liveness. gc_mark_task collects roots from three sources: persistent named tags (a flat tags-0 table, name → HashAndFormat), in-memory TempTags (RAII-refcounted, #[must_use], .leak() for pin-until-exit-of-process), and an injectable ProtectCb — a callback GC consults before each run that can add externally-known hashes or abort the run (ProtectOutcome::Abort; a flaky protection source skips the sweep rather than risking deletion).
  • Format-aware mark traversal — a non-raw root (HashSeq/ collection) contributes all reachable children to the live set via the bao hash stream. Protecting a collection protects everything reachable from it.
  • Sweep lists the whole store and batch-deletes (~100/batch) anything not live. list() being complete is therefore load-bearing for GC — carried into the backend trait contract (OQ-BL-02).
  • Write-safety against a concurrent sweep — the fs backend carries a separate DeleteSet transaction layer (ProtectHandle / mark-for-delete / protect-cancel / commit); TempTag holds a refcount so a blob put but not yet referenced cannot vanish mid-write. Temp tags are scoped (batch scope / process scope).
  • No namespaces in their model — flat named tags over a flat pool. Our namespace concept (OQ-BL-05) lands on their machinery as policy over what counts as a root: namespace → root manifest → (mark-walk) → {oids} is tags plus traversal, not new storage concepts.

rudolfs — git-lfs server; the composite-key + decorator-storage prior art

/workspace/rudolfs (v0.3.8, MIT, read-only reference; detailed notes: /workspace/@alkdev/alknet/docs/research/references/gitlfs/rudolfs-reference.md). Its storage layer (src/storage/) is the most direct ancestor of the composite-key + dispatch shape:

  • StorageKey = (Namespace, Oid) — a composite key over a flat CAS (Namespace = (org, project) strings, Oid = SHA-256). All backend operations take the composite key.
  • Storage trait — get/put/size/delete/list/ total_size/max_size/public_url/upload_url, with LFSObject = (len: u64, ByteStream) — the streaming-first put/get shape, the natural large-blob API and the template for the ops surface's fetch op.
  • Decorator composition — Verify ↔ Encrypted ↔ Cached(LRU → permanent) ↔ Retrying → s3/disk. The Cached decorator is the appfile shape in LRU form: "fast store in front of fallback" as a composable wrapper, an alternative dispatch policy to size thresholds. fanout() duplicates one stream into two lock-step copies (serve + persist) — rebuilt at our layer and validated as a broadcast seam (POC #3 finding A3).
  • Verify-as-decorator (OQ-BL-04 input) — streaming SHA-256 check on both put and get paths with auto-purge of a corrupted tier.
  • Footnote if at-rest encryption is ever considered — nonce derived from the oid: deterministic and correct, but key rotation invalidates every object and breaks dedup across keys.
  • THE ANTI-PATTERN (inverted, not inherited) — namespaces are physical: s3://{org}/{project}/{sha256} — identical content in two orgs is stored twice. It buys tenant isolation by paying the cross-tenant dedup that is our whole reason for pooled CAS. The reconciliation (OQ-BL-05): keep the composite-key API shape, put the namespace half in reference tables above the byte layer, keep physical storage flat (oid → bytes, backends namespace-blind).
  • list() is load-bearing — the S3 backend punts on list() (empty stream) and its delete() is a no-op, so it can never GC. A backend-trait-contract lesson: list/sweep support is a real requirement, not an optional extra (POC #1 finding 2 sharpens it: a malformed list is as deadly as an incomplete one).

gix-odb — git's own object database (alkgit's baseline)

/workspace/git-oxide/gix-odb (v0.84; read-only reference). The backend alkgit currently plans against; the hashing-algorithm baseline git actually uses (SHA-1/SHA-256), git's loose-object and packfile layout, and git's already-hash-addressed object model. This crate can sit under or beside gix-odb semantics because git objects are already content-addressed and self-verified by git's own model; the crate's value there is the large-blob story (git-lfs-shaped) and a better small-object backend than the proposed default, without fighting git's hash model. Audit conclusions carried into OQ-BL-04/OQ-BL-05: alternates/multi-pack-index give shared-read across packs with no dedup/GC/large-blob/streaming value-add (gix-lfs is an empty placeholder); gitoxide's pack-ID-stability machinery exists only because packs are immutable-with-stable-IDs, which a per-object appending pool avoids entirely. Note for the p2p case: git's alternates/object-pool mechanism (e.g. GitLab object pools for fork networks) is the prior art for cross-repo dedup among related repos — the pooled-CAS posture generalizes it to unrelated repos and to p2p replication.

alknet's appfile external-store probe — historical cross-check only

/workspace/@alkdev/alknet/docs/research/alknet-filesystem/ — written against an older iroh-blobs. Not a design input: alknet's code is old, was poorly planned, and has been decomposed into the alk* family; the substrate is alkcall. The re-verification ran 2026-10-02 as part of POC #3 (poc-largeblob-findings.md, final section): its architectural claims about iroh-blobs held up, its performance claims were asserted-never-measured and are superseded by POC #3's first-party benchmark. What remains relevant is only the shape lineage — the appfile pattern this crate now owns by its own validated design.

alkcall — the substrate (and the ops surface's foundation)

/workspace/@alkdev/alkcall (the call + channels RPC crate). It contains no transport code — no networking dependencies (README: "a pure protocol crate — no networking, no transport dependencies"); a Connection wraps an already-established byte stream that the caller provides (dialing/listening happens above it, per ADR-012's dial/call decoupling). What it owns is the wire/ops protocol: JSON op family with JSON Schema validation and typed error schemas (ADR-016), binary streaming via the channel_open marker (ADR-047) in both directions (Sub ADR-021 for verified fetch, Pub/ HandlerKind::Sink ADR-046 for put), ACL via AccessControl + ownership checks (ADR-011) under the ADR-017 privilege model, service discovery, and N-channel multiplexing; established data-path patterns (BiStream, two-pump pump_bidi ADR-050). The store layer itself stays alkcall-free (principle 1: substrate-agnostic); alkcall enters only at the ops module/sibling boundary — which is where transport, if p2p git ever needs a specific one, enters above that.

Open Questions

Numbering stable (OQ-BL-01..06); the final set is promoted into Phase 1's docs/architecture/open-questions.md with these states.

OQ-BL-01: Crate scope — store-only, or store + ops surface? — RESOLVED (shape + substrate)

Two separate questions were welded together in the original framing and are now settled separately.

The substrate question — settled: the ops surface rides alkcall. (This was previously carried as an open lean; it is now a justified decision rather than an inherited assumption.) The original draft's "undecided pending the first consumer's shape" conflated two things: iroh-blobs' actual weld (no Store trait — its message-actor store is inseparable from the irpc protocol machinery, the thing convention 7 rejects) and the presence of a protocol substrate under the ops layer, which is a different layer entirely and is family-standard. The inverted-point posture constrains the store layer only: the store grows no wire framing, op dispatch, ACL, or transport internals. It does not argue for an ops layer that reinvents them.

Why alkcall specifically, rather than an in-band op scheme:

  • Both op shapes exist there already — with validation. JSON ops carry OperationSpec + JSON Schema validation, typed error schemas (ADR-016), External/Internal visibility, and from_call discovery. Binary streaming ops are the channel_open marker pattern (ADR-047): an op declares "my stream is binary" and the channels layer allocates a data channel — the two directions are Sub (server→client stream, ADR-021: the verified-fetch case) and Pub/Sink (client→server stream, ADR-046: the put case). Blob ops are exactly this mixed family — a validated JSON control plane (have/need announcements, stat, offer/carry lengths) plus binary data channels (bulk bytes, broadcast fanout per POC #3 A3).
  • ACL is solved there — reviewed machinery we should not parallel. AccessControl (AND/OR scopes, resource_type/resource_id_path ownership checks per ADR-011) maps onto namespace gating directly (namespace = resource; put/fetch/delete sweeps map onto resource actions); ADR-017 fixes the privilege-escalation subtleties (internal calls switch authority context, never skip ACL) — a security model an in-band scheme would have to re-derive, unreviewed. Convention 6 already mandates this seam; an invented alternative would be a second authorization story in the family.
  • The store stays substrate-agnostic regardless. alkcall enters only at the ops module/sibling layer, behind the store's public surface — the same relationship alkhttp's handlers have to their storage. If the ops surface were ever re-homed, the store layer loses nothing (that independence is principle 1); the point is that the ops surface, when it exists, does not re-implement what the substrate provides — and neither layer grows transport, which has no home in either crate (alkcall's Connection is handed in, not dialed).

The crate-boundary question — reduced to placement only: given an alkcall-shaped op family, the remaining choice is where the registration/adapter lives: (i) feature-gated ops module in this crate (alkblobs/ops behind a feature flag; the alkcall dependency is feature-gated so the base crate stays lean), or (ii) a sibling crate. Lean: (i), matching the alk* feature-gate pattern (convention 8) — the op family is store-shaped (hashes in, streams out) rather than policy-shaped, so a sibling buys nothing until a second store-adjacent op family appears. Deciding input stays what alkgit's replicator wires first; the Phase 1 ADR records which placement shipped, not whether alkcall is used.

Residual shape note (rides Phase 1 regardless): the have/need + verified-fetch family is settled as the confirmed consumer; the streaming fetch handler shape is the broadcast-fanout seam POC #3 validated (finding A3), and the put/get shapes ride POC #3's LFSObject-shaped seams.

OQ-BL-02: Multi-backend dispatch — SETTLED (base); residuals named

Settled (POC #1 findings 1/2/6, POC #3 findings A1-A5):

  • The lean Backend trait contract works: opaque byte keys (typed Key converts at the store-layer boundary only), has/get/put/delete/list/name, and list() complete by contract — a malformed list is as deadly as an incomplete one (the redb key-vs-value trap; list correctness is only observable through GC).
  • Virgin-store read paths must treat a missing table as absent/empty.
  • Size-threshold dual dispatch (kv small / fs large) is a thin layer with deterministic per-digest routing (same content ⇒ same length ⇒ same backend). Default Phase 1 threshold: 128 KiB (tunable constructor parameter; measured crossover ~128-256 KiB, sqlite ~9-10× faster at 1-16 KiB — benchmark harness included in the POC crate, re-runnable on real media).
  • Streaming put seam: (Option<len>, stream); known-length one-pass hashing as the encouraged path, unknown-length stage-then-hash forced by the git preamble (finding A1). Stage-file hygiene: every failure path converges on stage-discard, every success path on commit-rename (finding A5).
  • Ops-layer fetch handler: broadcast fanout above the backends (finding A3), late joiners degrade to get-after-commit.

Residuals (Phase 1 ADR inputs, no POC gate): migration-between- backends policy; whether per-namespace backend config is ever needed; the unknown-length-vs-known-length dispatch asymmetry (finding A6: an unknown-length re-put of a small blob lands on fs — options (a) tier-correction sweep, (b) require len for small tier, (c) accept duplication, (d) bounded-memory pre-threshold buffering; (d) looks natural, weigh against batch-scope pinning).

OQ-BL-03: Hash abstraction — RESOLVED: canonical git-family hash

Resolved 2026-10-01; implementation-validated by POC #1 (byte-exact vs git hash-object, 2026-10-02). The original question — "trait, enum, or per-backend config for tolerating multiple hash algorithms?" — was based on an inherited premise: iroh-blobs uses BLAKE3, so we must. With that inheritance rejected (convention 7), the multi-hash framing dissolved; what the consumers force replaced it:

  • alkgit forces git's derivation — by protocol definition. Unavoidable.
  • The vfs/appfile/workspace case forces nothing — any cryptographically sound content hash works; consumers see addresses, not algorithms.
  • Nothing forces BLAKE3. Its virtues (merkle-native chunk trees, raw throughput) are irrelevant at our scales under the canonical digest.

The resolution, three tiers:

  1. Canonical: git-blob-sha-256 — git's oid derivation ("blob <len>\0" + content, SHA-256, uncompressed content) is the store's canonical algorithm. Git objects land in the pool as first-class entries (no indirection); one algorithm domain = one address space, so non-git consumers collide into the same pool by construction — cross-consumer dedup for free.
  2. Tolerated legacy: git-blob-sha-1 — existing git repos are overwhelmingly SHA-1; the store accepts them (protocol necessity; git's own hardened-SHA-1 threat model inherited verbatim).
  3. Conditional: BLAKE3 — demoted to an encoding consideration: if a bao-like chunk-tree transfer encoding is ever adopted (OQ-BL-04), BLAKE3 lives there, confined to the encoding layer; its verified whole-contents register into the pool as ordinary git-blob-sha-256 entries.

Security footnote: one algorithm across consumers is safe here — SHA-256 has no practical cross-protocol ambiguity at these input shapes, and the git-blob prefix domain-separates the derivation. The SHA-1 caveat is already git's accepted collision-hardened position.

Key-encoding mechanics (from POC #1 finding 5) — Phase 1 ADR: the tagged key encoding ([tag][algorithm byte][digest]) validated; algorithm identity must remain part of the key itself (a per-store single-algorithm config would not have survived the external-git-oid interop test); the single tag byte is probably needless (one key kind today); fixed-length enum-tagged keys sort cleanly for the sweep's live-set diffs. This encoding is the one-way door — byte layouts, length prefixes, wire representation are the ADR's content, and the wire-format question (OQ-BL-01) inherits whatever it settles.

OQ-BL-04: Verification and chunking — SETTLED (posture); encoding conditional

Settled posture:

  • Whole-file CAS is the default granularity; git objects (and any small-tier content) self-verify by reading whole and checking the canonical digest — put/get paths hash-check.
  • Per-range verification is impossible under the canonical digest (POC #3 finding A2, independently reproduced in iroh-blobs-eval.md): the preamble hashes the length, so no slice has any hash relationship to the whole. iroh gets range verification from the bao chunk tree, not from BLAKE3 the algorithm. Consequences: local range reads need no verification beyond a transport-integrity slice digest the caller carries out-of-band (the store's read_range returns one); git objects self-verify under git's model when read whole; small-tier range reads slice whole values in memory.
  • The chunk-tree encoding is a transfer-layer conditional, not a storage design. Its only genuine consumer is networked ranged fetch of large blobs (p2p git + LFS-shaped artifacts). If adopted (BLAKE3-confined per OQ-BL-03 tier 3), its verified whole-contents register as git-blob-sha-256 pool entries, so chunked and whole-file paths dedup in the one pool. Content-defined chunking (fastcdc/ restic-style) would be the shift-resistant choice for small-edit large binaries; fixed-size trees (iroh/bao) are not inherited as the default. Git objects never want chunking — they are unit blobs.
  • If chunks were ever materialized as themselves small blobs (the manifest-layer option), the store stays whole-file-only and chunks pool like any other content. Not designed; parked behind the same transfer-layer condition.

Phase 1 surface addendum: a cheap stat(digest) (length/type probe) — gix's header-only read is git's cheapest primitive, and the POC's LfsObject.len covers only the get side. stat(digest) + read_range + stream form the large-blob serving surface (also what packfile serving rides, OQ-BL-05).

OQ-BL-05: Pooling and GC — settled in shape; residual mechanics named

Settled decisions (2026-10-01 rounds; POC #1 findings 3/7/8; POC #3 pack analysis):

  1. One pooled CAS per node; repos/workspaces are sets of hash references (manifests/trees) rather than isolated stores — cross-repo dedup by construction.
  2. Physically flat, logically namespaced (the rudolfs inversion). Physical layer: hash → bytes, flat, {hex-prefix}/{hex-prefix}/{hash} sharding survives, backends never see the namespace. Logical layer: namespace → {hashes} reference tables giving GC roots, per- namespace sweeps, accounting, and the ACL boundary (for network ops, the namespace is the alkcall AccessControl resource — per OQ-BL-01's substrate settlement, with resource_id_path selecting the namespace).
  3. Tags are the root table; references stay above the crate. iroh-blobs' named Tag is the "an external consumer cares about this hash" concept — namespaces map onto it as namespace → root manifest tag → {oids}. Git refs are the names; LFS pointer files are the same shape (the pointer in the tree is the reference). Store stays structure-blind: it does not walk consumer-side manifests.
  4. GC mechanism: mark-and-sweep from registered roots — validated by POC #1 (exact sweep counts, abort semantics, delete-then-recover as re-put byte-identical under the same oid). Roots = registered namespace/root tags + put-path RAII pins + a protect callback with abort semantics (ProtectOutcome::Abort — protection-source errors skip the sweep rather than risk deletion). Liveness computation beyond "these roots exist" is the consumer's job, via the callback seam.
  5. Traversal ownership: lean (a) generalized protect callback + temp-tag pinning — the layer above computes liveness and hands hash sets to GC; the store stays structure-blind. Option (b) (store-native manifest traversal) stays parked unless a real consumer's live-set computation proves too heavy; option (c) (children-of callback walkthrough) superseded by the lean.
  6. Pack tension resolved in shape — git objects enter the pool as loose-equivalent kv entries (small tier, the common case); if packfile serving is ever wanted, packs are stored as large blobs and served via the store's range-read surface (git's own .idx does per-object offset lookup; local range serving is fine per finding A2). Dedup at pack granularity is weak, but pack-level dedup across unrelated repos was never the goal — object-level dedup lives in the loose tier. A per-object appending pool also avoids gitoxide's pack-ID-stability machinery entirely (there are no IDs to rebind; the address is the content).

Residuals (Phase 1 ADR inputs):

  • Namespace registry shape (namespaced root tags vs separate reference tables — POC #1 built in-memory NamespaceTables; table persistence is mechanical Phase 1 work).
  • GC scheduling (interval + injectable protection vs consumer-driven sweeps; incremental sweeps are an in-(a) option since the callback is injectable).
  • Concurrency — the sweep-vs-put race is a named Phase 1 requirement (POC #1 finding 3: the pin list read under a lock leaves a real race window between "list" and "delete"). Prior art named: iroh's DeleteSet/ProtectHandle transactions; the per-hash serialized-actor + snapshot-observation pattern (iroh-blobs-eval.md "re-borrowed conclusion"). Batch-scope pins (iroh's Batch::temp_tag) for multi-put writes (manifest writes are exactly this).
  • API-shape lesson to carry: the liveness seam must be live-shared — name it register_liveness_source (or have the source hold and register itself), never install_into (POC #1 finding 7: the "install" verb implied copy semantics and broke liveness).

OQ-BL-06: POC register — COMPLETE

# What Status Where
1 Backend-trait + dispatch shape (kv small / fs large); trait must include list() complete by contract (rudolfs anti-lesson) + temp-tag/pinning on the put path; the hash-enum abstraction with git-blob-sha-256's domain-separated preamble + git-sha-1 tolerance case Passed 2026-10-02 (8 findings) poc-trait-dispatch-findings.md; code: /workspace/alkblobs-trait-poc
2 Multi-hash store Absorbed into #1 (2026-10-01 hash round) — the canonical-hash resolution removed the "two families coexisting" question; the residual (preamble abstraction + SHA-1 tolerance) was POC #1's trait work Absorbed, validated by #1 —
3 Large-blob path (streaming LFSObject-shaped put/get, fanout seam, fs range reads) + the small-object sqlite-vs-fs micro-benchmark (the dual-belief anchor) Passed 2026-10-02 (6 findings A1-A6 + pack-tension analysis + alknet probe re-check) poc-largeblob-findings.md + iroh-blobs-eval.md; code: /workspace/alkblobs-largeblob-poc
4 Pooled CAS + GC: namespace reference tables over the flat pool; mark-and-sweep with protect-callback + TempTag pinning; delete-then-recover semantics Covered in miniature by #1 — namespace tables + sweep + recover validated single-threaded; the concurrency half is Phase 1 implementation work, not a POC gate poc-trait-dispatch-findings.md findings 3/7/8
5 Postgres as the kv engine (added post-convergence): inherited sqlite/fs/pg benchmark arms + write/read concurrency scale-out probes Passed 2026-10-02 (6 findings B1-B6: single-conn pg floor ~1 ms fsync-dominated, ~40× storage overhead; pg PUT scale-out ~12× at 12 conns / ~37k ops/s vs sqlite's ~1.2k WAL-serialized ceiling) poc-postgres-kv-findings.md; code: /workspace/alkblobs-postgres-poc
6 redb as the kv engine (added post-convergence, same standing as #5): inherited sqlite/fs arms + redb durability decomposition + scale-out probe Passed 2026-10-02 (6 findings C1-C6: the "2-7× over sqlite" claim inverted — sqlite ~430× over redb at crash-consistent puts; redb Immediate = 1 fdatasync/commit, 8-30 ms on this disk; reads ~530k/s but irrelevant; write scale-out flat ~42/s; ruled out at a durability-tier mismatch, not a benchmark quibble) poc-redb-kv-findings.md; code: /workspace/alkblobs-redb-poc
7 Postgres Large Objects as the fs tier's pg-lo engine (added 2026-10-03, per ADR-008 naming it the candidate fs-tier engine for REQ-2 fleets): LO write/read curves at the packfile regime, tx-scoped handle cost under pooling, pg_largeobject/vacuum posture under churn, crash-orphan behavior Passed 2026-10-03 (findings C1-C7: companion-table engine ~400 lines, contract gate passed 10/10 exact-count tests; durable put 60-65 MB/s ≥1 MiB ≈ fs's durable put and beating it below 1 MiB on this fsync-slow disk; cached gets ~70-180 MB/s single-stream / ~0.7 GB/s aggregate over 16 readers — 20-50× behind page-cache fs, the honest named delta; tx-scoped handles pool-compatible via descriptorless lo_get windows which become the shipped get shape; LO creation transactional — crash orphans structurally zero, sweep only for legacy bypassing the companion table; LO catalog pages churn-reused/never returned, autovacuum applies; lo_lseek64 discipline) poc-pglo-findings.md (spec: poc-pglo-spec.md); code: /workspace/alkblobs-pglo-poc

Sequencing outcome: #1 and #3 passed; #2 absorbed/validated under #1; #4's single-threaded core covered by #1 (the mechanism choice it gated — mark-and-sweep-with-protect-callback — is settled empirically; races are construction details with iroh's DeleteSet as named prior art). The POC register is complete — Phase 0 ended here; Phase 1 begins with the ADR backlog below. (#5 and #6 were added post-convergence 2026-10-02 as ADR-003 substitution-seam inputs, not Phase 0 gates: #5 — postgres can hold the kv contract but as a durability/ops posture change, the multi-client-replicator case where it scales and sqlite serializes; #6 — redb ruled out, its durability API cannot express the shipped crash-consistent-without-per-commit-fsync tier, see poc-redb-kv-findings.md. Together: sqlite's pin now has four-way triangulated evidence.)

POC placement conventions (inherited from alksocks/alktunnels): a POC that needs code from this repo runs in a worktree/branch (.worktrees/research/<task-id>/ per the SDD process); a self-contained POC runs as a standalone crate in the global workspace with findings written into docs/research/ here. Findings always land in docs/research/ regardless of where the code lives.

Post-convergence promotion (2026-10-02, Phase 1): #5's pg engine case and #6's redb guardrail became an accepted decision — ADR-007 (two kv engines: sqlite default + postgres feature-gated, behind one trait, constructor-selected; "two backends" counts tiers, not engines). An interim framing recorded in OQ-10 (postgres as an externally-owned "standing offer" opening on a first multi-tenant replicator deployment) was caught at review as circular hedging — the trigger was a fact only this crate could create — and dissolved on the same reasoning as the Schrödinger's-code rule's README corollary (the OQ-09 precedent). The POC evidence needed no strengthening; only the decision's recording did.

Convergence

Phase 0's final step per docs/sdd_process.md: converge on a recommended approach. This section is that convergence — the recommended approach, the evidence it stands on, and the ADR backlog Phase 1 opens with.

Build a content-addressed blob store keyed by the git-blob-sha-256 derivation (git-blob-sha-1 tolerated, BLAKE3 conditionally available only as a transfer-encoding leaf function): a lean Backend trait (opaque byte keys, complete-by-contract list()) with size-threshold dual dispatch — sqlite-like kv for small blobs, flat sharded fs for large (default threshold 128 KiB, constructor-tunable) — fronted by a store layer that owns typed keys, streaming put seam (Option<len>, stream) with stage-then-rename commit discipline, get/range-read surface (get, stat, read_range with out-of-band slice digests), RAII put-path pinning, and mark-and-sweep GC whose liveness sources register via a live-shared seam (protect callback with abort semantics) over a physically flat, logically namespaced pool. Everything structure-aware — manifests, refs, path→hash mapping, git semantics, replicator/gossip policy — lives above the crate; the network ops surface — have/need + verified fetch with broadcast fanout, JSON-shaped control ops plus binary data channels, ACL-gated — rides the alkcall substrate (OQ-BL-01), placed as a feature-gated ops module here or a sibling crate, recorded in Phase 1's boundary ADR.

Decision table (what was decided, on what evidence, what remains)

Area Decision Evidence
Canonical hash git-blob-sha-256 canonical; git-blob-sha-1 tolerated; BLAKE3 conditional (encoding-layer only) OQ-BL-03 resolution; POC #1 findings 4/5 (byte-exact vs git CLI, tagged-key coexistence)
Key encoding direction Algorithm identity inside the key; fixed-length enum-tagged keys; tag byte likely dropped POC #1 finding 5 (interop test survives per-store config)
Backend contract Lean Backend trait: opaque byte keys, has/get/put/delete/list/name, list() complete by contract, virgin-store reads are no-ops POC #1 findings 1/2 (redb traps); rudolfs S3 anti-lesson
Dispatch Size-threshold dual routing, deterministic per digest; default 128 KiB POC #1 finding 6; POC #3 A4 (sqlite ~9-10× at 1-16 KiB, crossover ~128-256 KiB)
Streaming Put seam (Option<len>, stream); known-length one-pass encouraged, unknown-length stage-then-hash; stage-hygiene invariants tested POC #3 A1/A5; iroh-blobs-eval (same shape, different reasons)
Ops fetch shape Broadcast fanout (one reader, store arm + subscriber arms), late-join = get-after-commit POC #3 A3; iroh get_blob_ranges_impl; rudolfs fanout()
Verification Whole-blob via canonical digest; range reads carry out-of-band slice digests; chunk-tree encoding is a transfer-layer conditional only POC #3 A2; iroh-blobs-eval §Verification
Pooling One pooled CAS per node; physically flat, logically namespaced (reference tables above the byte layer) Dedup discussion; rudolfs inversion; POC #1 namespace tests
GC Mark-and-sweep; roots = tags + RAII pins + protect callback (abort semantics); traversal ownership lean (a), consumer computes liveness iroh gc.rs verified read; POC #1 findings 3/7/8 (exact counts, abort, recover)
Git packs Loose-equivalent kv entries are the default storage; packs, if stored, are large blobs served by range reads POC #3 gix-odb analysis; finding A2 confirms local range serving is sound
Naming/API lessons register_liveness_source (live-shared), not install_into; stat(digest) in the surface POC #1 finding 7; POC #3 gix analysis

Phase 1 ADR backlog (open with these)

ADR candidate Scope
Key encoding (OQ-BL-03 residual) Byte layouts, length prefixes, algorithm discriminant, wire representation — the one-way door; must precede any wire consumer (OQ-BL-01)
Crate boundary & ops surface (OQ-BL-01 residual) Ops surface rides alkcall (settled — validation, ADR-016 error schemas, ADR-047 binary channel-open ops, ADR-011/017 ACL); remaining: feature-gated ops module (lean) vs sibling crate, recorded in the Phase 1 ADR
GC concurrency & scheduling (OQ-BL-05 residual) Sweep-vs-put race solution (DeleteSet-shaped transactions / per-hash actors), batch-scope pins, scheduling, namespace-table persistence
Dispatch policy residuals (OQ-BL-02 residual) Migration-between-backends; per-namespace backend config; unknown-length put asymmetry (finding A6 options)
Transfer encoding (OQ-BL-04 conditional) Whether/when to adopt a chunk-tree (BLAKE3-confined; CDC if chunking for edit-resistance) encoding for networked ranged fetch — deferred until that consumer exists
Phase 1 open-questions register Promote OQ-BL-01..06 (with the states above: OQ-BL-03 fully resolved; OQ-BL-01/02/04/05 partially resolved with named residuals; OQ-BL-06 complete) into docs/architecture/open-questions.md

Phase 0 plan (retrospective — complete)

  1. Write iroh-blobs-eval.md — done 2026-10-02 (POC #3's reading against checkout e82cbdc / v0.103.0).
  2. Re-verify alknet's probe — done 2026-10-02, historical cross-check only (poc-largeblob-findings.md final section; alknet's code is not a design input — the substrate is alkcall).
  3. Read gix-odb for the alkgit baseline — done 2026-10-02 (pack tension folded into OQ-BL-05; gix-lfs is a placeholder crate).
  4. POC #3 — done 2026-10-02, poc-largeblob-findings.md.
  5. Converge — done 2026-10-02: see §Convergence above. Inputs to Phase 1: the settled-approach list, the decision table, and the ADR backlog. Phase 1 opens with docs/architecture/ and the promoted open-questions register.