docs(architecture): Phase 1 spec tree — 7 specs, ADRs 001-006, OQ tracker

- overview/hashing/store-api/backends/gc/ops specs + README index,
  all draft status with YAML frontmatter and cross-referenced ADR/OQ
- ADR-001: substrate posture (store layer alkcall-free = family
  conformance, not deviation) + ops placement decided (feature-gated
  module, no consumer-waited deferral)
- ADR-002: canonical git-blob-sha-256 + key encoding — the declared
  wire-format ADR (one-way door)
- ADR-003: backend contract, sqlite/fs/mem tiers, pure-function
  dispatch, no migration, unknown-length buffering semantics pinned
- ADR-004: exactly two production backends; manifest/path-tree layers
  are consumer-side (the corrected two-vs-three-backends framing)
- ADR-005: pooled CAS, three liveness sources, delete windows with
  pin-before-publish + delete-time arbitration (race closed), no
  ambient scheduling
- ADR-006: whole-blob verification; chunk-tree encodings excluded by
  scoping, not deferred on a paused consumer
- open-questions.md: OQ-BL-01..06 resolved with ADR cross-refs; OQ-07
  /08 externally-owned, OQ-09 deferred(scope) with SDD tracker task
- AGENTS.md: status updated to Phase 1, gix-odb path corrected

Verification: cargo test / clippy -D warnings / fmt --check clean
This commit is contained in:
glm-5.3-flash committed 2026-10-01 14:45:32 +00:00
1 parent 7d9ac8a6ec
commit 4b5009d85a
16 files changed
+1953 -8

No files matched your search

+9 -8
View File
@@ -51,10 +51,10 @@ alk* family, inspired by iroh-blobs' store work but deliberately not a
fork of it. It sits downstream of alkcall (the call + channels substrate
its network operations ride, when they exist) and its first planned
consumer is alkgit (git object storage, where the iroh-blobs hashing
story collides with git's own object hashes). Status: **Phase 0
(Exploration)** — `docs/research/` is the working state of the repo;
`docs/architecture/` does not exist yet and no wire format, API shape,
or backend trait is decided. The conventions below apply to all work in
story collides with git's own object hashes). Status: **Phase 1
(Architecture)** — Phase 0 converged (see `docs/research/phase-0.md`);
`docs/architecture/` is the working state of the repo (specs, ADRs
001-006, open-questions tracker). The conventions below apply to all work in
`src/` and `tests/`. They mirror
`.opencode/agents/implementation-specialist.md` §Project Conventions and
are repeated here so they apply to every session.
@@ -152,16 +152,17 @@ corpus replay on stable) is expected — until then ignore fuzzing.
research findings, POC records, and `phase-0.md` (vision, prior art,
open questions, converged recommendation). Read it before non-trivial
work. The SDD process lives in `docs/sdd_process.md`.
- `docs/architecture/` — does not exist yet (Phase 1 output). When it
lands, the SDD process applies: specs describe WHAT, `decisions/`
ADRs explain WHY, `open-questions.md` tracks what's unresolved.
- `docs/architecture/` — exists (Phase 1 in progress: specs draft,
ADRs 001-006 accepted, open-questions tracker active). The SDD
process applies: specs describe WHAT, `decisions/` ADRs explain WHY,
`open-questions.md` tracks what's unresolved.
- Key prior art (read-only, in the global workspace):
- **iroh-blobs** — `/workspace/iroh-blobs` (fresh upstream checkout):
the shape inspiration, specifically its store work (kv + flat file
backends, bao verification, chunking). Its wire-surface choices
(tickets, postcard, BLAKE3-only) are the ones we are deliberately
diverging from.
- **gix-odb** — `/workspace/git-oxide/gix-odb`: git's own object
- **gix-odb** — `/workspace/gitoxide/gix-odb`: git's own object
database; the backend alkgit currently plans against and the
hashing-algorithm baseline git actually uses.
- **alknet's blobs research** —
+78
View File
@@ -0,0 +1,78 @@
---
status: draft
last_updated: 2026-10-01
---
# alkblobs — Architecture
## Current State
**Phase 1 (Architecture).** Phase 0 (Exploration) converged — see
`docs/research/phase-0.md` for the vision, the evidence trail, and the
POC register (all POCs passed or were absorbed; POC code lives in
standalone crates (`/workspace/alkblobs-trait-poc`,
`/workspace/alkblobs-largeblob-poc`); findings live in
`docs/research/`). No implementation exists yet (`src/lib.rs` is a
placeholder).
The converged Phase 0 posture has been promoted into the specs and ADRs
below. Two framing corrections from the convergence review are captured
there — not inherited as hedges:
- **ADR-001** makes the substrate posture explicit: the store layer
carries no alkcall dependency *by design* — which is conformance to
the family pattern, not deviation from it — and the ops surface rides
alkcall, feature-gated.
- **ADR-004** fixes the "two backends" framing: the crate ships exactly
two backends (kv small / fs large — both required by measured scale
economics); the path→hash mapping layer (the alkfs/alknet "vfs" shape)
is a consumer-layer concern and must never become a third backend.
## Architecture Documents
| Document | Scope | Status |
|---|---|---|
| [overview.md](overview.md) | Purpose, consumers, layer map, dependency posture | draft |
| [hashing-and-keys.md](hashing-and-keys.md) | Canonical hash, key encoding (the one-way door) | draft |
| [store-api.md](store-api.md) | Store facade: put/get/stat/range/pin/batch/sweep surface, errors, invariants | draft |
| [backends-and-dispatch.md](backends-and-dispatch.md) | Backend trait contract, shipped backends, size-threshold dispatch | draft |
| [gc-and-namespaces.md](gc-and-namespaces.md) | Pooled CAS, liveness seams, mark-and-sweep, delete windows | draft |
| [ops-surface.md](ops-surface.md) | alkcall-backed have/need + fetch/put ops, ACL mapping | draft |
| [open-questions.md](open-questions.md) | Centralized OQ tracker (incl. promoted Phase 0 register) | draft |
## Decision Records
| ADR | Decision | Status |
|---|---|---|
| [001](decisions/001-substrate-posture-and-ops-placement.md) | Substrate posture & ops-surface placement (conformance, not deviation) | Accepted |
| [002](decisions/002-canonical-hash-and-key-encoding.md) | Canonical hash + key encoding (the wire-format ADR; precedes any consumer) | Accepted |
| [003](decisions/003-backend-contract-and-dispatch.md) | Backend contract, shipped backends, dispatch policy | Accepted |
| [004](decisions/004-two-backends-no-third.md) | Exactly two backends; manifest layers stay above | Accepted |
| [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Pooled CAS, liveness seams, mark-and-sweep, delete windows | Accepted |
| [006](decisions/006-verification-posture-and-transfer-encoding.md) | Whole-blob verification; transfer encoding excluded by scoping | Accepted |
## Open Questions
Tracked in [open-questions.md](open-questions.md). The Phase 0 register
(OQ-BL-01..06) is promoted there with its resolutions; three questions
remain parked — OQ-07/OQ-08 externally-owned (alkgit seam mapping,
alkfs intake; carried for visibility, gating nothing here) and OQ-09
(ops namespace-visibility default, `deferred(scope)` until the first
embedded ops deployment). **Deferral policy (the "Schrödinger's code"
rule):** a *decision this crate needs before shipping* may not be
deferred on a dependency that is itself waiting for this crate to
exist. Two parking kinds remain legitimate: `deferred(scope)` on a
deciding fact that exists independently of this crate, and
externally-owned questions (how a consumer maps onto this crate) that
gate no decision here — full definitions in the header of
[open-questions.md](open-questions.md).
## Lifecycle
Spec docs: `draft` → `reviewed` → `stable` → `deprecated`. A doc moves
`reviewed` when every open question it references resolves and an
architecture review pass clears it; `stable` when implementation
verifies against it; `deprecated` when superseded (kept for reference).
ADRs use a separate status set (Accepted | Proposed | Deprecated |
Superseded), defined per ADR file. This tree is in `draft` pending the
first architecture review cycle.
+124
View File
@@ -0,0 +1,124 @@
---
status: draft
last_updated: 2026-10-01
---
# Backends and dispatch
## What this is
The physical storage layer: the `Backend` trait contract every backend
implements, the two shipped backends (kv, fs), and the size-threshold
dispatch that routes between them. The trait is the crate's second
stable seam (with the key encoding, ADR-002): backends must survive
digest-layer evolution, so the boundary stays dumb.
## The Backend trait contract (ADR-003)
- **Opaque byte keys, opaque byte values.** Backends never learn what a
digest is; typed `Key` converts at the store boundary (ADR-002). This
is also the persistence boundary — a kv row or an fs filename has no
schema-migration story, so the byte layout must be self-describing
(algorithm-tagged keys).
- **Methods:** `has` / `get` / `put` / `delete` / `list` / `name`.
Streaming put/get shapes ride the store layer's seams (store-api.md);
backends provide byte- or handle-level primitives underneath
(`get` on the fs tier yields a file handle, not a loaded buffer).
- **`list()` complete by contract.** A *malformed* list is as deadly as
an incomplete one: list correctness is only observable through GC
(POC #1 finding 2 — the redb key-vs-value trap produced a sweep that
deleted the wrong blobs, silently). Therefore:
- every backend implementation must prove list correctness through
sweep-outcome tests (the invariant's test gate lives in
store-api.md);
- an implementation that cannot enumerate (the rudolfs S3
anti-lesson) must not ship — GC is a structural requirement, not an
optional extra (ADR-005).
- **Virgin-store reads are no-ops:** read paths treat a missing table
as absent/empty (POC #1 finding 2).
- **Namespace-blind:** backends never see namespaces, tenants, or
reference structure (ADR-005 — the rudolfs inversion; physical
storage is `hash → bytes` flat).
## Shipped backends
Two, exactly — this is the complete set the problem requires; the
"third backend" fear is a category error fixed in ADR-004.
### kv backend (feature `kv`, default-on; sqlite)
Small blobs. Evidence (POC #3 finding A4, first-party measured):
sqlite is ~9-10× faster than fs at 1-16 KiB (the git small-blob regime
— most git objects, workspace files, manifests), with the crossover at
~128-256 KiB where fs stops paying the B-tree row rewrite and wins.
- Bounded reads; `has` as an EXISTS probe; prepared statements.
- WAL + `synchronous=NORMAL` as the shipped durability tier (matching
what the benchmark measured and what iroh's store ships).
- Read paths tolerate the no-tables-yet database (see contract).
### fs backend (feature `fs`, default-on)
Large blobs. Flat sharded layout: `{hex-prefix}/{hex-prefix}/{hash}`
sharding survives from iroh's conclusion (limits directory size on
huge pools); stage-then-commit-rename for the two-pass unknown-length
path; pread-based range reads (POC #3 finding A2 — local range serving
is sound, e.g. for packfiles).
### mem backend (feature `mem`, default-off)
`BTreeMap`-shaped ephemeral backend for tests and in-process
ephemerality. A testing/utility tier, never a production story.
## Size-threshold dispatch (ADR-003)
- **Routing is a pure function of content length.** Same content ⇒ same
length ⇒ same backend; re-puts are deterministic. No content ever
migrates between backends — the migration question existed only to
patch the unknown-length asymmetry, which ADR-003's pre-threshold
buffering eliminates by construction.
- **Default threshold: 128 KiB** (constructor-tunable). Midpoint of the
measured flat zone (POC #3 A4). Re-tuning per deployment media is a
constructor parameter, not an API change.
- **Get fall-through:** small-tier miss queries the large tier
(deterministic, cheap — a stat probe).
- Per-namespace or per-tenant backend configuration: **rejected**
(ADR-003 §Consequences — it would re-weld namespacing into the
physical layer, the rudolfs anti-pattern ADR-005 inverts).
## Where a *new* backend could come from
The trait is open to future implementations (network stores, S3-like
tiers), but nothing in the current consumer set requires one, and the
contract is deliberately hostile to half-implementations (complete
`list()`, GC-participating `delete`). Any future backend is a new ADR
carrying its own sweep-safety story. This crate's roadmap is not
blocked on one (see open-questions.md — alkfs intake may name needs
externally; OQ-08).
## Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| [003](decisions/003-backend-contract-and-dispatch.md) | Backend contract & dispatch | opaque keys, complete list, pure-function routing, no migration |
| [004](decisions/004-two-backends-no-third.md) | Two backends, no third | scope boundary against manifest-layer absorption |
| [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Namespace-blindness | backends see hashes only |
## Open Questions
- **OQ-08**: alkfs requirement intake may name storage requirements
(e.g., durability tiers, sync-friendly layouts) that touch this
layer — open, external owner (alkfs Phase 0).
## References
- `docs/research/poc-trait-dispatch-findings.md` findings 1/2/6
- `docs/research/poc-largeblob-findings.md` findings A2/A4 (+ the
re-runnable benchmark harness)
- `docs/research/iroh-blobs-eval.md` — the fs-layout conclusions borrowed
(sharding, crash ordering, inline thresholds rejected as weld)
- rudolfs notes — `list()` anti-lesson, decorator alternative noted and
not adopted (threshold dispatch chose the simpler policy;
ADR-003 §Context)
- ADR-003, ADR-004, ADR-005; [store-api.md](store-api.md) (the invariants
backends must satisfy)
@@ -0,0 +1,134 @@
# ADR-001: Substrate posture and ops-surface placement
## Status
Accepted
## Context
Phase 0 settled that a network ops surface for the blob store exists
and rides alkcall (OQ-BL-01). Two residuals were left:
1. **Placement** — feature-gated `ops` module in this crate vs a sibling
crate. Phase 0's lean for the module was gated on "deciding input:
what alkgit's replicator wires first" — a deferral pattern a review
of the convergence has rejected as Schrödinger's code: alkgit is
paused *because* this crate is being built; "what the paused consumer
does first" never becomes a deciding input. The deferral would never
collapse; the decision must stand on evidence in hand.
2. **Deviation framing** — the user-level concern that a crate whose
store layer carries no alkcall dependency *deviates from the family
pattern* (alktty, alktunnels, alksocks all depend on
`alkcall = 0.8.x`). If this is a deviation, its rationale must be
explicit rather than assumed; if it is not, the conformance claim
must also be explicit, or every future session re-litigates it.
Also loaded here: the user-level concern that the family pattern must
be stated so the two *coming network consumers* (alkgit, alkfs) — which
will certainly use alkcall — see this crate fitting under it, not
beside it.
## What is actually decided
### 1. The family pattern is "nothing invents wire/ops/ACL", not "everyone depends on alkcall"
The family pattern (conventions 4–6) is: wire framing, op dispatch, and
authorization are invented at most once in the family — by alkcall — and
consumed everywhere else. Whether a crate *depends* on alkcall is then a
function of whether it faces the network, not of family membership:
- alkcall itself: zero sibling deps (it *is* the substrate; the pattern's
home).
- alktty / alktunnels / alksocks: depend on alkcall because they define
network-facing protocol layers.
- alkblobs' **store layer**: does not face the network. It is a library
seam — the same posture as alkcall's own `Connection` (handed in, not
dialed) and alktty's `TtyBackend` (backend-injected). A store cannot
carry alkcall without violating convention 4 (substrate-agnostic by
construction): the store must not know whether bytes arrive over a
network.
- alkblobs' **ops surface**: faces the network; rides alkcall like every
other network-facing family member.
The store layer's alkcall-freeness is therefore **conformance, not
deviation**: it is the pattern applied to a layer that has no network
side. The deviation framing dissolves once layers are named — a crate
may contain both a substrate-free core and a substrate-backed shell, as
this one does. What would be an actual deviation — an ops surface that
invents its own framing/ACL to avoid the alkcall dependency — is
rejected.
### 2. Why alkcall for the ops surface (kept from Phase 0, now recorded)
Both op shapes the blob family needs exist there with validation: JSON
ops with schema validation + typed error schemas (alkcall ADR-016), and
binary streaming via `channel_open` (alkcall ADR-047) in both
directions (Sub ADR-021 = verified fetch; Pub/Sink ADR-046 = put). ACL
maps onto `AccessControl` (alkcall ADR-011) + the ADR-017 privilege
model; an in-band scheme would be a second authorization story
(convention 6). Not adopting iroh-blobs' wire surface ("tickets",
postcard, provider protocol) remains deliberate (convention 7) — the
choice is not "iroh's protocol vs nothing", it is "the family substrate
vs inventing one".
### 3. Placement: feature-gated ops module, decided (no deferral)
The ops surface ships as an `ops` module gated behind feature `ops`
(default-off; the alkcall dep is feature-gated so the base crate stays
lean — convention 8). Decided on evidence in hand:
- The op family is **store-shaped** (hashes in, streams out), not
policy-shaped; a sibling crate buys nothing until a second
store-adjacent op family appears.
- Convention 8 is established family precedent (alktunnels/alksocks
feature-gate their alkcall-facing layers the same way).
- Module→sibling promotion, if ever needed, is mechanically additive —
the types move, the consumers re-point; no stored bytes change. The
cost of being wrong is bounded and small.
Promotion criteria (documented, not a hedge): if a second store-adjacent
op family appears outside this crate's scope (e.g., an alkfs sync op
family that wants blob ops but not this crate's registration), the ops
module may be promoted to a sibling crate. That would be a new ADR — the
decision recorded here is where it ships *now*.
### 4. The store layer never grows wire concerns — even in ops code
The ops module is a consumer of the store facade, not a side door: all
ops go through put/get/stat/read_range/pin (ADR-005, ADR-006). There is
no ops-internal path touching backends directly. This is what makes the
claim "if the ops surface were re-homed, the store loses nothing" true
rather than aspirational.
## Consequences
**Positive**
- Family conformance is explicit and traceable; no future session
re-litigates "why doesn't the store use alkcall".
- The base crate builds lean and substrate-free; the ops story exists
without forcing an alkcall dependency onto non-network consumers.
- alkgit/alkfs integrate the store core and the ops shell à la carte,
with the authorization story unified in alkcall.
**Negative**
- The feature-gate is a configuration surface to preserve in CI
(default + all-features builds; see AGENTS conventions).
- Two-layer naming ("alkblobs core" vs "alkblobs ops") must be kept
tight in docs so consumers don't mistake the shell for the core.
**Neutral**
- alkgit/alkfs will still carry *their own* direct alkcall deps for
their protocol layers; nothing about ADR-001 constrains them.
## References
- `docs/research/phase-0.md` OQ-BL-01 (substrate settlement, promoted
into this ADR)
- `docs/sdd_process.md` (deferral policy — Schrödinger's-code rule)
- [overview.md](../overview.md); [ops-surface.md](../ops-surface.md)
- alkcall ADRs 011/012/016/017/021/046/047/050
- ADR-003 (backends stay substrate-free too); ADR-005 (facade-only ops
access); ADR-006 (verified fetch)
@@ -0,0 +1,125 @@
# ADR-002: Canonical hash and key encoding
## Status
Accepted
## Context
The store keys every entry under a content hash. Phase 0 (OQ-BL-03)
resolved the hash *policy* — canonical git-blob-sha-256, tolerated
git-blob-sha-1, BLAKE3 demoted — after the inherited premise "iroh uses
BLAKE3, so we must" was rejected (convention 7: borrow conclusions, not
welds; nothing in the consumer set forces BLAKE3: alkgit forces git's
derivation by protocol definition, and every other consumer sees
addresses, not algorithms).
The residual was the **key encoding mechanics** — the byte layout of
keys inside backends. This is the crate's declared one-way door
(AGENTS convention 5): key bytes persist in kv rows and fs filenames,
with no schema-migration story; and this is the wire-format ADR that
must exist before the first consumer (convention 7's gate on publishing
wire-relevant decisions). POC #1 validated the mechanics empirically:
byte-exact `git hash-object` interop (finding 4, both algorithms,
external-oid round-trip into the pool), and tagged-key coexistence of
both algorithms in one flat pool (finding 5).
## Decision
### Hash policy (three tiers)
1. **Canonical — `git-blob-sha-256`:** `H("blob <len>\0" + content)`,
SHA-256. The store's canonical algorithm; git objects are first-class
entries with zero indirection; one algorithm domain = one address
space, so non-git consumers dedup into the same pool by construction
(ADR-005's premise).
2. **Tolerated legacy — `git-blob-sha-1`:** same derivation, SHA-1 —
accepted for existing-repo interop (a protocol necessity; git's own
collision-hardened SHA-1 threat model inherited verbatim).
3. **Excluded — BLAKE3:** not present in this crate. It was an
iroh-blobs inheritance, not a requirement (its virtues — merkle-native
trees, raw throughput — are irrelevant here under the canonical
digest at our scales). If a chunk-tree transfer encoding is ever
adopted it is a *new* decision at a *new* layer (ADR-006 excludes it
from this crate's scope by scoping, not by hedge).
Security footnote (carried from OQ-BL-03): one algorithm across
consumers is safe here — SHA-256 has no practical cross-protocol
ambiguity at these input shapes; the preamble domain-separates the
derivation. SHA-1's caveat is git's accepted position, inherited.
### Key encoding (the one-way door's bytes)
Physical key layout inside all backends:
```
key := <algorithm byte> <digest bytes>
algorithm byte: 0x01 = git-blob-sha-256 (32-byte digest)
0x02 = git-blob-sha-1 (20-byte digest)
```
- **Algorithm identity is inside the key.** A per-store
single-algorithm configuration would not have survived POC #1's
external-oid interop test — external git oids and pooled workspace
content must be addressable from the same table.
- **The tag byte is dropped** (POC finding 5: exactly one key kind
exists; the algorithm byte is the load-bearing part). Honesty note:
the POC validated the *tagged* layout `[tag][algorithm][digest]`;
dropping the single tag byte is the decided simplification of that
layout (the finding's own reasoning — one key kind — is the evidence
the simplification rests on; the algorithm byte, the load-bearing
half, is the tested part).
- **Fixed length per algorithm** (32 or 20 bytes); total key lengths
33/21. Fixed-length enum-tagged keys sort cleanly in total order —
required for the sweep's live-set diffs (BTreeMap/range-scan walks).
- **Unknown-algorithm and truncated keys are rejects**, not guesses
(POC's length-checking test is the standard).
- **Backends are opaque to all of this:** they store `&[u8]` keys; only
the store core constructs/interprets typed keys (ADR-003).
Algorithm bytes are assigned from this ADR's registry; new algorithms
require a new ADR (another one-way-door byte).
### The preamble as abstraction
An algorithm = domain-separated derivation (its own input
preprocessing) + digest. The preamble `"blob <len>\0"` is not a
special case — POC #1's domain-separated-premise abstraction held with
no special-casing anywhere. Note the structural consequence: the length
is *inside* the hashed input, which is what forces the two put paths
(ADR-003) and caps verification at whole-blob (ADR-006) — both accepted
rather than worked around.
## Consequences
**Positive**
- Git oids (either format) address pool entries directly from external
producers — no mapping layer, byte-exact (tested against the git CLI,
not just vectors).
- One address space across all consumers; dedup is structural.
- Key bytes are self-describing; backend data survives digest-layer
evolution (opaque at the boundary).
**Negative**
- The encoding is frozen once a backend ships data (the door closes);
algorithm byte 0x00 is reserved and never assigned (a future encoding
revision would need a different key universe or a migration ADR).
- SHA-1 acceptance inherits git's threat model — documented, not
re-litigated here.
**Neutral**
- Variable-length length-prefixed keys were considered (POC finding 5)
and rejected for now: two algorithms make fixed-length simpler and
sort-stable. Extension to a third algorithm re-opens byte layout —
new ADR.
## References
- `docs/research/phase-0.md` OQ-BL-03 (resolution + encoding notes)
- `docs/research/poc-trait-dispatch-findings.md` findings 4/5
- [hashing-and-keys.md](../hashing-and-keys.md)
- ADR-003 (opaque byte keys; two put paths the preamble forces);
ADR-005 (one address space); ADR-006 (whole-blob verification)
@@ -0,0 +1,148 @@
# ADR-003: Backend contract, shipped backends, and dispatch policy
## Status
Accepted
## Context
OQ-BL-02 settled the *base* contract in Phase 0 (lean `Backend` trait,
opaque keys, complete `list()`, size-threshold dual dispatch) with
named residuals: migration-between-backends policy, per-namespace
backend config, and the unknown-length put asymmetry (POC #3 finding
A6: an unknown-length re-put of a small blob lands on fs, duplicating
content the same digest's known-length put would put in kv).
One framing correction also lands here: the "two backends is ugly" /
"three backends" framing that influenced early rounds — corrected to
its real content as ADR-004 (two backends are *required*, and the
manifest layer is downstream, not a backend). This ADR records the
physical layer; ADR-004 records the boundary.
Evidence in hand (no further POC or consumer-waiting needed):
- POC #1 findings 1/2/6 (trait seam, list traps, dispatch thinness),
- POC #3 findings A1/A4/A5/A6 (put shapes, benchmark, stage hygiene,
the asymmetry),
- rudolfs anti-lesson (S3 backend with no list can never GC).
## Decision
### Backend trait contract
- **Opaque byte keys (`&[u8]`), opaque byte values.** Typed `Key`
converts at the store boundary only (ADR-002). Backends must survive
digest-layer evolution — a kv row or fs filename has no migration
story, so opaque, self-describing bytes are the durable choice. (The
generic-over-`K` alternative was considered and rejected in POC #1
finding 1: with one key type, generics buy nothing and break the
persistence boundary.)
- **Methods:** `has`/`get`/`put`/`delete`/`list`/`name`. Streaming and
range shapes ride the store layer above (fs `get` yields a file
handle, not a loaded buffer).
- **`list()` complete by contract** — and *well-formed*: a malformed
list (wrong destructure, values-as-keys) is as deadly as an
incomplete one, and silent until a sweep deletes the wrong things
(POC #1 finding 2). Consequence: backend implementations are accepted
only with sweep-outcome tests proving list correctness.
- **Virgin-store reads are no-ops** (missing table = absent/empty).
- **Namespace-blind** (ADR-005): backends see hashes only.
### Shipped backends (two, default-on; plus a testing tier)
- **`kv` (default-on): sqlite** for small blobs — measured ~9-10×
faster than fs at 1-16 KiB (POC #3 A4, first-party benchmark),
the regime holding most git objects, workspace files, and manifests.
WAL + `synchronous=NORMAL` as shipped durability; prepared
statements; no-tables-yet tolerance.
- **`fs` (default-on)** for large blobs — flat
`{hex-prefix}/{hex-prefix}/{hash}` sharding (iroh's surviving
conclusion); stage-then-commit-rename; pread range reads (sound
locally per ADR-006).
- **`mem` (default-off):** ephemeral `BTreeMap` tier for tests and
in-process ephemerality. Never a production story.
Both production backends are default-on (batteries-included, the
alkgit feature-model precedent); `mem` is not.
### Dispatch policy
- **Size-threshold routing; default 128 KiB, constructor-tunable**
(midpoint of the measured ~128-256 KiB crossover zone; POC #3 A4).
- **Unknown-length puts buffer in memory up to the threshold; the
buffer's fate has three pinned semantics** (finding A6's option (d)):
1. **Threshold exceeded mid-stream** (non-restartable sources are
normal — network streams can't be replayed): the buffered prefix
flushes *into a fresh stage file*, and the stream continues
appending into it. There is no "restart the stream" requirement
anywhere.
2. **Verification runs over the eventual stored artifact as a
whole**: the memory buffer for the sub-threshold kv commit; the
complete staged bytes for the fs commit. "Two passes over staged
bytes" and "hash-check against the canonical derivation" are both
literally true under this — there is no stitched-input case left
ambiguous (the buffer is *in* the stage file by the time the
second pass runs).
3. **Stream error before commit**: no entry lands and no stage file
remains (the stage-hygiene invariant extends to the buffered
case: buffer discard on the failure path, exactly like stage
discard).
This is finding A6's option (d), decided over the alternatives:
- (a) tier-correction sweeps are moving machinery that re-creates
migration, and are unnecessary if routing never misplaces a blob;
- (b) rejecting unknown-length small puts is hostile to streaming
producers for no safety gain (dedup still holds under (d));
- (c) accepting duplicated content across tiers makes `list()` lie
about the pool (two entries for one digest) and complicates GC's
live-set diffs for no benefit.
Option (d) makes **routing a pure function of content length** — the
property that eliminates the asymmetry *and* the migration question
with it. Batch-scope pinning (POC #1 finding 3) is orthogonal — it
protects in-flight puts, it does not route them — and ships anyway
(ADR-005).
- **No migration between backends.** Content is immutable; its length
never changes; routing is pure. The migration question existed only
to patch the asymmetry and is *resolved by elimination*: there is
nothing to migrate. A future backend needing migration writes a new
ADR.
- **No per-namespace backend configuration.** It would re-weld
namespacing into the physical layer — the exact rudolfs anti-pattern
ADR-005 inverts. Tenancy differences that ever require physical
separation (a different pool per tenant) are a deployment/topology
concern (two stores), not a backend-config concern.
## Consequences
**Positive**
- The dispatch is deterministic and test-invariant: same content ⇒ same
backend, always — including across mixed put paths.
- GC's enumerate-everything substrate (`list()` unions) is contract-
enforced upfront rather than discovered against the rudolfs failure.
- The shipped default matches the measured economics on the git/
workspace regime, re-tunable per deployment without API change.
**Negative**
- Unknown-length puts of large content pay two passes (forced by the
preamble — ADR-002 §Consequences) plus staging I/O; bounded-memory
buffering adds a copy for sub-threshold unknown-length puts (cheap,
per the benchmark).
- sqlite's on-disk format enters the crate's surface durability story
(file compatibility across rusqlite/sqlite versions is upstream
sqlite's guarantee, not ours).
**Neutral**
- A network-backed backend (S3-like) remains possible via the same
trait but ships nothing here; its list/GC obligations are the
acceptance gate.
## References
- `docs/research/poc-trait-dispatch-findings.md` findings 1/2/6
- `docs/research/poc-largeblob-findings.md` findings A1/A4/A5/A6
- `docs/research/phase-0.md` OQ-BL-02
- [backends-and-dispatch.md](../backends-and-dispatch.md);
[store-api.md](../store-api.md)
- ADR-002 (key bytes), ADR-004 (boundary vs manifest layer), ADR-005
(namespace-blindness, pins)
@@ -0,0 +1,101 @@
# ADR-004: Exactly two backends; manifest layers stay above the crate
## Status
Accepted
## Context
An early Phase 0 framing — "using two backends is ugly" — shaped several
rounds before being edited away, but the correction it was reaching for
arrived only in the Phase 1 review, and needs recording so no future
round reabsorbs the concern incorrectly:
- There are **always** two physical backends in this problem: storage
for small content (fast, kv/sqlite) and the filesystem for large
content. This is not ugly, it is the measured shape of the problem
(POC #3 A4: crossover ~128-256 KiB; SQLite ~9-10× faster below it).
- The "two backends is ugly" sentiment was actually about a **third
layer** in the old alknet-filesystem research — the *virtual
filesystem / appfile* stack: path trees, branches, tombstones,
filename↔hash maps living *alongside* the blob store (in that research,
literally a co-equal SQLite store beside iroh-blobs).
If left implicit, this misframing has two failure modes: (a) this crate
absorbs a path-tree/manifest layer ("three backends"), growing tables,
schemas, and branching semantics that belong to consumers; (b) this
crate shapes itself such that a future vfs builder must hack around it.
alkfs is a pending consumer precisely of that manifest layer — this ADR
is the boundary it will build against.
## Decision
**This crate ships exactly two production backends and nothing more;
every manifest/namespace/path-tree layer is a consumer, by construction
and by contract.** (A non-production ephemeral `mem` tier exists for
tests — ADR-003; it is not a storage story and does not extend this
boundary.)
1. **Two production backends is the complete set.** kv/sqlite (small)
+ fs (large) cover the scale economics with measured evidence
(ADR-003); no other physical tier is needed by any current or
planned consumer's *blob* needs.
2. **The "third backend" is not this crate's layer.** The vfs/appfile
shape — path→hash mapping, branches/tombstones/deltas per workspace,
the alknet-filesystem lineage — is storage *above* the pool:
reference sets over hashes (ADR-005's namespace machinery), not a
Backend implementation. Consumers build it on their own storage
(alkfs's path tree, alkgit's refs).
3. **The store provides the seams the manifest layer needs, and nothing
that forces it to hack:**
- hash-addressed put/get with a verified, git-interop-compatible
addressing scheme (ADR-002) — manifests reference entries by the
same canonical addresses git uses;
- liveness registration so manifest-holds maps to GC roots without
the store learning manifest formats (ADR-005);
- `stat`/`read_range` — enough for serving path-resolved content and
large-file ranges;
- pinned batch puts — manifest-atomic multi-file writes are
protected in flight (ADR-005).
4. **Namespacing at the store level is logical only** (ADR-005):
reference tables above, flat bytes below. If a consumer wants
physical isolation per tenant/repo, that is two pool deployments (a
topology choice), never a namespace flag inside one.
## Consequences
**Positive**
- The crate's scope stays testable and small: two backends, one pool,
one hash domain. No schema/versioning surface for manifest formats —
consumers iterate on theirs freely.
- alkgit and alkfs build their (very different) mapping layers on one
shared pool without either bending the store; their manifests interop
at the address level by construction.
- The alknet-filesystem lineage is honored as *shape guidance* for
consumers (branches, write sessions, chain walks), with zero of its
storage-layer mechanics (SQLite path tables, honker wiring, CRDT
sync) leaking in here.
**Negative**
- Consumers each re-build the small amount of manifest plumbing they
need (a reference table and a root registration). This is deliberate:
the manifest needs of alkgit (git refs) and alkfs (path trees) are
different enough that a shared one would fit neither.
**Neutral**
- If a future consumer *did* want a first-party manifest layer, it
would be a sibling crate over this one (e.g. `alkfs` itself), not an
extension of this crate.
## References
- `docs/research/phase-0.md` (settled approach: pooled CAS; OQ-BL-05)
- `/workspace/@alkdev/alknet/docs/research/alknet-filesystem/poc-summary.md`
— the historical lineage this ADR explicitly bounds (not a design
input)
- ADR-003 (the two backends), ADR-005 (namespaces as reference sets)
- [overview.md](../overview.md) (consumer map: alkgit, alkfs);
[backends-and-dispatch.md](../backends-and-dispatch.md)
@@ -0,0 +1,201 @@
# ADR-005: Pooled CAS, liveness seams, and mark-and-sweep GC
## Status
Accepted
## Context
OQ-BL-05 settled the pooling/GC *shape* in Phase 0 (one pooled CAS;
mark-and-sweep from registered roots; lean traversal ownership) with
residual mechanics named: namespace registry shape, GC scheduling,
and the sweep-vs-put race. All were resolvable on evidence in hand —
the prior art was read and verified (iroh `gc.rs`, `delete_set.rs`), the
mechanism was POC-validated (POC #1 findings 3/7/8: exact sweep counts,
abort semantics, delete-then-recover as byte-identical re-put, live-
shared callback seam), and the residuals turned on which of the
validated shapes to adopt — not on unknown facts (the
Schrödinger's-code rule of the deferral policy applies: none of these
may wait on paused consumers).
Also loaded: the "physically flat, logically namespaced" inversion of
rudolfs (its `s3://{org}/{project}/{sha256}` physical namespaces pay
cross-tenant dedup — the whole reason for pooling — for tenant
isolation), and the packfile tension (resolved in Phase 0's shape:
loose-equivalent kv entries; packs, if ever stored, are large blobs
served by range reads).
## Decision
### One pooled CAS; namespaces are reference sets above it
- One pool per node; repos/workspaces/all consumers are sets of hash
references over it. Cross-repo dedup by construction (the property
both demanding consumers need; git `alternates` pools as the awkward
prior).
- **Physically flat** — backends see `hash → bytes`, never namespaces
(ties to ADR-003's no-per-namespace-config).
- **Logically namespaced** — a namespace is a consumer-held reference
set; the store-side counterpart is liveness registration only.
- **The store persists no root table.** Namespaced root tags vs separate
reference-table persistence (a Phase 0 residual) resolves as: *the
store holds no durable roots*; consumers keep durable references
(git refs, alkfs path trees) and hand liveness over per sweep via the
live-shared seam. In-memory namespace tables remain available for
ephemeral consumers. This also dissolves the hidden wrinkle of a
store-internal root table under fs-only deployments (where would its
durable bytes live?). Ephemeral consumers are covered by pins; a
deployment that needs store-durable roots registers them as
consumer-owned durable state and sweeps through the seam.
- **ACL boundary:** the namespace is the alkcall `AccessControl`
resource for network ops (resource_id_path; ADR-001/ops-surface.md).
### Marks: three liveness sources
1. **Registered liveness sources** — live-shared consumers (the seam
named `register_liveness_source`; the POC's finding 7 copy-semantics
failure is why the verb is pinned — "install" implies snapshot, and
snapshotting breaks consumer registration that happens after).
2. **Put-path pins (RAII)** — a `Pin` guard holds a refcount per key
until dropped or converted into a consumer reference; batch scope
available for multi-put writes (manifest writes are exactly this;
POC finding 3's named Phase 1 requirement, satisfied here). The
conversion ordering is pinned: a pin's drop and its replacement
reference's registration into a liveness source must be one
ordered step under the same arbitration lock the sweep consults —
a consumer may not release a pin before its registered source
observes the replacement (the local analog of the ops-layer pin
token contract).
3. **Protect callback** — consulted pre-sweep; may add known-live keys
or **abort the run** (typed `GcAborted`, nothing deleted). iroh's
`ProtectOutcome::Abort` conclusion adopted: a flaky protection
source skips the sweep rather than risk deletion.
### Sweep: mark → delete window → commit
- Enumerate the pool via backends' complete `list()` (ADR-003's
contract; its reason to exist).
- Compute the live set from all three sources; batch-delete the dead
(~100/batch; iroh's proven shape).
- **Delete windows** close the sweep-vs-put race (POC finding 3's named
Phase 1 requirement — the pin-map lock alone leaves a real window
between "list" and "delete"):
1. **mark:** enumerate the pool; compute the live set from all three
liveness sources; stage the candidate-dead set;
2. **delete window opens:** for each candidate, at *deletion time*
(not once upfront), arbitrate under the pin/liveness lock: if the
key is pinned or re-observed live, cancel it from the window;
otherwise proceed to delete;
3. **commit:** remaining candidates deleted in batches; window
closed.
The invariant holds because both halves of the race are now
ordered under the same lock:
- **Put commits pin before publish.** A put's `Pin` is acquired
*before* the entry becomes visible in its backend; an entry cannot
appear in a later sweep's `list()` without already carrying its
pin (or its liveness registration, whichever the consumer
converted the pin into) — publish and protection are one step,
and a put for an already-pool-existing entry (dedup) acquires no
*new pool entry* — the returned `Pin` refcounts the existing entry
under the same arbitration lock (a dedup call's returned Pin
protects exactly as a fresh put's does).
- **Deletes arbitrate under the pin lock at delete time.** A
deletion of key *k* cannot proceed while *k* carries a pin; since
any visible entry is pinned (invariant above), any deletion of a
visible entry happens only with no pin held — and a *concurrent*
re-put of *k* is either (a) blocked on the arbitration lock until
the deletion commits, after which the re-put lands as a fresh
pinned put (delete-then-recover semantics, validated), or (b)
ordered after, in which case its pin already protects it.
Net property, test-asserted: **a visible pool entry is never
deleted while liveness (pin or registered source) protects it, and
an in-flight put is never deleted by a sweep started before it
committed.** The arbitration can be cheap: the pin check is one
map lookup per candidate key; the delete window exists so the
reconciliation happens precisely at the delete boundary rather
than wholesale.
The iroh DeleteSet/ProtectHandle transactions and per-hash
serialized-actor pattern are the re-borrowed prior art this shape
generalizes (their serialized actor is exactly "arbitrate at
delete time"; our window batches the arbitration).
- `has`/`get` during a window see no torn state (immutable entries;
existence flips atomically per key).
- **Delete-then-recover is a re-put** (validated: byte-identical under
the same key); no tombstone layer exists.
- **Direct `delete(key)`:** refuses pinned keys with a typed error
(same arbitration); permitted for embedder correction flows on
unpinned entries. Outside-of-sweep deletion is unusual by posture —
most deletion should flow through sweeps.
### Traversal and scheduling
- **Lean (a) confirmed:** consumers compute liveness beyond "these
roots exist" and hand hash sets over; the store stays structure-blind
(what Phase 0 called "lean (a)": consumer-side traversal, vs option
(b) store-native traversal of consumer manifests). Store-native
traversal re-opens only on a *measured* consumer cost (an external
fact, not a parked hedge).
- **Explicit `sweep()`; no ambient timers.** The embedder owns cadence
(idle sweeps, interval sweeps, consumer-driven prunes — all one API).
- **Accounting:** pool-level totals via `list`+`stat`; per-namespace
accounting is computed by consumers from their reference sets.
### Packfiles (the Phase 0 pack tension, carried forward)
Git objects enter the pool as loose-equivalent kv entries (small tier
— the common case). If packfile serving is ever wanted by alkgit, packs
are stored as large blobs and served through `stat`/`read_range` (git's
`.idx` does per-object offset lookup; local range serving is sound —
ADR-006). This avoids gitoxide's pack-ID-stability machinery entirely
(there are no IDs to rebind; the address is the content). This restate-
as-decision carries no new open question; whether alkgit wants it is
OQ-07's.
## Consequences
**Positive**
- The dedup property the consumers exist for is structural, not
configured; pooling semantics can't be "off".
- Sweep safety is testable to exact counts (POC-proven) at the
architecture level: the invariant (delete-window re-observation) is
the test gate, independent of mechanism choice.
- No GC coupling in the store: a consumer without registered liveness
sources cannot have its content deleted (sweep aborts) — safe-by-
default.
**Negative**
- GC requires complete, well-formed `list()` from every backend
(enforced upstream, ADR-003) and delete-window machinery in the
sweep path — the most intricate state machine in the crate.
- Explicit sweep means an embedder that never schedules one accumulates
garbage (their choice; documented posture, not a defect).
- Pin/liveness misuse by consumers is runtime-observable (a sweep may
delete unreferenced content if a consumer deregisters early); the
seam's correctness contract is on consumers — documented in
store-api.md's invariants.
**Neutral**
- Cross-node GC coordination (p2p replicators) lives in the replicator
policy layer, above; the seam is all it needs.
## References
- `docs/research/phase-0.md` OQ-BL-05 (settled decisions + residuals,
all promoted here)
- `docs/research/poc-trait-dispatch-findings.md` findings 3/7/8
- `docs/research/iroh-blobs-eval.md` (gc.rs / delete_set.rs reads)
- rudolfs (the physical-namespace anti-pattern inverted);
gix-odb (alternates/pool prior art)
- ADR-002 (one address space), ADR-003 (list contract, pins), ADR-004
(manifest layers above), ADR-006 (verification limits inform recover
semantics)
- [gc-and-namespaces.md](../gc-and-namespaces.md);
[store-api.md](../store-api.md)
@@ -0,0 +1,106 @@
# ADR-006: Verification posture — whole-blob checks; transfer encodings excluded by scoping
## Status
Accepted
## Context
OQ-BL-04's settled Phase 0 posture: whole-file CAS as the default
granularity; per-range verification impossible under the canonical
digest (the preamble hashes the length — POC #3 finding A2,
independently reproduced in the iroh-blobs eval; no slice has any hash
relationship to the whole); the chunk-tree encoding framed as "a
transfer-layer conditional — deferred until a consumer exists".
That last framing is a deferral this crate cannot make under its own
deferral policy: "the consumer that would want networked verified
ranged fetch" (p2p git's replicator) is paused *because* this crate is
being built. A condition gated on it is Schrödinger's code — it never
collapses. The decision must stand on the evidence in hand, which is
complete for the *scoping* question even though nothing measures a
future consumer:
- What the consumers need today is verifiable **whole-blob** transfer:
fetch bytes, hash-check against the canonical digest the caller
carries. That is exactly verified-fetch (the `Sub` op with the digest
in the offer), works at any size, and needs no encoding — POC #3's
streaming paths are validated byte-exact against git.
- What would need a chunk tree is *per-chunk integrity during ranged
transfer of one large blob* — a property nothing in the current
consumer set requires, and which cannot be retrofitted onto the
canonical digest anyway (impossible under the preamble, finding A2;
iroh gets it from the bao *tree*, not BLAKE3 the algorithm).
- Storing chunk trees would add a second stored artifact per blob
(outboard encoding), a second write path, new GC surface, and a
second hash domain — speculative cost before any consumer names the
need.
Also settled here: local range reads need only transport-integrity
slice digests (out-of-band, carried by the caller — the store returns
one from `read_range`); git objects self-verify under git's own model
when read whole; small-tier range reads slice whole values in memory.
## Decision
1. **Verification is whole-blob.** Put and get paths hash-check against
the canonical derivation (ADR-002). `stat` probes length without
reading content.
2. **Range reads return `(slice, slice_digest)`** — the slice digest is
a transport-integrity check carried out-of-band by the caller (e.g.
the offering side of a fetch). The store performs no per-range
verification *because none is derivable* under the canonical digest;
local serving (packfiles, large-file reads) is sound on honest media
(finding A2).
3. **No chunk-tree, CDC, or otherwise partial-verification encoding is
adopted in this crate — by scoping, not by deferral.** The crate's
domain is the store; a transfer encoding is a network-layer artifact
and has no owner here. If a future networked consumer (a p2p
replicator above an un-paused alkgit, alkfs sync, a future ops
extension) names a requirement for verified ranged fetch of one
large blob, that requirement is met by a *new decision at a new
layer*: an optional encoding module (BLAKE3-confined per ADR-002's
tier 3 if a tree encoding is chosen; content-defined chunking if
edit-resistance is the goal) whose verified whole-contents register
back into the pool as ordinary canonical entries — dedup between
chunked and whole-file paths is preserved by registration, not by
shared structure.
4. **The verified-fetch op (ADR-001) carries whole-blob verification**
as its integrity story: digest in the offer, hash-check on receipt,
fanout above the store for concurrent subscribers (POC #3 finding
A3); late joiners degrade to post-commit verified `get`.
## Consequences
**Positive**
- The store stays single-granularity, single-hash-domain, single
write-path family — the simplest shape consistent with all
consumers' *current* needs; no speculative outboard/GC surface.
- Verified transfer exists at every blob size with zero extra machinery
(whole-blob digest check), which is what the confirmed ops family
serves.
**Negative**
- A future large-blob verified-ranged-fetch consumer pays a new module
and a second decision — deliberately deferred *cost*, not unmade
decision (the distinction this ADR records).
- Callers wanting mid-transfer corruption localization on very large
blobs get error-at-EOF, not error-at-chunk, without an encoding layer
(acceptable; named above).
**Neutral**
- BLAKE3's "large-blob encoding consideration" stays conditional and
out-of-crate; nothing here allocates an API slot for it.
## References
- `docs/research/phase-0.md` OQ-BL-04 (posture + surface addendum
promoted here and to ADR-003/store-api)
- `docs/research/poc-largeblob-findings.md` findings A2/A3
- `docs/research/iroh-blobs-eval.md` §Verification
- ADR-001 (ops surface; verified-fetch op), ADR-002 (digest + preamble
consequence), ADR-005 (range reads in packfile serving)
- [store-api.md](../store-api.md); [ops-surface.md](../ops-surface.md)
+136
View File
@@ -0,0 +1,136 @@
---
status: draft
last_updated: 2026-10-01
---
# Pooling, namespaces, and GC
## What this is
The layer that makes the store *pooled*: one content-addressed pool per
node, dedup across every consumer by construction, with garbage
collection that is safe against concurrent writes. The pool property is
the reason both demanding consumers exist (overview.md); GC is its
price, and this document specifies the mechanism the POCs validated
(ADR-005).
## The pooled CAS
- **One pool per node.** Repos, workspaces, and appfile stores are all
*sets of hash references* over the same CAS — never walled-off
per-repo stores. Git's `alternates`/object-pool mechanism (GitLab
object pools) is the same idea done awkwardly; the pool generalizes it
to unrelated repos and to p2p replication.
- **One address space.** Under the canonical hash (ADR-002), the same
file content has the same address whether it entered via a git oid or
a workspace manifest — cross-consumer dedup is free by construction,
not a feature.
- **Physically flat, logically namespaced** (the rudolfs inversion).
The byte layer is `hash → bytes`, backends namespace-blind; the
namespace is a *reference set* above the crate:
## Namespaces (logical, above the pool)
- A namespace is `namespace → set of root hashes` held by the consumer
(git refs *are* the names; an LFS pointer file in a tree *is* the
reference; an alkfs path tree *is* the manifest). The store never
walks consumer-side manifests — it stays structure-blind.
- The store-side counterpart is the **liveness seam**: consumers
register live-shared liveness sources (ADR-005, §Decision —
`register_liveness_source`; the POC's copy-semantics failure finding 7
is the pinning reason). Namespace tables may also exist purely
in-memory for ephemeral consumers; nothing in the store persists
namespaces — where roots live when a deployment runs fs-only is a
non-question by design: roots live with the consumer, whose references
are as durable as they need to be.
- ACL mapping for network ops: the namespace is the resource
(`resource_id_path` selects it) — see ops-surface.md.
## Mark-and-sweep GC (ADR-005)
Validated mechanism (POC #1 findings 3/7/8 — exact sweep counts, abort
semantics, delete-then-recover as byte-identical re-put):
- **Liveness sources (roots):**
1. consumer-registered liveness sources (live-shared, above),
2. put-path **pins** (RAII; batch scope available for multi-put
writes — manifest writes are exactly this),
3. the **protect callback** consulted before each sweep: it may add
externally-known hashes or **abort the run** (`GcAborted`; a
flaky protection source skips the sweep rather than risk
deletion — iroh's `ProtectOutcome::Abort` conclusion, adopted).
- **Sweep:** enumerate the whole pool (backends' complete `list()`,
ADR-003), compute the live set, batch-delete the dead
(batch-sized, ~100/batch, iroh's proven shape). Delete-then-recover
is a re-put — byte-identical under the same key (validated).
- **Traversal ownership: lean.** Liveness computation beyond "these
roots exist" is the consumer's job, handed over via the seam. The
store never learns manifest formats. If a real consumer's live-set
computation proves too heavy, store-native traversal is a *new* ADR
(documented constraint, not a parked hedge — the un-pause condition
is an external fact: a measured consumer cost).
## Delete windows (the sweep-vs-put race)
The race (POC #1 finding 3): a pin added between sweep-"list" and
sweep-"delete" is a lost pin — a blob deleted while in flight. The
architecture resolves it as a **delete window** with three phases —
mark (enumerate + live-set), per-key arbitration under the
pin/liveness lock *at delete time*, batched commit — resting on two
ordering invariants: **put commits pin before publish**, and **deletes
arbitrate under the pin lock at delete time** (full protocol + proof:
ADR-005 §Sweep; iroh `delete_set.rs`'s ProtectHandle/protect-cancel
shape and its serialized-actor pattern are the re-borrowed prior art —
their serialized actor is exactly "arbitrate at delete time"; the
window batches the arbitration).
Test-asserted property: **a visible pool entry is never deleted while
liveness (pin or registered source) protects it, and an in-flight put
is never deleted by a sweep started before it committed.** Direct
`delete(key)` refuses pinned keys (typed error; same arbitration).
`has()`/`get()` during a window observe either the old or the new
state; there is no torn observation (the pool's entries are immutable;
existence flips atomically per key).
## GC scheduling and scope (ADR-005)
- **Explicit `sweep()`; embedder owns cadence.** No background timers
in the crate. An embedder that wants interval sweeps adds them above
(one call). Consumer-driven sweeps (after a manifest prune, say) are
the same API.
- **Sweep scoping:** whole-pool is the base. Namespace-scoped sweeps
run by computing a narrower live set from that namespace's liveness
source; the pool remains shared (scoping is a liveness computation
choice, not a physical partition).
- **Accounting** (per-namespace sizes): computed from reference sets by
consumers; the store reports pool-level totals only (`list` +
`stat`). Physical-layer accounting would re-introduce namespace
physics — rejected with the per-namespace backend config
(ADR-003 §Consequences).
## Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Pooled CAS & GC | one pool, registered liveness, mark-and-sweep, delete windows, no ambient scheduling |
| [003](decisions/003-backend-contract-and-dispatch.md) | list() contract | complete lists are the GC substrate |
| [004](decisions/004-two-backends-no-third.md) | Structure-blindness | manifests/refs never sink into the store |
## Open Questions
None owned by this document. The p2p replicator's cross-node GC
coordination is a property of the replicator policy layer (above the
crate), governed by the same seam — it does not open a question here
until a consumer names a requirement (OQ-07/OQ-08 intake).
## References
- `docs/research/phase-0.md` OQ-BL-05 (settled decisions + residuals)
- `docs/research/poc-trait-dispatch-findings.md` findings 3/7/8
- `docs/research/iroh-blobs-eval.md` — gc.rs / delete_set.rs reads
(the re-borrowed conclusions)
- rudolfs — the physical-namespace anti-pattern (inverted)
- gix-odb — alternates/pool prior art for the p2p case
- ADR-005; [store-api.md](store-api.md) (Pin semantics, sweep API);
[ops-surface.md](ops-surface.md) (namespace-as-resource ACL)
+101
View File
@@ -0,0 +1,101 @@
---
status: draft
last_updated: 2026-10-01
---
# Hashing and keys
## What this is
The content-address layer: the canonical hash derivation every entry is
keyed under, the tolerated legacy algorithm, the typed `Key`/`Digest`
surface the store API speaks, and the physical byte encoding of keys
inside backends. This is the crate's one-way door: key byte layouts are
persisted in backend rows and filenames, survive digest-layer evolution,
and must precede any wire consumer (AGENTS convention 5; phase-0 OQ-BL-03
residual).
## The hash family
Three tiers (ADR-002 §Decision; resolved from OQ-BL-03 after the
"multi-hash" premise dissolved — nothing forces BLAKE3 once iroh-blobs'
inheritance is rejected):
1. **Canonical — `git-blob-sha-256`:** the git oid derivation
`H("blob <len>\0" + content)` with SHA-256, uncompressed content.
Chosen because alkgit *forces* git's derivation by protocol
definition, and every other consumer (workspaces, appfile) sees
addresses, not algorithms — so one algorithm domain makes the pool
one address space and cross-consumer dedup free by construction
(ADR-005).
2. **Tolerated legacy — `git-blob-sha-1`:** existing git repositories
are overwhelmingly SHA-1; the store accepts them under the same
preamble discipline. Git's own collision-hardened SHA-1 threat model
is inherited verbatim; the store adds nothing and weakens nothing.
3. **Not present — BLAKE3:** demoted out of the crate entirely. If a
chunk-tree transfer encoding is ever adopted, it is a new decision at
a new layer (ADR-006); its verified whole-contents would register as
ordinary `git-blob-sha-256` entries.
Security posture: one algorithm across consumers is safe here — SHA-256
has no practical cross-protocol ambiguity at these input shapes, and the
preamble domain-separates the derivation (ADR-002 §Consequences).
## Key encoding (the one-way door)
Physical key bytes inside backends (POC #1 finding 5, validated in the
interop test where an oid produced by *external git* addresses the same
pool entry the store puts):
- **Algorithm identity is part of the key itself** — a per-store
single-algorithm configuration cannot express git-sha-256 and
git-sha-1 coexisting in one flat pool, which the external-oid test
proves is required.
- **Fixed-length, enum-tagged:** `1 algorithm byte + digest bytes`
(32 for sha-256, 20 for sha-1). The POC's extra "kind" tag byte is
dropped — there is exactly one key kind. Algorithm bytes are assigned
from a single registry constant in ADR-002 so external producers can
compute them.
- **Sort stability:** fixed-length enum-tagged keys sort cleanly in a
total order — required for the sweep's live-set diffs (list() walks
are plain range scans; ADR-005).
- **Opaque to backends:** the `Backend` trait speaks opaque byte keys
(ADR-003); only the store core constructs and interprets typed keys.
## The Key/Digest surface (WHAT the API exposes)
- `Digest` — an algorithm + digest pair; construction via the canonical
derivation from content (known length) or from a pre-computed pair
(interop: `from_git_oid_hex` shape validated in POC #1 finding 4).
- `Key` — the typed wrapper; `as_bytes()` / `from_bytes()` are the only
conversions across the backend boundary. Malformed key bytes
(truncated, unknown algorithm) are rejects, never guesses — the POC's
length-checking test is the standard.
- Verification semantics ride the whole-blob posture (ADR-006): put and
get paths hash-check against the canonical derivation; there is no
per-range hash relationship under this derivation (preamble includes
the length), which is the constraint ADR-006 documents rather than
works around.
## Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| [002](decisions/002-canonical-hash-and-key-encoding.md) | Canonical hash + key encoding | three-tier hash policy; algorithm-in-key byte layout; the declared wire-format ADR |
## Open Questions
None owned by this document. The interop surface alkgit ultimately
needs (oid formats, hex/binary conventions at *its* boundary) is OQ-07
— alkgit's question, not a key-encoding question.
## References
- `docs/research/phase-0.md` OQ-BL-03 (resolution + key-encoding
mechanics)
- `docs/research/poc-trait-dispatch-findings.md` findings 4/5
(byte-exact git CLI validation; tagged-key coexistence)
- gix-odb (`/workspace/gitoxide/gix-odb`) — the consumer-side baseline
for object reads/writes
- ADR-002; ADR-003 (opaque byte keys at the backend boundary); ADR-006
(what verification is possible under this derivation)
+196
View File
@@ -0,0 +1,196 @@
---
status: draft
last_updated: 2026-10-01
---
# Open Questions
Centralized tracker. Theme sections hold OQs; cross-theme status lives
in the Deferred/Blocked index below.
**Deferral policy (the Schrödinger's-code rule).** A *decision this
crate needs before shipping* may not be deferred on a dependency that
is itself waiting for this crate to exist. alkgit is paused mid-Phase-1
*to build* this core; alkfs is pending *that shared base*. "Deciding
input: what alkgit wires first" never collapses into an input, so it is
not one. Two kinds of parked question remain legitimate:
- **deferred(scope)** — the deciding fact exists independently of this
crate (e.g., alkfs Phase 0 outcomes); the crate proceeds on its own
ADRs meanwhile.
- **externally-owned questions** — questions about *how a consumer maps
onto this crate*, which are not decisions this crate needs before
shipping and are not this document's to decide; they are carried here
for visibility, owned by the consumer's own process, and collapse
when that consumer acts (including on an API of this crate that will
exist by then — that is acceptable *only* because the question gates
no decision here; it would be Schrödinger's code if it gated one).
Invalid for either kind: a blocker that is a decision this crate must
make before shipping. Phase 0's residual list was swept under this
rule: every residual either resolved into an ADR (the evidence was
already in hand) or re-owned as below.
**Index of active OQs:** OQ-07, OQ-08 (externally-owned, carried for
visibility), OQ-09 (deferred(scope)). Promoted Phase 0 questions
OQ-BL-01..06 are recorded here with their resolutions for
traceability.
---
## Theme: consumer integration
### OQ-07: How alkgit's object-storage seam consumes alkblobs
- **Origin**: [overview.md](overview.md), [gc-and-namespaces.md](gc-and-namespaces.md)
- **Status**: externally-owned (alkgit) — carried here for visibility;
not a decision this crate must make before shipping
- **Priority**: high (it is why this crate exists) — but not a blocker:
alkblobs builds by its own ADRs; this question maps the *consumer on*.
- **Owner**: alkgit (external)
- **Question**: how do alkgit's reviewed `GitRefs`/`GitPackGen`/`GitPackIngest`
traits (its backend.md, ADR-018) sit over alkblobs — gix-odb-over-pool,
alkblobs-behind-those-traits, or a mixed composition (e.g. refs in the
pool, pack generation still gix)? Which tier serves the loose tier;
does packfile serving (ADR-005) get wanted at all?
- **Resolution**: owned by alkgit's architecture process; it is
answerable on paper against this document tree (backend.md + these
ADRs) whenever alkgit's process acts, and gates nothing here. This
crate's obligations toward it are already fixed: git-oid addressing
(ADR-002), liveness seam (ADR-005), serving surface (ADR-006/store-api).
- **Cross-references**: OQ-08, ADR-002, ADR-005, ADR-006
### OQ-08: alkfs requirement intake
- **Origin**: [overview.md](overview.md), [backends-and-dispatch.md](backends-and-dispatch.md)
- **Status**: externally-owned (alkfs) — carried here for visibility;
alkfs Phase 0 has not run
- **Priority**: medium (alkfs Phase 0 pending; nothing here blocks on it)
- **Owner**: alkfs (external)
- **Question**: what does alkfs's Phase 0 name as requirements on the
shared base — durability tiers, sync/replication seams the store does
not expose, manifest-layer shapes beyond ADR-004's boundary (e.g. the
appfile write-session shape as a consumer concern)?
- **Resolution**: owned by alkfs's Phase 0 (independently obtainable —
it needs no alkblobs artifact). The store's posture toward an unknown
consumer is deliberately permissive: open trait + addresses +
liveness seams (ADR-003/004/005) without baking an alkfs shape in.
- **Cross-references**: OQ-07, ADR-004
## Theme: ops surface
### OQ-09: Default visibility of namespaces in network ops
- **Origin**: [ops-surface.md](ops-surface.md)
- **Status**: deferred(scope)
- **Priority**: low (a policy default, one line to set; needs a real
deployment's posture to set it against)
- **Owner**: this crate — decided with the first embedded `ops`
deployment
- **Question**: are namespaces closed-by-default (an embedder must
grant read/write per namespace) or open-by-default (public read,
write gated)? The `AccessControl` machinery supports either posture;
the default shapes embedder onboarding and accidental-exposure risk.
- **Impacts**: one line in the ops module's ACL defaults at
implementation; ships after the decision, so it gates no
architecture output.
- **Resolution**: not yet decidable — the wrong default is a real risk
(an open-by-default store accidentally exposing content), and the
first *embedded deployment* of the ops surface is the deciding fact
(a deployment that can only exist once the ops module ships —
permitted under the rule above only because the decision gates
nothing before shipping; the module ships closed-by-default as the
documented interim posture).
- **Blocked on**: first embedded ops deployment (tracked in
`tasks/architecture/oq-09-ops-visibility-tracker.md`, the external-
trigger task per the SDD deferred-OQ convention).
- **Cross-references**: ADR-001, [ops-surface.md](ops-surface.md)
---
## Theme: promoted Phase 0 register (resolutions recorded)
### OQ-BL-01: Crate scope — store-only, or store + ops surface?
- **Origin**: `docs/research/phase-0.md`
- **Status**: resolved
- **Resolution**: store-only core; ops surface exists as a
feature-gated alkcall-backed module (placement *decided*, not hedged —
Schrödinger's-code rule applied: no waiting on paused consumers).
Full rationale: [ADR-001](decisions/001-substrate-posture-and-ops-placement.md)
(family-pattern conformance record; placement; no-store-layer-wire
invariant).
- **Cross-references**: ADR-001; OQ-09 (the one residual policy line)
### OQ-BL-02: Multi-backend dispatch
- **Origin**: `docs/research/phase-0.md`
- **Status**: resolved
- **Resolution**: lean `Backend` contract + two shipped backends +
pure-function size routing (pre-threshold buffering for unknown-length
puts); migration eliminated by construction; per-namespace backend
config rejected. Full rationale:
[ADR-003](decisions/003-backend-contract-and-dispatch.md).
- **Cross-references**: ADR-003, ADR-004
### OQ-BL-03: Hash abstraction
- **Origin**: `docs/research/phase-0.md`
- **Status**: resolved
- **Resolution**: canonical git-blob-sha-256; git-blob-sha-1 tolerated;
BLAKE3 excluded from the crate. Key bytes: algorithm-in-key,
fixed-length enum-tagged, tag byte dropped. The declared wire-format
ADR. Full rationale:
[ADR-002](decisions/002-canonical-hash-and-key-encoding.md).
- **Cross-references**: ADR-002, ADR-006
### OQ-BL-04: Verification and chunking
- **Origin**: `docs/research/phase-0.md`
- **Status**: resolved
- **Resolution**: whole-blob verification; range reads carry
out-of-band slice digests; chunk-tree/CDC encodings excluded by
scoping (not deferred — the transfer-encoding consumer cannot be
waited on; re-entry is a new decision at a new layer if ever named).
Full rationale:
[ADR-006](decisions/006-verification-posture-and-transfer-encoding.md).
- **Cross-references**: ADR-006, ADR-001
### OQ-BL-05: Pooling and GC
- **Origin**: `docs/research/phase-0.md`
- **Status**: resolved
- **Resolution**: one pooled CAS; flat bytes + logical namespaces;
store persists no root table; mark-and-sweep with three liveness
sources (registered sources / RAII + batch pins / protect callback
with abort); delete windows close the sweep-vs-put race; explicit
sweep, embedder-owned cadence; packfile serving rides large-blob
range reads if ever wanted. Full rationale:
[ADR-005](decisions/005-pooled-cas-and-gc-mechanism.md).
- **Cross-references**: ADR-005, ADR-003
### OQ-BL-06: POC register
- **Origin**: `docs/research/phase-0.md`
- **Status**: resolved (complete: #1, #3 passed; #2 absorbed; #4
covered in miniature, concurrency half specified as architecture in
ADR-005)
- **Resolution**: register complete; Phase 0 ended; evidence trail in
`docs/research/`.
- **Cross-references**: all ADRs above
---
## Parked index (deciding fact / owner — externally-owned rows are not blockers)
| OQ | Status | Deciding fact / owner |
|---|---|---|
| OQ-09 | deferred(scope) | first embedded ops deployment; tracker task `tasks/architecture/oq-09-ops-visibility-tracker.md` |
| OQ-07 | externally-owned | alkgit's architecture process (answerable on paper anytime; gates nothing here) |
| OQ-08 | externally-owned | alkfs Phase 0 intake |
Per the SDD deferred-OQ convention, OQ-09 carries a tracker task under
`tasks/architecture/` (`[external-trigger, deferred-oq]`); OQ-07/OQ-08
are owned by other repos' processes and are not alkblobs tracker tasks
(their outcome arrives *through* their owners, not through any artifact
this repo creates).
+167
View File
@@ -0,0 +1,167 @@
---
status: draft
last_updated: 2026-10-01
---
# Ops surface (alkcall-backed network operations)
## What this is
The network-facing operation family over the store: have/need
announcements, verified fetch, and streaming put — riding the alkcall
substrate behind a **feature-gated `ops` module in this crate**
(feature `ops`, default-off; ADR-001). The base crate stays lean and
alkcall-free; enabling `ops` adds the alkcall dependency and registers
the op family.
## Why alkcall (ADR-001, the short form)
The op family is exactly alkcall's mixed-shape sweet spot (validated
machinery, not parallel invention):
- **JSON control plane** — have/need announcements, `stat`-shaped
probes, offer/carry lengths: alkcall `OperationSpec` + JSON Schema
validation (alkcall ADR-016) + typed error schemas + `from_call`
discovery + External/Internal visibility.
- **Binary data plane** — bulk bytes one way (fetch: `Sub`,
alkcall ADR-021) or the other (put: `Pub`/`Sink`, alkcall ADR-046),
allocated via the `channel_open` marker (alkcall ADR-047);
established pump patterns (`pump_bidi`, alkcall ADR-050).
- **ACL is solved there** — `AccessControl` + ownership checks
(alkcall ADR-011) under the ADR-017 privilege model map onto
namespace gating directly; an in-band invented scheme would be a
second, unreviewed authorization story in the family (AGENTS
convention 6).
Producer/consumer vocabulary throughout; no transport enters either
layer — the caller (an embedder, alkgit's replicator, a future alkfs
sync) dials and hands an established `Connection` in (alkcall is a
pure protocol crate; ADR-012 there).
## Op family (WHAT is exposed)
**Namespaces on the wire.** Every ops payload carries an explicit
`namespace` field — the consumer-scoped identifier the embedder's
registry recognizes (syntax: non-empty UTF-8 string, embedder-validated;
the *minting* of namespaces is the embedder's act — e.g. alkgit's
registry maps repo ids to namespaces, alkfs maps workspace roots). The
store core never sees this string (backends are namespace-blind,
ADR-005); it selects the liveness source and the ACL resource. The ops
module's job is to bind `namespace → the embedder-registered liveness
source` for pin conversion, and to present the namespace as the
alkcall ACL resource.
JSON control ops (`Visibility` per ADR-001 §Decision):
- **`blobs/stat`** `{namespace, digest}` → `{len}` — probe before
offering; ACL-gated like fetch (read action on the namespace)
- **`blobs/have`** `{namespace, digests[]}` → `{present[]}` — the
have half of have/need set diffing (hashes only — content never
traverses this op); read-gated: an existence probe over arbitrary
digests is a discovery surface, so it is gated exactly like fetch,
not public
- **`blobs/delete`** `{namespace, digests[]}` — operator machinery;
`Visibility::Internal`, evaluated under the internal authority
context per alkcall ADR-017 (internal calls switch authority context,
never skip ACL): requires the embedder's operator/admin authority
context, not a namespace grant. The `namespace` field exists for
audit/triage scoping (which slice of the pool the deletion targets),
not for gating.
Binary channel ops (registered via the `channel_open` marker):
- **`blobs/fetch`** (`Sub` — the server→client streaming op shape)
— `{namespace, digest, ranges?}` in;
verified bytes out: the consumer hash-checks each received chunk/whole against
the carried digest (ADR-006). Broadcast fanout above the store:
one reader, store arm + subscriber arms (POC #3 finding A3); late
joiners degrade to a normal post-commit `get` — identical bytes under
CAS. Slow-subscriber policy (drop-and-late-join) is ops-layer
policy, not a store concern.
- **`blobs/put`** (`Sink` — the client→server streaming op shape)
— `{namespace, digest?, len?}` offer +
byte stream in; server verifies against the canonical derivation
before commit. A known-length offer is the encouraged path (one-
pass); unknown-length rides the store's pre-threshold buffering path
(ADR-003). **Need half of have/need**: a fetch miss *is* the need
announcement — the consumer computes its need set by diffing
`blobs/have` results and fetches the absent digests; no separate
need op exists (the diff is caller-side; the wire carries only
concrete fetch requests).
**Pin hand-over for remote puts (the cross-wire contract).** The
server-side put handler owns the `Pin` guard; the offering side's
reference registration happens in the *remote* process, invisible to
the server. The hand-over contract:
1. the put lands pinned (ADR-005) — the entry cannot be swept while
the handler holds the pin;
2. the `blobs/put` response returns a **pin token** (opaque handle);
the putter can later confirm its registration landed by observing
the digest present via `blobs/have`;
3. the embedder's registered liveness source is the *actual* root of
record: the contract on the embedder (documented on the ops module)
is that it registers the namespace's roots such that a digest
referenced by a namespace's manifest is either (a) already
registered before the caller's next sweep could run on the serving
node, or (b) protected by holding the pin token until its
registration is confirmed (`blobs/have` returns present after
registration). In short: **the remote putter may not rely on an
unconfirmed put surviving the next sweep — it holds its pin token
until its registration shows up in `blobs/have`.** A `Pin`
conversion on the server side (token → registered liveness source)
is the mechanism the embedder plugs its registry into.
Placement of the op *registration* (which embedders wire where they
want them exposed) matches the alkgit ops pattern (its
`git/repo/*` call ops): the module exports spec/handler pairs; the
embedder assembles.
## ACL mapping
Resources and actions, evaluated by alkcall's `AccessControl` machinery
at op entry — this module never re-implements an authorization check
(convention 6; ADR-001):
- **Namespace = resource.** `resource_type: "blob-namespace"`,
`resource_id_path` selecting the payload's `namespace` field;
actions map onto read/put/manage:
- `blobs/fetch`, `blobs/stat`, `blobs/have` — the **read** action;
unlisted namespaces (no ACL entries) deny by default (closed
posture until OQ-09 decides a different default)
- `blobs/put` — the **write** action
- `blobs/delete` — not namespace-gated at all; internal authority
context only (alkcall ADR-017: internal switches context, never
skips ACL)
## Store-facing relationship
The ops module is a *consumer* of the store core (store-api.md), not a
side channel: fetch handlers call `get`/`read_range`/`stat`; put
handlers call `put`; both honor pins and delete windows. There is no
ops-internal path around the store facade (ADR-001 §Consequences —
this is what makes "re-homing the ops surface later loses nothing"
true).
## Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| [001](decisions/001-substrate-posture-and-ops-placement.md) | Substrate posture & placement | alkcall-backed, feature-gated ops module; conformance-not-deviation rationale |
| [006](decisions/006-verification-posture-and-transfer-encoding.md) | Verification | verified-fetch via carried digests; no chunk trees |
## Open Questions
- **OQ-09**: default visibility of namespaces in network ops
(open-by-default vs closed-by-default) — deferred(scope), decided at
the first embedded ops deployment; ships closed-by-default meanwhile.
## References
- `docs/research/phase-0.md` OQ-BL-01 (the substrate settlement + its
justification, promoted verbatim into ADR-001 §Context)
- alkcall docs/architecture — ADR-011/012/016/017/021/046/047/050 (the
machinery this module rides)
- `docs/research/poc-largeblob-findings.md` finding A3 (fanout seam)
- ADR-001; [store-api.md](store-api.md); [gc-and-namespaces.md](gc-and-namespaces.md)
(namespace-as-resource)
+131
View File
@@ -0,0 +1,131 @@
---
status: draft
last_updated: 2026-10-01
---
# Overview
## What this crate is
`alkblobs` is content-addressed blob storage in the alk* family: a
put/get/verify store keyed under one canonical content hash (the
git-family derivation; legacy SHA-1-tolerant), with pluggable backends
(small blobs in a kv store, large blobs on a filesystem fallback), split
from any wire/ops-protocol layer. Network ops ride the alkcall substrate
behind a feature-gated `ops` module (ADR-001).
## Why it exists — the extracted shared core
Two downstream projects were converging on the same storage stack, so it
is being built once, here:
- **alkgit** was paused mid-Phase-1 (its architecture is reviewed,
`docs/architecture/` in the alkgit repo) to stop duplicated
storage-layer work. Its object backend needs git's own oid derivation,
a better small-object backend than its proposed default file layout,
and a git-lfs-shaped large-blob story. Under the canonical-hash
resolution (ADR-002) git objects are first-class pool entries with
zero indirection.
- **alkfs** (Phase 0 pending) needs the appfile/workspace shape: small
content in a fast kv store, large content on the filesystem, path→hash
mapping above. The path-tree/manifest layer is alkgit's and alkfs's,
not this crate's (ADR-004).
Both demanding consumers (p2p git replicators, agent workspaces) reduce
to the same shape: **a pooled CAS per node**, where repos and workspaces
are *sets of hash references*, not walled-off object stores — cross-repo
dedup by construction (ADR-005). That pooling property is why the
storage layer was pulled out of both consumers rather than built twice
or three times.
## Layer map
```
┌────────────────────────────────────────────────────────────┐
│ consumers (above this crate) │
│ alkgit object storage · alkfs path trees/workspaces │
│ refs, manifests, path→hash maps, git semantics │
└────────────┬───────────────────────────────────────────────┘
▼
┌────────────────────────────────────────────────────────────┐
│ ops surface (feature "ops"; ADR-001) │
│ have/need + verified fetch + put; JSON control ops, │
│ binary data channels; ACL via alkcall AccessControl │
└────────────┬───────────────────────────────────────────────┘
▼
┌────────────────────────────────────────────────────────────┐
│ store core (always on) │
│ hashing & keys (ADR-002) · put/get/stat/range/pin seams │
│ mark-and-sweep GC + liveness registration (ADR-005) │
│ size-threshold dispatch (ADR-003) │
└───────┬───────────────────┬───────────────────────────────┘
▼ ▼
kv backend fs backend (ADR-003: exactly two)
transport: none, anywhere. alkgit/alkfs or the embedder dials;
an alkcall Connection is handed in at the ops boundary.
```
## Dependency posture
| Layer | Dependencies | Notes |
|---|---|---|
| base crate (store core + backends) | `tokio`, `thiserror`, `parking_lot`, plus `rusqlite` behind the `kv` feature and nothing behind `fs`/`mem` | no alkcall, no serialization frameworks |
| `ops` feature | adds `alkcall` (feature-gated, default-off) | JSON op specs + binary channels + `AccessControl`; ADR-001 |
| consumer crates | depend on alkblobs (and on alkcall directly when they speak ops) | alkgit, alkfs |
Backends are feature-gated; both `kv` and `fs` ship default-on
(batteries-included, the alkgit feature-model pattern); the `ops`
feature is default-off. Wasm: the backends are not a wasm story; a wasm
client is a consumer of the `ops` surface (which rides alkcall, itself
wasm-clean), never an in-process embedder (ADR-001 §Consequences).
## Non-goals (each enforced by an ADR)
- No wire framing, op dispatch, ACL, or transport inside the store layer
(ADR-001).
- No adopted iroh-blobs wire surface — no tickets, no postcard, no
provider protocol (AGENTS convention 7; ADR-001 §Context).
- No namespaces-as-physical-storage; the pool is flat, namespaces are
consumer-level reference sets (ADR-005).
- No manifest/path-tree/vfs tables; consumers never implement a Backend
to model paths (ADR-004).
- No chunk-tree/CDC storage encoding; whole-file CAS is the store's only
granularity (ADR-006).
- No per-repo or per-namespace walled stores — one pooled CAS
(ADR-005).
## Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| [001](decisions/001-substrate-posture-and-ops-placement.md) | Substrate posture & ops placement | store stays substrate-free; ops ride alkcall feature-gated |
| [002](decisions/002-canonical-hash-and-key-encoding.md) | Canonical hash + key encoding | git-blob-sha-256 canonical; algorithm in-key; one-way door |
| [003](decisions/003-backend-contract-and-dispatch.md) | Backend contract & dispatch | two backends, size-threshold routing, no migration |
| [004](decisions/004-two-backends-no-third.md) | Two backends, no third | manifest layers are consumer-side |
| [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Pooled CAS & GC | flat pool, registered liveness, delete windows |
| [006](decisions/006-verification-posture-and-transfer-encoding.md) | Verification posture | whole-blob verification; chunk trees excluded by scoping |
## Open Questions
Tracked in [open-questions.md](open-questions.md). Ones affecting the
crate level:
- **OQ-07**: how alkgit's object-storage seam consumes alkblobs
(externally-owned: alkgit's architecture process)
- **OQ-08**: alkfs requirement intake (externally-owned: alkfs Phase 0)
- **OQ-09**: default namespace visibility in network ops
(deferred(scope) — decided at the first embedded `ops` deployment;
ships closed-by-default meanwhile)
## References
- `docs/research/phase-0.md` — the converged Phase 0 record; inputs to
every ADR above
- `docs/research/poc-trait-dispatch-findings.md`,
`docs/research/poc-largeblob-findings.md`,
`docs/research/iroh-blobs-eval.md` — the evidence trail
- `docs/sdd_process.md` — the process governing this tree
- `/workspace/@alkdev/alkgit/docs/architecture/backend.md` — the
reviewed consumer-side seam that OQ-07 maps onto alkblobs
- AGENTS.md — family conventions (conventions 4–8 gate the decisions)
+156
View File
@@ -0,0 +1,156 @@
---
status: draft
last_updated: 2026-10-01
---
# Store API
## What this is
The store core's public surface: the typed facade above the backends
that consumers (alkgit, alkfs, the ops module) program against. It owns
typed keys and hashing (ADR-002), size-threshold dispatch (ADR-003),
the put-pinning seam, and the liveness/sweep seam (ADR-005) — and
nothing else: no paths, no manifests, no wire (ADR-004, ADR-001).
## Public surface
### Write path
- **`put(len: Option<u64>, stream)` → `(Key, Pin)`** — the streaming
put seam (POC #3 finding A1). Two internal paths behind one signature:
- *Known length* (encouraged, documented as such on the API): the
preamble hashes before content flows; one pass, no staging needed
for hashing. Covers git objects, pre-staged files (`stat`-derived
length), network offers, `Content-Length`-bearing uploads.
- *Unknown length*: content buffers in memory up to the dispatch
threshold; on commit the buffer is the whole stored artifact
(kv tier) — or, on mid-stream threshold overflow, the buffer
flushes into a stage file and the stream continues into it
(fs tier — the streaming pass cannot include the preamble, so the
derivation pass restarts over preamble + staged bytes; forced by
the preamble being inside the hashed input, not a design choice).
Verification runs over the eventual stored artifact as a whole;
a stream error on either path leaves no entry and no stage file.
Eliminates the A6 dispatch asymmetry by construction: routing is
a pure function of content length (ADR-003).
- **Stage-file hygiene invariant** (POC #3 finding A5, codified): every
failure path converges on stage-discard; every success path converges
on commit-rename. No early returns that bypass cleanup; this is an
asserted invariant with dedicated exact-count tests, not review
hygiene.
- **`Pin` / `Batch`** — RAII guard on the put path (ADR-005): the
entry is liveness-protected until the guard drops or the caller
converts it into a consumer-side reference (registered liveness, a
root tag). **`Batch`** groups multi-put writes under one batch-scoped
pin (manifest writes are exactly this); a batch's pins drop together
on batch commit or drop (ADR-005 §Decision).
### Read path
- **`get(key) → Option<(len, stream)>`** — small tier returns a
bounded read; large tier streams a file handle. Dispatch
fall-through: a small-tier miss queries the large tier (POC #1
finding 6).
- **`stat(key) → Option<EntryMeta>`** — cheap length/type probe
(phase-0 OQ-BL-04 addendum; gix's header-only read is git's cheapest
primitive and the packfile-serving surface rides this). `EntryMeta`
carries `len: u64` and the key's algorithm (the "type" from the
store's perspective is the hash algorithm — there is no other type at
this layer).
- **`read_range(key, range) → (slice, slice_digest)`** — slice plus an
out-of-band slice digest: SHA-256 over the slice bytes, fixed by
convention so callers and ops handlers agree without negotiation
(per-range verification against the canonical digest is impossible —
ADR-006). Range reads fall through tiers exactly as `get` does
(small-tier miss slices the large tier's value). Local range serving
(packfiles) is the consumer of this.
- **`get_stream` / fanout** — the ops module's fetch handler is a
broadcast above the store: one reader, store arm + subscriber
arms (POC #3 finding A3); late joiners degrade to normal verified
`get` after commit. The store exposes the seam; the fanout policy
lives in the ops layer (ADR-001).
### Lifecycle
- **`has(key)`, `delete(key)`** — deletion participates in ADR-005's
delete windows; direct deletes are permitted but refuse protected
keys — pinned or re-observed live via registered sources (the same
per-key arbitration sweeps apply, typed error) — and are unusual by
posture; most deletion flows through sweeps.
- **`list() → stream of keys`** — whole-pool enumeration; complete by
contract (see backends doc for why list correctness is load-bearing).
- **`register_liveness_source(...)`** — the live-shared GC seam
(ADR-005, §Decision; the POC's "install implies copy" failure
(finding 7) is why the verb/name is pinned here).
- **`sweep() → SweepReport`** — explicit mark-and-sweep; embedders own
the cadence (ADR-005, no ambient timers). A sweep with no registered
liveness sources aborts without deleting — the safe default.
## Error model
`thiserror`; no panics in library code; no `unwrap`/`expect` outside
tests (AGENTS convention 2). Distinguished failure families:
- `Missing` — key absent (get/stat miss)
- `Verification` — put/get hash-check failure (content ≠ key)
- `Io(String)` — backend media failure (stringly because
`std::io::Error` is not stable across versions)
- `GcAborted` — a protection source failed; nothing deleted (ADR-005)
- `KeyInvalid` — malformed key bytes at the boundary (hashing doc)
Virgin-store semantics: read paths on a fresh store see absent/empty,
never "table does not exist" errors (POC #1 finding 2 — the redb
lesson generalizes to any kv engine).
## Concurrency posture
- Async I/O throughout; `tokio::sync` for lifecycle correlation;
`parking_lot` for short-held internal locks (AGENTS convention 3);
poisoned locks degrade via `unwrap_or_else(|e| e.into_inner())`.
- Blocking file work lives in `spawn_blocking` inside backend impls
(the alkgit trait-execution pattern); the store never blocks the
executor.
- The sweep-vs-put window is an architectural mechanism — delete
windows — specified in ADR-005, not an implementation note.
## Invariants (the test gate)
1. Whole-blob put/get round-trips byte-identical under the canonical
derivation, both tiers, both put paths (interop with real git
remains the source of truth, per POC #1's lesson about hardcoded
vectors).
2. Dedup: putting identical content twice (same or different path) is
one pool entry.
3. Stage hygiene: any failure mid-put leaves zero stage files; any
success leaves exactly one committed entry.
4. Sweep safety: with correct liveness registered, sweep counts are
exact; with aborting sources, sweep deletes nothing.
5. `list()` correctness is observable only through GC — list-related
tests assert through sweep outcomes (POC #1 finding 2's lesson
codified).
## Design Decisions
| ADR | Decision | Summary |
|---|---|---|
| [002](decisions/002-canonical-hash-and-key-encoding.md) | Hashing & keys | derivation + encoding this surface is built on |
| [003](decisions/003-backend-contract-and-dispatch.md) | Dispatch | routing is a pure function of length; pre-threshold buffering |
| [005](decisions/005-pooled-cas-and-gc-mechanism.md) | Pools & GC | pins, liveness seams, delete windows, sweep semantics |
| [006](decisions/006-verification-posture-and-transfer-encoding.md) | Verification | whole-blob checks; slice digests out-of-band |
## Open Questions
None owned by this document beyond the cross-references above.
## References
- `docs/research/poc-largeblob-findings.md` findings A1/A5/A6,
benchmark table
- `docs/research/poc-trait-dispatch-findings.md` findings 1–3, 7
- ADR-003 (put-path buffering), ADR-005 (pin/Batch/Pin lifecycle),
ADR-006 (range-read semantics), ADR-002 (keys)
- [ops-surface.md](ops-surface.md) — the fanout/fetch consumer of the
read path
- [gc-and-namespaces.md](gc-and-namespaces.md) — the lifecycle seam
details
@@ -0,0 +1,40 @@
---
id: oq-09-ops-visibility-tracker
name: "[external-trigger] first embedded ops deployment (OQ-09 unblock)"
status: pending
depends_on: []
scope: single
risk: trivial
impact: component
level: research
tags: [external-trigger, deferred-oq]
---
## Description
Tracker task (not actionable work): watches for the arrival of the
first embedded deployment of the `ops` surface. Its arrival unblocks
[OQ-09](../../docs/architecture/open-questions.md) (default namespace
visibility in network ops), at which point OQ-09 transitions
`deferred(scope)` → `open` and the deciding work is scheduled.
## Acceptance Criteria
- [ ] N/A — this task completes only when the external condition
arrives (first embedded ops deployment exists); it carries no
work itself.
## References
- docs/architecture/open-questions.md (OQ-09 — the `Blocked on` field
references this task id)
- docs/architecture/ops-surface.md
## Notes
> Per SDD deferred-OQ convention: `risk: trivial`, `level: research`,
> no implementation content.
## Summary
> Awaiting the external condition.