Phase 0 third pass: fork fan-out workload, dedup economics, cache ownership
- New OQ-FS-17: the multi-agent fork fan-out (many sessions forking the same workspace on the shared dev server — directories, docker containers, or remote mounts) as the workload that makes content-addressed dedup load-bearing, with the honest economics: file-level granularity already captures the fork win (unchanged files dedup identically); chunk-level pays only on large in-place-edited files and costs per-chunk metadata - Inline threshold recorded as backend-dependent (SQLite's ~100KB observation; postgres/KV differ) — a configuration knob over a storage profile, not a constant - Two named objectives recorded in the working direction: efficient many-repo git storage (OQ-FS-12) + SFTP-capable VFS for remote mounting (OQ-FS-16), with the container fan-out variant noted - rudolfs precision: verified it genuinely has an S3 backend (rusoto_s3); the bucket/namespace key, disk-cache-permanent split, and swappable permanent backend are its actual patterns; cache-ownership question (store-layer decorator vs door-layer cache) parked in OQ-FS-17 - POC #4 feeds OQ-FS-17; survey list gains the dedup-economics item - AGENTS.md: new convention 13 (caches key on content hashes, never on paths), conventions 14-18 renumbered with cross-references fixed, rudolfs added to the third-party reference list, fork fan-out added to the quick reference Verification: docs-only repo — OQ-FS-01..17 contiguous, every referenced workspace path exists, convention cross-references checked after renumbering
This commit is contained in:
1 parent
23a39cc9c0
commit
2f5f0ad8ed
2 files changed
+140
-26
No files matched your search
@@ -187,9 +187,19 @@ implementation agents.
|
||||
channels through `AccessControl` for free via
|
||||
`ChannelCore::register_openable`, and treat per-path/per-bucket
|
||||
policy as a Phase 1 design question (OQ). Multi-tenancy isolation
|
||||
(the POC's `bucket` concept) is a where-clause, not an afterthought.
|
||||
(the POC's `bucket` concept) is a where-clause, not an afterthought;
|
||||
bucket semantics across the alkblobs/alkfs seam are part of the
|
||||
seam's spec, not an afterthought.
|
||||
|
||||
13. **Backend/storage trait shapes are one-way doors** — the seam the
|
||||
13. **Caches key on content hashes, never on paths** — content under a
|
||||
hash is immutable, so content-hash-keyed caches never invalidate on
|
||||
bytes; only the path/branch-head → hash *mapping* goes stale, and
|
||||
that is cheap to refresh. Any cache (store-layer decorator per
|
||||
rudolfs, mount-door cache, session prefetch) follows this rule.
|
||||
Path-keyed caches reintroduce coherence problems content addressing
|
||||
exists to remove.
|
||||
|
||||
14. **Backend/storage trait shapes are one-way doors** — the seam the
|
||||
protocol crate exposes to storage implementers (the alkgit
|
||||
`GitRefs`/`GitPackGen`/`GitPackIngest` trait-family precedent,
|
||||
`docs/architecture/backend.md` there) is a contract once consumers
|
||||
@@ -199,17 +209,17 @@ implementation agents.
|
||||
the sealed/`pub(crate)` surface is exactly what turns "external
|
||||
store" into "fork".
|
||||
|
||||
14. **Feature flags** — substrate backends and heavy dependencies are
|
||||
15. **Feature flags** — substrate backends and heavy dependencies are
|
||||
feature-gated if the need arises. The base crate should compile lean
|
||||
(no SQLite, no gix, no kernel-FS access unless the feature is on).
|
||||
Verify both `cargo test` (default) and `cargo test --all-features`
|
||||
pass if features are added.
|
||||
|
||||
15. **Naming** — Rust standard: `snake_case` for functions/variables/
|
||||
16. **Naming** — Rust standard: `snake_case` for functions/variables/
|
||||
modules, `PascalCase` for types/traits, `SCREAMING_SNAKE_CASE` for
|
||||
constants.
|
||||
|
||||
16. **Module structure** — one module per file under `src/`, re-exported
|
||||
17. **Module structure** — one module per file under `src/`, re-exported
|
||||
from `src/lib.rs`. Public API surface is `lib.rs` re-exports. The
|
||||
expected shape (pending Phase 0/1 pinning): storage engine modules
|
||||
(path tree, blob store, write sessions, GC), backend traits +
|
||||
@@ -218,17 +228,19 @@ implementation agents.
|
||||
modules are feature-gated and never imported from the
|
||||
protocol/adapter/client modules.
|
||||
|
||||
17. **Upstream posture** — alkcall, alktunnels, and the alk* crates are
|
||||
18. **Upstream posture** — alkcall, alktunnels, and the alk* crates are
|
||||
ours to shape: file asks early and land them there rather than
|
||||
working around them locally (the alktunnels E-01/E-02 precedent —
|
||||
filed from Phase 0, landed within a day). Third-party crates
|
||||
(iroh-blobs, sqlite/rusqlite, russh-sftp, gix, automerge, honker)
|
||||
filed from Phase 0, landed within a day). Third-party crates
|
||||
(iroh-blobs, sqlite/rusqlite, russh-sftp, gix, automerge, honker,
|
||||
rudolfs)
|
||||
are NOT ours and never will be: wrap, extract, or fork deliberately
|
||||
per the alksocks precedent (OQ-SK-04 → ADR-013: extraction as owned
|
||||
code with provenance notices, differential tests against the
|
||||
reference checkout, no silent absorption) — and only when the carried
|
||||
changes pay for themselves. `/workspace/iroh-blobs`,
|
||||
`/workspace/russh-sftp`, etc. are read-only reference checkouts.
|
||||
`/workspace/russh-sftp`, `/workspace/rudolfs`, etc. are read-only
|
||||
reference checkouts.
|
||||
|
||||
## Verification Commands
|
||||
|
||||
@@ -324,7 +336,10 @@ Recorded so sessions don't relitigate settled leanings — see
|
||||
blob store beneath the path-tree mapping; repo shape (one repo, two
|
||||
crates vs two repos) is an open sub-question.
|
||||
- **Anchor workload: remote mounting** — sshfs/vanilla-SFTP replacement;
|
||||
per-syscall round trips are the problem to beat.
|
||||
per-syscall round trips are the problem to beat. Second named
|
||||
workload: the multi-agent fork fan-out (many sessions forking the
|
||||
same workspace — containers or remote mounts), which is what makes
|
||||
content-addressed dedup load-bearing (OQ-FS-17).
|
||||
- These are **leanings, not decisions**: every one is an open OQ
|
||||
(OQ-FS-01..16) until the Phase 0 convergence; don't mark anything
|
||||
(OQ-FS-01..17) until the Phase 0 convergence; don't mark anything
|
||||
"settled" in code or docs until the checklist starts ticking.
|
||||
+114
-15
@@ -57,18 +57,25 @@ converge, not decisions):**
|
||||
only identical files; git itself dedups at object granularity (packs
|
||||
delta-compress within a pack, but the address space is per-object);
|
||||
chunking is the storage-layer generalization (OQ-FS-04 — measure
|
||||
whether the workloads pay before committing).
|
||||
whether the workloads pay before committing). The workload that makes
|
||||
dedup load-bearing is the **multi-agent fork fan-out** (OQ-FS-17):
|
||||
many agent sessions forking the same workspace, each ultimately
|
||||
making small changes — at file granularity that alone already wins
|
||||
(unchanged files dedup identically); chunking adds wins only on
|
||||
large files edited in place.
|
||||
- **The two-crate split: `alkblobs` + `alkfs`.** A bucket-style
|
||||
content-addressed blob store (the S3 shape) beneath the path-tree
|
||||
mapping (the alknet-filesystem POC's layer 2). alkgit can consume the
|
||||
blob store without a filesystem; the FS layer keeps its own crate
|
||||
(OQ-FS-01).
|
||||
- **The anchor workload is remote mounting.** The concrete use case:
|
||||
the dev-server workspace mounted on a local Linux machine under
|
||||
`remote/` — today over vanilla SFTP (sshfs), which is slow in known
|
||||
ways (a round trip per syscall, stat storms, full inode work per
|
||||
stat). An appfile-indexed VFS fixes the server side; caching and
|
||||
pipelining close the rest (OQ-FS-16).
|
||||
content-addressed blob store (the S3/rudolfs shape — and rudolfs
|
||||
really does have an S3 backend, see §Prior art) beneath the
|
||||
path-tree mapping (the alknet-filesystem POC's layer 2). alkgit can
|
||||
consume the blob store without a filesystem; the FS layer keeps its
|
||||
own crate (OQ-FS-01).
|
||||
- **Two named objectives, two anchor workloads.** (1) Store potentially
|
||||
many git repos efficiently (alkgit's un-pause condition — OQ-FS-12).
|
||||
(2) A virtual filesystem supporting an SFTP door for remote mounting
|
||||
(OQ-FS-16) — including the container fan-out: the same workspace
|
||||
mounted into potentially many docker containers on the same host
|
||||
(many consumers, one producer). The dedup economics (OQ-FS-17) is
|
||||
what makes the aggregate sane when every session is a fork.
|
||||
|
||||
**Why this crate exists.** The alknet mono-repo's filesystem research
|
||||
(three POC iterations, 24 passing tests) proved a viable three-layer
|
||||
@@ -209,7 +216,7 @@ does not start from zero. It inherits:
|
||||
~120-160 lines of bao-tree glue must be re-implemented per store. The
|
||||
probe's cautionary half: `TempTags`/`TempTagScope` are sealed
|
||||
`pub(crate)` — sealed surface is exactly what turns "external store"
|
||||
into "fork" (AGENTS.md convention 13's cautionary example).
|
||||
into "fork" (AGENTS.md convention 14's cautionary example).
|
||||
- **The consumer shapes this crate must serve.** alkgit's backend trait
|
||||
family (`docs/architecture/backend.md` there: `GitRefs`, `GitPackGen`,
|
||||
`GitPackIngest` — streaming pack generation/ingestion, CAS ref
|
||||
@@ -305,7 +312,7 @@ consumable).
|
||||
(OQ-FS-02/03).
|
||||
- **Sealed internals** (`TempTags`, the `BaoTreeSender`) — the probe's
|
||||
fork-vs-PR calculus. Third-party upstream, not ours, never will be
|
||||
(AGENTS.md convention 17).
|
||||
(AGENTS.md convention 18).
|
||||
|
||||
### git + git-lfs — the shape to optimize for
|
||||
|
||||
@@ -373,13 +380,14 @@ https://sqlite.org/appfileformat.html — the insight that anchored the
|
||||
POC. honker (`/workspace/honker`, `honker-core` 0.2.4) supplies
|
||||
notify/locks/queues as SQL functions on *your* connection — the
|
||||
transactional-outbox property. Caveats: honker is third-party (not ours,
|
||||
convention 17), single-machine by explicit design ("two servers writing
|
||||
convention 18), single-machine by explicit design ("two servers writing
|
||||
the same .db over NFS is not a Honker deployment strategy"), and its
|
||||
scope in alkfs is an OQ (notify-on-commit may be the only load-bearing
|
||||
piece — locks/queues/scheduler may be excess).
|
||||
|
||||
### rudolfs — the git-lfs server reference
|
||||
|
||||
`/workspace/rudolfs` (local checkout) +
|
||||
`/workspace/@alkdev/alknet/docs/research/references/gitlfs/
|
||||
rudolfs-reference.md`: `StorageKey = (Namespace, Oid)` tenant isolation;
|
||||
the decorator composition `Verify ↔ Encrypted ↔ Cached ↔ Retrying(Disk →
|
||||
@@ -388,6 +396,18 @@ client and cache simultaneously. The cache/permanent split and the
|
||||
fanout pattern are the production-shape references for alkfs's blob
|
||||
layer and its serving read path.
|
||||
|
||||
**Two precise notes (2026-09-23):** (1) rudolfs *genuinely has an S3
|
||||
backend* (`rusoto_s3` in its Cargo.toml; local disk cache in front of
|
||||
permanent S3 storage) — so "bucket-style storage like S3" in the
|
||||
alkblobs direction is grounded in rudolfs' actual shape, not a loose
|
||||
analogy: the bucket/namespace key, the disk-cache-permanent split, and
|
||||
a swappable permanent backend are all rudolfs patterns. (2) The
|
||||
decorator chain is the compositional answer to "local cache in front of
|
||||
remote storage" — the same shape OQ-FS-16's mount caching wants, at the
|
||||
*store* layer rather than the door layer. Which layer owns the cache
|
||||
(door-local content cache vs store-layer decorator) is an OQ-FS-17
|
||||
sub-question.
|
||||
|
||||
### Anti-prior-art (what NOT to carry over)
|
||||
|
||||
- **The alknet mono-repo's scope.** The filesystem research is inherited
|
||||
@@ -793,6 +813,77 @@ known slowness:
|
||||
This OQ is the *consumer pull* behind the whole crate — if the mount
|
||||
use case stops mattering, OQ-FS-08's urgency drops with it.
|
||||
|
||||
#### OQ-FS-17: The multi-agent fork fan-out — dedup economics and cache ownership
|
||||
|
||||
**Added 2026-09-23.** The second named workload, and the one that makes
|
||||
dedup *load-bearing* rather than a nicety. The concrete deployment: one
|
||||
dedicated dev server (the user's OVH box, not a container), with many
|
||||
agent sessions working on parts of the same codebase — each session a
|
||||
fork of (essentially) the same workspace, whether as separate
|
||||
directories, per-session docker containers mounting the workspace, or
|
||||
remote mounts from local machines. In a traditional setup that
|
||||
aggregate is a lot of copied files; with content addressing the
|
||||
aggregate is (mostly) shared pointers plus each session's actual
|
||||
changes. The POC already proved the file-granularity half
|
||||
(`content_is_deduped_across_branches`: same bytes → same link, shared
|
||||
across branches for free).
|
||||
|
||||
**The dedup economics, honestly stated:**
|
||||
|
||||
- **File-level granularity (the git/POC shape) already captures the
|
||||
fork fan-out win.** When N agents fork a workspace and each edits a
|
||||
handful of files, the unchanged files are byte-identical — they dedup
|
||||
at whole-file identity. The aggregate cost is then: one copy per
|
||||
unique file version + metadata rows. That is the dominant win, and it
|
||||
is the cheapest design (the POC's tested shape, git's own shape).
|
||||
- **Chunk-level granularity adds wins only where files are large and
|
||||
edited in place** (datasets, logs, generated artifacts, big
|
||||
notebooks, binary assets) — and pays for it in per-chunk metadata
|
||||
(storage, write-path hashing/insert cost, GC walk cost) — the same
|
||||
trade iroh-blobs' inline-vs-outboard split wrestled with, and the
|
||||
reason the appfile rationale ("BLOBs < ~100KB inline faster")
|
||||
*applies to chunks under certain sizes* rather than to files. The
|
||||
right inline threshold is backend-dependent (SQLite's ~100KB
|
||||
observation; postgres and KV stores have different sweet spots) — so
|
||||
the threshold must be a configuration knob over an abstracted
|
||||
storage profile, not a constant (OQ-FS-02's backend knob).
|
||||
- **The hybrid threshold policy (OQ-FS-04's probable landing) is thus
|
||||
anchored by two real workloads:** git-object + small-file traffic
|
||||
(whole-file regime) and the mount/fan-out traffic (where chunking
|
||||
pays only above a size threshold). POC #4 measures where that
|
||||
crossover is; until then "chunk-level dedup is the working goal"
|
||||
means *for large content*, not universally.
|
||||
- **Fork fan-out has a second cost center the dedup doesn't cover:
|
||||
write-session bookkeeping.** Many concurrent sessions forking the
|
||||
same base means many live write sessions + branch rows; the POC's
|
||||
last-close-wins semantics is per-path safe, but aggregate churn
|
||||
(temp branches, orphaned sessions after agent crashes) makes OQ-FS-07's
|
||||
GC/reaping story a first-class requirement for this workload, not a
|
||||
background nicety.
|
||||
|
||||
**Cache ownership (the rudolfs-pattern question):** with many consumers
|
||||
reading shared content (docker containers mounting the same workspace;
|
||||
local machines mounting remote), where the read cache lives matters:
|
||||
|
||||
- **Store-layer decorator** (rudolfs' `Cached ↔ Retrying(Disk →
|
||||
Permanent)` shape): one cache per process, serves all local
|
||||
consumers — right for containers sharing a host.
|
||||
- **Door/door-client cache** (OQ-FS-16's mount-side caching): right
|
||||
for remote mounts where the RTT is the enemy.
|
||||
- Likely both, at different layers (host-local store cache + remote
|
||||
mount-door cache). Content addressing is what makes layered caches
|
||||
coherent cheaply — invalidation is per-hash-never-invalidates (bytes
|
||||
under a hash never change) + branch-head refresh for *which* hashes
|
||||
are current. The design note: caches should key on content hashes,
|
||||
never on paths.
|
||||
|
||||
**Sub-questions for Phase 0:** realistic dedup-ratio measurement on
|
||||
actual agent workloads (POC #4's bench); inline threshold per storage
|
||||
profile; the cache-layer split above; whether docker-mount fan-out
|
||||
wants a local unix-socket/in-process door shape (many containers, one
|
||||
producer process exposing alkfs via the local loopback shape —
|
||||
OQ-FS-10's door list grows a "container door" entry).
|
||||
|
||||
## Candidate POC register (proposals — none run yet)
|
||||
|
||||
POC placement conventions (inherited from alktunnels/alksocks): a POC
|
||||
@@ -806,7 +897,7 @@ standalone crate in the global workspace with findings written into
|
||||
| 1 | Engine skeleton: path tree + inline-blob store in one SQLite file | One-file metadata+blob-metadata (OQ-FS-02) holds up under the POC's 15-test suite ported over; single-transaction durability ordering (principle 3) is expressible | OQ-FS-02, OQ-FS-05 |
|
||||
| 2 | Write path at pack scale | Chunked write sessions handle 100MB+ streams (spill-to-file regime) with correct durability ordering and crash-abort semantics | OQ-FS-05, OQ-FS-13 |
|
||||
| 3 | alkblobs store skeleton (sha1/sha256 identity) | Own-store direction (OQ-FS-03): algorithm-tagged hash ids, bucket isolation, inline-vs-spill layout, manifest-of-chunks for large content — measure against the probe's store-seam map as the API checklist | OQ-FS-03, OQ-FS-04 |
|
||||
| 4 | Chunking probe (paper + micro-bench) | Whole-file vs chunk manifest: dedup gains on realistic workloads (agent workspaces, datasets, logs — git objects are single-chunk and ride whole-file), write-path impact, metadata size — enough to decide the format one-way door or at least stage it | OQ-FS-04 |
|
||||
| 4 | Chunking probe (paper + micro-bench) | Whole-file vs chunk manifest: dedup gains on realistic workloads (agent fork fan-out, datasets, logs — git objects are single-chunk and ride whole-file), inline-threshold crossover per storage profile, write-path impact, metadata size — enough to decide the format one-way door or at least stage it | OQ-FS-04, OQ-FS-17 |
|
||||
| 5 | Serving protocol spike | SFTP-verb protocol over channels: per-file-channel vs multiplexed-handles session model, open cost at directory-walk and git-clone fan-out, framing sketch (not a wire ADR) | OQ-FS-08, OQ-FS-09 |
|
||||
| 6 | alkgit seam spike | `GitPackGen`/`GitPackIngest` shaped against an alkblobs-backed engine — does the trait family survive contact; git oids as native addresses | OQ-FS-12 |
|
||||
| 7 | GC walk prototype | Reachability-mark GC over path rows + refs: correctness on shared content across branches, cost model, orphaned-session reaping | OQ-FS-07 |
|
||||
@@ -866,6 +957,11 @@ To evaluate (research-specialist queue, feeding the OQs named):
|
||||
baseline's "why it's slow" decomposition, verified against real
|
||||
measurements from POC #8 rather than folklore) (OQ-FS-10,
|
||||
OQ-FS-16; defer design, record constraints).
|
||||
- **Dedup/chunking economics** — measured dedup ratios on real agent
|
||||
workspaces (the fork fan-out), inline-threshold sweet spots per
|
||||
storage backend (SQLite ~100KB claim verified; postgres/KV
|
||||
equivalents), CDC chunker benchmark under workspace-like data
|
||||
(OQ-FS-17, OQ-FS-02).
|
||||
|
||||
## Convergence checklist (what Phase 0 must produce)
|
||||
|
||||
@@ -910,6 +1006,9 @@ To evaluate (research-specialist queue, feeding the OQs named):
|
||||
- [ ] OQ-FS-16 (remote mount) — the anchor workload's requirements
|
||||
written down (verb-set implications, caching shape); sequencing
|
||||
(SFTP floor first vs native protocol tier) decided
|
||||
- [ ] OQ-FS-17 (fork fan-out) — dedup-ratio measured on real agent
|
||||
workloads; inline-threshold-as-knob adopted or rejected; cache
|
||||
layer split decided; container-door shape recorded
|
||||
- [ ] Targeted POCs run + findings in `docs/research/`
|
||||
- [ ] Converge: recommended approach written up, ready to hand to the
|
||||
Architect for Phase 1
|
||||
Reference in new issue
Block a user