diff --git a/AGENTS.md b/AGENTS.md index 87a3d0f..368d283 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -187,9 +187,19 @@ implementation agents. channels through `AccessControl` for free via `ChannelCore::register_openable`, and treat per-path/per-bucket policy as a Phase 1 design question (OQ). Multi-tenancy isolation - (the POC's `bucket` concept) is a where-clause, not an afterthought. + (the POC's `bucket` concept) is a where-clause, not an afterthought; + bucket semantics across the alkblobs/alkfs seam are part of the + seam's spec, not an afterthought. -13. **Backend/storage trait shapes are one-way doors** — the seam the +13. **Caches key on content hashes, never on paths** — content under a + hash is immutable, so content-hash-keyed caches never invalidate on + bytes; only the path/branch-head → hash *mapping* goes stale, and + that is cheap to refresh. Any cache (store-layer decorator per + rudolfs, mount-door cache, session prefetch) follows this rule. + Path-keyed caches reintroduce coherence problems content addressing + exists to remove. + +14. **Backend/storage trait shapes are one-way doors** — the seam the protocol crate exposes to storage implementers (the alkgit `GitRefs`/`GitPackGen`/`GitPackIngest` trait-family precedent, `docs/architecture/backend.md` there) is a contract once consumers @@ -199,17 +209,17 @@ implementation agents. the sealed/`pub(crate)` surface is exactly what turns "external store" into "fork". -14. **Feature flags** — substrate backends and heavy dependencies are +15. **Feature flags** — substrate backends and heavy dependencies are feature-gated if the need arises. The base crate should compile lean (no SQLite, no gix, no kernel-FS access unless the feature is on). Verify both `cargo test` (default) and `cargo test --all-features` pass if features are added. -15. **Naming** — Rust standard: `snake_case` for functions/variables/ +16. **Naming** — Rust standard: `snake_case` for functions/variables/ modules, `PascalCase` for types/traits, `SCREAMING_SNAKE_CASE` for constants. -16. **Module structure** — one module per file under `src/`, re-exported +17. **Module structure** — one module per file under `src/`, re-exported from `src/lib.rs`. Public API surface is `lib.rs` re-exports. The expected shape (pending Phase 0/1 pinning): storage engine modules (path tree, blob store, write sessions, GC), backend traits + @@ -218,17 +228,19 @@ implementation agents. modules are feature-gated and never imported from the protocol/adapter/client modules. -17. **Upstream posture** — alkcall, alktunnels, and the alk* crates are +18. **Upstream posture** — alkcall, alktunnels, and the alk* crates are ours to shape: file asks early and land them there rather than working around them locally (the alktunnels E-01/E-02 precedent — - filed from Phase 0, landed within a day). Third-party crates - (iroh-blobs, sqlite/rusqlite, russh-sftp, gix, automerge, honker) + filed from Phase 0, landed within a day). Third-party crates + (iroh-blobs, sqlite/rusqlite, russh-sftp, gix, automerge, honker, + rudolfs) are NOT ours and never will be: wrap, extract, or fork deliberately per the alksocks precedent (OQ-SK-04 → ADR-013: extraction as owned code with provenance notices, differential tests against the reference checkout, no silent absorption) — and only when the carried changes pay for themselves. `/workspace/iroh-blobs`, - `/workspace/russh-sftp`, etc. are read-only reference checkouts. + `/workspace/russh-sftp`, `/workspace/rudolfs`, etc. are read-only + reference checkouts. ## Verification Commands @@ -324,7 +336,10 @@ Recorded so sessions don't relitigate settled leanings — see blob store beneath the path-tree mapping; repo shape (one repo, two crates vs two repos) is an open sub-question. - **Anchor workload: remote mounting** — sshfs/vanilla-SFTP replacement; - per-syscall round trips are the problem to beat. + per-syscall round trips are the problem to beat. Second named + workload: the multi-agent fork fan-out (many sessions forking the + same workspace — containers or remote mounts), which is what makes + content-addressed dedup load-bearing (OQ-FS-17). - These are **leanings, not decisions**: every one is an open OQ - (OQ-FS-01..16) until the Phase 0 convergence; don't mark anything + (OQ-FS-01..17) until the Phase 0 convergence; don't mark anything "settled" in code or docs until the checklist starts ticking. \ No newline at end of file diff --git a/docs/research/phase-0.md b/docs/research/phase-0.md index e7c70e2..59a9271 100644 --- a/docs/research/phase-0.md +++ b/docs/research/phase-0.md @@ -57,18 +57,25 @@ converge, not decisions):** only identical files; git itself dedups at object granularity (packs delta-compress within a pack, but the address space is per-object); chunking is the storage-layer generalization (OQ-FS-04 — measure - whether the workloads pay before committing). + whether the workloads pay before committing). The workload that makes + dedup load-bearing is the **multi-agent fork fan-out** (OQ-FS-17): + many agent sessions forking the same workspace, each ultimately + making small changes — at file granularity that alone already wins + (unchanged files dedup identically); chunking adds wins only on + large files edited in place. - **The two-crate split: `alkblobs` + `alkfs`.** A bucket-style - content-addressed blob store (the S3 shape) beneath the path-tree - mapping (the alknet-filesystem POC's layer 2). alkgit can consume the - blob store without a filesystem; the FS layer keeps its own crate - (OQ-FS-01). -- **The anchor workload is remote mounting.** The concrete use case: - the dev-server workspace mounted on a local Linux machine under - `remote/` — today over vanilla SFTP (sshfs), which is slow in known - ways (a round trip per syscall, stat storms, full inode work per - stat). An appfile-indexed VFS fixes the server side; caching and - pipelining close the rest (OQ-FS-16). + content-addressed blob store (the S3/rudolfs shape — and rudolfs + really does have an S3 backend, see §Prior art) beneath the + path-tree mapping (the alknet-filesystem POC's layer 2). alkgit can + consume the blob store without a filesystem; the FS layer keeps its + own crate (OQ-FS-01). +- **Two named objectives, two anchor workloads.** (1) Store potentially + many git repos efficiently (alkgit's un-pause condition — OQ-FS-12). + (2) A virtual filesystem supporting an SFTP door for remote mounting + (OQ-FS-16) — including the container fan-out: the same workspace + mounted into potentially many docker containers on the same host + (many consumers, one producer). The dedup economics (OQ-FS-17) is + what makes the aggregate sane when every session is a fork. **Why this crate exists.** The alknet mono-repo's filesystem research (three POC iterations, 24 passing tests) proved a viable three-layer @@ -209,7 +216,7 @@ does not start from zero. It inherits: ~120-160 lines of bao-tree glue must be re-implemented per store. The probe's cautionary half: `TempTags`/`TempTagScope` are sealed `pub(crate)` — sealed surface is exactly what turns "external store" - into "fork" (AGENTS.md convention 13's cautionary example). + into "fork" (AGENTS.md convention 14's cautionary example). - **The consumer shapes this crate must serve.** alkgit's backend trait family (`docs/architecture/backend.md` there: `GitRefs`, `GitPackGen`, `GitPackIngest` — streaming pack generation/ingestion, CAS ref @@ -305,7 +312,7 @@ consumable). (OQ-FS-02/03). - **Sealed internals** (`TempTags`, the `BaoTreeSender`) — the probe's fork-vs-PR calculus. Third-party upstream, not ours, never will be - (AGENTS.md convention 17). + (AGENTS.md convention 18). ### git + git-lfs — the shape to optimize for @@ -373,13 +380,14 @@ https://sqlite.org/appfileformat.html — the insight that anchored the POC. honker (`/workspace/honker`, `honker-core` 0.2.4) supplies notify/locks/queues as SQL functions on *your* connection — the transactional-outbox property. Caveats: honker is third-party (not ours, -convention 17), single-machine by explicit design ("two servers writing +convention 18), single-machine by explicit design ("two servers writing the same .db over NFS is not a Honker deployment strategy"), and its scope in alkfs is an OQ (notify-on-commit may be the only load-bearing piece — locks/queues/scheduler may be excess). ### rudolfs — the git-lfs server reference +`/workspace/rudolfs` (local checkout) + `/workspace/@alkdev/alknet/docs/research/references/gitlfs/ rudolfs-reference.md`: `StorageKey = (Namespace, Oid)` tenant isolation; the decorator composition `Verify ↔ Encrypted ↔ Cached ↔ Retrying(Disk → @@ -388,6 +396,18 @@ client and cache simultaneously. The cache/permanent split and the fanout pattern are the production-shape references for alkfs's blob layer and its serving read path. +**Two precise notes (2026-09-23):** (1) rudolfs *genuinely has an S3 +backend* (`rusoto_s3` in its Cargo.toml; local disk cache in front of +permanent S3 storage) — so "bucket-style storage like S3" in the +alkblobs direction is grounded in rudolfs' actual shape, not a loose +analogy: the bucket/namespace key, the disk-cache-permanent split, and +a swappable permanent backend are all rudolfs patterns. (2) The +decorator chain is the compositional answer to "local cache in front of +remote storage" — the same shape OQ-FS-16's mount caching wants, at the +*store* layer rather than the door layer. Which layer owns the cache +(door-local content cache vs store-layer decorator) is an OQ-FS-17 +sub-question. + ### Anti-prior-art (what NOT to carry over) - **The alknet mono-repo's scope.** The filesystem research is inherited @@ -793,6 +813,77 @@ known slowness: This OQ is the *consumer pull* behind the whole crate — if the mount use case stops mattering, OQ-FS-08's urgency drops with it. +#### OQ-FS-17: The multi-agent fork fan-out — dedup economics and cache ownership + +**Added 2026-09-23.** The second named workload, and the one that makes +dedup *load-bearing* rather than a nicety. The concrete deployment: one +dedicated dev server (the user's OVH box, not a container), with many +agent sessions working on parts of the same codebase — each session a +fork of (essentially) the same workspace, whether as separate +directories, per-session docker containers mounting the workspace, or +remote mounts from local machines. In a traditional setup that +aggregate is a lot of copied files; with content addressing the +aggregate is (mostly) shared pointers plus each session's actual +changes. The POC already proved the file-granularity half +(`content_is_deduped_across_branches`: same bytes → same link, shared +across branches for free). + +**The dedup economics, honestly stated:** + +- **File-level granularity (the git/POC shape) already captures the + fork fan-out win.** When N agents fork a workspace and each edits a + handful of files, the unchanged files are byte-identical — they dedup + at whole-file identity. The aggregate cost is then: one copy per + unique file version + metadata rows. That is the dominant win, and it + is the cheapest design (the POC's tested shape, git's own shape). +- **Chunk-level granularity adds wins only where files are large and + edited in place** (datasets, logs, generated artifacts, big + notebooks, binary assets) — and pays for it in per-chunk metadata + (storage, write-path hashing/insert cost, GC walk cost) — the same + trade iroh-blobs' inline-vs-outboard split wrestled with, and the + reason the appfile rationale ("BLOBs < ~100KB inline faster") + *applies to chunks under certain sizes* rather than to files. The + right inline threshold is backend-dependent (SQLite's ~100KB + observation; postgres and KV stores have different sweet spots) — so + the threshold must be a configuration knob over an abstracted + storage profile, not a constant (OQ-FS-02's backend knob). +- **The hybrid threshold policy (OQ-FS-04's probable landing) is thus + anchored by two real workloads:** git-object + small-file traffic + (whole-file regime) and the mount/fan-out traffic (where chunking + pays only above a size threshold). POC #4 measures where that + crossover is; until then "chunk-level dedup is the working goal" + means *for large content*, not universally. +- **Fork fan-out has a second cost center the dedup doesn't cover: + write-session bookkeeping.** Many concurrent sessions forking the + same base means many live write sessions + branch rows; the POC's + last-close-wins semantics is per-path safe, but aggregate churn + (temp branches, orphaned sessions after agent crashes) makes OQ-FS-07's + GC/reaping story a first-class requirement for this workload, not a + background nicety. + +**Cache ownership (the rudolfs-pattern question):** with many consumers +reading shared content (docker containers mounting the same workspace; +local machines mounting remote), where the read cache lives matters: + +- **Store-layer decorator** (rudolfs' `Cached ↔ Retrying(Disk → + Permanent)` shape): one cache per process, serves all local + consumers — right for containers sharing a host. +- **Door/door-client cache** (OQ-FS-16's mount-side caching): right + for remote mounts where the RTT is the enemy. +- Likely both, at different layers (host-local store cache + remote + mount-door cache). Content addressing is what makes layered caches + coherent cheaply — invalidation is per-hash-never-invalidates (bytes + under a hash never change) + branch-head refresh for *which* hashes + are current. The design note: caches should key on content hashes, + never on paths. + +**Sub-questions for Phase 0:** realistic dedup-ratio measurement on +actual agent workloads (POC #4's bench); inline threshold per storage +profile; the cache-layer split above; whether docker-mount fan-out +wants a local unix-socket/in-process door shape (many containers, one +producer process exposing alkfs via the local loopback shape — +OQ-FS-10's door list grows a "container door" entry). + ## Candidate POC register (proposals — none run yet) POC placement conventions (inherited from alktunnels/alksocks): a POC @@ -806,7 +897,7 @@ standalone crate in the global workspace with findings written into | 1 | Engine skeleton: path tree + inline-blob store in one SQLite file | One-file metadata+blob-metadata (OQ-FS-02) holds up under the POC's 15-test suite ported over; single-transaction durability ordering (principle 3) is expressible | OQ-FS-02, OQ-FS-05 | | 2 | Write path at pack scale | Chunked write sessions handle 100MB+ streams (spill-to-file regime) with correct durability ordering and crash-abort semantics | OQ-FS-05, OQ-FS-13 | | 3 | alkblobs store skeleton (sha1/sha256 identity) | Own-store direction (OQ-FS-03): algorithm-tagged hash ids, bucket isolation, inline-vs-spill layout, manifest-of-chunks for large content — measure against the probe's store-seam map as the API checklist | OQ-FS-03, OQ-FS-04 | -| 4 | Chunking probe (paper + micro-bench) | Whole-file vs chunk manifest: dedup gains on realistic workloads (agent workspaces, datasets, logs — git objects are single-chunk and ride whole-file), write-path impact, metadata size — enough to decide the format one-way door or at least stage it | OQ-FS-04 | +| 4 | Chunking probe (paper + micro-bench) | Whole-file vs chunk manifest: dedup gains on realistic workloads (agent fork fan-out, datasets, logs — git objects are single-chunk and ride whole-file), inline-threshold crossover per storage profile, write-path impact, metadata size — enough to decide the format one-way door or at least stage it | OQ-FS-04, OQ-FS-17 | | 5 | Serving protocol spike | SFTP-verb protocol over channels: per-file-channel vs multiplexed-handles session model, open cost at directory-walk and git-clone fan-out, framing sketch (not a wire ADR) | OQ-FS-08, OQ-FS-09 | | 6 | alkgit seam spike | `GitPackGen`/`GitPackIngest` shaped against an alkblobs-backed engine — does the trait family survive contact; git oids as native addresses | OQ-FS-12 | | 7 | GC walk prototype | Reachability-mark GC over path rows + refs: correctness on shared content across branches, cost model, orphaned-session reaping | OQ-FS-07 | @@ -866,6 +957,11 @@ To evaluate (research-specialist queue, feeding the OQs named): baseline's "why it's slow" decomposition, verified against real measurements from POC #8 rather than folklore) (OQ-FS-10, OQ-FS-16; defer design, record constraints). +- **Dedup/chunking economics** — measured dedup ratios on real agent + workspaces (the fork fan-out), inline-threshold sweet spots per + storage backend (SQLite ~100KB claim verified; postgres/KV + equivalents), CDC chunker benchmark under workspace-like data + (OQ-FS-17, OQ-FS-02). ## Convergence checklist (what Phase 0 must produce) @@ -910,6 +1006,9 @@ To evaluate (research-specialist queue, feeding the OQs named): - [ ] OQ-FS-16 (remote mount) — the anchor workload's requirements written down (verb-set implications, caching shape); sequencing (SFTP floor first vs native protocol tier) decided +- [ ] OQ-FS-17 (fork fan-out) — dedup-ratio measured on real agent + workloads; inline-threshold-as-knob adopted or rejected; cache + layer split decided; container-door shape recorded - [ ] Targeted POCs run + findings in `docs/research/` - [ ] Converge: recommended approach written up, ready to hand to the Architect for Phase 1 \ No newline at end of file