Files
glm-5.3-flash 2f5f0ad8ed Phase 0 third pass: fork fan-out workload, dedup economics, cache ownership
- New OQ-FS-17: the multi-agent fork fan-out (many sessions forking the
  same workspace on the shared dev server — directories, docker
  containers, or remote mounts) as the workload that makes
  content-addressed dedup load-bearing, with the honest economics:
  file-level granularity already captures the fork win (unchanged files
  dedup identically); chunk-level pays only on large in-place-edited
  files and costs per-chunk metadata
- Inline threshold recorded as backend-dependent (SQLite's ~100KB
  observation; postgres/KV differ) — a configuration knob over a storage
  profile, not a constant
- Two named objectives recorded in the working direction: efficient
  many-repo git storage (OQ-FS-12) + SFTP-capable VFS for remote
  mounting (OQ-FS-16), with the container fan-out variant noted
- rudolfs precision: verified it genuinely has an S3 backend (rusoto_s3);
  the bucket/namespace key, disk-cache-permanent split, and swappable
  permanent backend are its actual patterns; cache-ownership question
  (store-layer decorator vs door-layer cache) parked in OQ-FS-17
- POC #4 feeds OQ-FS-17; survey list gains the dedup-economics item
- AGENTS.md: new convention 13 (caches key on content hashes, never on
  paths), conventions 14-18 renumbered with cross-references fixed,
  rudolfs added to the third-party reference list, fork fan-out added
  to the quick reference

Verification: docs-only repo — OQ-FS-01..17 contiguous, every referenced
workspace path exists, convention cross-references checked after
renumbering
2026-09-23 09:56:12 +00:00

19 KiB

AGENTS.md

Operating instructions for opencode agents working in this repo. opencode auto-loads this file as instructions, overriding the built-in defaults for this project. Custom agents in .opencode/agents/ inherit these rules unless their own prompts say otherwise.

Git Workflow

Commit and push when reasonable. When a change is complete and verified, commit and push to origin/main without asking. This overrides the built-in default of "only commit when explicitly asked."

The workflow:

  1. Make the change
  2. Verify (once src/ exists): cargo test, cargo clippy --all-targets -- -D warnings, cargo fmt --check, cargo doc --no-deps if docs changed. While the repo is docs-only (Phase 0), verification is a careful re-read of the changed docs and a check that every referenced path/ADR/OQ id exists
  3. Inspect git status and git diff before staging — stage only the intended files, never secrets
  4. Write a concise commit message matching the repo style (see git log --oneline -10). For multi-point changes, use a summary line plus a body with bullet points and a verification block.
  5. git push origin main
  6. Report the commit hash and the verification summary

Exceptions — do not commit or push without asking:

  • The change is exploratory / speculative (you're not sure the user wants it kept)
  • The user is actively reviewing the diff and may ask for changes
  • The change touches a wire format or a trait shape that backends or consumers implement (one-way doors — see "Wire formats are stable" and "Backend/storage trait shapes are one-way doors" below; once consumers exist, those signatures are wire-stable contracts)
  • You'd be force-pushing, amending a published commit, creating an empty commit, or skipping hooks

Never commit secrets, keys, or credentials. If a commit fails or hooks reject it, fix the issue and create a new commit — do not amend the failed one.

Git identity is preconfigured (glm-5.3-flash <glm-5.3-flash@alk.dev>). Do not change git config, skip hooks, or use git commit -i.

Project Conventions (Rust / virtual filesystem crate)

This is the alkfs repo — a content-addressed, branch-aware virtual filesystem for the alk family. The working direction (pending Phase 0 convergence): a two-crate split — alkblobs (bucket-style content-addressed blob store: appfile layout, git-compatible sha1/sha256 identity, chunk-level dedup for large content) beneath alkfs (the path-tree mapping + branches + write sessions, plus a producer/consumer serving protocol on alkcall channels) — so that alkgit's object storage, the coming alksftp, and the alknet filesystem vision all compose on the same substrate. The anchor workload is remote mounting: a dev-server workspace mounted on a local Linux machine (today via sshfs/vanilla SFTP, whose per-syscall round trips are the problem to beat). It sits in the alk* family: alkcall (call + channels — the substrate), alktty, alktunnels, alksocks (the protocol-crate siblings), alkgit (the first in-family storage consumer). The conventions below apply to all work in src/ and tests/ once they exist. They mirror .opencode/agents/implementation-specialist.md §Project Conventions and are repeated here so they apply to every session, not just spawned implementation agents.

  1. No comments in code unless the user explicitly asks. This is a project-wide convention. Doc comments (///, //!) are fine and expected on public API. Inline // comments only when the user asks or when a non-obvious safety/correctness constraint would otherwise be missed (e.g., durability ordering constraints — "the blob bytes must be durable before the path-tree row that names them commits, or a crash orphans the reference").

  2. Error handling — thiserror for library error types. No panics in library code. No unwrap() or expect() outside tests. If you reach for unwrap, the error path wasn't specified — stop and decide what should actually happen. For poisoned RwLock/Mutex, use unwrap_or_else(|e| e.into_inner()) so a panic in one operation does not cascade to other operations. Filesystem and storage errors get faithful, typed mapping — never collapse a durability/corruption error into a generic one (a silent-corruption bug in a filesystem crate is the worst-case failure mode).

  3. tokio is the async runtime — all I/O is async. Use the wasm-clean tokio subset (rt, sync, io-util, macros, time) — do NOT use features = ["full"]. CPU-heavy storage work (hashing, pack/pao assembly, fsync-heavy batches) goes through tokio::task::spawn_blocking, the alkgit POC-2 pattern. Use tokio::sync primitives (oneshot, mpsc) for lifecycle correlation; parking_lot for short-held internal locks.

  4. WASM target is load-bearing by default, with an honest escape hatch — the family default (alktty/alktunnels/alksocks) is a wasm-clean protocol layer. A filesystem crate has heavier native gravity (real FS I/O, SQLite); the expected shape is a wasm-clean protocol/ops layer with the storage engine native-only behind features — mirroring how alksocks keeps socket I/O behind local. Whether even the protocol layer stays wasm-clean is a Phase 0 question, not a settled convention; decide deliberately and record it. Run cargo check --target wasm32-unknown-unknown once the split exists.

  5. Wire formats are stable — no wire format exists yet; when the first one is specced (the serving ALPN, its channel open-op params, any chunk/framing on the data path), it is a one-way door: decide via ADR before the first consumer exists, then additive-only. The ALPN string (provisional alk/fs; final naming per alkcall ADR-004/006) and the params-object shape follow the alkcall ADR-039 precedent (params is ALPN-specific, interpreted by the open handler).

  6. Producer/consumer, not server/client — both sides of a channels connection can initiate. A producer exposes filesystem resources (registers openable channels via ChannelCore::register_openable); a consumer opens them and speaks the file protocol inside. Both sides can be both simultaneously — connection direction is independent of service direction. Avoid "server" and "client" framing in docs and API names; use "producer" and "consumer," or "accept side" / "connect side" for the connection-establishment half specifically. See alkcall ADR-022, ADR-037.

  7. Substrate-agnostic by construction — the protocol layer must not know whether the far side is a channels BiStream, an in-process handle, a local loopback, or a door adapter (alksftp, FUSE, a sync client). The file-protocol state machine is generic over T: AsyncRead + AsyncWrite + Unpin (the fast-socks5/alksocks genericity precedent); substrate-specific types are confined to feature-gated modules injected at the assembly layer (alktty TtyBackend inversion-point pattern).

  8. No forced host-FS access — the defining requirement of this crate, the analogue of alksocks' "never binds a port." The virtual filesystem manages its own content store and never reads, writes, or mounts the host filesystem unless explicitly configured to. Bridges to the real world are explicit, optional, feature-gated assembly shapes: a local blob-file store, an SFTP door (alksftp), a mount adapter, a directory-sync tool. The engine itself is storage-agnostic about what backs its blob layer (filesystem, SQLite, remote).

  9. Metadata/content split is the load-bearing architecture — small structured state (path edges, branches, snapshots, refs, cached sizes) lives in a transactional store; content bytes live in a content-addressed blob layer (dedup by hash, branch sharing for free). This is the alknet-filesystem POC's three-layer conclusion and the same "shape" as git + gitlfs and iroh-blobs' inline/outboard split. Path-tree operations are O(path edges), never O(bytes) — rename is O(1) on edges regardless of file size. Do not let byte-scale concerns leak into the metadata layer or vice versa. See docs/research/phase-0.md.

  10. Content addressing and durability ordering — content is identified by its hash, not its path, with git-compatible identity (sha1/sha256, algorithm-tagged): git objects are natively addressable in the blob layer, never mapped through a foreign hash. Identical content is shared across branches/paths by construction; chunk-level dedup applies to large content (git objects are effectively single-chunk and ride whole-file identity). The durability contract is the crash-safety story: content must be durable before any metadata naming it commits; a crash mid-write leaves the old version intact and visible (the POC's branch-on-write/merge-on-close property). Across the alkblobs/alkfs seam the contract is the put/get API (what a returned handle promises about durability is part of that seam's spec). Exact chunking/manifest choices are Phase 0 OQs; the invariant (no torn versions visible, no orphaned-name commits) is not.

  11. Vendored core types come from alkcall — Connection, ProtocolHandler, BiStream, BidiStreamSource, AuthContext, Identity, IdentityProvider, AccessControl, OwnershipProvider, HandlerError, StreamError come from alkcall::core. Do not vendor copies into this crate. alkcall is ours and co-developed — breaking changes are expected at this major-zero stage; find and fix issues upstream rather than working around them. Pin deliberately and bump deliberately. The establishment surface (alkcall ADR-049 + amendments) and the identity seam (CF-005/CF-006) are load-bearing for the serving protocol, same as the alktty/alktunnels/alksocks pattern.

  12. Access control — a filesystem is arbitrary read/write by nature: the open gate and path-scope policy are the security boundary, the same posture as alksocks' arbitrary egress. Scope-gate file opens (an alkfs-shaped scope following alktty's TTY_OPEN_SCOPE / alksocks' SOCKS5_OPEN_SCOPE precedent), wire producer openable channels through AccessControl for free via ChannelCore::register_openable, and treat per-path/per-bucket policy as a Phase 1 design question (OQ). Multi-tenancy isolation (the POC's bucket concept) is a where-clause, not an afterthought; bucket semantics across the alkblobs/alkfs seam are part of the seam's spec, not an afterthought.

  13. Caches key on content hashes, never on paths — content under a hash is immutable, so content-hash-keyed caches never invalidate on bytes; only the path/branch-head → hash mapping goes stale, and that is cheap to refresh. Any cache (store-layer decorator per rudolfs, mount-door cache, session prefetch) follows this rule. Path-keyed caches reintroduce coherence problems content addressing exists to remove.

  14. Backend/storage trait shapes are one-way doors — the seam the protocol crate exposes to storage implementers (the alkgit GitRefs/GitPackGen/GitPackIngest trait-family precedent, docs/architecture/backend.md there) is a contract once consumers exist: keep traits small, orthogonal, and substrate-blind; never leak gix/iroh/SQLite types across them. The alknet-filesystem probe (alknet-blobs-external-store-probe.md) is the cautionary example — the sealed/pub(crate) surface is exactly what turns "external store" into "fork".

  15. Feature flags — substrate backends and heavy dependencies are feature-gated if the need arises. The base crate should compile lean (no SQLite, no gix, no kernel-FS access unless the feature is on). Verify both cargo test (default) and cargo test --all-features pass if features are added.

  16. Naming — Rust standard: snake_case for functions/variables/ modules, PascalCase for types/traits, SCREAMING_SNAKE_CASE for constants.

  17. Module structure — one module per file under src/, re-exported from src/lib.rs. Public API surface is lib.rs re-exports. The expected shape (pending Phase 0/1 pinning): storage engine modules (path tree, blob store, write sessions, GC), backend traits + feature-gated implementations, and the producer/consumer protocol modules mirroring the alktty/alktunnels/alksocks structure. Backend modules are feature-gated and never imported from the protocol/adapter/client modules.

  18. Upstream posture — alkcall, alktunnels, and the alk* crates are ours to shape: file asks early and land them there rather than working around them locally (the alktunnels E-01/E-02 precedent — filed from Phase 0, landed within a day). Third-party crates (iroh-blobs, sqlite/rusqlite, russh-sftp, gix, automerge, honker, rudolfs) are NOT ours and never will be: wrap, extract, or fork deliberately per the alksocks precedent (OQ-SK-04 → ADR-013: extraction as owned code with provenance notices, differential tests against the reference checkout, no silent absorption) — and only when the carried changes pay for themselves. /workspace/iroh-blobs, /workspace/russh-sftp, /workspace/rudolfs, etc. are read-only reference checkouts.

Verification Commands

Run these before committing (once src/ exists). All must pass.

cargo test                                    # full suite
cargo clippy --all-targets -- -D warnings
cargo fmt --check
cargo doc --no-deps                           # if docs changed
cargo test --all-features                     # if features are added
cargo check --target wasm32-unknown-unknown   # if the wasm posture applies (see convention 4)
cargo publish --dry-run --allow-dirty         # before a release

While the repo is docs-only (Phase 0), there is no build to verify — verification means re-reading changed docs and checking that every referenced path, ADR, and OQ id exists.

Architecture Context

  • docs/research/phase-0.md — the Phase 0 (Exploration) document and the current state of this repo: vision, prior art, open questions (OQ-FS-NN), and the POC register. Read it before non-trivial work. The SDD process lives in docs/sdd_process.md (Phase 0 in progress; docs/architecture/ does not exist yet — do not create it; that is Phase 1).
  • The ancestor research:
    • alknet-filesystem POCs — /workspace/@alkdev/alknet/docs/research/alknet-filesystem/ (poc-summary.md — the three-iteration POC: SQLite path tree + iroh-blobs + honker, branch-on-write/merge-on-close, automerge sync; alknet-blobs-external-store-probe.md — the iroh-blobs Command-actor probe). This crate is the decomposition-era continuation of that research on the published alk* substrate.
    • alkgit — /workspace/@alkdev/alkgit: the first in-family storage consumer (currently paused on this problem). Its backend trait family (docs/architecture/backend.md — GitRefs, GitPackGen, GitPackIngest over gix-odb) is the shape alkfs must serve efficiently: streaming pack generation/ingestion, CAS ref transactions, budgeted resources. Its docs/research/ holds the gitoxide and git-protocol surveys.
  • The substrate and sibling crates:
    • alkcall — /workspace/@alkdev/alkcall (v0.8.0, crates.io). The substrate: call protocol + channels multiplexing. This crate will consume alkcall::core types and the channels ChannelCore/ChannelClient/register_openable_with_establisher surface, same as alktty/alktunnels/alksocks.
    • alktty — /workspace/@alkdev/alktty: the first producer/consumer protocol crate; the backend inversion point, wasm-clean default, scope-gating, and feature-gated backend precedents.
    • alktunnels — /workspace/@alkdev/alktunnels: -L/-R/-D tunnels on channels; the Phase 0 findings format and POC placement conventions (docs/research/phase-0-findings.md) this repo inherits.
    • alksocks — /workspace/@alkdev/alksocks: the SOCKS5 sibling; the closest Phase 0 template (docs/research/phase-0.md shape, OQ numbering, upstream-posture convention).
    • alkvault — /workspace/@alkdev/alkvault: secure secret handling; the family rule is metadata stores hold vault references, never plaintext secrets (alkgit vision principle 3).
  • Key upstream ADRs that inform this crate's design (alkcall numbers unless noted):
    • ADR-035 — channels pure channel multiplexing (8-byte header); the file-session rides inside a channel's BiStream
    • ADR-050 — two-pump shutdown-on-completion (channels::pump_bidi); use it, do not hand-roll
    • ADR-049 (amendments) — the establishment phase; register_openable_with_establisher; refused sessions are typed channel:open_failed call errors, never phantom channels
    • ADR-039 — params is ALPN-specific; ADR-037 — channel lifecycle ops on channel 0; ADR-034 — the channels wire format (one-way door); ADR-036 — channel 0 is pre-negotiated alk/call
    • ADR-042 — hub relay (terminate-and-re-produce); ADR-051 — ChannelRelay/HubLegTemplate
    • Ledger CF-005/CF-006 — the per-call opener identity seam on the open-op hooks
  • If a TODO references a design direction that a later ADR has decided against, the TODO is stale — remove it and align with the ADR. Do not implement the rejected design.

Phase 0 working direction (quick reference)

Recorded so sessions don't relitigate settled leanings — see docs/research/phase-0.md §Working direction for full context:

  • Git-compatible identity (sha1/sha256), not BLAKE3 — the ancestor POCs' BLAKE3 was an iroh-blobs artifact; iroh-blobs is evaluated, not adopted (findings transfer, code doesn't).
  • Chunk-level dedup is the working goal for large content; git objects ride whole-file identity (single-chunk regime).
  • Two-crate split (working direction): alkblobs + alkfs — the blob store beneath the path-tree mapping; repo shape (one repo, two crates vs two repos) is an open sub-question.
  • Anchor workload: remote mounting — sshfs/vanilla-SFTP replacement; per-syscall round trips are the problem to beat. Second named workload: the multi-agent fork fan-out (many sessions forking the same workspace — containers or remote mounts), which is what makes content-addressed dedup load-bearing (OQ-FS-17).
  • These are leanings, not decisions: every one is an open OQ (OQ-FS-01..17) until the Phase 0 convergence; don't mark anything "settled" in code or docs until the checklist starts ticking.