- .gitignore matching sibling repos (target/, node_modules/, .worktrees/, Cargo.lock) - AGENTS.md on the alkcall/alksocks template: git workflow, 17 project conventions adapted for a virtual-filesystem crate (no forced host-FS access, metadata/content split, durability ordering, storage-trait one-way doors), verification commands (docs-only posture for Phase 0), and architecture context pointing at the ancestor research and family ADRs - docs/research/phase-0.md: initial Phase 0 draft — vision/scope sketch, guiding principles, what's already settled, prior art (alknet-filesystem POCs, iroh-blobs + external-store probe, git/git-lfs, alkgit backend seam, russh-sftp, SQLite/honker, rudolfs), 15 open questions (OQ-FS-01..15) grouped by theme, a 7-entry candidate POC register (proposals, none run), survey list, and an unconverged checklist Verification: docs-only repo — every referenced path, ADR id, and OQ id checked to exist (workspace paths, alkcall ADRs 034-051, alkgit research files); OQ numbering OQ-FS-01..15 complete with no gaps
41 KiB
status: draft last_updated: 2026-09-23 (initial draft from the setup discussion — unconverged by design: every OQ is open, the POC register holds proposals not results, and the prior-art pass is partial)
alkfs — Phase 0 (Exploration)
This document captures Phase 0 (Exploration) for the alkfs crate: vision,
guiding principles, prior art, open questions (OQ-FS-01..NN), and the
candidate POC register. Phase 0's objective per docs/sdd_process.md:
capture vision and guiding principles; research options; validate
approaches; converge on a recommended approach. It is the input to Phase 1
(Architecture), where the Architect will produce docs/architecture/
specs, ADRs, and the open-questions tracker.
Drafted 2026-09-23 from the initial setup discussion. The crate is the storage-engine member of the alk* family: alkcall (call + channels — the substrate), alktty/alktunnels/alksocks (the protocol-crate siblings), alkgit (the first in-family storage consumer, currently paused on exactly the problem this crate exists to solve).
Honest framing: this Phase 0 is earlier in its lifecycle than alksocks' was at the equivalent commit. There are more unknowns here than in any sibling so far — a filesystem is both a storage engine (hashing, chunking, durability, GC) and a serving protocol (verbs, sessions, wire framing), and almost none of those decisions have empirical ground yet. The POC register below is proposals, not results. Expect this document to be rewritten several times before convergence.
Vision and guiding principles
One sentence: a content-addressed, branch-aware virtual filesystem for the alk family — a local storage engine (path-tree metadata + content-addressed blobs) with a producer/consumer serving protocol on alkcall channels, so that alkgit's object storage, the coming alksftp door, and the alknet filesystem vision all compose on the same substrate — and it never touches the host filesystem unless explicitly configured to.
Why this crate exists. The alknet mono-repo's filesystem research (three POC iterations, 24 passing tests) proved a viable three-layer architecture but lived inside a project with excessive scope. The alk* family has since been decomposed and published (alkcall, alktty, alktunnels, alkvault; alksocks and alkgit in flight). alkgit paused because its storage problem — content-addressed blob storage with branch/CAS semantics — is not a git-specific problem, and solving it inside alkgit would leave alksftp and the alknet filesystem rewrite to solve it again. The same basic "shape" recurs across the family: git + git-lfs is small-object CAS metadata over big-blob content storage; iroh-blobs' FsStore is inline-small/outboard-large; the alknet-filesystem POC is SQLite path edges over a blob store. alkfs is that shape, made a crate.
Scope sketch (unconverged — this is the shape Phase 0 must confirm or cut):
- Storage engine half — the path tree (paths, branches, snapshots, tombstones, cached sizes) in a transactional metadata store; content bytes in a content-addressed blob layer; write sessions (branch-on-write/merge-on-close); GC. Fully local, no channels required — a consumer embeds the engine in-process.
- Serving protocol half — the producer/consumer pair on alkcall
channels: a producer registers openable filesystem channels
(
ChannelCore::register_openable); a consumer opens them and speaks the file protocol inside. Thealk/fs-shaped ALPN resource, the same "ALPN as a service" shape as alktty/alktunnels/alksocks. - Doors stay out — alksftp (SFTP door), FUSE/mount adapters, sync tools are separate crates/features that consume alkfs, per the family rule (alkgit vision: doors are family infrastructure, not crate contents).
Guiding principles, inherited from the alk* family plus the filesystem research:
- No forced host-FS access — the defining requirement. The analogue of alksocks' "never binds a port": the engine manages its own content store and never reads, writes, or mounts the host filesystem unless explicitly configured to. Every bridge to the real world (a blob-file store, an SFTP door, a mount, a sync tool) is an explicit, optional, feature-gated assembly shape. The engine itself is storage-agnostic about what backs its blob layer (filesystem files, SQLite, remote).
- Metadata/content split is load-bearing. Small structured state (path edges, branch pointers, tombstones, sizes) lives in a transactional store; content bytes live under their hash in the blob layer. Path operations are O(path edges), never O(bytes) — rename is O(1) regardless of file size; identical content dedups across branches and paths by construction. This is the alknet-filesystem POC's core conclusion and the same shape as git + git-lfs and iroh-blobs' inline/outboard split.
- Content addressing with a durability contract. Content is identified by its hash, not its path. The invariant that must survive every design choice: content is durable before any metadata row naming it commits; a crash mid-write leaves the old version intact and visible; no torn version is ever readable; no metadata commit orphans its content. Exact hashing/chunking is an OQ; the invariant is not.
- Branch-aware from day one. Fossil-style branches (a branch is a parent chain + per-branch overrides + tombstones), branch-on-write/ merge-on-close for the write path. This is what gives free multi-agent/multi-fork content sharing and the POSIX concurrent-reader-sees-old-version property — and it is the property git's object model needs (content shared across refs by hash).
- "ALPN as a service," not "a server." The filesystem is a produced, ACL-scoped resource on a channels connection. The protocol never opens host files or binds anything; door-shaped conveniences (local loopback front door, in-process consumer handle) are optional assembly-layer shapes on either side. Producer/consumer vocabulary, never server/client.
- Substrate-agnostic protocol, native engine. The file-protocol
state machine is generic over
T: AsyncRead + AsyncWrite + Unpin(fast-socks5/alksocks genericity precedent); the storage engine is native-only behind features (real FS I/O, SQLite). Whether the protocol layer itself stays wasm-clean is an OQ, not an assumption. - Multi-tenancy is a where-clause, not an afterthought. The POC's
bucketcolumn is the isolation unit; identity → bucket mapping is an assembly-layer auth question; per-path/per-bucket policy is a Phase 1 design question. A filesystem is arbitrary read/write by nature — the open gate and path-scope policy are the security boundary. - Bounded resources on every serving path. alkgit's ADR-009 posture (wall-clock, size, round limits) is the consumer shape this crate must serve; the serving protocol inherits alkcall's backpressure and channel-cap invariants rather than re-deriving them.
What is already settled
The foundation is POC-validated and (upstream) ADR-pinned; this crate does not start from zero. It inherits:
- The channels substrate — demux→Connection→handler→mux, the
establishment story (
register_openable_with_establisher, alkcall ADR-049; refused sessions are typedchannel:open_failedcall errors, never phantom channels), the identity seam (CF-005/CF-006: per-call opener identity on the open-op hooks), ACL-for-free viaregister_openable, the two-pump data plane (channels::pump_bidi, ADR-050 — use it, do not hand-roll), hub relay (ChannelRelay/HubLegTemplate, ADR-051). All production-validated by alktty/alktunnels/alksocks. - The metadata/content split, POC-proven. alknet-filesystem POC iteration 1: SQLite path tree over a blob store, 8 tests — branch inheritance, tombstones, O(1) rename, bucket isolation, honker notify-atomic-with-mutation.
- The write path, POC-proven. Iteration 2: branch-on-write/ merge-on-close reconciles "BLAKE3 must hash the complete file" with "writes arrive in chunks, possibly out of order" — concurrent readers see the old version until close() commits atomically; crash/abort leaves the old version intact. 7 tests.
- The distributed-sync shape, POC-proven but scope-uncommitted. Iteration 3: path tree as an automerge CRDT synced over QUIC; 9 tests; metadata and content sync separately; concurrent-root-map initialization is a real constraint (root structures must be created by one node and synced before others write). Whether any of this is alkfs v1 scope is OQ-FS-14.
- The iroh-blobs store-seam findings (the external-store probe). The
alknet-blobs probe traced iroh-blobs 0.103's full
Commandactor surface: there is noStoretrait —Storeis a concrete struct over anirpc::Client;MemStore/FsStoreare actors; a third store is a third actor speaking the pubCommandenum. Blockers are a ~4-line visibility PR (Store::from_sender,ref_from_sender,ApiClient, theScopefield); GC (run_gc) is pub and reusable; ~120-160 lines of bao-tree glue must be re-implemented per store. The probe's cautionary half:TempTags/TempTagScopeare sealedpub(crate)— sealed surface is exactly what turns "external store" into "fork" (AGENTS.md convention 13's cautionary example). - The consumer shapes this crate must serve. alkgit's backend trait
family (
docs/architecture/backend.mdthere:GitRefs,GitPackGen,GitPackIngest— streaming pack generation/ingestion, CAS ref transactions, budgeted resources) and the russh-sftpHandlertrait (the POC's near-1:1 SFTP↔PathTree mapping table). "Optimize for git and sftp-like" is the working posture. - The SQLite-as-application-file-format insight. BLOBs < ~100KB are faster inline in SQLite than as filesystem files; atomic transactions over path-tree metadata; the schema is the documentation. iroh-blobs' FsStore discovered the same shape independently (redb inline tables + filesystem files). The inline-small/outboard-large hybrid is the metadata/content split at storage granularity.
- The family conventions — AGENTS.md in this repo (error handling, tokio subset, feature-gating, upstream posture, one-way doors) and the alkcall ADR set listed there.
Prior art
alknet-filesystem POCs — the direct ancestor
/workspace/@alkdev/alknet/docs/research/alknet-filesystem/
(poc-summary.md, alknet-blobs-external-store-probe.md), POC crates at
/workspace/alknet-filesystem-poc and /workspace/alknet-fs-sync-poc.
The three-layer conclusion (SQLite path tree + iroh-blobs content +
honker coordination) and the write-path resolution
(branch-on-write/merge-on-close) are this crate's inheritance, not prior
art to re-litigate — but every dependency choice in those POCs is open
again:
- SQLite path tree (
rusqlite0.39, bundled):buckets/branches/paths/tombstones+write_sessions/write_chunks; recursive-CTE chain walk resolves(bucket, branch, path)→ content link in sub-millisecond time; only deltas are stored per branch. - iroh-blobs
MemStoreas the content layer (deliberately notFsStore— no redb, no fsync rabbit hole in the POC). - honker-core for notify-in-transaction (the transactional-outbox
pattern built in):
SELECT notify(...)inside the same transaction as a path-tree mutation; watchers wake on commit, not poll. - What did NOT transfer and why — the POC's remaining unknowns (FsStore/redb coexistence, SFTP wiring, GC/tags, chain-depth perf, snapshot semantics) are this Phase 0's OQ list; the probe resolved the two-database question by reframing it (external actor, own SQLite file), and the redb question dissolves entirely if alkfs owns its blob metadata in its own store.
iroh-blobs — the nearest CAS store (evaluated; not adopted as-is)
/workspace/iroh-blobs (v0.100 checkout) + published 0.103. DESIGN.md
is the best available treatment of the blob-store trade space (files are
hard: the bitfield/data/outboard write-ordering problem). What it gives:
BLAKE3 content addressing, bao-tree verified streaming (per-chunk
integrity with an outboard), dedup by construction, partial-range reads,
tags + mark-sweep GC, a hybrid store (small inline in redb, large as
filesystem files) — the same metadata/content split as principle 2.
The known frictions (why the user called it "a pita" for git):
- The transfer protocol is welded to iroh/QUIC. iroh-blobs' network
story is an iroh
ProtocolHandlerspeaking its own ALPN over QUIC. For alkfs this is irrelevant-by-design — content transfer rides alkcall channels — but it means iroh-blobs is at most a local store dependency, never the transfer layer. (The store side itself is transport-independent — that is what the Command-actor probe proved.) - Whole-file naming vs incremental writes. BLAKE3/bao names a blob
by the root hash of its complete byte sequence; you cannot compute
the address until all bytes exist. Partial data can live under temp
tags (
BlobStatus::Partial) but cannot be pinned by its final hash. The POC's branch-on-write/merge-on-close is one reconciliation; chunk lists / Merkle-DAG naming is the other (OQ-FS-04). git escapes the problem structurally: objects are small and independently hashed; packs are the exception, and pack generation streams to a consumer that hashes on arrival. - Two-database coexistence (redb for blob metadata + SQLite for the path tree): two WALs, two fsync paths, two crash-recovery stories — the POC summary's unknown #1, reframed by the probe ("neither fork nor coexist: a third store actor with its own single SQLite file"), and dissolvable entirely if alkfs owns blob metadata in its own store (OQ-FS-02/03).
- Sealed internals (
TempTags, theBaoTreeSender) — the probe's fork-vs-PR calculus. Third-party upstream, not ours, never will be (AGENTS.md convention 17).
git + git-lfs — the shape to optimize for
The user-level framing this crate must honor: "the same basic shape exists with git + gitlfs as with how iroh's blobs handled the mix of small and large blobs." git itself is the small-object CAS: every object individually hashed, refs are names over hashes, content shared across refs/branches for free, rename is metadata-only. git-lfs is the large-blob sidecar: content-addressed pointer rows in the tree, bytes in a separate content store (rudolfs is the family's reference server for that half — namespace/bucket isolation, LRU cache decorator, fanout streaming). The alkfs analogue: the metadata layer is git's object/ref shape; the blob layer is the LFS store; small content can inline into the metadata store exactly as iroh-blobs inlines small blobs into redb. Whether alkfs's engine API makes alkgit's gix-odb backend an implementation detail of alkfs (or leaves gix-odb native and serves only the general-FS shape) is OQ-FS-12 — the highest-stakes scope question in this document.
alkgit — the first consumer
/workspace/@alkdev/alkgit — paused on the storage problem. Its backend
doc (docs/architecture/backend.md) pins the seam: five traits
(GitRegistry, GitRegistryStore, GitRefs, GitPackGen,
GitPackIngest), kept small/orthogonal/substrate-blind after the
alknet-blobs sealed-surface lesson. The properties alkfs must serve
efficiently: streaming pack generation (O(counts) memory, blocking-thread
friendly), streaming pack ingestion (budgeted, fsck/connectivity report),
CAS ref transactions, registry records with grants. Its research
(docs/research/gitoxide.md, poc2-findings.md) carries the gix-odb
API contract notes — including the durability shape (pack bytes durable
before the ref transaction that names them commits) that mirrors
principle 3 exactly.
russh-sftp / alksftp — the door shape
/workspace/russh-sftp (read-only reference): the Handler trait is a
near-1:1 mirror of POSIX syscalls as SFTP packets, and the POC mapped
every op onto the PathTree + WriteSession API (open→resolve or
WriteSession::open; write→write_chunk with offset as key; close→hash+
merge+notify; readdir/list_dir; stat→indexed row — cheaper than a real
FS fstat; rename→O(1) edge move). The client's File already implements
AsyncRead + AsyncSeek + AsyncWrite with pipelined writes. alksftp (the
coming family crate) would speak this verb set; whether alkfs's serving
protocol adopts SFTP-shaped verbs natively or defines its own framing
that alksftp translates is OQ-FS-08. The alknet-tty async-IO adapter
patterns (AsyncWrite-over-mpsc, kill-on-drop) are reusable for the write
bridge.
SQLite (application file format) + honker
https://sqlite.org/appfileformat.html — the insight that anchored the
POC. honker (/workspace/honker, honker-core 0.2.4) supplies
notify/locks/queues as SQL functions on your connection — the
transactional-outbox property. Caveats: honker is third-party (not ours,
convention 17), single-machine by explicit design ("two servers writing
the same .db over NFS is not a Honker deployment strategy"), and its
scope in alkfs is an OQ (notify-on-commit may be the only load-bearing
piece — locks/queues/scheduler may be excess).
rudolfs — the git-lfs server reference
/workspace/@alkdev/alknet/docs/research/references/gitlfs/ rudolfs-reference.md: StorageKey = (Namespace, Oid) tenant isolation;
the decorator composition Verify ↔ Encrypted ↔ Cached ↔ Retrying(Disk → S3); LRU cache in front of permanent storage; fanout() streaming to
client and cache simultaneously. The cache/permanent split and the
fanout pattern are the production-shape references for alkfs's blob
layer and its serving read path.
Anti-prior-art (what NOT to carry over)
- The alknet mono-repo's scope. The filesystem research is inherited as findings, not as a code path; the crate stays small per the decomposition discipline (alksocks' "socks + channels, nothing else" resolution is the template — alkfs's version is OQ-FS-01).
- iroh-blobs' transfer protocol as the content path. Content transfer rides alkcall channels, full stop; an iroh/QUIC side-protocol would fork the family's transport story.
- iroh-blobs' FsStore partial-file lifecycle as a requirement. Its
749-line
bao_file.rspartial-blob-on-filesystem machinery exists because redb/filesystem is its substrate; a SQLite-inline store (POC write_chunks) or an alkfs-owned blob store has no need for it. - gix-odb's layout as alkfs's layout. Loose-objects + packfile directory layout is git's on-disk contract; alkfs is a general engine. If alkgit keeps gix-odb native, alkfs serves the general shape and git storage stays in alkgit (OQ-FS-12).
- NFS-style shared-file multi-node assumptions. honker's own warning applies to the whole engine: single-machine local state first; distribution is an explicit layer (OQ-FS-14), not an ambient property.
Open Questions
These are the design questions Phase 0 must resolve (or explicitly
defer) before the architecture spec. Numbered OQ-FS-01.. so they can be
referenced, tracked, and promoted into docs/architecture/ open-questions.md in Phase 1. Status is open unless stated; the point
of this document is to hold the reasoning without forcing premature
decisions. Grouped by theme; numbering is stable across reorganizations
(do not renumber — findings files and commits reference these ids).
Core architecture
OQ-FS-01: Crate scope — what lives in alkfs?
The component split (§Vision) is a sketch, not a decision. Questions to converge on:
- One crate with engine + serving protocol (feature-gated apart), or engine and protocol as separate crates (the alkcall registry/protocol split precedent)? The engine must be embeddable in-process without channels (alkgit's use case); the protocol must be usable without the default storage engine (an embedder with its own backend).
- Is there a backend trait seam at all (the alkgit
GitRefs-family pattern: traits in-crate, implementations behind features), or is the storage engine simply the implementation with the protocol generic over a small handle type? The sealed-surface lesson (alknet-blobs-external-store-probe.md) says: decide this before consumers exist; it is a one-way door. - What is explicitly out: doors (alksftp, FUSE, mount), sync tooling, UI/registry applications, the alknet rewrite's client stack.
OQ-FS-02: Metadata store substrate — and is blob metadata in the same file?
The POC used rusqlite (bundled) for the path tree and it worked well
(schema-as-documentation, recursive CTEs, BLOB columns at chunk
granularity). Open sub-questions:
- SQLite/rusqlite vs redb vs something else for the path tree. The POC evidence favors SQLite (honker integration, transactional outbox, sub-ms resolves); redb buys a lighter footprint but loses SQL and the honker seam.
- The one-file-vs-two question: if alkfs owns the blob layer too, blob metadata (hash index, sizes, refcounts/liveness) can live in the same SQLite file as the path tree — one WAL, one commit boundary, and the durability ordering of principle 3 becomes a single transaction instead of a cross-store fsync dance. The probe's external-actor shape (iroh-blobs with its own store) vs an alkfs-owned store decides whether this is even on the table. Small content can then inline as BLOBs (the appfileformat sweet spot) — possibly making the "blob store" a layer inside one database with file-backed spill-over for large content (the iroh-blobs FsStore hybrid shape).
- Concurrency model: single writer connection + reader pool? WAL
settings (
synchronous=NORMALvsFULL) as a durability-vs-throughput knob? What is the honest fsync contract per operation class?
OQ-FS-03: Blob substrate — reuse iroh-blobs, own the store, or hybrid?
Three live options, shaped by the probe:
- (a) Consume iroh-blobs as a store (implement the
Commandactor externally per the probe; 4-line upstream PR; ~120-160 lines of bao glue) — gets BLAKE3+bao verified streaming, GC, tags for free; pays the iroh-blobs/irpc dependency weight, the sealed-TempTagsfallback (self-managed liveness), and still needs the metadata-store coexistence story (OQ-FS-02). Also carriesbao_treeversion-tracking as the maintenance surface. - (b) Own the blob store. BLAKE3 (
blake3crate) + bao (bao_treecrate) directly, alkfs-owned storage layout — full control of durability ordering (principle 3 becomes enforceable in one transaction with OQ-FS-02's single-file shape), no sealed surfaces, no iroh-blobs dependency. Cost: re-derive verified-streaming import/ export (the glue the probe measured), GC, partial states — the pieces iroh-blobs already solved, now ours to keep correct. - (c) Hybrid/no-bao start. Whole-file BLAKE3 + a simple file-per-blob or SQLite-inline store first, add verified streaming (bao outboards) when a consumer needs range verification. git's own objects are small; packs are written-once-then-verified-by-consumer. Honest question: does any alkfs v1 consumer actually need per-chunk bao verification on the read path, or is end-to-end TLS + hash-check-on-close sufficient?
The deciding inputs: the alkgit workload (pack-sized blobs, streaming), the sftp workload (range reads, seek), and how much of iroh-blobs' surface survives contact with "no iroh/QUIC transfer."
Write path, branching, GC
OQ-FS-04: Hashing and chunking — whole-file addresses vs chunk DAGs
The load-bearing storage-format question (a one-way door once content exists in the wild):
- Whole-file BLAKE3 (+ bao outboard for verified streaming). Simple, git-shaped, matches the POC and alkgit's objects; but the address doesn't exist until the write completes (staging required — the branch-on-write answer), no partial availability under a final name, no cross-file chunk dedup (two 1GB files sharing a 500MB prefix dedup zero bytes).
- Chunk-list / Merkle-DAG addressing (content-defined or fixed-size chunks; the file "name" is a root hash over chunk hashes, iroh-blobs style): dedup at chunk granularity, partial availability, resumable writes under a stable name, cheap append; but bigger metadata, a chunker to pin down (CDC boundary-shift sensitivity), and whole-file identity becomes "same root hash" rather than "same bytes hash" (root hash equality still implies byte equality with bao-style trees, at hash-of-chunk-hashes cost).
- Hybrid (the git-lfs/iroh shape): small content inline/whole-hash; large content chunked. Threshold policy is a knob; the format split is the one-way door.
Sub-questions: CDC (rollsum/buzhash) vs fixed-size vs BLAKE3's native
chunking (bao_tree's chunk size); where chunk indices live (metadata
store rows vs outboard files); whether chunk dedup across files is a
v1 requirement or a defer-able nicety; interaction with OQ-FS-05 (a
chunk-DAG write can commit incrementally, weakening the branch-on-write
staging argument).
OQ-FS-05: Write path semantics beyond the POC
Branch-on-write/merge-on-close is POC-proven for the happy path. What Phase 0 must still answer:
- Durability details — per-chunk transaction commit (POC shape) vs
batched WAL;
synchronouslevel; whether fsync of spilled large content must precede the path-row commit (principle 3) and how that is ordered when content lives outside the DB file; crash-mid-close behavior. - Concurrency — the POC left honker named locks unwired; concurrent writers to the same path need a real policy (last-close-wins is session-safe but not POSIX; O_EXCL-ish create semantics; advisory byte-range locks for the sftp door?).
- Streaming large writes — pack generation and 1GB+ uploads: the POC's per-chunk SQLite rows are proven at ~32KB×1MB; the pack-size regime needs a spill story (OQ-FS-02's file-backed overflow) and a backpressure story inherited from alkcall channels.
- Sparse/out-of-order/ftruncate — SFTP pipelines out-of-order (proven); sparse files and explicit truncation/extend are unaddressed.
OQ-FS-06: Branch/snapshot/commit model
The POC has Fossil-style branches (name → parent chain + deltas) but no explicit commit/snapshot op — writes to a branch are immediately visible on that branch. Open questions:
- Is a snapshot a new branch (Fossil's model — maps naturally to the chain walk) or a recorded point-in-time within a branch (git's model)? What do alkgit (refs + CAS) and the sftp door each need?
- Refs as a first-class concept: alkfs stores
refsrows naming commits (branch heads), enabling CAS ref transactions (alkgit'sGitRefs::applyshape) over the same metadata store. Is the ref namespace git-shaped from day one, or a later addition? - Merge semantics across branches: the POC has none (branch-on-write collapses trivially at close). Real merges (three-way? automerge-CRDT for the metadata layer per POC 3?) are presumably out of v1 — say so explicitly rather than leaving it implied.
- Chain-depth performance: the recursive CTE degrades on deep chains (POC unknown #6); resolve-cache/materialized-view shape if needed — perf probe before spec, or accept and document.
OQ-FS-07: GC and liveness
Content is shared across branches/paths by hash, so naive per-path tags (Poc unknown #5's iroh-tags shape) break sharing — deleting a path must not collect content another path/branch still names. The candidate shape: liveness = reachability walk (mark from all live branch heads
- refs + write sessions; sweep unreachable) — a metadata-store query over path rows, no separate tag table needed if blob liveness rows are refcounted or mark-computed at GC time. Sub-questions: incremental vs stop-the-world GC; tombstone retention (content behind tombstones is unreachable-but-recent — grace windows); orphaned write-session chunk reaping (crash cleanup, the POC's known orphans); cross-bucket sharing (never — buckets are tenants).
Serving protocol
OQ-FS-08: The file protocol over channels — verbs, sessions, framing
The biggest unspecced surface, and the crate's first real wire format (one-way door — ADR before the first consumer):
- Verb set. SFTP-shaped (the russh-sftp Handler mapping is proven
1:1 against the engine) vs a leaner custom set (open/read/write/
close/stat/readdir/rename/remove + extended ops). Working posture:
SFTP-shaped verbs, SFTP-versioned semantics where it matters
(handles, pipelined writes,
SeekFrom::Endcheapness). - Session model. One channel per file session (alksocks' one-channel- per-CONNECT shape) vs one control channel multiplexing many file handles (the alktty demux shape). A directory traversal or a git clone opens many files; per-file channels spend channel IDs and opens (the 256-cap arithmetic, OQ-SK-09's twin); a multiplexed control channel reintroduces the demux trade alksocks resolved for SOCKS5 (whose RFC structure genuinely has one command per connection — a file protocol has no such constraint). This trade has different physics here than it had there; no inherited answer.
- Framing on the data path. If sessions are multiplexed, requests need IDs and payloads need framing (length-prefix + request-id — the one-way door). If per-session, the payload may be near-pass-through. Chunk sizes, backpressure (bounded buffers are inherited), and flow-control for pipelined writes all live here.
- Read-path streaming. Range reads (offset+len), EOF semantics, verified-streaming vs trust-the-store; zero-copy ambitions explicitly out (channels are the substrate, not splice(2)).
- Establishment.
register_openable_with_establisherwith a scope gate (OQ-FS-10); refused opens are typed errors with per-verb error mapping (SFTP status codes at the door, faithful engine errors in protocol — convention 2).
OQ-FS-09: ALPN, params, and the resource model
Provisional alk/fs (final naming per alkcall ADR-004/006 convention).
Open sub-questions (mostly Phase 1 ADRs with Phase 0 ground):
- The params-object shape (ADR-039 precedent: ALPN-specific, interpreted by the open handler) — what identifies a produced filesystem resource? (bucket? branch? scope string? limits?)
- In-process and loopback consumers: the engine half must be directly
embeddable (no channels) — the protocol layer is one substrate among
several, generic over
T(principle 6). The in-process shape is also the test seam, the alksocks POC pattern. - Hub relay story: a filesystem resource traverses relays terminate-and-re-produce like any produced resource; nothing alkfs-specific expected — verify during POC rather than assume.
OQ-FS-10: Doors — alksftp, mounts, sync (out of crate, but shapes the API)
Family rule: doors are separate crates/features. Phase 0's job is to
keep the protocol layer door-ready: the verb set (OQ-FS-08) must map
onto russh-sftp's Handler without translation loss; a FUSE/mount door
needs POSIX-ish error mapping and inode-ish handle semantics — do not
design for FUSE in v1, but record what it would demand so the protocol
doesn't preclude it. Directory-sync tooling (alkfs ↔ host dir) is the
explicit, feature-gated host-FS bridge if/when wanted — never ambient.
OQ-FS-11: Multi-tenancy, identity, and path-scope policy
Buckets are the isolation unit (POC-proven free). Open questions: the
ALKFS_OPEN_SCOPE-shaped scope gate (alktty/alksocks precedent) — does
scope name buckets, path prefixes, or both; identity → bucket mapping
at the assembly layer (vault references, never plaintext secrets —
family rule); per-path/per-bucket policy as a producer-side policy
callback vs declarative ACL (alksocks OQ-SK-05's dialer-refuses
resolution is the evidence-backed template); write-vs-read grant
separation (a read-only bucket mount is the obvious first split).
Consumer fit
OQ-FS-12: alkgit fit — the question that un-pauses alkgit
The first consumer and the reason this crate moved ahead of alkgit. The decision tree:
- (a) alkgit keeps gix-odb native; alkfs serves the general FS. Lowest integration risk; but then alkfs doesn't solve the problem that paused alkgit, and the family carries two storage stories.
- (b) alkgit's backend traits get an alkfs-backed implementation.
GitPackGenstreams a pack out of alkfs content (pack assembly is alkgit-side; the blob bytes come from alkfs reads);GitPackIngeststreams a pack into alkfs (each unpacked object lands as content; refs land as metadata rows);GitRefsbecomes CAS transactions over alkfs refs. gix-odb's on-disk layout disappears; git semantics ride alkfs's content addressing. This is the "optimize for git" path and the probable answer — but it makes alkgit's workload (pack-scale blobs, O(counts) memory, streaming) the benchmark OQ-FS-13 must meet. - (c) gix-odb on top of alkfs (an odb backend where loose objects/packs are alkfs files) — keeps gix-odb's pack machinery but treats alkfs as the disk. Probably the worst of both: gix-odb assumes filesystem semantics alkfs won't natively provide (mmap, file locks).
Sub-questions: does git's object-id namespace (sha1/sha256) coexist with alkfs's content addresses (path rows mapping git-oids → alkfs hashes, dedup at the alkfs layer); CAS ref transaction shape; whether pack ingestion budget enforcement (alkgit ADR-009) maps onto alkfs write sessions.
OQ-FS-13: Performance shape and budgets
No numbers exist yet. What Phase 0 should collect (empirically, via the POC register) so Phase 1 can pin budgets: path-resolve cost at realistic branch depths (POC: sub-ms in-memory; chains of 10+?); small-file throughput (the inline-BLOB sweet spot claim vs reality at 4KB metadata ops); large-file write throughput through chunked sessions; pack gen/ingest streaming rates over a channel vs native gix-odb; memory bounds under the alkcall channel caps. The alkgit ADR-009 limits shape is the consumer contract to serve.
Posture and deferred
OQ-FS-14: Distributed sync / multi-node — in scope at all?
POC 3 proved automerge-CRDT sync of the path tree (9 tests, LWW conflicts, the concurrent-root-init constraint) with content fetched separately. But sync is a product-layer concern that may not belong in the engine crate: the family pattern is composition (alkcall relay for transport, automerge as a third-party sync layer, honker queues as the local outbox). Working posture: deferred — out of v1 scope, recorded so the metadata schema doesn't preclude it (branch/paths rows are already CRDT-encodable per POC 3's document structure). Revisit when a consumer needs multi-node.
OQ-FS-15: WASM posture
The expected shape (convention 4): wasm-clean protocol/ops layer; native
storage engine behind features (real FS/SQLite never compiles to
wasm32-unknown-unknown). Open: does any consumer actually want the
protocol layer in wasm (a sandboxed alkfs consumer, mirroring alksocks'
sandboxed-SOCKS-consumer story)? If none is in sight, the honest answer
may be alkhttp-style native-only (the "wasm-clean is preferred, not
mandatory" precedent) — but the protocol layer being wasm-clean by
construction costs little if the engine split exists anyway. Decide
deliberately once the split (OQ-FS-01) exists; verify with
cargo check --target wasm32-unknown-unknown.
Candidate POC register (proposals — none run yet)
POC placement conventions (inherited from alktunnels/alksocks): a POC
needing this repo's code runs in a worktree (.worktrees/research/ <task-id>/ per docs/sdd_process.md); a self-contained POC runs as a
standalone crate in the global workspace with findings written into
docs/research/ here. Findings always land in docs/research/.
| # | Candidate | Hypothesis to validate | Feeds |
|---|---|---|---|
| 1 | Engine skeleton: path tree + inline-blob store in one SQLite file | One-file metadata+blob-metadata (OQ-FS-02) holds up under the POC's 15-test suite ported over; single-transaction durability ordering (principle 3) is expressible | OQ-FS-02, OQ-FS-05 |
| 2 | Write path at pack scale | Chunked write sessions handle 100MB+ streams (spill-to-file regime) with correct durability ordering and crash-abort semantics | OQ-FS-05, OQ-FS-13 |
| 3 | Blob substrate bake-off (narrow) | (a) iroh-blobs external actor vs (b) own blake3/bao store vs (c) inline-hybrid: measure integration cost, GC story, sealed-surface fallout against the probe's predictions | OQ-FS-03 |
| 4 | Chunking probe (paper + micro-bench) | Whole-file vs chunk-DAG: dedup gains on realistic workloads (git packs, agent workspaces), write-path impact, metadata size — enough to decide the format one-way door or at least stage it | OQ-FS-04 |
| 5 | Serving protocol spike | SFTP-verb protocol over channels: per-file-channel vs multiplexed-handles session model, open cost at directory-walk and git-clone fan-out, framing sketch (not a wire ADR) | OQ-FS-08, OQ-FS-09 |
| 6 | alkgit seam spike | GitPackGen/GitPackIngest shaped against an alkfs-backed engine on the POC's API — does the trait family survive contact; where do git oids map |
OQ-FS-12 |
| 7 | GC walk prototype | Reachability-mark GC over path rows + refs: correctness on shared content across branches, cost model, orphaned-session reaping | OQ-FS-07 |
Order is indicative, not committed: #1/#2 de-risk the engine core; #3/#4 resolve the storage-format one-way doors; #5/#6 are only worth running once #1 exists (the protocol needs an engine to serve). The register exists so POC work has home ids; expect entries to be split, merged, or dropped as OQs sharpen.
Survey / prior-art list
Read/verified (cited above):
- alknet-filesystem POCs —
/workspace/@alkdev/alknet/docs/research/alknet-filesystem/ poc-summary.md,alknet-blobs-external-store-probe.md; POC crates/workspace/alknet-filesystem-poc,/workspace/alknet-fs-sync-poc. - iroh-blobs —
/workspace/iroh-blobs(v0.100 checkout +DESIGN.md), published 0.103 in cargo cache (probe source); the family's reference notes at/workspace/@alkdev/alknet/docs/research/references/iroh/ iroh-blobs/(overview, key types, transfer protocol, storage). - alkgit —
docs/architecture/backend.md(the trait seam),docs/research/{vision,gitoxide,git-protocol,poc2-findings}.md. - russh-sftp —
/workspace/russh-sftp(server/handler.rs, clientFilepipelining). - SQLite appfileformat — https://sqlite.org/appfileformat.html; honker
—
/workspace/honker(honker-core0.2.4). - rudolfs (git-lfs server) —
/workspace/@alkdev/alknet/docs/research/references/gitlfs/ rudolfs-reference.md. - alkcall — ADRs 034/035/036/037/039/042/049/050/051, ledger
CF-005/CF-006 (
docs/architecture/decisions/); alktty/alktunnels/ alksocks Phase 0 docs as the format/template precedents.
To evaluate (research-specialist queue, feeding the OQs named):
- Chunking/CDC prior art — FastCDC, rollsum (bup), rsync-style
boundaries;
bao_tree's BLAKE3 chunking as the default candidate (OQ-FS-04). - CAS filesystems — Plan 9 fossil+venti (the branch/CAS ancestry the POC name-checks), IPFS unixfs/iros DAG shapes, OST/lessfs-style inline-vs-outboard hybrids (OQ-FS-04, OQ-FS-06).
- git storage internals — gix-odb pack/idx layout and
gix-pack::data::inputstreaming (alkgit research has the survey; re-read for OQ-FS-12's object-id/namespace questions). - SQLite durability posture — WAL +
synchronous=NORMALvsFULLunder the commit-boundary contract; SQLite-as-blob-store limits (OQ-FS-02, OQ-FS-05). - FUSE/mount doors — rust bindings shape, POSIX semantics alkfs would need to fake (OQ-FS-10; defer design, record constraints).
Convergence checklist (what Phase 0 must produce)
- Vision + guiding principles captured (this doc, §Vision — the scope sketch needs confirmation or cutting, OQ-FS-01)
- Prior-art pass complete: alknet-filesystem POCs re-read and mapped (done); iroh-blobs surface re-verified against the probe's 0.103 findings; git/git-lfs shape pinned against alkgit's backend contract; CDC/chunking survey (OQ-FS-04 input); CAS-filesystem survey
- OQ-FS-01 (crate scope) — converged: engine/protocol split, trait seam or not, explicit non-goals
- OQ-FS-02 (metadata store) — converged: substrate, one-file-vs-two (blob metadata placement), durability knob
- OQ-FS-03 (blob substrate) — converged: consume iroh-blobs vs own the store vs hybrid (POC #3)
- OQ-FS-04 (hashing/chunking) — converged: the storage-format one-way door (POC #4)
- OQ-FS-05 (write path) — durability/concurrency/spill semantics pinned (POC #2)
- OQ-FS-06 (branch/snapshot/refs) — model chosen; ref namespace decided; merge explicitly in-or-out of v1
- OQ-FS-07 (GC) — liveness shape chosen; retention/reaping policy pinned (POC #7)
- OQ-FS-08 (file protocol) — verbs, session model, framing sketch; the wire ADR itself is Phase 1's (one-way door, first consumer in sight) (POC #5)
- OQ-FS-09 (ALPN/params) — naming + params ground collected; ADR in Phase 1
- OQ-FS-10 (doors) — protocol kept door-ready; nothing designed in
- OQ-FS-11 (tenancy/policy) — scope-string + policy shape settled
- OQ-FS-12 (alkgit fit) — the (a)/(b)/(c) decision made; alkgit un-pause plan written (POC #6)
- OQ-FS-13 (budgets) — numbers collected; budget table drafted
- OQ-FS-14 (sync) — explicitly deferred or scoped with evidence
- OQ-FS-15 (wasm) — decided once the engine/protocol split exists
- Targeted POCs run + findings in
docs/research/ - Converge: recommended approach written up, ready to hand to the Architect for Phase 1