From 8692f9748e31ae9176d9b691b71f5cd43598edd5 Mon Sep 17 00:00:00 2001 From: "glm-5.2" Date: Sat, 11 Jul 2026 08:26:56 +0000 Subject: [PATCH] =?UTF-8?q?docs(research):=20flesh=20out=20alknet-channels?= =?UTF-8?q?=20=E2=80=94=20hub=20motivation,=20channel=20open=20negotiation?= =?UTF-8?q?,=20channel=20manager=20internals?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds three sections filling the conceptual gaps in the phase-0 findings: - Hub Motivation: The Multi-Transport Collapse — diagnoses the O(protocols × transports × spokes) mess the hub crate faces and shows how channels collapses it to one connection per leg with channel-by-channel byte forwarding, reusing the call protocol's auth/forwarded-for model. - Channel Open Negotiation — concrete channel/open, channel/close, channel/control, channel/resources operation payloads with field tables, error codes as CallError strings, bidirectional open semantics, and the end-to-end ACL flow for browser→hub→spoke. No new wire framing; all four operations register on the call protocol's existing OperationRegistry. - Channel Manager and Connection Internals — the ChannelsAdapter/ ChannelManager split, ChannelManager state sketch, ChannelsAdapter::handle read/demux loop with channel 0 preinstall, the channel/open handler closure threading into OperationRegistry, ChannelConnection as a Connection via the existing from_stream path, the hub relay as ChannelManager-to-ChannelManager byte pumping, and the boundary (ChannelManager holds no handlers, no ALPN parsing, no auth, no transport coupling). Also adds four new open questions (OQ-CH-08 through OQ-CH-11) on resource staleness, responder-to-initiator lifecycle, typed destructure ownership, and hub relay channel_id remapping; two new POC stretch goals (hub relay, channel/resources); and updates the opening to note the revision scope. --- .../alknet-channels/phase-0-findings.md | 743 +++++++++++++++++- 1 file changed, 742 insertions(+), 1 deletion(-) diff --git a/docs/research/alknet-channels/phase-0-findings.md b/docs/research/alknet-channels/phase-0-findings.md index 4c888be..f2705d8 100644 --- a/docs/research/alknet-channels/phase-0-findings.md +++ b/docs/research/alknet-channels/phase-0-findings.md @@ -1,6 +1,6 @@ --- status: draft -last_updated: 2026-07-10 +last_updated: 2026-07-11 --- # alknet-channels — Phase 0 Research Findings @@ -20,6 +20,14 @@ crate's chunk format already solves sub-stream multiplexing for terminal sessions — this document generalizes that pattern into a universal channel multiplexer that the call protocol can orchestrate. +The 2026-07-11 revision adds the hub motivation (§Hub Motivation), the +concrete channel-open negotiation protocol (§Channel Open Negotiation), and +the channel-manager/connection internals (§Channel Manager and Connection +Internals) — driven by the realization that the hub crate's multi-transport +complexity is what channels specifically collapses, and that the call +protocol's auth/operation-overlay model carries over to channel lifecycle +with no new machinery. + ## Vision Recap `alknet-channels` is a **multiplexing proxy** — a `ProtocolHandler` on @@ -98,6 +106,143 @@ bidirectional — either side can open a channel to the other, just like the call protocol's operation overlay where each side populates what operations they can call. +## Hub Motivation: The Multi-Transport Collapse + +This design was forced by the hub crate. A hub is the architectural role +(ADR-029, ADR-034) that bridges peers and browsers: it terminates QUIC from +spokes, terminates WebTransport/WSS from browsers, and may itself be a spoke +of an upstream hub. The hub's job is to **route and relay**, not to +re-implement every protocol's framing per transport. + +Without channels, the hub quickly becomes a mess: + +### The mess, concretely + +A hub serving a browser over WebTransport and a spoke over QUIC wants to +offer the browser three things on the same logical session: + +1. Call operations on the spoke (e.g., `docker/container/list` — JSON). +2. A TTY session on the spoke (interactive exec — raw bytes, ADR-052 chunk + format). +3. An SSH session to the spoke (raw bytes, SSH binary protocol). + +Today each of these is a **separate ALPN** on a **separate connection**, and +each transport multiplexes differently: + +| Need | Today's mechanism | Transport | Multiplexing | +|------|-------------------|-----------|--------------| +| JSON ops | `alknet/call` ALPN | QUIC stream or WSS | one op per bidi stream | +| TTY | `alknet/tty` ALPN | QUIC stream or WSS | TTY chunk format *within* one bidi stream | +| SSH | `alknet/ssh` ALPN (future) | QUIC stream or WSS | SSH multiplexes *within* one bidi stream | + +So the hub faces a matrix: **3 needs × 2 transports × N spokes**, and the +multiplexing layer is different in each cell. The hub would have to: + +- Maintain a `quinn::Connection` per spoke AND a WebTransport session per + browser, AND decide which one carries which ALPN. +- For a browser→spoke TTY, the hub can't just "forward a stream" — the + browser's WSS stream carries `alknet/tty` chunks, and the spoke's QUIC + stream also carries `alknet/tty` chunks, so the hub *can* relay bytes — + but it had to negotiate two separate `alknet/tty` connections (one per + leg) and correlate them by out-of-band state. There's no "one session that + spans the relay." +- For a browser→spoke SSH, same problem with a different ALPN and a + different internal multiplexer. +- The call protocol that *orchestrates* all this lives on a **third** ALPN + (`alknet/call`), so the hub is juggling three connections per (browser, + spoke) pair, each with its own framing, and the correlation across them is + implicit. +- Every new protocol (tunnel, future file-transfer, future git-pack) adds + **another column** to the matrix and **another ALPN** the hub must + register, dispatch, and relay. + +This is the "PITA to maintain a bunch of connections" problem: the hub's +complexity is **O(protocols × transports × spokes)**, when it should be +**O(spokes)**. + +### What channels collapses + +With `alknet/channels`, the hub's per-spoke and per-browser state is **one +channels connection each**, and everything else is routing: + +``` +Browser ──WebTransport──► Hub ──QUIC──► Spoke + alknet/channels alknet/channels + ┌─────────────┐ ┌─────────────┐ + │ ch0: call │ │ ch0: call │ + │ ch1: tty │ relay │ ch1: tty │ + │ ch2: ssh │ ◄─────► │ ch2: ssh │ + │ ch3: tunnel │ │ ch3: tunnel │ + └─────────────┘ └─────────────┘ +``` + +The hub's relay logic is **channel-by-channel byte forwarding**: read chunks +for `(channel_id)` off the browser's channels connection, write the same +payload (with the spoke's `channel_id` substituted) onto the spoke's +channels connection. The hub does **not** parse `alknet/tty` chunks, does +**not** understand SSH's internal multiplexer, does **not** run a separate +`CallAdapter` per leg — it forwards opaque `(channel_id, stream_type, +payload)` tuples and lets the endpoints at each end do the protocol work. + +The collapse is at three levels: + +1. **One connection per leg, not one per protocol.** The browser holds one + WebTransport session to the hub; the hub holds one QUIC connection to the + spoke. All three needs (call, TTY, SSH) ride as channels on those two + connections. The hub correlates by channel, not by ALPN-per-connection. +2. **One multiplexing model, not three.** Connection-level (ALPN router), + stream-level (QUIC native), and sub-stream-level (TTY chunks) collapse + into one: channels chunks. The hub's relay code is one loop, not three. +3. **The call protocol orchestrates from inside.** Channel 0 is `alknet/call` + on both legs. The browser calls `channel/open` on its channel 0; the hub + forwards that call to the spoke's channel 0 (via `from_call`); the spoke + allocates the channel and returns the ID; the hub opens a matching channel + on the browser's side and bridges them. The hub never runs a handler for + `alknet/tty` or `alknet/ssh` — it only runs `alknet/channels` (the relay) + and `alknet/call` (for its own hub-level operations like routing and + resource lookup). + +### Why the auth model reuses cleanly + +The call protocol's `OperationContext` (identity, scopes, capabilities, +ownership) already gates every operation. `channel/open` is just another +operation on channel 0, so the same `OperationContext` flows through: + +- The browser's `channel/open` carries the browser's identity (bearer token + resolved by the hub per ADR-043). +- The hub forwards via `from_call`, which populates `forwarded_for` from the + hub's `OperationContext.identity` (ADR-032 §3) — exactly the kernel/user- + land + forwarded-for model from ADR-050. +- The spoke's `AccessControl::check` sees the hub as the caller and the + browser as `forwarded_for`. The spoke authorizes the hub (its direct + peer); the hub's "who is this for" is its own app state, carried as + `forwarded_for`. + +No new auth machinery. The hub doesn't authenticate channels — it +authenticates the call operations that *open* channels, which it already +does. The channels layer inherits the call protocol's ACL by being one of +its operation types. + +### What the hub *does* still own + +Channels does not eliminate the hub's responsibilities; it relocates them: + +- **Routing**: which spoke serves `container:abc123`? That's the hub's + resource registry / ownership store (ADR-050), queried via call operations + on channel 0. Channels doesn't touch this. +- **ACL at the hub**: does this browser's identity have `channel:open` scope + for `alknet/ssh` to `spoke-X`? That's the call protocol's + `AccessControl::check` on the `channel/open` operation, run by the hub's + `CallAdapter` before it forwards. Channels doesn't touch this either. +- **Relay lifecycle**: when a browser disconnects, the hub tears down the + spoke-side channels (and vice versa). This is `channel/close` on each + channel, or a transport-level close that the channels layer observes. + +What the hub *no longer* owns: per-protocol framing parsers, per-protocol +relay loops, per-ALPN connection management, and the correlation state +across multiple connections per (browser, spoke) pair. Those move into the +channels layer, which is transport-agnostic and protocol-agnostic. + ## The Wire Format ### Chunk header @@ -240,6 +385,549 @@ recommendation is to allow it but not encourage it — the primary use case is one level of multiplexing. Recursive composition is a natural consequence of the `Connection` abstraction, not a feature to design for. +## Channel Open Negotiation + +Channel lifecycle is orchestrated by the call protocol on channel 0. The +call protocol is JSON-only — it carries `channel/open`, +`channel/close`, `channel/control`, and `channel/resources` as **new +operation types** in the existing `OperationRegistry`, dispatched through +the existing `OperationContext` (identity, scopes, capabilities, ownership, +`forwarded_for`). No wire-format change to `EventEnvelope`, no new carriage +type. The channels layer registers these operations on the call protocol's +`OperationRegistry` at assembly time. + +This is where the call protocol's auth model is reused: `channel/open` is +gated by `AccessControl::check` exactly like any other operation, and the +operation's `AccessControl` can declare `required_scopes` (e.g. +`channel:open:alknet/tty`), `resource_type` + `resource_action` (e.g. +`container` + `tty`), or `required_scopes_any` for multi-scope gates. The +channels layer doesn't re-implement auth — it leans on the +`OperationRegistry::invoke` path that already runs the check before the +handler runs. + +### `channel/open` — request a channel + +Either side sends a `call.requested` on channel 0 with operation name +`channel/open`: + +```json +{ + "type": "call.requested", + "id": "req-abc", + "payload": { + "operation": "channel/open", + "input": { + "alpn": "alknet/tty", + "stream_types": [0, 1, 2, 3], + "params": { + "backend": "docker", + "cmd": ["bash"], + "container": "abc123" + }, + "direction": "initiator-to-responder" + } + } +} +``` + +Fields: + +| field | type | meaning | +|-------|------|---------| +| `alpn` | string | The ALPN the channel will carry, e.g. `alknet/tty`, `alknet/ssh`, `alknet/tunnel`. The responder looks this up in its `HandlerRegistry`. | +| `stream_types` | `[u8]` | Which sub-stream types this channel will use. Declared at open time so both sides size their reassembly buffers and know which `(channel_id, stream_type)` pairs are valid. E.g. `[0, 1, 2, 3]` for TTY, `[0, 1]` for a raw tunnel. | +| `params` | object | ALPN-specific parameters passed to the handler. For `alknet/tty` this is the `NegotiateRequest` (backend, cmd, env, size). For `alknet/tunnel` this is the target resource (`{"resource": "container:abc123", "port": 5432}`). For `alknet/ssh` this is the SSH auth/method hints. The channels layer does not interpret `params` — it hands the JSON to the handler's allocate entry point. | +| `direction` | string | `initiator-to-responder` (the initiator wants the responder to expose a resource) or `responder-to-initiator` (the initiator wants to expose a resource *to* the responder — the "worker exposes, hub consumes" case). See "Bidirectional open" below. | + +The `direction` field is what makes channel open match the call protocol's +operation-overlay symmetry: just as either side can populate operations the +other can call, either side can expose resources the other can open +channels to. `initiator-to-responder` is the common case ("open me a TTY on +your docker container"). `responder-to-initiator` is the reverse ("I'm a +worker exposing my local Postgres as a tunnel target; hub, you can open a +channel to it") — which is how a hub consumes a worker's resources without +the worker initiating. + +### `channel/open` — response + +The responder validates ACL, looks up the ALPN in `HandlerRegistry`, and — +crucially — **allocates the `channel_id` before returning** (DP-1: server- +assigned). The response is a normal `call.responded`: + +```json +{ + "type": "call.responded", + "id": "req-abc", + "payload": { + "output": { + "channel_id": 7, + "stream_types": [0, 1, 2, 3] + } + } +} +``` + +| field | type | meaning | +|-------|------|---------| +| `channel_id` | u32 | The server-assigned channel ID. Both sides now know to route chunks with this `channel_id` to the new channel. | +| `stream_types` | `[u8]` | The *negotiated* set — the responder may narrow the initiator's requested set (e.g. refuse stderr for a backend that doesn't produce it). The intersection of requested and supported. | + +Once the initiator receives the response, data may flow on +`(channel_id=7, stream_type=*)`. The responder's handler is already +allocated and reading from its reassembled streams. + +### `channel/open` — error cases + +Errors use the call protocol's existing `CallError` shape, dispatched as +`call.error`. The channels layer defines a small set of error codes: + +| code | meaning | retryable | +|------|---------|-----------| +| `channel:unknown_alpn` | The `alpn` is not in the responder's `HandlerRegistry`. | false | +| `channel:forbidden` | `AccessControl::check` denied the open (missing scope, not the resource owner, `forwarded_for` not trusted). | false | +| `channel:allocation_failed` | The handler's allocate failed (e.g. `DockerTtyBackend::allocate` couldn't start the exec). The `details` carry the backend's error message. | true (often transient) | +| `channel:invalid_params` | The `params` JSON didn't satisfy the ALPN's expectations (e.g. missing `cmd` for TTY). | false | +| `channel:too_many_channels` | The responder hit its per-connection channel limit (OQ-CH-06). | false | +| `channel:stream_type_unavailable` | The responder's handler can't provide one of the requested `stream_types`. The `details` carry the supported set. | false | + +These are new `CallError.code` strings, not new framing. The call protocol +already carries `code`/`message`/`retryable`/`details`. + +### `channel/close` — tear down a channel + +Either side sends: + +```json +{ + "operation": "channel/close", + "input": { "channel_id": 7, "reason": "exit" } +} +``` + +The responder (the side that *didn't* send the close) drains its +reassembled streams for `channel_id`, signals EOF to the handler, and +returns `call.responded` with `{ "closed": true }`. The `channel_id` is now +eligible for reuse (OQ-CH-04). `reason` is a free-form string for +observability — `"exit"`, `"cancel"`, `"error"` — and is not semantically +required. + +The ordering invariant (ADR-055 for TTY): the channel's **data chunks** must +be written and flushed before the `channel/close` operation is sent on +channel 0. This is an implementation constraint on the side closing — the +channels layer's close handler must observe the data-channel pump complete +before issuing the call operation. For TTY this is the exit-chunk-is-last +invariant carried forward. + +### `channel/control` — out-of-band control on channel 0 + +For control that doesn't need ordering relative to data (resize, signal, +keepalive), a call operation on channel 0: + +```json +{ + "operation": "channel/control", + "input": { + "channel_id": 7, + "stream_type": 3, + "message": { "type": "resize", "cols": 80, "rows": 24 } + } +} +``` + +The channels layer routes `message` to the handler's control handle for +`channel_id`. This is the "call protocol for orchestration" half of DP-4. +The `message` JSON is ALPN-specific; the channels layer doesn't interpret +it — it hands it to the handler's control entry point, same as `params` on +open. + +### `stream_type 3` — in-band control on the data channel + +For control that **does** need ordering relative to data (EOF before exit, +flush before close), a chunk with `stream_type = 3` on the data channel +itself: + +``` +[channel_id: u32 be][stream_type: 0x03][length: u32 be][json payload] +``` + +This is the "stream_type 3 for data-ordered control" half of DP-4. The +handler parses the JSON; the channels layer just reassembles and delivers +it in-order with the data. The TTY crate's `ControlMessage::Eof` is the +canonical example — it must arrive after the last stdin chunk, which is +guaranteed by chunk ordering within `(channel_id, stream_type)`, not by a +call-protocol round-trip. + +### `channel/resources` — populate the resource overlay + +This is the conceptual gap the doc is closing. The call protocol already +has `services/list` and `services/list-peers` for discovering **operations** +each side exposes. Channels needs the equivalent for **resources** — what +ALPNs each side is willing to open channels for, and with what constraints. + +A side calls `channel/resources` on channel 0 to ask the other side what it +exposes: + +```json +{ + "operation": "channel/resources", + "input": {} +} +``` + +Response: + +```json +{ + "output": { + "resources": [ + { + "alpn": "alknet/tty", + "backends": ["docker", "local"], + "access": { "required_scopes": ["tty:open"] } + }, + { + "alpn": "alknet/ssh", + "access": { "required_scopes": ["ssh:open"] } + }, + { + "alpn": "alknet/tunnel", + "targets": ["container:*", "service:postgres"], + "access": { "required_scopes_any": ["tunnel:open", "admin"] } + } + ] + } +} +``` + +| field | type | meaning | +|-------|------|---------| +| `alpn` | string | The ALPN this side will accept `channel/open` for. | +| `backends` / `targets` | `[string]` | ALPN-specific enumeration of what's available — TTY backends, tunnel target resource patterns. The channels layer doesn't interpret these; they're for the *initiator* to know what `params` to send. | +| `access` | object | A preview of the `AccessControl` that `channel/open` will check. Lets the initiator fail fast (e.g. a browser without `tty:open` doesn't bother trying to open a TTY channel). This is advisory — the real check happens on `channel/open`. | + +This is the resource-discovery analogue of `services/list`. It's the +mechanism by which both sides populate what resources they expose, matching +the bidirectional symmetry of the operation overlay. The hub, after +`from_call` discovers the spoke's operations, also calls +`channel/resources` to discover what channels the spoke will accept — and +the hub aggregates that into what it exposes to the browser. + +### Bidirectional open — who initiates + +Channel open is bidirectional. Two cases: + +1. **Initiator wants responder's resource** (`direction: + initiator-to-responder`): the browser opens a TTY channel on a spoke. + The initiator sends `channel/open`; the responder allocates the handler + and returns the ID. This is the common case. + +2. **Initiator exposes a resource to the responder** (`direction: + responder-to-initiator`): a worker wants the hub to be able to open a + tunnel channel *to* the worker's local service. Here the worker sends + `channel/open` with `direction: responder-to-initiator` — it's saying + "I'm making myself available; when *you* want to connect, use this + channel." The semantics: the channel is created, but the worker's handler + is the *server* side of the ALPN, and the hub is the *client* side. The + `params` describe what the worker is exposing (e.g. + `{"target": "service:postgres", "port": 5432}`), not what it's asking + for. + + This is the mirror of the call protocol's operation overlay: a worker + registers `bash/exec` (the hub can call it); a worker opens a + `responder-to-initiator` channel (the hub can use it). In both cases the + worker is the server, the hub is the client, and the *worker* initiates + the registration/open because it's the one that knows what it has. + +### ACL flow end-to-end + +Putting it together, a browser opening a TTY channel to a spoke through a +hub: + +1. Browser's channel 0 → hub's channel 0: `channel/open` + `{ alpn: "alknet/tty", params: { backend: "docker", cmd: ["bash"], container: "abc123" } }`. + The browser's identity is a bearer token (ADR-043). +2. Hub's `CallAdapter` runs `AccessControl::check` on `channel/open` with + the browser's identity. If the browser lacks `channel:open:alknet/tty` + or the hub's policy forbids it → `channel:forbidden`. +3. Hub forwards to spoke via `from_call`: the hub's `forwarded_for` handler + constructs a `call.requested` with the hub as caller and the browser as + `forwarded_for` (ADR-032 §3). The spoke receives + `channel/open` with `caller = hub`, `forwarded_for = browser`. +4. Spoke's `CallAdapter` runs `AccessControl::check` with the hub as caller + (the spoke authorizes the hub, its direct peer — ADR-050). The spoke's + ownership store (ADR-050) verifies the hub (or the `forwarded_for` + browser, if the spoke's policy says so) owns `container:abc123`. +5. Spoke allocates the channel via `TtyAdapter` / `DockerTtyBackend`, returns + `channel_id`. +6. Hub opens a matching channel on the browser's side (it's now the + *responder* for the browser leg, *initiator* for the spoke leg) and + bridges them: read chunks off browser channel, write onto spoke channel + (with `channel_id` remapped), and vice versa. + +The hub ran **zero** protocol-specific auth. It ran `channel/open`'s +`AccessControl::check` (call-protocol machinery) and forwarded. The channels +layer inherited the auth model by being a call-protocol operation. + +## Channel Manager and Connection Internals + +This section fills the second conceptual gap: what state the channels layer +holds, how `ChannelsAdapter::handle` is wired, and how the `channel/open` +handler threads back into the `CallAdapter`'s `OperationRegistry`. The +guiding constraint is that **the channels layer is a re-framing proxy** — it +converts between "one transport stream carrying N channels" (the wire) and +"N independent `AsyncRead + AsyncWrite` handles" (what handlers see) — and +it does **no protocol work itself**. + +### The two halves: ChannelsAdapter and ChannelManager + +The channels crate has two internal components, split by responsibility: + +1. **`ChannelsAdapter`** — implements `ProtocolHandler` for + `alknet/channels`. Its `handle()` receives one `Connection` (the + transport), reads 9-byte chunk headers off the single bidi stream, and + routes each chunk to the `ChannelManager`. It is the *read/demux* half. + It does not know what ALPNs exist or what a handler is — it just splits + streams. + +2. **`ChannelManager`** — the shared state both halves touch. It holds the + map of `channel_id → ChannelState`, the `HandlerRegistry` reference, and + the `OperationRegistry` reference. It is the *reassemble/allocate* half. + It is what the `channel/open` operation handler closes over. + +The split mirrors the TTY crate's `ChunkReader`/`ChunkWriter` + adapter +pattern, generalized: the adapter no longer drives one session — it drives +N channels, and the channel-0 session is special only in that it's pre- +allocated. + +### ChannelManager state + +```rust +pub struct ChannelManager { + /// channel_id → per-channel state. Channel 0 is pre-inserted at construction. + channels: Mutex>, + /// The handler registry for looking up ALPNs on channel/open. + handlers: Arc, + /// The call protocol's operation registry, so channel/open etc. can be + /// registered at assembly time. The ChannelsAdapter holds a clone. + call_ops: Arc, + /// Next server-assigned channel_id. Monotonic; wraps at u32::MAX. + next_id: AtomicU32, + /// Per-channel reassembly buffer cap (DP-5). Default 1 MiB. + buffer_cap: usize, + /// Per-connection channel limit (OQ-CH-06). Default 256. + max_channels: usize, +} + +struct ChannelState { + /// The ALPN this channel carries, for routing and observability. + alpn: String, + /// Reassembly buffers per active stream_type. Each is a bounded + /// channel feeding a `SendStream`/`RecvStream` returned to the handler. + streams: HashMap, + /// The handler task driving this channel (TtyAdapter::drive_session, + /// SshAdapter::handle, etc.). Dropping this aborts the channel. + handler_task: JoinHandle<()>, + /// Which stream_types are active (from the open negotiation). + stream_types: Vec, +} +``` + +The `ChannelManager` is `Clone` (cheap — it's an `Arc` internally) so that +the `ChannelsAdapter`, the `channel/open` operation handler, and any relay +logic can all hold a handle to it. + +### ChannelsAdapter::handle — the read/demux loop + +```rust +#[async_trait] +impl ProtocolHandler for ChannelsAdapter { + fn alpn(&self) -> &'static [u8] { b"alknet/channels" } + + async fn handle(&self, connection: Connection, auth: &AuthContext) -> Result<(), HandlerError> { + // One bidi stream carries all channels. + let (send, recv) = connection.accept_bi().await?; + // Channel 0 is pre-negotiated as alknet/call. Construct its + // reassembly buffers and hand the reassembled Connection to the + // CallAdapter (looked up in the registry, same as every other ALPN). + self.manager.preinstall_channel_0(send, recv, auth).await?; + // Now read chunks off the transport and route them. + self.manager.run_demux_loop().await + } +} +``` + +The `preinstall_channel_0` step is the only special case: it constructs the +reassembly buffers for `channel_id = 0`, wraps them as a `Connection` (via +`Connection::from_stream`, which already exists in `crates/alknet-core/src/ +types.rs`), and hands that `Connection` to the `CallAdapter` — exactly as if +`alknet/call` had been the top-level ALPN. The `CallAdapter` is none the +wiser: it calls `accept_bi()` on its `Connection`, gets one bidi stream (the +channel-0 reassembled stream), and runs its dispatch loop. The call +protocol's `EventEnvelope` frames ride on `stream_type = 0` of channel 0. + +The `run_demux_loop` reads 9-byte headers, looks up `channel_id` in +`channels`, and pushes the payload into the right `ReassemblyBuffer` for +`(channel_id, stream_type)`. If the buffer is full (DP-5), the loop +**stops reading that channel's chunks** until the consumer drains — this is +the backpressure mechanism. Other channels keep flowing. This is the one +place the channels layer does flow control, and it's deliberately simple. + +### The channel/open handler — threading into OperationRegistry + +The `channel/open` (and `channel/close`, `channel/control`, +`channel/resources`) operations are registered on the call protocol's +`OperationRegistry` at **assembly time**, before the endpoint starts. The +handler closures close over a `ChannelManager` clone: + +```rust +// At assembly time, after both adapters are constructed: +let channel_ops = ChannelOperations::new(manager.clone()); +channel_ops.register_on(&mut call_registry)?; +``` + +Where `ChannelOperations::register_on` inserts four +`HandlerRegistration`s (one per operation) into the `OperationRegistry`. +The `channel/open` handler is the interesting one: + +```rust +fn channel_open_handler(mgr: ChannelManager) -> Handler { + Arc::new(move |input: Value, ctx: OperationContext| { + let mgr = mgr.clone(); + Box::pin(async move { + let req: ChannelOpenRequest = serde_json::from_value(input)?; + // 1. ACL is already checked by OperationRegistry::invoke before + // this handler runs — AccessControl::check on the operation. + // 2. Look up the ALPN in the HandlerRegistry. + let handler = mgr.handlers.get(req.alpn.as_bytes()) + .ok_or_else(|| channel_error("channel:unknown_alpn"))?; + // 3. Allocate the channel_id (server-assigned, DP-1). + let channel_id = mgr.next_id.fetch_add(1, Relaxed); + // 4. Construct the reassembly buffers for the negotiated + // stream_types and wrap as a Connection. + let conn = mgr.build_channel_connection(channel_id, &req.stream_types); + // 5. Spawn the handler task — same as TtyAdapter::handle does + // today, but on a ChannelConnection instead of a QUIC conn. + let task = tokio::spawn(async move { + let _ = handler.handle(conn, &auth_from_ctx(&ctx)).await; + }); + // 6. Record the channel state. + mgr.insert_channel(channel_id, ChannelState { /* ... */ }); + // 7. Return the channel_id. + ResponseEnvelope::ok(ctx.request_id, json!({ + "channel_id": channel_id, + "stream_types": req.stream_types, + })) + }) + }) +} +``` + +The key insight: **step 5 is identical to what `TtyAdapter::handle` does +today** (`crates/alknet-tty/src/adapter.rs:130`) — `tokio::spawn` a +`drive_session` task per bidi stream. The only difference is the `Connection` +passed in is a `ChannelConnection` (backed by reassembly) rather than a +quinn connection. The handler doesn't care. + +### ChannelConnection as a Connection + +`ChannelConnection` is a new `ConnectionKind` variant, or — to avoid +touching `alknet-core` — a `Connection` constructed via the existing +`Connection::from_stream` path, where the "stream" is a reassembly buffer +pair. The reassembly buffer exposes `AsyncRead` (drained by the handler's +`RecvStream`) and the write side exposes `AsyncWrite` (the handler's +`SendStream`, which the channels layer reads from and re-chunks onto the +transport with the right `channel_id`). + +The flow for a handler writing to its `SendStream`: + +``` +handler writes bytes + → SendStream::poll_write (AsyncWrite on the reassembly write-half) + → channels layer's per-channel pump reads from the write-half + → frames as [channel_id][stream_type][length][payload] onto the transport +``` + +And for reading: + +``` +transport arrives with a chunk for (channel_id, stream_type) + → ChannelsAdapter::run_demux_loop routes payload to ReassemblyBuffer + → handler's RecvStream::poll_read (AsyncRead on the reassembly read-half) yields bytes +``` + +The `stream_type` is chosen by *which* `SendStream`/`RecvStream` the handler +writes to. A TTY handler gets four handles (stdin=0, stdout=1, stderr=2, +control=3); a tunnel handler gets two (data-in=0, data-out=1). The +`ChannelConnection::accept_bi()` call returns the pair for the *next* +expected stream_type, or — more concretely — the handler is handed a typed +struct (`TtyChannel { stdin, stdout, stderr, control }`) rather than a +generic `Connection`, because the channel's stream_types are known at open +time. + +This is a slight divergence from "every handler sees a `Connection`": TTY +wants four named handles, not a `Connection` you call `accept_bi` on four +times. The resolution (carried into Phase 1): the `ChannelConnection` +*implements* the `Connection` interface (for recursion and generic +handlers), **and** can be destructured into typed sub-stream handles for +handlers that know their ALPN's shape. The typed destructure is a +convenience layer over the same reassembly buffers; it's not a separate +abstraction. + +### The hub relay: ChannelManager-to-ChannelManager + +For the hub use case (§Hub Motivation), the hub holds *two* `ChannelManager` +instances per (browser, spoke) pair — one for the browser leg, one for the +spoke leg — and a relay task per channel that bridges them: + +```rust +// For channel_id=7 on browser side, channel_id=12 on spoke side: +tokio::spawn(async move { + let (b_send, b_recv) = browser_mgr.open_channel_stream(7, stream_type).await; + let (s_send, s_recv) = spoke_mgr.open_channel_stream(12, stream_type).await; + tokio::join!( + pump(b_recv, s_send), // browser → spoke + pump(s_recv, b_send), // spoke → browser + ); +}); +``` + +The relay reads opaque bytes off one `ChannelManager`'s reassembled stream +and writes them onto the other's write-half, which re-chunks them with the +other leg's `channel_id`. The relay does not parse the bytes — it doesn't +know if they're TTY chunks, SSH frames, or tunnel data. The channels layer +on each end does the chunk↔stream conversion; the relay just moves bytes +between two `AsyncRead + AsyncWrite` pairs. + +This is why the hub's complexity collapses: the relay is **one pump +function** applied per channel, not per (protocol × transport) cell. The +`ChannelManager` is the uniform interface on both ends. + +### What is NOT in the ChannelManager + +To keep the boundary clean, the `ChannelManager` deliberately does **not** +hold: + +- **No `ProtocolHandler` implementations.** It holds a `HandlerRegistry` + reference for ALPN lookup, but it doesn't *be* a handler. The handlers + live in their crates and register on the same registry. +- **No ALPN-specific parsing.** It does not parse `NegotiateRequest` JSON, + SSH binary frames, or tunnel target strings. It hands `params` JSON to + the handler and gets back a handler task; it hands `stream_type 3` JSON + to the handler's control handle. The `ChannelManager` is ALPN-blind. +- **No auth state.** Auth lives in the `OperationContext` that the call + protocol passes to `channel/open`. The `ChannelManager` doesn't check + scopes or ownership — that's `AccessControl::check` in + `OperationRegistry::invoke`, run before the `channel/open` handler is + called. +- **No transport coupling.** The `ChannelManager` talks to the transport + only through the `ChannelsAdapter`'s read loop and the per-channel write + pumps, both of which use `AsyncRead + AsyncWrite`. QUIC, TCP+TLS, + WebTransport, SSH channel — all look the same. + +This is what makes the channels layer WASM-compatible (§WASM Compatibility) +and transport-agnostic (§Transport Agnosticism): the `ChannelManager` is +pure byte routing with no platform or protocol dependencies. + ## Relationship to Existing Crates ### alknet-tty @@ -641,6 +1329,54 @@ sent. This is an implementation constraint, not a design change. call protocol, TTY session inside a channel. Stretch: SSH connection inside a channel, tunnel inside a channel, recursive composition. +- **OQ-CH-08 (channel/resources staleness)**: `channel/resources` is a + snapshot of what a side exposes at call time. Resources can appear and + disappear (a docker container starts/stops, a worker connects/disconnects). + Does the channels layer need a `channel/resources/subscribe` streaming + operation (analogous to a subscription), or is polling `channel/resources` + sufficient for v1? The call protocol already has streaming + `OperationType::Subscription`; reusing it here is natural but adds a + long-lived stream on channel 0 per interested peer. Recommendation for + Phase 1: poll for v1, add subscription if staleness bites. + +- **OQ-CH-09 (responder-to-initiator channel lifecycle)**: in the + `responder-to-initiator` open case (worker exposes, hub consumes), who + drives the data first? The channel is allocated on the worker's side, but + the hub is the client of the ALPN — does the hub's handler start pumping + immediately, or does the worker's handler wait for the hub to write first? + For TTY the server writes the negotiation response first; for a tunnel the + client connects first. This is ALPN-specific and probably doesn't need a + channels-layer rule, but the `direction` field's effect on which side's + `ChannelConnection` is the "server" vs "client" of the ALPN needs to be + pinned down in Phase 1. + +- **OQ-CH-10 (ChannelConnection: typed destructure vs generic Connection)**: + the doc proposes that `ChannelConnection` implements the `Connection` + trait (for recursion and generic handlers) **and** can be destructured + into typed sub-stream handles (e.g. `TtyChannel { stdin, stdout, stderr, + control }`). Is the typed destructure a `ChannelConnection` method, a + separate trait, or an ALPN-specific constructor in the handler crate? + This affects whether `alknet-channels` needs to know about TTY's + `stream_type` semantics or whether the TTY crate does the destructure + itself given a `ChannelConnection`. Recommendation: the TTY crate + destructures — `alknet-channels` exposes `(channel_id, stream_type) → + (SendStream, RecvStream)` accessors and the handler crate maps those to + its typed names. + +- **OQ-CH-11 (hub relay channel_id remapping)**: when the hub bridges a + browser channel (`channel_id=7`) to a spoke channel (`channel_id=12`), + the relay rewrites the `channel_id` field in each chunk header. But + `channel/control` and `channel/close` operations on channel 0 carry + `channel_id` in their *JSON payload*, not in the chunk header. Does the + hub's call-protocol forwarding (via `from_call`) need to rewrite + `channel_id` inside the `call.requested` payload too? This is a + relay-level concern: the hub is terminating channel 0 on both legs (it + runs its own `CallAdapter`), so it likely translates `channel/open` from + the browser into a *new* `channel/open` on the spoke (not a raw forward), + and the spoke's returned `channel_id` is mapped back to the browser's + side. Phase 1 must specify whether the hub translates or transparently + forwards, and how the `channel_id` mapping is maintained. + ## Recommended Approach ### Crate @@ -711,6 +1447,11 @@ channel types, orchestrated by the call protocol.** Minimum scope: Stretch goals: 5. Two different channel types (TTY + tunnel) on the same connection 6. WASM build of the chunk splitting/recombining logic +7. Hub relay: two `ChannelManager` instances bridged by a byte pump, + validating that `channel/open` on leg A translates to `channel/open` on + leg B and the `channel_id` remapping works end-to-end (OQ-CH-11) +8. `channel/resources` discovery — one side lists its exposed ALPNs and the + other opens a channel based on the response The POC can be built as an extension to the existing `alknet-tty-poc` or as a standalone POC in `/workspace/alknet-channels-poc/`.