docs(research): flesh out alknet-channels — hub motivation, channel open negotiation, channel manager internals

Adds three sections filling the conceptual gaps in the phase-0 findings:

- Hub Motivation: The Multi-Transport Collapse — diagnoses the
  O(protocols × transports × spokes) mess the hub crate faces and shows how
  channels collapses it to one connection per leg with channel-by-channel
  byte forwarding, reusing the call protocol's auth/forwarded-for model.

- Channel Open Negotiation — concrete channel/open, channel/close,
  channel/control, channel/resources operation payloads with field tables,
  error codes as CallError strings, bidirectional open semantics, and the
  end-to-end ACL flow for browser→hub→spoke. No new wire framing; all four
  operations register on the call protocol's existing OperationRegistry.

- Channel Manager and Connection Internals — the ChannelsAdapter/
  ChannelManager split, ChannelManager state sketch, ChannelsAdapter::handle
  read/demux loop with channel 0 preinstall, the channel/open handler
  closure threading into OperationRegistry, ChannelConnection as a
  Connection via the existing from_stream path, the hub relay as
  ChannelManager-to-ChannelManager byte pumping, and the boundary
  (ChannelManager holds no handlers, no ALPN parsing, no auth, no transport
  coupling).

Also adds four new open questions (OQ-CH-08 through OQ-CH-11) on resource
staleness, responder-to-initiator lifecycle, typed destructure ownership,
and hub relay channel_id remapping; two new POC stretch goals (hub relay,
channel/resources); and updates the opening to note the revision scope.
This commit is contained in:
glm-5.2 committed 2026-07-11 08:26:56 +00:00
1 parent eb9ea506f6
commit 8692f9748e
1 file changed
+742 -1
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-10
last_updated: 2026-07-11
---
# alknet-channels — Phase 0 Research Findings
@@ -20,6 +20,14 @@ crate's chunk format already solves sub-stream multiplexing for terminal
sessions — this document generalizes that pattern into a universal channel
multiplexer that the call protocol can orchestrate.
The 2026-07-11 revision adds the hub motivation (§Hub Motivation), the
concrete channel-open negotiation protocol (§Channel Open Negotiation), and
the channel-manager/connection internals (§Channel Manager and Connection
Internals) — driven by the realization that the hub crate's multi-transport
complexity is what channels specifically collapses, and that the call
protocol's auth/operation-overlay model carries over to channel lifecycle
with no new machinery.
## Vision Recap
`alknet-channels` is a **multiplexing proxy** — a `ProtocolHandler` on
@@ -98,6 +106,143 @@ bidirectional — either side can open a channel to the other, just like the
call protocol's operation overlay where each side populates what operations
they can call.
## Hub Motivation: The Multi-Transport Collapse
This design was forced by the hub crate. A hub is the architectural role
(ADR-029, ADR-034) that bridges peers and browsers: it terminates QUIC from
spokes, terminates WebTransport/WSS from browsers, and may itself be a spoke
of an upstream hub. The hub's job is to **route and relay**, not to
re-implement every protocol's framing per transport.
Without channels, the hub quickly becomes a mess:
### The mess, concretely
A hub serving a browser over WebTransport and a spoke over QUIC wants to
offer the browser three things on the same logical session:
1. Call operations on the spoke (e.g., `docker/container/list` — JSON).
2. A TTY session on the spoke (interactive exec — raw bytes, ADR-052 chunk
format).
3. An SSH session to the spoke (raw bytes, SSH binary protocol).
Today each of these is a **separate ALPN** on a **separate connection**, and
each transport multiplexes differently:
| Need | Today's mechanism | Transport | Multiplexing |
|------|-------------------|-----------|--------------|
| JSON ops | `alknet/call` ALPN | QUIC stream or WSS | one op per bidi stream |
| TTY | `alknet/tty` ALPN | QUIC stream or WSS | TTY chunk format *within* one bidi stream |
| SSH | `alknet/ssh` ALPN (future) | QUIC stream or WSS | SSH multiplexes *within* one bidi stream |
So the hub faces a matrix: **3 needs × 2 transports × N spokes**, and the
multiplexing layer is different in each cell. The hub would have to:
- Maintain a `quinn::Connection` per spoke AND a WebTransport session per
browser, AND decide which one carries which ALPN.
- For a browser→spoke TTY, the hub can't just "forward a stream" — the
browser's WSS stream carries `alknet/tty` chunks, and the spoke's QUIC
stream also carries `alknet/tty` chunks, so the hub *can* relay bytes —
but it had to negotiate two separate `alknet/tty` connections (one per
leg) and correlate them by out-of-band state. There's no "one session that
spans the relay."
- For a browser→spoke SSH, same problem with a different ALPN and a
different internal multiplexer.
- The call protocol that *orchestrates* all this lives on a **third** ALPN
(`alknet/call`), so the hub is juggling three connections per (browser,
spoke) pair, each with its own framing, and the correlation across them is
implicit.
- Every new protocol (tunnel, future file-transfer, future git-pack) adds
**another column** to the matrix and **another ALPN** the hub must
register, dispatch, and relay.
This is the "PITA to maintain a bunch of connections" problem: the hub's
complexity is **O(protocols × transports × spokes)**, when it should be
**O(spokes)**.
### What channels collapses
With `alknet/channels`, the hub's per-spoke and per-browser state is **one
channels connection each**, and everything else is routing:
```
Browser ──WebTransport──► Hub ──QUIC──► Spoke
alknet/channels alknet/channels
┌─────────────┐ ┌─────────────┐
│ ch0: call │ │ ch0: call │
│ ch1: tty │ relay │ ch1: tty │
│ ch2: ssh │ ◄─────► │ ch2: ssh │
│ ch3: tunnel │ │ ch3: tunnel │
└─────────────┘ └─────────────┘
```
The hub's relay logic is **channel-by-channel byte forwarding**: read chunks
for `(channel_id)` off the browser's channels connection, write the same
payload (with the spoke's `channel_id` substituted) onto the spoke's
channels connection. The hub does **not** parse `alknet/tty` chunks, does
**not** understand SSH's internal multiplexer, does **not** run a separate
`CallAdapter` per leg — it forwards opaque `(channel_id, stream_type,
payload)` tuples and lets the endpoints at each end do the protocol work.
The collapse is at three levels:
1. **One connection per leg, not one per protocol.** The browser holds one
WebTransport session to the hub; the hub holds one QUIC connection to the
spoke. All three needs (call, TTY, SSH) ride as channels on those two
connections. The hub correlates by channel, not by ALPN-per-connection.
2. **One multiplexing model, not three.** Connection-level (ALPN router),
stream-level (QUIC native), and sub-stream-level (TTY chunks) collapse
into one: channels chunks. The hub's relay code is one loop, not three.
3. **The call protocol orchestrates from inside.** Channel 0 is `alknet/call`
on both legs. The browser calls `channel/open` on its channel 0; the hub
forwards that call to the spoke's channel 0 (via `from_call`); the spoke
allocates the channel and returns the ID; the hub opens a matching channel
on the browser's side and bridges them. The hub never runs a handler for
`alknet/tty` or `alknet/ssh` — it only runs `alknet/channels` (the relay)
and `alknet/call` (for its own hub-level operations like routing and
resource lookup).
### Why the auth model reuses cleanly
The call protocol's `OperationContext` (identity, scopes, capabilities,
ownership) already gates every operation. `channel/open` is just another
operation on channel 0, so the same `OperationContext` flows through:
- The browser's `channel/open` carries the browser's identity (bearer token
resolved by the hub per ADR-043).
- The hub forwards via `from_call`, which populates `forwarded_for` from the
hub's `OperationContext.identity` (ADR-032 §3) — exactly the kernel/user-
land + forwarded-for model from ADR-050.
- The spoke's `AccessControl::check` sees the hub as the caller and the
browser as `forwarded_for`. The spoke authorizes the hub (its direct
peer); the hub's "who is this for" is its own app state, carried as
`forwarded_for`.
No new auth machinery. The hub doesn't authenticate channels — it
authenticates the call operations that *open* channels, which it already
does. The channels layer inherits the call protocol's ACL by being one of
its operation types.
### What the hub *does* still own
Channels does not eliminate the hub's responsibilities; it relocates them:
- **Routing**: which spoke serves `container:abc123`? That's the hub's
resource registry / ownership store (ADR-050), queried via call operations
on channel 0. Channels doesn't touch this.
- **ACL at the hub**: does this browser's identity have `channel:open` scope
for `alknet/ssh` to `spoke-X`? That's the call protocol's
`AccessControl::check` on the `channel/open` operation, run by the hub's
`CallAdapter` before it forwards. Channels doesn't touch this either.
- **Relay lifecycle**: when a browser disconnects, the hub tears down the
spoke-side channels (and vice versa). This is `channel/close` on each
channel, or a transport-level close that the channels layer observes.
What the hub *no longer* owns: per-protocol framing parsers, per-protocol
relay loops, per-ALPN connection management, and the correlation state
across multiple connections per (browser, spoke) pair. Those move into the
channels layer, which is transport-agnostic and protocol-agnostic.
## The Wire Format
### Chunk header
@@ -240,6 +385,549 @@ recommendation is to allow it but not encourage it — the primary use case is
one level of multiplexing. Recursive composition is a natural consequence of
the `Connection` abstraction, not a feature to design for.
## Channel Open Negotiation
Channel lifecycle is orchestrated by the call protocol on channel 0. The
call protocol is JSON-only — it carries `channel/open`,
`channel/close`, `channel/control`, and `channel/resources` as **new
operation types** in the existing `OperationRegistry`, dispatched through
the existing `OperationContext` (identity, scopes, capabilities, ownership,
`forwarded_for`). No wire-format change to `EventEnvelope`, no new carriage
type. The channels layer registers these operations on the call protocol's
`OperationRegistry` at assembly time.
This is where the call protocol's auth model is reused: `channel/open` is
gated by `AccessControl::check` exactly like any other operation, and the
operation's `AccessControl` can declare `required_scopes` (e.g.
`channel:open:alknet/tty`), `resource_type` + `resource_action` (e.g.
`container` + `tty`), or `required_scopes_any` for multi-scope gates. The
channels layer doesn't re-implement auth — it leans on the
`OperationRegistry::invoke` path that already runs the check before the
handler runs.
### `channel/open` — request a channel
Either side sends a `call.requested` on channel 0 with operation name
`channel/open`:
```json
{
"type": "call.requested",
"id": "req-abc",
"payload": {
"operation": "channel/open",
"input": {
"alpn": "alknet/tty",
"stream_types": [0, 1, 2, 3],
"params": {
"backend": "docker",
"cmd": ["bash"],
"container": "abc123"
},
"direction": "initiator-to-responder"
}
}
}
```
Fields:
| field | type | meaning |
|-------|------|---------|
| `alpn` | string | The ALPN the channel will carry, e.g. `alknet/tty`, `alknet/ssh`, `alknet/tunnel`. The responder looks this up in its `HandlerRegistry`. |
| `stream_types` | `[u8]` | Which sub-stream types this channel will use. Declared at open time so both sides size their reassembly buffers and know which `(channel_id, stream_type)` pairs are valid. E.g. `[0, 1, 2, 3]` for TTY, `[0, 1]` for a raw tunnel. |
| `params` | object | ALPN-specific parameters passed to the handler. For `alknet/tty` this is the `NegotiateRequest` (backend, cmd, env, size). For `alknet/tunnel` this is the target resource (`{"resource": "container:abc123", "port": 5432}`). For `alknet/ssh` this is the SSH auth/method hints. The channels layer does not interpret `params` — it hands the JSON to the handler's allocate entry point. |
| `direction` | string | `initiator-to-responder` (the initiator wants the responder to expose a resource) or `responder-to-initiator` (the initiator wants to expose a resource *to* the responder — the "worker exposes, hub consumes" case). See "Bidirectional open" below. |
The `direction` field is what makes channel open match the call protocol's
operation-overlay symmetry: just as either side can populate operations the
other can call, either side can expose resources the other can open
channels to. `initiator-to-responder` is the common case ("open me a TTY on
your docker container"). `responder-to-initiator` is the reverse ("I'm a
worker exposing my local Postgres as a tunnel target; hub, you can open a
channel to it") — which is how a hub consumes a worker's resources without
the worker initiating.
### `channel/open` — response
The responder validates ACL, looks up the ALPN in `HandlerRegistry`, and —
crucially — **allocates the `channel_id` before returning** (DP-1: server-
assigned). The response is a normal `call.responded`:
```json
{
"type": "call.responded",
"id": "req-abc",
"payload": {
"output": {
"channel_id": 7,
"stream_types": [0, 1, 2, 3]
}
}
}
```
| field | type | meaning |
|-------|------|---------|
| `channel_id` | u32 | The server-assigned channel ID. Both sides now know to route chunks with this `channel_id` to the new channel. |
| `stream_types` | `[u8]` | The *negotiated* set — the responder may narrow the initiator's requested set (e.g. refuse stderr for a backend that doesn't produce it). The intersection of requested and supported. |
Once the initiator receives the response, data may flow on
`(channel_id=7, stream_type=*)`. The responder's handler is already
allocated and reading from its reassembled streams.
### `channel/open` — error cases
Errors use the call protocol's existing `CallError` shape, dispatched as
`call.error`. The channels layer defines a small set of error codes:
| code | meaning | retryable |
|------|---------|-----------|
| `channel:unknown_alpn` | The `alpn` is not in the responder's `HandlerRegistry`. | false |
| `channel:forbidden` | `AccessControl::check` denied the open (missing scope, not the resource owner, `forwarded_for` not trusted). | false |
| `channel:allocation_failed` | The handler's allocate failed (e.g. `DockerTtyBackend::allocate` couldn't start the exec). The `details` carry the backend's error message. | true (often transient) |
| `channel:invalid_params` | The `params` JSON didn't satisfy the ALPN's expectations (e.g. missing `cmd` for TTY). | false |
| `channel:too_many_channels` | The responder hit its per-connection channel limit (OQ-CH-06). | false |
| `channel:stream_type_unavailable` | The responder's handler can't provide one of the requested `stream_types`. The `details` carry the supported set. | false |
These are new `CallError.code` strings, not new framing. The call protocol
already carries `code`/`message`/`retryable`/`details`.
### `channel/close` — tear down a channel
Either side sends:
```json
{
"operation": "channel/close",
"input": { "channel_id": 7, "reason": "exit" }
}
```
The responder (the side that *didn't* send the close) drains its
reassembled streams for `channel_id`, signals EOF to the handler, and
returns `call.responded` with `{ "closed": true }`. The `channel_id` is now
eligible for reuse (OQ-CH-04). `reason` is a free-form string for
observability — `"exit"`, `"cancel"`, `"error"` — and is not semantically
required.
The ordering invariant (ADR-055 for TTY): the channel's **data chunks** must
be written and flushed before the `channel/close` operation is sent on
channel 0. This is an implementation constraint on the side closing — the
channels layer's close handler must observe the data-channel pump complete
before issuing the call operation. For TTY this is the exit-chunk-is-last
invariant carried forward.
### `channel/control` — out-of-band control on channel 0
For control that doesn't need ordering relative to data (resize, signal,
keepalive), a call operation on channel 0:
```json
{
"operation": "channel/control",
"input": {
"channel_id": 7,
"stream_type": 3,
"message": { "type": "resize", "cols": 80, "rows": 24 }
}
}
```
The channels layer routes `message` to the handler's control handle for
`channel_id`. This is the "call protocol for orchestration" half of DP-4.
The `message` JSON is ALPN-specific; the channels layer doesn't interpret
it — it hands it to the handler's control entry point, same as `params` on
open.
### `stream_type 3` — in-band control on the data channel
For control that **does** need ordering relative to data (EOF before exit,
flush before close), a chunk with `stream_type = 3` on the data channel
itself:
```
[channel_id: u32 be][stream_type: 0x03][length: u32 be][json payload]
```
This is the "stream_type 3 for data-ordered control" half of DP-4. The
handler parses the JSON; the channels layer just reassembles and delivers
it in-order with the data. The TTY crate's `ControlMessage::Eof` is the
canonical example — it must arrive after the last stdin chunk, which is
guaranteed by chunk ordering within `(channel_id, stream_type)`, not by a
call-protocol round-trip.
### `channel/resources` — populate the resource overlay
This is the conceptual gap the doc is closing. The call protocol already
has `services/list` and `services/list-peers` for discovering **operations**
each side exposes. Channels needs the equivalent for **resources** — what
ALPNs each side is willing to open channels for, and with what constraints.
A side calls `channel/resources` on channel 0 to ask the other side what it
exposes:
```json
{
"operation": "channel/resources",
"input": {}
}
```
Response:
```json
{
"output": {
"resources": [
{
"alpn": "alknet/tty",
"backends": ["docker", "local"],
"access": { "required_scopes": ["tty:open"] }
},
{
"alpn": "alknet/ssh",
"access": { "required_scopes": ["ssh:open"] }
},
{
"alpn": "alknet/tunnel",
"targets": ["container:*", "service:postgres"],
"access": { "required_scopes_any": ["tunnel:open", "admin"] }
}
]
}
}
```
| field | type | meaning |
|-------|------|---------|
| `alpn` | string | The ALPN this side will accept `channel/open` for. |
| `backends` / `targets` | `[string]` | ALPN-specific enumeration of what's available — TTY backends, tunnel target resource patterns. The channels layer doesn't interpret these; they're for the *initiator* to know what `params` to send. |
| `access` | object | A preview of the `AccessControl` that `channel/open` will check. Lets the initiator fail fast (e.g. a browser without `tty:open` doesn't bother trying to open a TTY channel). This is advisory — the real check happens on `channel/open`. |
This is the resource-discovery analogue of `services/list`. It's the
mechanism by which both sides populate what resources they expose, matching
the bidirectional symmetry of the operation overlay. The hub, after
`from_call` discovers the spoke's operations, also calls
`channel/resources` to discover what channels the spoke will accept — and
the hub aggregates that into what it exposes to the browser.
### Bidirectional open — who initiates
Channel open is bidirectional. Two cases:
1. **Initiator wants responder's resource** (`direction:
initiator-to-responder`): the browser opens a TTY channel on a spoke.
The initiator sends `channel/open`; the responder allocates the handler
and returns the ID. This is the common case.
2. **Initiator exposes a resource to the responder** (`direction:
responder-to-initiator`): a worker wants the hub to be able to open a
tunnel channel *to* the worker's local service. Here the worker sends
`channel/open` with `direction: responder-to-initiator` — it's saying
"I'm making myself available; when *you* want to connect, use this
channel." The semantics: the channel is created, but the worker's handler
is the *server* side of the ALPN, and the hub is the *client* side. The
`params` describe what the worker is exposing (e.g.
`{"target": "service:postgres", "port": 5432}`), not what it's asking
for.
This is the mirror of the call protocol's operation overlay: a worker
registers `bash/exec` (the hub can call it); a worker opens a
`responder-to-initiator` channel (the hub can use it). In both cases the
worker is the server, the hub is the client, and the *worker* initiates
the registration/open because it's the one that knows what it has.
### ACL flow end-to-end
Putting it together, a browser opening a TTY channel to a spoke through a
hub:
1. Browser's channel 0 → hub's channel 0: `channel/open`
`{ alpn: "alknet/tty", params: { backend: "docker", cmd: ["bash"], container: "abc123" } }`.
The browser's identity is a bearer token (ADR-043).
2. Hub's `CallAdapter` runs `AccessControl::check` on `channel/open` with
the browser's identity. If the browser lacks `channel:open:alknet/tty`
or the hub's policy forbids it → `channel:forbidden`.
3. Hub forwards to spoke via `from_call`: the hub's `forwarded_for` handler
constructs a `call.requested` with the hub as caller and the browser as
`forwarded_for` (ADR-032 §3). The spoke receives
`channel/open` with `caller = hub`, `forwarded_for = browser`.
4. Spoke's `CallAdapter` runs `AccessControl::check` with the hub as caller
(the spoke authorizes the hub, its direct peer — ADR-050). The spoke's
ownership store (ADR-050) verifies the hub (or the `forwarded_for`
browser, if the spoke's policy says so) owns `container:abc123`.
5. Spoke allocates the channel via `TtyAdapter` / `DockerTtyBackend`, returns
`channel_id`.
6. Hub opens a matching channel on the browser's side (it's now the
*responder* for the browser leg, *initiator* for the spoke leg) and
bridges them: read chunks off browser channel, write onto spoke channel
(with `channel_id` remapped), and vice versa.
The hub ran **zero** protocol-specific auth. It ran `channel/open`'s
`AccessControl::check` (call-protocol machinery) and forwarded. The channels
layer inherited the auth model by being a call-protocol operation.
## Channel Manager and Connection Internals
This section fills the second conceptual gap: what state the channels layer
holds, how `ChannelsAdapter::handle` is wired, and how the `channel/open`
handler threads back into the `CallAdapter`'s `OperationRegistry`. The
guiding constraint is that **the channels layer is a re-framing proxy** — it
converts between "one transport stream carrying N channels" (the wire) and
"N independent `AsyncRead + AsyncWrite` handles" (what handlers see) — and
it does **no protocol work itself**.
### The two halves: ChannelsAdapter and ChannelManager
The channels crate has two internal components, split by responsibility:
1. **`ChannelsAdapter`** — implements `ProtocolHandler` for
`alknet/channels`. Its `handle()` receives one `Connection` (the
transport), reads 9-byte chunk headers off the single bidi stream, and
routes each chunk to the `ChannelManager`. It is the *read/demux* half.
It does not know what ALPNs exist or what a handler is — it just splits
streams.
2. **`ChannelManager`** — the shared state both halves touch. It holds the
map of `channel_id → ChannelState`, the `HandlerRegistry` reference, and
the `OperationRegistry` reference. It is the *reassemble/allocate* half.
It is what the `channel/open` operation handler closes over.
The split mirrors the TTY crate's `ChunkReader`/`ChunkWriter` + adapter
pattern, generalized: the adapter no longer drives one session — it drives
N channels, and the channel-0 session is special only in that it's pre-
allocated.
### ChannelManager state
```rust
pub struct ChannelManager {
/// channel_id → per-channel state. Channel 0 is pre-inserted at construction.
channels: Mutex<HashMap<u32, ChannelState>>,
/// The handler registry for looking up ALPNs on channel/open.
handlers: Arc<HandlerRegistry>,
/// The call protocol's operation registry, so channel/open etc. can be
/// registered at assembly time. The ChannelsAdapter holds a clone.
call_ops: Arc<OperationRegistry>,
/// Next server-assigned channel_id. Monotonic; wraps at u32::MAX.
next_id: AtomicU32,
/// Per-channel reassembly buffer cap (DP-5). Default 1 MiB.
buffer_cap: usize,
/// Per-connection channel limit (OQ-CH-06). Default 256.
max_channels: usize,
}
struct ChannelState {
/// The ALPN this channel carries, for routing and observability.
alpn: String,
/// Reassembly buffers per active stream_type. Each is a bounded
/// channel feeding a `SendStream`/`RecvStream` returned to the handler.
streams: HashMap<u8, ReassemblyBuffer>,
/// The handler task driving this channel (TtyAdapter::drive_session,
/// SshAdapter::handle, etc.). Dropping this aborts the channel.
handler_task: JoinHandle<()>,
/// Which stream_types are active (from the open negotiation).
stream_types: Vec<u8>,
}
```
The `ChannelManager` is `Clone` (cheap — it's an `Arc` internally) so that
the `ChannelsAdapter`, the `channel/open` operation handler, and any relay
logic can all hold a handle to it.
### ChannelsAdapter::handle — the read/demux loop
```rust
#[async_trait]
impl ProtocolHandler for ChannelsAdapter {
fn alpn(&self) -> &'static [u8] { b"alknet/channels" }
async fn handle(&self, connection: Connection, auth: &AuthContext) -> Result<(), HandlerError> {
// One bidi stream carries all channels.
let (send, recv) = connection.accept_bi().await?;
// Channel 0 is pre-negotiated as alknet/call. Construct its
// reassembly buffers and hand the reassembled Connection to the
// CallAdapter (looked up in the registry, same as every other ALPN).
self.manager.preinstall_channel_0(send, recv, auth).await?;
// Now read chunks off the transport and route them.
self.manager.run_demux_loop().await
}
}
```
The `preinstall_channel_0` step is the only special case: it constructs the
reassembly buffers for `channel_id = 0`, wraps them as a `Connection` (via
`Connection::from_stream`, which already exists in `crates/alknet-core/src/
types.rs`), and hands that `Connection` to the `CallAdapter` — exactly as if
`alknet/call` had been the top-level ALPN. The `CallAdapter` is none the
wiser: it calls `accept_bi()` on its `Connection`, gets one bidi stream (the
channel-0 reassembled stream), and runs its dispatch loop. The call
protocol's `EventEnvelope` frames ride on `stream_type = 0` of channel 0.
The `run_demux_loop` reads 9-byte headers, looks up `channel_id` in
`channels`, and pushes the payload into the right `ReassemblyBuffer` for
`(channel_id, stream_type)`. If the buffer is full (DP-5), the loop
**stops reading that channel's chunks** until the consumer drains — this is
the backpressure mechanism. Other channels keep flowing. This is the one
place the channels layer does flow control, and it's deliberately simple.
### The channel/open handler — threading into OperationRegistry
The `channel/open` (and `channel/close`, `channel/control`,
`channel/resources`) operations are registered on the call protocol's
`OperationRegistry` at **assembly time**, before the endpoint starts. The
handler closures close over a `ChannelManager` clone:
```rust
// At assembly time, after both adapters are constructed:
let channel_ops = ChannelOperations::new(manager.clone());
channel_ops.register_on(&mut call_registry)?;
```
Where `ChannelOperations::register_on` inserts four
`HandlerRegistration`s (one per operation) into the `OperationRegistry`.
The `channel/open` handler is the interesting one:
```rust
fn channel_open_handler(mgr: ChannelManager) -> Handler {
Arc::new(move |input: Value, ctx: OperationContext| {
let mgr = mgr.clone();
Box::pin(async move {
let req: ChannelOpenRequest = serde_json::from_value(input)?;
// 1. ACL is already checked by OperationRegistry::invoke before
// this handler runs — AccessControl::check on the operation.
// 2. Look up the ALPN in the HandlerRegistry.
let handler = mgr.handlers.get(req.alpn.as_bytes())
.ok_or_else(|| channel_error("channel:unknown_alpn"))?;
// 3. Allocate the channel_id (server-assigned, DP-1).
let channel_id = mgr.next_id.fetch_add(1, Relaxed);
// 4. Construct the reassembly buffers for the negotiated
// stream_types and wrap as a Connection.
let conn = mgr.build_channel_connection(channel_id, &req.stream_types);
// 5. Spawn the handler task — same as TtyAdapter::handle does
// today, but on a ChannelConnection instead of a QUIC conn.
let task = tokio::spawn(async move {
let _ = handler.handle(conn, &auth_from_ctx(&ctx)).await;
});
// 6. Record the channel state.
mgr.insert_channel(channel_id, ChannelState { /* ... */ });
// 7. Return the channel_id.
ResponseEnvelope::ok(ctx.request_id, json!({
"channel_id": channel_id,
"stream_types": req.stream_types,
}))
})
})
}
```
The key insight: **step 5 is identical to what `TtyAdapter::handle` does
today** (`crates/alknet-tty/src/adapter.rs:130`) — `tokio::spawn` a
`drive_session` task per bidi stream. The only difference is the `Connection`
passed in is a `ChannelConnection` (backed by reassembly) rather than a
quinn connection. The handler doesn't care.
### ChannelConnection as a Connection
`ChannelConnection` is a new `ConnectionKind` variant, or — to avoid
touching `alknet-core` — a `Connection` constructed via the existing
`Connection::from_stream` path, where the "stream" is a reassembly buffer
pair. The reassembly buffer exposes `AsyncRead` (drained by the handler's
`RecvStream`) and the write side exposes `AsyncWrite` (the handler's
`SendStream`, which the channels layer reads from and re-chunks onto the
transport with the right `channel_id`).
The flow for a handler writing to its `SendStream`:
```
handler writes bytes
→ SendStream::poll_write (AsyncWrite on the reassembly write-half)
→ channels layer's per-channel pump reads from the write-half
→ frames as [channel_id][stream_type][length][payload] onto the transport
```
And for reading:
```
transport arrives with a chunk for (channel_id, stream_type)
→ ChannelsAdapter::run_demux_loop routes payload to ReassemblyBuffer
→ handler's RecvStream::poll_read (AsyncRead on the reassembly read-half) yields bytes
```
The `stream_type` is chosen by *which* `SendStream`/`RecvStream` the handler
writes to. A TTY handler gets four handles (stdin=0, stdout=1, stderr=2,
control=3); a tunnel handler gets two (data-in=0, data-out=1). The
`ChannelConnection::accept_bi()` call returns the pair for the *next*
expected stream_type, or — more concretely — the handler is handed a typed
struct (`TtyChannel { stdin, stdout, stderr, control }`) rather than a
generic `Connection`, because the channel's stream_types are known at open
time.
This is a slight divergence from "every handler sees a `Connection`": TTY
wants four named handles, not a `Connection` you call `accept_bi` on four
times. The resolution (carried into Phase 1): the `ChannelConnection`
*implements* the `Connection` interface (for recursion and generic
handlers), **and** can be destructured into typed sub-stream handles for
handlers that know their ALPN's shape. The typed destructure is a
convenience layer over the same reassembly buffers; it's not a separate
abstraction.
### The hub relay: ChannelManager-to-ChannelManager
For the hub use case (§Hub Motivation), the hub holds *two* `ChannelManager`
instances per (browser, spoke) pair — one for the browser leg, one for the
spoke leg — and a relay task per channel that bridges them:
```rust
// For channel_id=7 on browser side, channel_id=12 on spoke side:
tokio::spawn(async move {
let (b_send, b_recv) = browser_mgr.open_channel_stream(7, stream_type).await;
let (s_send, s_recv) = spoke_mgr.open_channel_stream(12, stream_type).await;
tokio::join!(
pump(b_recv, s_send), // browser → spoke
pump(s_recv, b_send), // spoke → browser
);
});
```
The relay reads opaque bytes off one `ChannelManager`'s reassembled stream
and writes them onto the other's write-half, which re-chunks them with the
other leg's `channel_id`. The relay does not parse the bytes — it doesn't
know if they're TTY chunks, SSH frames, or tunnel data. The channels layer
on each end does the chunk↔stream conversion; the relay just moves bytes
between two `AsyncRead + AsyncWrite` pairs.
This is why the hub's complexity collapses: the relay is **one pump
function** applied per channel, not per (protocol × transport) cell. The
`ChannelManager` is the uniform interface on both ends.
### What is NOT in the ChannelManager
To keep the boundary clean, the `ChannelManager` deliberately does **not**
hold:
- **No `ProtocolHandler` implementations.** It holds a `HandlerRegistry`
reference for ALPN lookup, but it doesn't *be* a handler. The handlers
live in their crates and register on the same registry.
- **No ALPN-specific parsing.** It does not parse `NegotiateRequest` JSON,
SSH binary frames, or tunnel target strings. It hands `params` JSON to
the handler and gets back a handler task; it hands `stream_type 3` JSON
to the handler's control handle. The `ChannelManager` is ALPN-blind.
- **No auth state.** Auth lives in the `OperationContext` that the call
protocol passes to `channel/open`. The `ChannelManager` doesn't check
scopes or ownership — that's `AccessControl::check` in
`OperationRegistry::invoke`, run before the `channel/open` handler is
called.
- **No transport coupling.** The `ChannelManager` talks to the transport
only through the `ChannelsAdapter`'s read loop and the per-channel write
pumps, both of which use `AsyncRead + AsyncWrite`. QUIC, TCP+TLS,
WebTransport, SSH channel — all look the same.
This is what makes the channels layer WASM-compatible (§WASM Compatibility)
and transport-agnostic (§Transport Agnosticism): the `ChannelManager` is
pure byte routing with no platform or protocol dependencies.
## Relationship to Existing Crates
### alknet-tty
@@ -641,6 +1329,54 @@ sent. This is an implementation constraint, not a design change.
call protocol, TTY session inside a channel. Stretch: SSH connection
inside a channel, tunnel inside a channel, recursive composition.
- **OQ-CH-08 (channel/resources staleness)**: `channel/resources` is a
snapshot of what a side exposes at call time. Resources can appear and
disappear (a docker container starts/stops, a worker connects/disconnects).
Does the channels layer need a `channel/resources/subscribe` streaming
operation (analogous to a subscription), or is polling `channel/resources`
sufficient for v1? The call protocol already has streaming
`OperationType::Subscription`; reusing it here is natural but adds a
long-lived stream on channel 0 per interested peer. Recommendation for
Phase 1: poll for v1, add subscription if staleness bites.
- **OQ-CH-09 (responder-to-initiator channel lifecycle)**: in the
`responder-to-initiator` open case (worker exposes, hub consumes), who
drives the data first? The channel is allocated on the worker's side, but
the hub is the client of the ALPN — does the hub's handler start pumping
immediately, or does the worker's handler wait for the hub to write first?
For TTY the server writes the negotiation response first; for a tunnel the
client connects first. This is ALPN-specific and probably doesn't need a
channels-layer rule, but the `direction` field's effect on which side's
`ChannelConnection` is the "server" vs "client" of the ALPN needs to be
pinned down in Phase 1.
- **OQ-CH-10 (ChannelConnection: typed destructure vs generic Connection)**:
the doc proposes that `ChannelConnection` implements the `Connection`
trait (for recursion and generic handlers) **and** can be destructured
into typed sub-stream handles (e.g. `TtyChannel { stdin, stdout, stderr,
control }`). Is the typed destructure a `ChannelConnection` method, a
separate trait, or an ALPN-specific constructor in the handler crate?
This affects whether `alknet-channels` needs to know about TTY's
`stream_type` semantics or whether the TTY crate does the destructure
itself given a `ChannelConnection`. Recommendation: the TTY crate
destructures — `alknet-channels` exposes `(channel_id, stream_type) →
(SendStream, RecvStream)` accessors and the handler crate maps those to
its typed names.
- **OQ-CH-11 (hub relay channel_id remapping)**: when the hub bridges a
browser channel (`channel_id=7`) to a spoke channel (`channel_id=12`),
the relay rewrites the `channel_id` field in each chunk header. But
`channel/control` and `channel/close` operations on channel 0 carry
`channel_id` in their *JSON payload*, not in the chunk header. Does the
hub's call-protocol forwarding (via `from_call`) need to rewrite
`channel_id` inside the `call.requested` payload too? This is a
relay-level concern: the hub is terminating channel 0 on both legs (it
runs its own `CallAdapter`), so it likely translates `channel/open` from
the browser into a *new* `channel/open` on the spoke (not a raw forward),
and the spoke's returned `channel_id` is mapped back to the browser's
side. Phase 1 must specify whether the hub translates or transparently
forwards, and how the `channel_id` mapping is maintained.
## Recommended Approach
### Crate
@@ -711,6 +1447,11 @@ channel types, orchestrated by the call protocol.** Minimum scope:
Stretch goals:
5. Two different channel types (TTY + tunnel) on the same connection
6. WASM build of the chunk splitting/recombining logic
7. Hub relay: two `ChannelManager` instances bridged by a byte pump,
validating that `channel/open` on leg A translates to `channel/open` on
leg B and the `channel_id` remapping works end-to-end (OQ-CH-11)
8. `channel/resources` discovery — one side lists its exposed ALPNs and the
other opens a channel based on the response
The POC can be built as an extension to the existing `alknet-tty-poc` or as
a standalone POC in `/workspace/alknet-channels-poc/`.