Files
glm-5.3-flash d22cecd317 docs: ADR-049/050 in indexes; changelog link refs; ported-ADR range 001..048
- architecture README: index rows for ADR-049 (establishment phase) and
  ADR-050 (pump_bidi); stale ADR-001..045 range labels corrected
- channel-operations.md: design-decision table + references list the
  two new ADRs
- CHANGELOG: missing link references for 0.4.0/0.4.1/0.5.0/0.6.0
- AGENTS.md: ADR range 001..050
2026-09-07 09:03:10 +00:00

24 KiB

status, last_updated
status last_updated
draft 2026-08-12

channel-operations.md — Channel Lifecycle on the Call Protocol

Channel lifecycle is orchestrated by the call protocol on channel 0 (ADR-036). Under the unified model (ADR-047), channel/open dissolves into per-ALPN ops (channels/<alpn>/sub, channels/<alpn>/pub), each with its own access_control, input_schema, resource_id_path, and a channel_open marker. The generic ops channel/close, channel/control, channel/resources/subscribe stay (keyed by channel_id). All go through the existing OperationContext / AccessControl::check path — no new auth machinery, no new framing.

The open ops (per-ALPN, ADR-047)

Each openable ALPN registers its own ops on the call OperationRegistry at assembly time, named channels/<alpn>/sub and/or channels/<alpn>/pub:

Op OperationType Initiator Responder Stream direction
channels/<alpn>/sub Sub consumer (subscribes) producer (streams) server → client
channels/<alpn>/pub Pub producer (publishes) consumer (receives) client → server

The channels/<alpn>/... path segment is the ALPN with the alk/ prefix stripped (ADR-047 §"Negative"). An ALPN may register one or both ops; separate op types → separate ACLs. The op spec carries the channel_open: Option<ChannelOpenSpec> marker (ADR-047 §2) — the dispatch hint that tells the channels layer "this op's stream is binary, allocate a channel for it."

channels/<alpn>/sub — subscribe to a binary stream

Request (call.requested on channel 0):

{
  "operation": "channels/tty/sub",
  "input": { "backend": "docker", "cmd": ["bash"], "container": "abc123" }
}

The input is ALPN-specific params (the former params field, now the op's input_schema). For alk/tty this is the NegotiateRequest; for alk/tunnel this is the target resource. The channels layer does not interpret input.

Response (call.responded):

{
  "output": { "channel_id": 7 }
}
field type meaning
channel_id u32 Allocated by the connection owner (ADR-047 §5). In the Sub case, the responder (producer). Both sides route chunks with this ID to the new channel.

channel_id allocation (ADR-047 §5): connection-owner allocates. The side that holds the ChannelManager for that connection allocates the channel_id via a monotonic AtomicU32 (next_id.fetch_add(1, Relaxed)) and returns it in the response. In the Sub case the responder owns the connection and allocates; in the Pub case the initiator owns and allocates. One round-trip before data flows — the same round-trip the call protocol makes for every operation. All current channel types (TTY, tunnel, SSH) already require a negotiation round-trip, so the open round-trip is not additive latency.

Error codes (CallError.code strings):

code meaning retryable
channel:forbidden AccessControl::check denied the open false
channel:allocation_failed Handler allocate failed (e.g., backend couldn't start) true (often transient)
channel:too_many_channels Per-connection or per-identity channel limit hit (ADR-040, ADR-041) false
channel:open_failed The establishment phase failed (ADR-049) — details: { reason, message }, reason ∈ dial_failed / unknown_resource / resource_shortage / handler_error / timeout false
channel:no_channels_session The op was invoked outside a channels session (no ChannelManager — ADR-047 §2) false

channel:unknown_alpn and channel:invalid_params (ADR-037) are gone — an unregistered ALPN is an ordinary NOT_FOUND (op not registered); a bad input is ordinary schema rejection.

The establishment phase (ADR-049)

An openable ALPN may register an establisher alongside its open handler (ChannelCore::register_openable_with_establisher). The establisher is the awaited preparation step ADR-047 §3 described ("validate params, consult ownership, prepare the backend, return a channel plan"): the wrapper awaits it bounded (the dispatch deadline when the op carries one, else the registration's override or the 10s ESTABLISHMENT_TIMEOUT — the earlier of the two), before replying and before spawning the pump handler. On establisher failure (error or deadline) the wrapper tears down the just-allocated channel (demux sender, opener-ledger entry, policy.on_close un-increment — the same atomic ledger-take gate every teardown path runs) and replies channel:open_failed with details: { reason, message } (the table row above). The SSH contract holds consumer-visibly: a failed open never returns a channel_id.

Registrations without an establisher behave exactly as before (the open op cannot fail post-allocation; establishment work inside the handler is invisible to the open reply). The bound applies only to the establisher — the spawned pump handler's lifetime is governed by the existing teardown machinery, unchanged. Head-of-line safety: the serving loop spawns Once invocations as independent tasks, so a slow establisher on one open op does not block other calls on channel 0.

channels/<alpn>/pub — publish a binary stream

OperationType::Pub (ADR-046). The initiator publishes a stream of call.published events to the responder; the responder's SinkHandler consumes them and returns a single call.responded. The channel_open marker on the op spec tells the channels layer the stream is binary — the call.published chunks carry binary payloads on a data channel (not JSON on channel 0).

The hub-as-proxy pattern (ADR-042, ADR-047 §1) composes on top: a worker publishes to the hub, the hub owns the resource, and consumers subscribe from the hub. The broker (topic registry, Pub↔Sub matching) is a hub concern, not an alkcall concern (ADR-047 §1, Gap B).

The generic ops (keyed by channel_id)

channel/close — tear down a channel

{
  "operation": "channel/close",
  "input": { "channel_id": 7, "reason": "exit" }
}

The responder (the side that didn't send the close) drains its reassembled stream for channel_id, signals EOF to the handler, and returns { "closed": true }. The channel_id is eligible for reuse after the drain completes (ADR-040 — monotonic IDs with wrap-around, not a free-list). reason is free-form for observability — not semantically required.

Per-identity quota decrement (ADR-047 §7). The decrement is keyed by the opener (from the per-connection opener ledger), not the closer. The decrement is called from every teardown path — close received, close sent locally, handler exit, connection drop — not just channel/close. The ledger entry is removed atomically with its decrement.

REQ-CH-06: exit-chunk-before-close ordering. The channel's data chunks MUST be written and flushed before the channel/close operation is sent on channel 0. The side closing must observe the data-channel pump complete before issuing the call operation. For TTY this is the exit-chunk-is-last invariant (ADR-055) carried forward — the exit control message rides on TTY's STREAM_CTRL_OUT (stream_type 4, inside TTY's 5-byte payload format); for tunnels it is the last data byte before close. This invariant crosses two channels (the data channel and channel 0), so the channels layer owns the ordering guarantee.

channel/control — out-of-band control on channel 0

For control that doesn't need ordering relative to data (resize, signal, keepalive):

{
  "operation": "channel/control",
  "input": {
    "channel_id": 7,
    "message": { "type": "resize", "cols": 80, "rows": 24 }
  }
}

The channels layer routes message to the handler's control handle for channel_id. The message JSON is ALPN-specific; the channels layer does not interpret it.

channel/resources/subscribe — live resource discovery

This is a Sub operation (ADR-021), not a polled Query. The call protocol has StreamingHandler / invoke_streaming (implemented and tested). The first consumer (the hub aggregating worker resources) needs live updates when workers connect/disconnect or containers start/stop.

{
  "operation": "channel/resources/subscribe",
  "input": {}
}

The responder registers a StreamingHandler that emits a ResponseEnvelope whenever the resource set changes. Each event:

{
  "output": {
    "resources": [
      {
        "alpn": "alk/tty",
        "backends": ["docker", "local"]
      },
      {
        "alpn": "alk/tunnel",
        "targets": ["container:*", "service:postgres"]
      }
    ]
  }
}
field type meaning
alpn string The ALPN this side accepts open ops for.
backends / targets [string] ALPN-specific enumeration of what's available. The channels layer doesn't interpret these; they're for the initiator to know what input to send.

The access preview (ADR-037) is dropped (ADR-047 §6) — it's on the op spec, available via services/schema. Carrying a preview in a different shape invites staleness; the spec is the authority.

The stream emits an initial snapshot immediately, then subsequent events on any change. The stream is long-lived; the subscriber cancels by dropping the subscription (ADR-020 abort cascade applies).

A channel/resources (non-subscribe, Query) operation is NOT provided. The subscription's initial snapshot serves the poll use case (subscribe, read the first event, cancel). Providing both would be redundant and would pressure consumers toward the stale-poll path.

Control-message division (DP-4 — pinned)

Control path When Examples
Call operations on channel 0 (channel/control, channel/close) Control that doesn't need ordering relative to data, or lifecycle events resize, signal, keepalive, close
Data-ordered bytes on the data channel's BiStream (handler-internal framing) Control that MUST be ordered relative to data EOF before exit, flush before close

The TTY crate's exit-chunk-is-last invariant (ADR-055) is the canonical example of data-ordered control — it rides on TTY's STREAM_CTRL_OUT (stream_type 4, inside TTY's 5-byte payload format) because it must arrive after the last data on TTY's stdout stream_type, guaranteed by TTY's per-stream_type chunk ordering within its own 5-byte format, not by a call-protocol round-trip. The channel/close operation that follows is on channel 0 and is ordered after the data pump completes (REQ-CH-06).

The control-message division is handler-internal. Under ADR-035, the channels layer has no stream_type concept — it carries the handler's framing transparently in the payload. TTY's STREAM_CTRL_IN (stream_type 3) and STREAM_CTRL_OUT (stream_type 4) are stream_types in TTY's 5-byte format (ADR-052, amended by Phase 7), not channels-layer concepts. The channels layer routes by channel_id only; the handler owns its sub-stream multiplexing on the BiStream it receives.

ACL flow (end-to-end)

A browser opening a TTY channel to a spoke through a hub (ADR-042):

  1. Browser's channel 0 → hub's channel 0: channels/tty/sub { backend: "docker", cmd: ["bash"], container: "abc123" }. The browser's identity is a bearer token (ADR-004).
  2. Hub's CallAdapter runs AccessControl::check on channels/tty/sub with the browser's identity. If denied → channel:forbidden.
  3. Hub forwards to spoke via from_call: the hub's forwarded_for handler constructs a call.requested with the hub as caller and the browser as forwarded_for (ADR-026 §3). The spoke receives channels/tty/sub with caller = hub, forwarded_for = browser.
  4. Spoke's CallAdapter runs AccessControl::check with the hub as caller (the spoke authorizes the hub — ADR-011). The spoke's ownership store verifies the hub owns container:abc123 (ADR-050 §4c — forwarded_for is metadata, not authority; the spoke sees the hub as the owner).
  5. Spoke's ChannelCore allocates channel_id via next_id.fetch_add (ADR-047 §5 — connection-owner allocates; the spoke is the responder for Sub, so it allocates), spawns TtyAdapter on the channel's BiStream with the docker backend, records opener (hub) in the ledger (ADR-047 §7), returns {channel_id}.
  6. Hub receives the spoke's {channel_id}, opens a matching channel on the browser's side (hub is the responder for the browser leg), records the channel_id mapping browser_id ↔ spoke_id, returns {channel_id: browser_id} to the consumer.
  7. Hub byte-forwards between browser_id and spoke_id with 4-byte channel_id rewrite (ADR-042 unchanged).

The hub ran zero protocol-specific auth. It ran channels/tty/sub's AccessControl::check (call-protocol machinery) and forwarded. The relay contract from ADR-042 holds unchanged in shape; only the op name changed (from generic channel/open to per-ALPN channels/tty/sub), and the channel_open marker (not prefix-matching) is how the hub recognizes and translates channel-open ops.

Per-identity channel cap (ADR-041, ADR-047 §7)

A channel slot is a resource. The cap on how many channels an identity may hold open is a quota check on that resource — parallel to OwnershipProvider::owns (ADR-011) for spawned resources. Same primitive, different resource. The cap is a peer concern, not a hub-specific concern: any accepting peer (worker or hub) enforces the cap on its inbound channels, just as it enforces AccessControl::check on the open op. The cap is also symmetric — both sides of a channels connection enforce their cap on the other's channels.

Why the cap is not in the channels layer

ChannelManager (ADR-039) is auth-blind by design — no auth state, no identity, no scopes. That decision is load-bearing (it is what makes the channels layer WASM-compatible, transport-agnostic, and ALPN-blind). So the per-identity cap lives in channels-call, where the identity is already on OperationContext (the same place AccessControl::check runs). The channels layer (channels-core) is unchanged. See ADR-041 §"Why the channels layer cannot hold the cap".

The channels-layer per-connection max_channels = 256 (ADR-040) is a per-connection memory bound (limits one connection's reassembly-buffer cost), not a DoS defense. A peer can open an unbounded number of transport connections, so a per-connection cap is not a per-peer DoS defense. The per-identity DoS defense is the cap documented here; see ADR-041 for the corrected DoS-defense framing.

The ChannelLifecyclePolicy trait

/// Per-identity channel lifecycle policy. Consulted by the open-op
/// wrapper (after `AccessControl::check`, before allocation) and on
/// every teardown path (after deallocation). Both have the identity
/// via `OperationContext` (for open) or the opener ledger (for
/// teardown — ADR-047 §7).
pub trait ChannelLifecyclePolicy: Send + Sync + 'static {
    /// Before channel allocation. Deny with `channel:too_many_channels`
    /// when the identity is over its cap. The identity is the direct
    /// caller (the peer that opened this channels connection);
    /// `forwarded_for` is metadata and is NOT consulted (ADR-026).
    fn check_open(&self, identity: &Identity) -> Result<(), ChannelError>;

    /// After channel deallocation. Decrement the per-identity count.
    /// Called on every teardown path (close received, close sent
    /// locally, handler exit, connection drop), keyed by the opener
    /// from the per-connection ledger (ADR-047 §7) — not the closer.
    fn on_close(&self, opener: &Identity);
}

Default: PerIdentityChannelPolicy::new(256)

The default constructor enforces 256 per identity out of the box — no "NoOp default + wire it later." A channels-accepting peer that constructs ChannelOperations::new(manager) with no policy argument gets PerIdentityChannelPolicy::new(256). The default is secure; opt-outs are explicit:

  • PerIdentityChannelPolicy::new(cap) — shared per-identity state (HashMap<PeerId, usize> + cap), constructed once per accepting peer and shared (via Arc) across every channels connection that peer accepts. The sharing is what makes the cap per-identity, not per-connection.
  • PerIdentityChannelPolicy::with_per_identity_caps(mapping) — per-peer-role variant: HashMap<PeerId, usize> overrides the default cap for specific peers. Used by a spoke that serves a high-fan-out hub (the hub peer's cap is set higher than a worker peer's cap — see "Relay consequence" below).
  • NoCap — no cap. Explicit opt-out for tests, POCs, and trusted single-peer deployments. Not the default.

The policy is constructed once and passed to ChannelOperations at registration time:

let policy = Arc::new(PerIdentityChannelPolicy::new(256));
let channel_ops = ChannelOperations::new(manager, policy);
channel_ops.register_on(&mut call_registry)?;

Enforcement point: between AccessControl::check and allocation

The open-op wrapper (ADR-047 §3 — the ChannelCore wrapper around the ALPN's open handler) gains the policy check after ACL and before next_id.fetch_add:

  1. ACL is already checked by OperationRegistry::invoke (the existing AccessControl::check path — unchanged).
  2. NEW: policy.check_open(&op_ctx.identity)? — deny with channel:too_many_channels if over cap.
  3. Allocate the channel_id via next_id.fetch_add(1, Relaxed) (ADR-047 §5 — connection-owner allocates).
  4. Record the opener in the per-connection ledger (ADR-047 §7).
  5. Construct the ChannelBidiStreamSource, spawn the handler, record the ChannelState (unchanged).
  6. Return the channel_id.

Every teardown path (close received, close sent locally, handler exit, connection drop) walks the ledger and calls policy.on_close(opener) per open channel, removing the ledger entry atomically with its decrement (ADR-047 §7).

Relay consequence: the spoke caps the hub, not the browser

When the hub relays a browser's channel to a spoke (ADR-042), the spoke sees the hub as the direct caller. forwarded_for carries the browser's identity as metadata (ADR-026 — forwarded_for is not authority; AccessControl::check never reads it). The channel cap follows the same shape: the spoke's ChannelLifecyclePolicy is consulted with the hub's identity, not the browser's. The spoke asks "does the hub have access to open another channel?" and the hub's quota on the spoke reflects the aggregate of all relayed channels. The hub's per-browser caps are the hub's own concern (enforced on the browser leg by the hub's own policy), not the spoke's.

Deployment consequence: a spoke that serves a hub relaying for many browsers must set the hub peer's cap higher than a worker peer's cap, or the spoke denies legitimate relayed channels when the hub's aggregate count exceeds a worker-sized cap. This is a per-peer-role policy, set by the spoke via with_per_identity_caps. The architecture provides the mechanism; the deployment sets the numbers. This is not a flaw — it is the same shape as any per-peer ACL.

Recursive channels do not bypass the cap

A recursive alk/channels-inside-alk/channels channel runs a new ChannelsAdapter with a new ChannelManager. If the same ChannelLifecyclePolicy is wired into the inner ChannelOperations, the inner channels are counted against the same identity. Recursion is not a bypass; the 13-byte-per-chunk overhead is the documented cost (ADR-035), and the cap behavior is unchanged. Recursive channels are an edge case for edge cases and not specced further.

Hub relay contract (ADR-042 — summary)

The hub translates, not transparently forwards:

  1. Call-protocol layer (channel 0): translate. The hub terminates channel 0 on both legs. channels/<alpn>/sub from the browser → hub's AccessControl::check → hub re-issues channels/<alpn>/sub on the spoke leg with forwarded_for → spoke returns its channel_id → hub maps browser-id ↔ spoke-id.
  2. Data-channel layer: byte-forward with channel_id rewrite. The relay reads chunks for browser_id, rewrites the channel_id field to spoke_id, writes onto the spoke's channels connection — and vice versa. The relay does not parse the payload.

channel/control operations on channel 0 carry channel_id in their JSON payload; the hub's CallAdapter translates these too (rewrites channel_id in the payload). The relay does not touch channel/control — it's a call operation, translated, not byte-forwarded.

The hub never runs a handler for alk/tty, alk/ssh, or alk/tunnel. It runs alk/channels (the relay) and alk/call (for its own hub-level operations + translation). The channel_open marker (ADR-047 §2) is how the hub recognizes a channel-open op during from_call discovery (ADR-047 §1, Gap C) — the from_call relay wrapper wraps marked ops with relay machinery.

Design Decisions

All design decisions are documented as ADRs in decisions/.

ADR Decision Summary
037 Channel Lifecycle Operations The generic ops; direction pinned (amended by ADR-047 — channel/open dissolves; direction removed)
047 Openable ALPNs Are Operations Per-ALPN open ops; channel_open marker; ChannelCore wrapper; opener ledger
036 Channel 0 Pre-Negotiated Channel 0 = alk/call
042 Hub Relay Translate channel 0, byte-forward data channels
041 Per-Identity Channel Cap 256 per PeerId, enforced via ChannelLifecyclePolicy in channels-call (amended by ADR-047 §7 — opener ledger, every teardown path)
035 Pure Channel Multiplexing No stream_types; handler owns sub-mux
021 StreamingHandler The machinery channel/resources/subscribe uses
046 Pub Operation Type The Pub/Sub primitives the per-ALPN open ops build on
049 Channel-Open Establishment Phase The establisher hook; typed channel:open_failed; Establishment.plan (amendment 2); the JoinHandle lifetime contract
050 pump_bidi Two-Pump Helper The data-plane helper openable-ALPN handlers await inline
026 Forwarded-For Identity The auth chain for hub-relayed opens (and why the cap is per direct-caller, not per forwarded_for)
011 Dynamic Resource Ownership The parallel — a channel slot is a resource, the cap is a quota check; resource_id_path works again under per-ALPN ops

References

  • ADR-047: openable ALPNs are operations (the unifying ADR — per-ALPN open ops, channel_open marker, ChannelCore wrapper, opener ledger)
  • ADR-049: channel-open establishment phase (the establisher hook, channel:open_failed, the plan payload)
  • ADR-050: pump_bidi (the two-pump helper)
  • ADR-037: channel lifecycle operations (amended by ADR-047)
  • ADR-041: per-identity channel cap (amended by ADR-047 §7 — opener ledger, every teardown path)
  • ADR-042: hub relay (the translate contract)
  • ADR-046: Pub operation type (the Pub/Sub primitives)
  • /workspace/@alkdev/alknet/docs/research/call-channels-unification/ findings.md — the research that surfaced the unification and the gaps