docs(arch): draft alknet-docker architecture specs (ADRs 058-063, OQs 048-051)

Greenfield architecture spec set for the alknet-docker crate — a thin,
single-host bollard wrapper exposing docker container/image operations
as call-protocol ops on the shared alknet/call ALPN, plus a
DockerTtyBackend (impl TtyBackend) behind a tty feature for interactive
terminal sessions into containers over alknet/tty.

Six ADRs:
- 058: docker ops on alknet/call (no separate ALPN; raw-carriage handoff
  dissolved by alknet-tty extraction — interactive attach moved to
  alknet/tty via DockerTtyBackend, no carriage field on call.requested)
- 059: bollard 0.21 (verified current on crates.io) + feature selection
  (http+pipe+time; no ssl/ssh/websocket/buildkit)
- 060: container resource model (ADR-050 application) — alknet.managed/
  alknet.owner labels, list owned_only flag, hosted-services operator
  role via static-resource fallback, handler-driven revoke with
  autonomous-death tolerance, resource-action vocabulary
- 061: DockerTtyBackend in alknet-docker behind tty feature (attach vs
  exec mode; POC drive_attach_raw as reference)
- 062: Docker client + OwnershipStore injection via closure capture
  (not Capabilities, not OperationContext — matches from_openapi pattern)
- 063: exit code on terminal call.responded for non-interactive exec
  (call.completed stays empty, ADR-012 unchanged)

Four spec docs (crates/docker/): README, overview, docker-operations,
docker-tty-backend. Four deferred-scope OQs (048-051): network/volume
ops, buildkit, system events subscription, create options surface.

Updates the tty-backend.md Backend implementations table (DockerTtyBackend
row now specced) and the architecture README (doc table, ADR table,
current-state paragraph, OQ count).

Grounded in the alknet-docker POC (docs/research/alknet-docker/poc-summary.md)
which validated the hard parts; the remaining lifecycle ops are mechanical
bollard wrapping. Reviewed by architecture-reviewer subagent; criticals
(the docker_client injection model conflicting with the Capabilities
contract) resolved via ADR-062/063 before commit.
This commit is contained in:
glm-5.2 committed 2026-07-08 09:29:22 +00:00
1 parent 5873fa22d6
commit 1ce11717d8
17 files changed
+3196 -9

No files matched your search

+15 -3
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-06
last_updated: 2026-07-08
---
# Alknet Architecture
@@ -24,12 +24,14 @@ The alknet-call crate is **implemented and reviewed** — both the server-side c
**alknet-tty specs drafted.** The alknet-tty crate (terminal session protocol handler — `ProtocolHandler` on `alknet/tty`, two-carriage wire format with a raw chunk codec + JSON control channel, backend-agnostic via a `TtyBackend` trait) now has architecture specs: [crates/tty/](crates/tty/) (overview, tty-wire, tty-backend, tty-adapter, tty-local) and six ADRs — [ADR-052](decisions/052-alknet-tty-wire-format-and-two-carriage.md) (wire format: `alknet/tty` ALPN, JSON negotiation frame then raw chunks, fixed channel set 0-3, control as JSON), [ADR-053](decisions/053-ttybackend-trait-and-ttyhandle.md) (`TtyBackend` trait + `TtyHandle`; `exit_code` as a `Future`; backends need not be natively async — REQ-TTY-01 from the local-PTY POC; `TtyControlHandle` newtype for `Clone`-ability — `Clone` is not object-safe), [ADR-054](decisions/054-local-tty-backend-sibling-crate.md) (`alknet-tty-local` sibling crate behind a `local` feature re-export; PTY vs pipe per-session; the runner pattern preserved), [ADR-055](decisions/055-exit-code-on-control-chunk.md) (exit code on a stream_type 3 control chunk; "exit chunk is last" invariant; adapter owns the ordering), [ADR-056](decisions/056-backend-cleanup-on-session-cancel.md) (backend cleanup contract: dropping the `exit_code` future on session cancel MUST kill the session target — closes the orphaned-process gap the local-PTY POC surfaced for the waiter thread), [ADR-057](decisions/057-alknet-tty-no-alknet-call-dep.md) (alknet-tty does not depend on alknet-call — the negotiation framing is self-contained; the earlier "reuse `FrameFramedReader`" claim was unsound because the utility is welded to `EventEnvelope` deserialization). The specs are grounded in the alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`) and the alknet-tty POC (`/workspace/alknet-tty-poc/`, built 2026-07-05), which validated the wire format, the control channel, the local-PTY bridge, and the signal-delivery contract (REQ-TTY-02). The docker and SSH backends are future crates that implement the `TtyBackend` trait — out of scope for this spec set, but the trait shape is committed so they can be built against it.
**alknet-docker specs drafted.** The alknet-docker crate (docker operations on the shared `alknet/call` ALPN + `DockerTtyBackend` behind a `tty` feature) now has architecture specs: [crates/docker/](crates/docker/) (overview, docker-operations, docker-tty-backend) and six ADRs — [ADR-058](decisions/058-alknet-docker-on-alknet-call.md) (docker ops register on `alknet/call`, not a separate `alknet/docker` ALPN; the raw-carriage handoff the POC struggled with is dissolved by the alknet-tty extraction — interactive attach moved to `alknet/tty` via `DockerTtyBackend`, no `carriage` field on `call.requested`), [ADR-059](decisions/059-bollard-021-dependency-and-features.md) (bollard 0.21, verified current on crates.io; features `http`+`pipe`+`time`, no `ssl`/`ssh`/`websocket`/`buildkit` — single-host by construction, fleet is a call-protocol concern), [ADR-060](decisions/060-container-resource-model-and-label-namespace.md) (ADR-050 application to bollard: `alknet.managed`/`alknet.owner` labels; `list` `owned_only` flag; hosted-services operator role via the static-resource fallback; handler-driven `revoke` on `remove` with autonomous-death tolerance), [ADR-061](decisions/061-docker-tty-backend-in-alknet-docker.md) (`DockerTtyBackend` in alknet-docker behind a `tty` feature, not a sibling crate; attach vs exec mode; the POC's `drive_attach_raw` as the reference), [ADR-062](decisions/062-docker-client-injection-via-closure-capture.md) (the `Docker` client + `OwnershipStore` are closure-captured at registration time, not read from `OperationContext` and not smuggled through `Capabilities` — `Capabilities` is for secret material only per ADR-014; matches the `from_openapi` pattern), [ADR-063](decisions/063-exit-code-on-terminal-call-responded.md) (non-interactive exec puts `{ "exitCode": N, "terminal": true }` on a final `call.responded` before `call.completed` — `call.completed` stays empty, ADR-012 unchanged). The specs are grounded in the alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`, `/workspace/alknet-docker-poc/`), which validated the hard parts (interactive attach, logs subscription, exec with exit code); the remaining lifecycle operations are mechanical bollard wrapping. The two use cases — disposable dev containers (coordinator-spawned, ownership-recorded) and long-running hosted services (operator-managed, static-resource fallback, per `/workspace/system/dev1/docker.md`) — both work through one `AccessControl` model (ADR-050/060). The `DockerTtyBackend` fills the `TtyBackend` row the alknet-tty spec left open. Four OQs (048–051) track deferred scope: network/volume ops, buildkit, system events subscription, and the full `CreateContainerOptions` surface (deferred to v1 implementation).
## Architecture Documents
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Workspace-level overview, crate graph, shared types, design principles |
| [open-questions.md](open-questions.md) | draft | OQ index — theme-grouped tables + Deferred/Blocked section; per-OQ files in [`questions/`](open-questions/questions/) |
| [open-questions.md](open-questions.md) | draft | OQ index — theme-grouped tables + Deferred/Blocked section; per-OQ files in [`questions/`](questions/) |
| [crates/core/README.md](crates/core/README.md) | draft | alknet-core crate index |
| [crates/core/core-types.md](crates/core/core-types.md) | draft | ProtocolHandler, HandlerError, Connection, BiStream, StreamError |
| [crates/core/endpoint.md](crates/core/endpoint.md) | draft | ALPN router, HandlerRegistry, accept loop, shutdown |
@@ -52,6 +54,10 @@ The alknet-call crate is **implemented and reviewed** — both the server-side c
| [crates/tty/tty-backend.md](crates/tty/tty-backend.md) | draft | `TtyBackend` trait, `TtyParams`, `TtyHandle`, `TtyControl` — the backend inversion point. Carries REQ-TTY-01 (backends need not be natively async) |
| [crates/tty/tty-adapter.md](crates/tty/tty-adapter.md) | draft | `TtyAdapter` (`ProtocolHandler` on `alknet/tty`): session lifecycle, three-pump bidirectional driver, negotiation errors, exit-chunk ordering (ADR-055), access control |
| [crates/tty/tty-local.md](crates/tty/tty-local.md) | draft | `alknet-tty-local` sibling crate: `LocalTtyBackend` via `portable_pty` (PTY) and `std::process::Command` (pipe/runner). Carries REQ-TTY-02 (signal forwarding to the process group) |
| [crates/docker/README.md](crates/docker/README.md) | draft | alknet-docker crate index |
| [crates/docker/overview.md](crates/docker/overview.md) | draft | Crate purpose, two-role design (call ops + DockerTtyBackend), dependencies, ALPN, label namespace, feature gates, assembly-layer wiring |
| [crates/docker/docker-operations.md](crates/docker/docker-operations.md) | draft | Operation surface: lifecycle (Query/Mutation), logs/exec/pull (Subscription via StreamingHandler), access control (ADR-050/060), label namespace, teardown coupling |
| [crates/docker/docker-tty-backend.md](crates/docker/docker-tty-backend.md) | draft | `DockerTtyBackend` (impl `TtyBackend`): attach vs exec mode, `TtyHandle` field mapping, `TtyControl` → bollard resize/signal, `exit_code` Drop-kill (ADR-056) |
| [crates/vault/README.md](crates/vault/README.md) | stable | alknet-vault crate index |
| [crates/vault/mnemonic-derivation.md](crates/vault/mnemonic-derivation.md) | stable | BIP39, SLIP-0010, BIP-0032, derivation paths, key types |
| [crates/vault/encryption.md](crates/vault/encryption.md) | stable | AES-256-GCM, EncryptedData, key versioning, salt (Phase B reserved) |
@@ -119,10 +125,16 @@ The alknet-call crate is **implemented and reviewed** — both the server-side c
| [055](decisions/055-exit-code-on-control-chunk.md) | Exit Code on a Control Chunk (the Last Chunk Before Stream Close) | Accepted |
| [056](decisions/056-backend-cleanup-on-session-cancel.md) | Backend Cleanup on Session Cancel (Drop of `exit_code` Kills) | Accepted |
| [057](decisions/057-alknet-tty-no-alknet-call-dep.md) | alknet-tty Does Not Depend on alknet-call (Self-Contained Negotiation Framing) | Accepted |
| [058](decisions/058-alknet-docker-on-alknet-call.md) | alknet-docker Registers on `alknet/call` (No Separate ALPN) | Accepted |
| [059](decisions/059-bollard-021-dependency-and-features.md) | bollard 0.21 Dependency and Feature Selection | Accepted |
| [060](decisions/060-container-resource-model-and-label-namespace.md) | Container Resource Model and Label Namespace | Accepted |
| [061](decisions/061-docker-tty-backend-in-alknet-docker.md) | DockerTtyBackend in alknet-docker | Accepted |
| [062](decisions/062-docker-client-injection-via-closure-capture.md) | Docker Client and OwnershipStore Injection via Closure Capture | Accepted |
| [063](decisions/063-exit-code-on-terminal-call-responded.md) | Exit Code on a Terminal `call.responded` for Non-Interactive Exec | Accepted |
## Open Questions
Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (47 OQs across 15 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](open-questions/questions/) (`NNN-slug.md`, mirroring the ADR convention).
Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (51 OQs across 16 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](questions/) (`NNN-slug.md`, mirroring the ADR convention).
## Document Lifecycle
+166
View File
@@ -0,0 +1,166 @@
---
status: draft
last_updated: 2026-07-08
---
# alknet-docker
The docker operations crate for the ALPN-as-service architecture: a
thin, single-host bollard wrapper that exposes docker container and
image operations as call-protocol operations on the shared
`alknet/call` ALPN, plus (behind a `tty` feature) a `DockerTtyBackend`
implementing `alknet-tty`'s `TtyBackend` trait for interactive terminal
sessions into containers.
## What
alknet-docker does two things:
1. **Call operations.** A set of `OperationSpec`-registered operations
(`docker/container/*`, `docker/image/*`) on the shared
`alknet/call` ALPN, mapping bollard's docker API to the call
protocol's `Query` / `Mutation` / `Subscription` dispatch paths.
Lifecycle (create/start/stop/remove/list/inspect) is `Query`/
`Mutation`; logs, non-interactive exec, and image pull are
`Subscription` (streaming via `StreamingHandler`, ADR-049). The
operations declare `AccessControl` against the ADR-050 container-
as-resource model. Decided in
[ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md).
2. **`DockerTtyBackend`** (behind the `tty` feature). An
`impl TtyBackend` (ADR-053) wrapping `bollard::attach_container()`
and `bollard::exec::start_exec` with `tty: true`, for interactive
terminal sessions into containers over `alknet/tty`. This is the
`TtyBackend` row the alknet-tty spec left open ("future, out of
scope here"). Decided in
[ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md).
The two use cases the crate serves (per the user's brief and the system
docs):
- **Disposable dev containers** (the common case by volume) — a
coordinator spawns a container for an implementation agent or an
isolated env. alknet-docker's `create` records ownership
(`OwnershipStore::record`); the coordinator owns the container for
its lifetime; `remove` revokes. Interactive sessions into the
container go through `DockerTtyBackend` on `alknet/tty`.
- **Long-running hosted services** (the production server case) —
rarely-changing services (reverse-proxy, postgres, redis, gitea on
dev1, per `/workspace/system/dev1/docker.md`) created by an operator
via `docker compose`, not via alknet. alknet-docker's operations
(start/stop/inspect/logs) manage them; the operator role reaches them
via the static-resource fallback (ADR-060 §3).
## Documents
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Crate purpose, two-role design (call ops + DockerTtyBackend), dependencies, ALPN, label namespace, feature gates, assembly-layer wiring |
| [docker-operations.md](docker-operations.md) | draft | The operation surface: lifecycle (Query/Mutation), logs/exec/pull (Subscription via StreamingHandler), access control, label namespace, teardown coupling |
| [docker-tty-backend.md](docker-tty-backend.md) | draft | `DockerTtyBackend` (impl `TtyBackend`): attach vs exec mode, `TtyHandle` field mapping, `TtyControl` → bollard resize/signal, `exit_code` Drop-kill (ADR-056) |
## Applicable ADRs
| ADR | Title | Relevance |
|-----|-------|-----------|
| [058](../../decisions/058-alknet-docker-on-alknet-call.md) | alknet-docker Registers on `alknet/call` | Docker ops are call-protocol operations, not a separate ALPN; raw-carriage dissolved by alknet-tty extraction |
| [059](../../decisions/059-bollard-021-dependency-and-features.md) | bollard 0.21 Dependency and Feature Selection | Version pin (0.21, verified current); features `http`+`pipe`+`time`, no `ssl`/`ssh`/`websocket`/`buildkit` |
| [060](../../decisions/060-container-resource-model-and-label-namespace.md) | Container Resource Model and Label Namespace | ADR-050 application: `alknet.managed`/`alknet.owner` labels; `list` owned_only flag; hosted-services static-resource fallback; handler-driven revoke + autonomous-death tolerance |
| [061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) | DockerTtyBackend in alknet-docker | Backend in alknet-docker behind `tty` feature; attach/exec mode split; POC `drive_attach_raw` as reference |
| [003](../../decisions/003-crate-decomposition.md) | Crate Decomposition | alknet-docker depends on alknet-core + alknet-call (ops) and alknet-tty (tty feature); no handler-depends-on-handler violation |
| [012](../../decisions/012-call-protocol-stream-model.md) | Call Protocol Stream Model | The wire format docker ops use (`EventEnvelope`, no `carriage` field) |
| [017](../../decisions/017-call-protocol-client-and-adapter-contract.md) | Call Protocol Client and Adapter Contract | `from_call` re-export — the proxy pattern's mechanism for hub→worker docker management |
| [023](../../decisions/023-operation-error-schemas.md) | Operation Error Schemas | Docker ops declare domain error codes (CONTAINER_NOT_FOUND, IMAGE_NOT_FOUND, etc.) |
| [024](../../decisions/024-operation-registry-layering.md) | Operation Registry Layering | Docker ops register in the curated Layer 0 at startup |
| [029](../../decisions/029-peer-graph-routing-model.md) | Peer-Graph Routing Model | Head→worker docker management via `PeerRef` / `invoke_peer` |
| [032](../../decisions/032-forwarded-for-identity.md) | Forwarded-For Identity | End-user identity as metadata when a coordinator proxies docker ops |
| [049](../../decisions/049-streaming-handler-for-subscriptions.md) | Streaming Handler for Subscriptions | `StreamingHandler` for logs/exec/pull; exit code on final `call.responded` |
| [050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Dynamic Resource Ownership | Containers as `AccessControl` resources; the model ADR-060 applies |
| [052](../../decisions/052-alknet-tty-wire-format-and-two-carriage.md) | alknet-tty Wire Format | The `alknet/tty` wire format `DockerTtyBackend` sessions use |
| [053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) | TtyBackend Trait and TtyHandle | The trait `DockerTtyBackend` implements |
| [055](../../decisions/055-exit-code-on-control-chunk.md) | Exit Code on a Control Chunk | The exit-chunk ordering `DockerTtyBackend`'s `exit_code` feeds into |
| [056](../../decisions/056-backend-cleanup-on-session-cancel.md) | Backend Cleanup on Session Cancel | `DockerTtyBackend`'s `exit_code` future `Drop` kills the container/exec |
| [014](../../decisions/014-secret-material-flow-and-capability-injection.md) | Secret Material Flow and Capability Injection | `Capabilities` is for secret material only — the `Docker` handle is not secret (ADR-062) |
| [062](../../decisions/062-docker-client-injection-via-closure-capture.md) | Docker Client and OwnershipStore Injection via Closure Capture | Closure capture at registration; `Capabilities` for secrets only; matches `from_openapi` pattern |
| [063](../../decisions/063-exit-code-on-terminal-call-responded.md) | Exit Code on a Terminal `call.responded` for Non-Interactive Exec | `{ "exitCode": N, "terminal": true }` on final `call.responded`; `call.completed` stays empty |
## Relevant Open Questions
| OQ | Title | Status | Relevance |
|----|-------|--------|-----------|
| OQ-048 | Network and volume operation surface | deferred(scope) | Network/volume CRUD deferred; v1 is containers + images |
| OQ-049 | Image build (buildkit) scope | deferred(scope) | `buildkit` feature deferred; v1 has `image/pull` + `image/list` + `image/inspect` |
| OQ-050 | Docker system events subscription | deferred(scope) | `docker/system/events` subscription for stale-ownership cleanup deferred |
| OQ-051 | Container create options surface | deferred(scope) | Full `CreateContainerOptions` (mounts, port bindings, networks) surface deferred to v1 implementation |
## Key Design Principles
1. **Single-host, bollard-specific.** alknet-docker talks to one local
docker daemon over `/var/run/docker.sock`. The fleet case (multiple
daemons on different machines) is a call-protocol concern — a
`CallClient` per remote daemon, each running alknet-docker locally
— not a bollard-feature concern. No `ssl`/`ssh` features, no remote
daemon over TLS. See [overview.md](overview.md) and
[ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md).
2. **Call operations on `alknet/call`, not a separate ALPN.** Docker
ops are ordinary call-protocol operations with `OperationSpec`s,
`AccessControl`, and service discovery. They inherit `from_call`
re-export and peer routing. The one operation that didn't fit
(interactive attach) moved to `alknet/tty` via `DockerTtyBackend`.
No `carriage` field on `call.requested`, no parallel dispatch
surface. See [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md).
3. **Containers are runtime-spawned resources (ADR-050).** `create`
records ownership; `remove` revokes; `exec`/`stop`/`inspect` check
ownership via `OperationSpec.resource_id_path`. The
hosted-services case (operator-managed, pre-existing containers)
works through the static-resource fallback. Two labels
(`alknet.managed`, `alknet.owner`) mark alknet-spawned containers
for the `list` filter and the cross-check. See
[docker-operations.md](docker-operations.md) and
[ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md).
4. **Interactive terminal is a tty concern, not a docker concern.**
`DockerTtyBackend` implements `alknet-tty`'s `TtyBackend` trait,
behind a `tty` feature. The wire format, the session lifecycle, and
the exit-chunk ordering live in alknet-tty; the backend produces
bollard-backed handles. This is the same inversion as
`LocalTtyBackend` (local process) and `SshTtyBackend` (SSH). See
[docker-tty-backend.md](docker-tty-backend.md) and
[ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md).
5. **bollard 0.21, verified current.** The POC used a local 0.21
checkout; the crate depends on published 0.21 from crates.io
(verified as latest). Features: `http` + `pipe` (default, local
daemon) + `time` (log timestamps). No `websocket` (reliable attach
only), no `ssl`/`ssh` (fleet is call-protocol), no `buildkit`
(deferred). See [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md).
6. **Mechanical mapping, no feasibility risk.** The POC validated the
hard parts (raw carriage attach, logs subscription, exec with exit
code). The remaining lifecycle operations (create/start/stop/remove/
list/inspect) are mechanical bollard wrapping — `Query`/`Mutation`
with single `call.responded` responses, "the boring case" (POC §"What
the POC Does NOT Validate" #4). See
[docker-operations.md](docker-operations.md).
## References
- `docs/research/alknet-docker/poc-summary.md` — the POC that validated
the hard parts (interactive attach, logs subscription, exec with exit
code) and surfaced the open unknowns this spec set resolves
- `/workspace/alknet-docker-poc/` — the POC source (`src/ops.rs`
`DockerOps`, `src/raw.rs` chunk codec, `src/frame.rs` EventEnvelope
mirror, `tests/integration.rs` 6 tests against a live daemon)
- `/workspace/bollard/` — bollard 0.21.0 source (the local checkout the
POC used; identical API surface to the published 0.21)
- `/workspace/@alkdev/dispatch/` — the dispatch POC (bollard 0.18,
`dispatch.managed=true` labels, SSH-tunnel fleet model — the prior
art this crate generalizes and the friction it removes)
- `/workspace/system/dev1/docker.md` — the production hosted-services
use case (reverse-proxy, postgres, redis, gitea on dev1)
- `/workspace/@alkdev/reverse-proxy/deploy/docker-compose.yml` — the
reverse-proxy's docker setup (operator-created, not alknet-spawned)
- `docs/architecture/crates/tty/` — the alknet-tty spec
(`DockerTtyBackend` is the row `tty-backend.md` left open)
@@ -0,0 +1,537 @@
---
status: draft
last_updated: 2026-07-08
---
# alknet-docker — Operations
The operation surface alknet-docker registers on the shared
`alknet/call` `OperationRegistry`: lifecycle operations (Query/
Mutation), streaming operations (Subscription via `StreamingHandler`),
access control (ADR-050/060 application), the label namespace, and
teardown coupling. This document specifies what an implementer builds
against. The ALPN decision (shared `alknet/call`, no separate
`alknet/docker`) is in [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md).
## What
alknet-docker exports a `register_docker_ops()` function (or a
`DockerOps` registration bundle) that the assembly layer calls to add
docker operations to the shared `OperationRegistryBuilder`. Each
operation is a `HandlerRegistration` with:
- An `OperationSpec` (name, namespace, op_type, schemas,
`access_control`, `resource_id_path` per ADR-050).
- A `HandlerKind` (`Once` for Query/Mutation, `Stream` for
Subscription — ADR-049).
- `provenance: Local` (assembly-registered, can compose).
- A `CompositionAuthority` (the scopes the docker ops run under when
composed — e.g., `container:list`, `container:exec`).
- Empty `Capabilities` (local bollard, no outbound credentials).
The operations are dispatched by the existing `CallAdapter` through
`invoke()` / `invoke_streaming()`. There is no docker-specific
dispatch code in alknet-call.
## Why
The POC (`docs/research/alknet-docker/poc-summary.md`) validated the
hard parts (interactive attach, logs subscription, exec with exit
code) and noted the remaining lifecycle operations are "mechanical
bollard wrapping, no feasibility risk... `Query`/`Mutation` operations
with single `call.responded` responses, the boring case" (POC §"What
the POC Does NOT Validate" #4). This document specifies the boring
case + the validated streaming cases as a concrete operation surface.
The operations declare `AccessControl` against the ADR-050
container-as-resource model, applied to bollard's API via ADR-060
(label namespace, `list` filter, hosted-services fallback, teardown
coupling). No new auth model is invented; the existing model is
applied.
## Architecture
### Operation Surface (v1 scope)
| Operation | Type | bollard method | `resource_id_path` | Notes |
|-----------|------|---------------|-------------------|-------|
| `docker/container/list` | Query | `list_containers` | (none — list case) | `owned_only` input flag (ADR-060 §2) |
| `docker/container/inspect` | Query | `inspect_container` | `$.containerId` | |
| `docker/container/create` | Mutation | `create_container` | (none — creates the resource) | Records ownership (ADR-060 §5) |
| `docker/container/start` | Mutation | `start_container` | `$.containerId` | |
| `docker/container/stop` | Mutation | `stop_container` | `$.containerId` | |
| `docker/container/remove` | Mutation | `remove_container` | `$.containerId` | Revokes ownership (ADR-060 §4) |
| `docker/container/restart` | Mutation | `restart_container` | `$.containerId` | |
| `docker/container/logs` | Subscription | `logs` | `$.containerId` | StreamingHandler; `follow` input flag |
| `docker/container/exec` | Subscription | `create_exec` + `start_exec` + `inspect_exec` | `$.containerId` | StreamingHandler; `tty:false` only; exit code on final `call.responded` |
| `docker/image/list` | Query | `list_images` | (none) | |
| `docker/image/pull` | Subscription | `create_image` | (none) | StreamingHandler; progress events |
| `docker/image/inspect` | Query | `inspect_image` | (none) | |
**Out of scope for v1 (deferred):**
- Network operations (`docker/network/*`) — OQ-048.
- Volume operations (`docker/volume/*`) — OQ-048.
- Image build (`docker/image/build` with buildkit) — OQ-049.
- System events (`docker/system/events`) — OQ-050.
- Full `CreateContainerOptions` surface (mounts, port bindings,
networks) — OQ-051. v1 `create` accepts the image, command, env,
labels, and name; the full options surface is a v1 implementation
refinement.
- Interactive exec / attach (`tty: true`) — not a call operation;
`DockerTtyBackend` on `alknet/tty` (ADR-061).
### Lifecycle Operations (Query / Mutation)
The lifecycle operations are the "boring case": a single bollard
async method call, a single `call.responded` with the JSON result
(or `call.error` on failure). The handler is a `Handler` (not
`StreamingHandler`), registered as `HandlerKind::Once`.
The `start`, `stop`, and `restart` operations share the `inspect`
shape (a `Mutation` with `resource_id_path: "$.containerId"`,
`container:<action>` scope, a single bollard method call, a single
`call.responded`). The representative example below is `inspect`;
the others differ only in the bollard method and the declared
scope.
```rust
// docker/container/inspect — representative lifecycle op.
// The handler closure captures the bollard Docker client by
// Arc::clone at registration time (ADR-062); it does not read the
// client from OperationContext.
let docker_clone = docker.clone(); // Arc<bollard::Docker>
let handler: Handler = Arc::new(move |input, ctx| {
let docker = docker_clone.clone();
Box::pin(async move {
let container_id = input["containerId"].as_str()
.ok_or_else(|| /* INVALID_INPUT */)?;
match docker.inspect_container(container_id, None::<()>).await {
Ok(info) => ResponseEnvelope::ok(to_json_value(info)),
Err(bollard::Error::DockerResponseServerError { status_code: 404, .. }) =>
ResponseEnvelope::error(CallError {
code: "CONTAINER_NOT_FOUND".into(),
message: format!("container {} not found", container_id),
retryable: false, details: None,
}),
Err(e) => ResponseEnvelope::error(CallError {
code: "DOCKER_ERROR".into(),
message: e.to_string(),
retryable: false, details: None,
}),
}
})
});
```
The handler extracts the container ID from the input, calls the
bollard method, and maps the result to a `ResponseEnvelope`. The
bollard error is mapped to a declared domain error code
(`CONTAINER_NOT_FOUND`, `DOCKER_ERROR`) per ADR-023. The
`OperationSpec.error_schemas` declare these codes so clients get
typed error enums.
### Handler Injection (ADR-062)
The `Docker` client and `OwnershipStore` are **closure-captured** at
registration time, not read from `OperationContext` or smuggled
through `Capabilities`. This matches the established `from_openapi`
pattern (ADR-017): non-secret shared state via closure capture,
secret material via `Capabilities`. A bollard `Docker` handle to a
local unix socket is not secret material — it must not go through
`Capabilities` (ADR-014's contract: `Capabilities` is for outbound
secret material only). The `register_docker_ops` function takes
`Arc<Docker>` + `Arc<dyn OwnershipStore>` + the label config + the
`CompositionAuthority`, and constructs each handler closure with the
handles it needs captured by `Arc::clone`. The `Capabilities` passed
to each registration is empty (`Capabilities::new()`) — local bollard
needs no API key. See [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md)
for the full decision and the rationale for why `Capabilities` and
`OperationContext` extension are the wrong channels.
### `docker/container/create` — ownership recording
`create` is the spawn event (ADR-060 §5). The handler:
1. Parses the input (image, command, env, labels, name, and v1-limited
`CreateContainerOptions` — OQ-051 for the full surface).
2. Injects the alknet labels: `alknet.managed=true`,
`alknet.owner=<caller_peer_id>`. The caller's peer ID comes from
`ctx.identity` (`Identity.id`, ADR-030). For a composed call (the
coordinator composing `create`), this is the coordinator's
identity, not the end user's — the proxy pattern (ADR-050 §3).
3. Calls `bollard::create_container()`.
4. On success, calls `OwnershipStore::record(identity, "container",
container_id)` (ADR-050 §1, ADR-060 §5).
5. Returns the `ContainerCreateResponse` (with the container ID) as
`call.responded`.
The label injection is the handler's job, not bollard's — bollard's
`CreateContainerOptions.labels` accepts a map; the handler merges the
alknet labels with any caller-provided labels (caller labels win on
conflict for non-`alknet.*` keys; `alknet.*` keys are reserved and
overwritten by the handler to prevent a caller spoofing ownership).
### `docker/container/remove` — ownership revocation
`remove` is the teardown event (ADR-060 §4). The handler:
1. Calls `bollard::remove_container()`.
2. On success, calls `OwnershipStore::revoke("container", container_id)`.
3. Returns `call.responded` with an empty/ok result.
Autonomous container death (a `--rm` exit, external `docker rm`,
daemon restart) is tolerated — stale ownership entries are not
promptly cleaned up (no reaper); they're inert (a reused container ID
gets a fresh `record` on its next `create`). See ADR-060 §4.
### `docker/container/list` — scope-gate + optional result-filter
`list` is the ADR-050 #4a "list case": `resource_type: "container"`,
no `resource_id_path`. The input accepts an `owned_only: bool` flag
(default `false`):
- `owned_only: false` — `bollard::list_containers()` with no label
filter; returns all containers the daemon sees. The scope check
gates the call (caller needs `container:list`). This is the
hosted-services case (operator lists all).
- `owned_only: true` — `bollard::list_containers()` with a label
filter (`label: alknet.owner=<caller_peer_id>`); returns only the
caller's containers. This is the disposable-dev-container case
(coordinator lists its own).
The result-filter is bollard-side (the label filter is in the
`ListContainersOptions`), not handler-side — bollard filters at the
daemon. The handler doesn't call `OwnershipProvider::owned_resources`
and filter in Rust; it pushes the filter to bollard via the label.
This is more efficient (the daemon filters before sending) and
correct (the label and the ownership store agree for alknet-spawned
containers; the cross-check is ADR-060 §1's secondary signal).
### Streaming Operations (Subscription)
The streaming operations (logs, exec, image pull) are `Subscription`
operations using `StreamingHandler` (ADR-049). The handler returns a
`Stream<ResponseEnvelope>`; the dispatcher's `pump_stream` writes
each `Ok(value)` as `call.responded`, an `Err` as `call.error`
(terminal), and natural stream end as `call.completed`.
#### `docker/container/logs`
Maps `bollard::container::logs()` (returns
`Stream<Item = Result<LogOutput, Error>>`) to a stream of
`call.responded` frames. The POC validated this path (POC target 2:
`docker_logs_subscription_pumps_frames_and_completes`).
```rust
// The StreamingHandler closure captures the bollard Docker client
// by Arc::clone at registration time (ADR-062). ResponseStream is
// the type alias Pin<Box<dyn Stream<Item = ResponseEnvelope> + Send>>
// (ADR-049, operation-registry.md §"Handler").
let docker_clone = docker.clone();
let handler: StreamingHandler = Arc::new(move |input, ctx| {
let docker = docker_clone.clone();
Box::pin(async_stream::stream! {
let container_id = input["containerId"].as_str().unwrap();
let follow = input["follow"].as_bool().unwrap_or(false);
let options = LogsOptionsBuilder::default()
.follow(follow)
.stdout(true).stderr(true)
.timestamps(true) // ADR-059 §4: time feature
.build();
let mut stream = docker.logs(container_id, Some(options));
while let Some(result) = stream.next().await {
match result {
Ok(LogOutput::StdOut { message }) |
Ok(LogOutput::Console { message }) => yield ResponseEnvelope::ok(json!({
"stream": "stdout",
"timestamp": /* from LogOutput */,
"text": /* message bytes as UTF-8 */
})),
Ok(LogOutput::StdErr { message }) => yield ResponseEnvelope::ok(json!({
"stream": "stderr",
"timestamp": /* ... */,
"text": /* ... */
})),
Ok(_) => yield ResponseEnvelope::ok(json!({})), // other variants
Err(e) => yield ResponseEnvelope::error(CallError {
code: "DOCKER_ERROR".into(),
message: e.to_string(),
retryable: false, details: None,
}),
}
}
// stream end → call.completed (the dispatcher's pump_stream
// writes call.completed on natural stream end — ADR-049)
}) as ResponseStream
});
```
Each `LogOutput` becomes a `call.responded` with `stream`
(stdout/stderr), `timestamp` (from the `time` feature, ADR-059 §4),
and `text` (the log line, as UTF-8 — the POC's single-`text`-field
refinement separates timestamp and text). Stream end (the container
exits for `follow: true`, or historical logs drain for `follow: false`)
→ `call.completed` (the dispatcher's `pump_stream` handles this on
natural stream end).
#### `docker/container/exec` (non-interactive, `tty: false`)
Maps `bollard::create_exec` + `start_exec` + `inspect_exec` to a
stream of `call.responded` frames, with the exit code on a final
`call.responded` before `call.completed`. The POC validated this
path (POC target 3: `docker_exec_streams_output_and_exit_code`).
The `tty` field in the input MUST be `false` (or absent). If `tty:
true`, the handler returns a single `ResponseEnvelope::error` with
code `INVALID_INPUT` and a message directing the caller to
`alknet/tty` for interactive exec (ADR-058 §3). This is the
one-operation-per-carriage invariant: the call operation is captured
output; the tty session is interactive.
The handler:
1. `create_exec(container_id, CreateExecOptions { cmd, env, tty: false, ... })`
→ exec ID.
2. `start_exec(exec_id, None)` → `StartExecResults::Attached { output, input }`.
3. Pump `output` (a `Stream<LogOutput>`) as `call.responded` frames
(stdout/stderr separated, since `tty: false`).
4. After the output stream ends, `inspect_exec(exec_id)` →
`ExecInspectResponse { exit_code, ... }`.
5. Emit a final `call.responded` with `{ "exitCode": N, "terminal": true }`.
6. Stream end → `call.completed` (the exit-code `call.responded` is
the last one before `completed`).
The exit code rides on a `call.responded`, not on `call.completed`.
The `terminal: true` flag marks this `call.responded` as the final
value before `call.completed`. The full shape decision (why
`terminal: true` on `call.responded` and not an exit code on
`call.completed` or a `call.exit` event) is in
[ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md).
#### `docker/image/pull`
Maps `bollard::image::create_image()` (returns
`Stream<Item = Result<CreateImageInfo, Error>>`) to a stream of
`call.responded` frames with progress events. Same shape as logs:
each progress event → `call.responded`, stream end → `call.completed`.
No exit code (image pull has no exit code; success is stream end
without error).
### Access Control
The operations declare `AccessControl` against the ADR-050 model,
applied via ADR-060. Three patterns:
#### Specific-container operations (`exec`, `inspect`, `start`, `stop`, `remove`, `restart`, `logs`)
```rust
OperationSpec {
name: "docker/container/exec",
access_control: AccessControl {
required_scopes: vec!["container:exec".into()],
resource_type: Some("container".into()),
resource_action: Some("exec".into()),
..
},
resource_id_path: Some("$.containerId".into()),
..
}
```
The dispatcher extracts `containerId` from the input via
`resource_id_path` and passes it to `AccessControl::check`, which
consults the `OwnershipProvider` (ADR-050 §2). The check passes if:
- the caller owns the container (the ownership store says so — the
`create` path recorded it), OR
- the caller's static `Identity.resources["container"]` includes an
action that subsumes the required action (the operator-role
fallback — ADR-060 §3, e.g., `container:manage` ⊇ `container:exec`).
#### The `list` case (`container/list`)
```rust
OperationSpec {
name: "docker/container/list",
access_control: AccessControl {
required_scopes: vec!["container:list".into()],
resource_type: Some("container".into()),
// no resource_action — the list case
..
},
// no resource_id_path — the list case (ADR-050 #4a)
..
}
```
The scope check gates the call. The `owned_only` input flag selects
the result-filter (label-based, bollard-side). See ADR-060 §2.
#### The `create` case (`container/create`)
```rust
OperationSpec {
name: "docker/container/create",
access_control: AccessControl {
required_scopes: vec!["container:create".into()],
// no resource_type — create spawns the resource; the
// ownership is recorded after the bollard call succeeds.
// The scope check gates; the ownership is a side effect.
..
},
// no resource_id_path — the container doesn't exist yet
..
}
```
`create` has no `resource_type` (the resource doesn't exist yet);
the scope check gates. The ownership recording is a handler side
effect (ADR-060 §5), not an ACL check.
#### Image operations (`image/list`, `image/pull`, `image/inspect`)
```rust
OperationSpec {
name: "docker/image/pull",
access_control: AccessControl {
required_scopes: vec!["image:pull".into()],
// no resource_type — images are not runtime-spawned resources
// in the ADR-050 sense (they're pulled, not spawned with
// per-caller ownership). Scope-gate only.
..
},
..
}
```
Images are not runtime-spawned resources with per-caller ownership
— they're shared daemon state. The scope check gates; no ownership
provider consultation.
### Error Schemas (ADR-023)
Each operation declares its domain error codes in
`OperationSpec.error_schemas`:
| Code | Operations | HTTP status (for to_openapi projection) | Description |
|------|-----------|----------------------------------------|-------------|
| `CONTAINER_NOT_FOUND` | inspect, start, stop, remove, restart, logs, exec | 404 | The container ID doesn't exist on the daemon |
| `IMAGE_NOT_FOUND` | image/inspect | 404 | The image doesn't exist locally |
| `IMAGE_PULL_FAILED` | image/pull | 502 | The pull failed (registry error, network) |
| `DOCKER_ERROR` | all | 500 | A bollard error not covered by a specific code |
| `INVALID_INPUT` | exec (tty:true) | 400 | `tty: true` on `docker/container/exec` — use `alknet/tty` |
The `INVALID_INPUT` for `tty: true` on exec carries a message
directing the caller to `alknet/tty` for interactive exec. This is
the one-operation-per-carriage invariant's user-visible surface
(ADR-058 §3).
### Composition and Peer Routing
Docker operations compose through the standard call-protocol
mechanisms:
- **`from_call` re-export (ADR-017).** A hub that wants to manage
docker on a worker imports the worker's `docker/container/*`
operations via `from_call` and re-registers them locally. The hub's
clients call the re-exported operations; the hub is the direct
caller to the worker; the worker's ownership store sees the hub as
the owner (the proxy pattern, ADR-050 §3). The end user's identity
rides as `forwarded_for` (ADR-032).
- **Peer routing (ADR-029).** A hub with multiple worker connections
composes `docker/container/exec` via `invoke_peer(PeerRef::Specific(
worker_peer_id), ...)`. The `PeerCompositeEnv` routes to the named
worker's sub-overlay. This is the head→worker fan-out primitive.
- **Local composition.** A coordinator handler that spawns a
container and then execs into it composes `docker/container/create`
then `docker/container/exec` through `OperationEnv::invoke()`. The
coordinator's `CompositionAuthority` has `container:create` +
`container:exec` scopes (static, ADR-022); the ownership check
passes because `create` recorded the coordinator as the owner.
The `docker/container/exec` with `tty: true` rejection is a
composition boundary: a coordinator composing exec must use `tty:
false` (captured output); interactive exec requires an `alknet/tty`
session, which is a different protocol (the coordinator opens a tty
connection, not a call operation). This is the same boundary as
"SSH exec" (call op) vs "SSH PTY session" (alknet-tty).
## Constraints
- **No `carriage` field on `call.requested`.** The call protocol's
wire format (ADR-012) is `EventEnvelope`-only on `alknet/call`.
The one operation that needed raw carriage (interactive attach)
moved to `alknet/tty` via `DockerTtyBackend` (ADR-058 §2, ADR-061).
A `docker/container/exec` call with `tty: true` is rejected with
`INVALID_INPUT`.
- **`resource_id_path` is required for specific-container ops.** The
dispatcher extracts the container ID from the input via the JSON
pointer; `AccessControl::check` receives it. Operations without
`resource_id_path` (`list`, `create`, image ops) don't reference a
specific container and don't consult the ownership provider for a
specific resource.
- **`create` records ownership; `remove` revokes; lifecycle ops do
neither.** Only the spawn (`create`) and the teardown (`remove`)
touch the ownership store. `start`/`stop`/`restart` act on an
existing container whose ownership was recorded at create time
(or which pre-exists and is reached via the static fallback).
- **Stale ownership entries are tolerated.** Autonomous container
death leaves a stale entry; no reaper cleans it promptly. The
entry is inert (a reused container ID gets a fresh `record`). A
`docker/system/events` subscription for prompt cleanup is deferred
(OQ-050).
- **`owned_only` on `list` is the caller's choice.** Default
`false` (returns all); `true` filters to the caller's containers.
The default and the two cases are decided in ADR-060 §2; the scope
check gates regardless of the flag.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Docker ops on `alknet/call` | [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md) | No separate ALPN; raw-carriage dissolved; `tty:true` exec rejected |
| Container resource model + labels | [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) | `alknet.managed`/`alknet.owner` labels; `list` `owned_only`; hosted-services fallback; teardown coupling |
| Streaming handler for subscriptions | [ADR-049](../../decisions/049-streaming-handler-for-subscriptions.md) | `StreamingHandler` for logs/exec/pull; exit code on final `call.responded` |
| Dynamic resource ownership | [ADR-050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Containers as `AccessControl` resources; `resource_id_path` |
| Operation error schemas | [ADR-023](../../decisions/023-operation-error-schemas.md) | `CONTAINER_NOT_FOUND`, `IMAGE_NOT_FOUND`, `DOCKER_ERROR`, `INVALID_INPUT` |
| Peer-graph routing | [ADR-029](../../decisions/029-peer-graph-routing-model.md) | Head→worker docker management |
| Forwarded-for identity | [ADR-032](../../decisions/032-forwarded-for-identity.md) | End-user identity as metadata in the proxy pattern |
| bollard 0.21 + features | [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) | `time` feature for log timestamps |
| Docker client + OwnershipStore injection | [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md) | Closure capture at registration (not `Capabilities`, not `OperationContext`); matches `from_openapi` pattern |
| Exit code on terminal `call.responded` | [ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md) | `{ "exitCode": N, "terminal": true }` on final `call.responded`; `call.completed` stays empty |
## Open Questions
See [open-questions.md](../../open-questions.md) for full details.
- **OQ-048** (deferred(scope)): Network and volume operation surface.
- **OQ-049** (deferred(scope)): Image build (buildkit) scope.
- **OQ-050** (deferred(scope)): Docker system events subscription.
- **OQ-051** (deferred(scope)): Container create options surface.
## References
- [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md) — the
ALPN decision (shared `alknet/call`, no `carriage` field)
- [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md)
— the resource model and label namespace this surface applies
- [ADR-049](../../decisions/049-streaming-handler-for-subscriptions.md)
— `StreamingHandler` for logs/exec/pull
- [ADR-050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md)
— the ownership model
- [ADR-023](../../decisions/023-operation-error-schemas.md) — declared
domain error codes
- [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md)
— the Docker client + OwnershipStore injection model (closure
capture, not `Capabilities`)
- [ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md)
— the exec exit-code frame shape (`terminal: true`)
- `docs/research/alknet-docker/poc-summary.md` — the POC (validated
logs subscription, exec with exit code; lifecycle ops are "the
boring case")
- `/workspace/alknet-docker-poc/src/ops.rs` — `DockerOps`:
`drive_logs`, `drive_exec` (the streaming-handler reference)
- `/workspace/bollard/src/container.rs` (`list_containers` :245,
`create_container` :296, `inspect_container` :777, `logs` :928),
`/workspace/bollard/src/exec.rs` (`create_exec` :172, `start_exec`
:225, `inspect_exec` :315), `/workspace/bollard/src/image.rs`
(`create_image` :120, `list_images` :66, `inspect_image` :177)
@@ -0,0 +1,482 @@
---
status: draft
last_updated: 2026-07-08
---
# alknet-docker — DockerTtyBackend
The `DockerTtyBackend` is alknet-docker's `impl TtyBackend` (ADR-053)
for interactive terminal sessions into docker containers over
`alknet/tty`. It wraps `bollard::attach_container()` (attach mode) and
`bollard::exec::start_exec` with `tty: true` (exec mode), producing a
`TtyHandle` the `TtyAdapter` pumps. This document specifies the
backend's allocation, control, and cancel-cleanup. The crate
placement is decided in
[ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md).
## What
`DockerTtyBackend` is the `TtyBackend` implementation the
`TtyAdapter` dispatches to when a client opens an `alknet/tty` session
with `backend: "docker"` in the negotiation frame. The adapter selects
the backend by the `backend` field, calls `allocate()`, and pumps the
resulting `TtyHandle`'s fields bidirectionally using the alknet-tty
wire format (ADR-052). The backend produces handles; the adapter owns
the wire format.
```rust
// behind the `tty` feature in alknet-docker
pub struct DockerTtyBackend {
docker: bollard::Docker,
label_prefix: String, // "alknet" by default; for the ownership cross-check
}
#[async_trait]
impl TtyBackend for DockerTtyBackend {
async fn allocate(&self, params: &TtyParams) -> Result<TtyHandle, TtyError>;
fn resource_id(&self, params: &TtyParams) -> Option<(&'static str, String)>;
}
```
The backend is constructed by the assembly layer with the bollard
client (shared with the call operations) and the label prefix (shared
with ADR-060's label namespace), then registered in the
`TtyAdapter`'s backend map under the `"docker"` key.
## Why
The alknet-tty spec ([tty-backend.md](../tty/tty-backend.md) §"Backend
implementations") listed the `DockerTtyBackend` location as "future,
out of scope here." This spec fills that row: the backend lives in
alknet-docker (ADR-061), behind the `tty` feature, and wraps the
bollard attach/exec API the POC validated.
The POC's `drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`)
is the seed of `allocate()`. The POC proved the mechanism (raw chunks
on a bidi stream after a JSON request); the `TtyBackend` trait
extracts the bollard interaction into a backend that produces handles,
while the `TtyAdapter` (in alknet-tty) owns the wire format and
session lifecycle. The backend is the inversion point
([ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md)).
## Architecture
### Backend Params
The `DockerTtyBackend` deserializes its params from
`TtyParams.backend_params` (the opaque `serde_json::Map` the adapter
passes through verbatim, per ADR-053 §"Backend params are opaque"):
```rust
#[derive(Deserialize)]
struct DockerBackendParams {
/// The container to attach to or exec in.
container: String,
/// "attach" (attach to the running primary process) or "exec"
/// (start a new process with a PTY). Default: "exec".
#[serde(default = "default_mode")]
mode: String,
}
fn default_mode() -> String { "exec".into() }
```
- `container` — the container ID or name. Required. The backend
validates it exists (the bollard call will fail if not; the error
maps to `TtyError::Backend`).
- `mode` — `"attach"` or `"exec"`. `"attach"` attaches to the
container's primary process (the running process, pid 1 or the
`ENTRYPOINT`); `"exec"` starts a new process in the container with a
PTY. Default `"exec"` (the common case — a new shell). For `exec`
mode, `TtyParams.cmd` is the command vector; for `attach` mode,
`cmd` is ignored (the primary process is already running).
The adapter does not interpret these fields — it passes the
negotiation frame's backend-specific JSON through as
`backend_params`, and the backend deserializes its own
strongly-typed struct (ADR-053).
### `allocate()` — attach mode
Wraps `bollard::container::attach_container()` (the reliable
HTTP-upgrade-to-TCP path, per the POC and
[ADR-059](../../decisions/059-bollard-021-dependency-and-features.md)
§3 — not the websocket path).
```rust
async fn allocate_attach(&self, params: &TtyParams, container: &str) -> Result<TtyHandle, TtyError> {
let options = AttachContainerOptionsBuilder::default()
.stdout(true).stderr(true).stdin(true)
.stream(true).logs(false)
.build();
let AttachContainerResults { output, input } = self.docker
.attach_container(container, Some(options))
.await
.map_err(|e| TtyError::Backend { message: e.to_string() })?;
// output: Stream<LogOutput> → TtyHandle.stdout (Stream<Bytes>)
// PTY mode (tty: true on the container) merges stdout/stderr;
// bollard's LogOutput on a TTY returns StdOut only, so stderr
// is None. (tty-backend.md §DockerTtyBackend notes this.)
let stdout = Box::pin(output.map(|r| match r {
Ok(LogOutput::StdOut { message }) | Ok(LogOutput::Console { message }) =>
message, // Bytes
Ok(_) => bytes::Bytes::new(),
Err(_) => bytes::Bytes::new(),
}));
// input: AsyncWrite → TtyHandle.stdin (Box<dyn AsyncWrite + Send + Unpin>)
let stdin: Box<dyn tokio::io::AsyncWrite + Send + Unpin> = Box::new(input);
// exit_code: wait for the container's process to exit, then read
// the exit status from inspect_container. This future resolves
// independently of the output stream — the adapter (ADR-055)
// awaits BOTH the stdout pump's EOF AND this future's resolve
// before sending the exit chunk, so the backend does not need to
// drain stdout itself. For a container that exits with a non-zero
// status, the State.exit_code carries it. The Drop of this future
// (cancel) kills the container (ADR-056) — see "Cancel-Cleanup".
let docker = self.docker.clone();
let container_for_exit = container.to_string();
let exit_code: BoxFuture<'static, Result<i32, TtyError>> = Box::pin(async move {
// Poll the container until it exits. wait_container returns a
// stream of WaitContainerResponse; the first (and usually
// only) one carries the exit code. Alternatively, poll
// inspect_container until State.running is false. The
// adapter's stdout pump drains the output stream in
// parallel; this future resolves when the process exits.
let mut wait_stream = docker.wait_container(&container_for_exit,
None::<WaitContainerOptions<String>>);
if let Some(result) = wait_stream.next().await {
match result {
Ok(resp) => Ok(resp.status_code.unwrap_or(-1) as i32),
Err(e) => Err(TtyError::WaitFailed { message: e.to_string() }),
}
} else {
Ok(-1)
}
});
Ok(TtyHandle {
stdin,
stdout,
stderr: None, // PTY mode merges
exit_code,
control: Some(TtyControlHandle::new(Arc::new(DockerControl {
docker: self.docker.clone(),
container: container.to_string(),
mode: "attach".into(),
}))),
})
}
```
The output stream maps `LogOutput` → `Bytes` (the adapter pumps these
as stdout chunks, stream_type 1). PTY mode (`tty: true` on the
container) merges stdout/stderr into `StdOut`, so `TtyHandle.stderr` is
`None` — the adapter pumps only stdout chunks. This matches the
`tty-backend.md` sketch: "stdout/stderr are merged when `tty: true`
(bollard's `LogOutput` on a TTY exec returns `StdOut` only), so
`TtyHandle.stderr` is `None` for the PTY case."
### `allocate()` — exec mode
Wraps `bollard::exec::create_exec` + `start_exec` with `tty: true`:
```rust
async fn allocate_exec(&self, params: &TtyParams, container: &str) -> Result<TtyHandle, TtyError> {
let cmd = &params.cmd; // argv[0] + args
let config = CreateExecOptions {
cmd: Some(cmd.clone()),
env: Some(params.env.clone()),
attach_stdout: Some(true),
attach_stderr: Some(true),
attach_stdin: Some(true),
tty: Some(true), // PTY mode
..Default::default()
};
let exec = self.docker.create_exec(container, config).await
.map_err(|e| TtyError::Backend { message: e.to_string() })?;
let exec_id = exec.id;
let StartExecResults::Attached { output, input } = self.docker
.start_exec(&exec_id, None).await
.map_err(|e| TtyError::Backend { message: e.to_string() })?;
// Same mapping as attach mode: output → stdout, input → stdin,
// stderr None (PTY mode merges).
let stdout = /* output → Stream<Bytes> */;
let stdin: Box<dyn AsyncWrite + Send + Unpin> = Box::new(input);
// exit_code: after output stream ends, inspect_exec for the exit code
let docker = self.docker.clone();
let exec_id = exec_id.clone();
let exit_code: BoxFuture<'static, Result<i32, TtyError>> = Box::pin(async move {
// Wait for output stream end (the future resolves after the
// pump drains output — the adapter's exit-chunk ordering
// awaits both stdout EOF and exit_code resolve, ADR-055).
// Then inspect_exec:
let inspect = docker.inspect_exec(&exec_id).await
.map_err(|e| TtyError::WaitFailed { message: e.to_string() })?;
Ok(inspect.exit_code.unwrap_or(-1))
});
Ok(TtyHandle { stdin, stdout, stderr: None, exit_code,
control: Some(TtyControlHandle::new(Arc::new(DockerControl {
docker: self.docker.clone(),
container: container.to_string(),
mode: "exec".into(),
exec_id: Some(exec_id),
}))) })
}
```
The exit-code future for exec mode follows the POC's pattern (POC
target 3): pump the output stream, then `inspect_exec` for the exit
code. The adapter's exit-chunk ordering (ADR-055) awaits both the
stdout pump's EOF and the `exit_code` future's resolve before sending
the exit chunk — so the `exit_code` future must not resolve until the
output stream has ended. For exec mode, the output stream end is the
signal that the exec process exited; `inspect_exec` after that returns
the exit code.
### `resource_id()` — ownership check delegation
```rust
fn resource_id(&self, params: &TtyParams) -> Option<(&'static str, String)> {
let p: DockerBackendParams = serde_json::from_value(
serde_json::Value::Object(params.backend_params.clone())
).ok()?;
Some(("container", p.container))
}
```
The adapter calls this at negotiation and checks
`OwnershipProvider::owns(identity, "container", container_id, "tty")`
if an ownership provider is wired (ADR-050, ADR-060). The backend owns
the extraction; the adapter doesn't parse docker-specific JSON. The
`"tty"` action is a distinct resource action from `"exec"` (the call
operation's action) — a caller authorized to `docker/container/exec`
(captured, non-interactive) is not automatically authorized for an
interactive tty session; the `tty` action is a separate scope. This
matches the "interactive terminal ≠ captured command" boundary
(ADR-058 §3).
### `DockerControl` — `TtyControl` impl
```rust
struct DockerControl {
docker: bollard::Docker,
container: String,
mode: String, // "attach" or "exec"
exec_id: Option<String>, // Some for exec mode (for resize_exec)
}
impl TtyControl for DockerControl {
fn resize(&self, cols: u16, rows: u16, pixel_width: u16, pixel_height: u16) {
let docker = self.docker.clone();
let container = self.container.clone();
let exec_id = self.exec_id.clone();
let mode = self.mode.clone();
tokio::spawn(async move {
let _ = match mode.as_str() {
"attach" => docker.resize_container_tty(&container,
ResizeContainerTtyOptions { width: cols, height: rows, .. }).await,
"exec" if exec_id.is_some() => docker.resize_exec(
exec_id.as_ref().unwrap(),
ResizeExecOptions { height: rows, width: cols, .. }).await,
_ => Ok(()),
};
});
}
fn signal(&self, name: &str) {
// docker has no per-exec signal API in bollard's stable surface.
// Best-effort: kill_container with the mapped signal.
// Unknown signals fall back to SIGKILL (TtyControl::signal contract).
let docker = self.docker.clone();
let container = self.container.clone();
let sig = signal_name_to_bollard(name); // SIGINT, SIGTERM, etc.
tokio::spawn(async move {
let _ = docker.kill_container(&container,
Some(KillContainerOptions { signal: sig })).await;
});
}
}
```
- `resize()` — `bollard::container::resize_container_tty()` (attach
mode) or `bollard::exec::resize_exec()` (exec mode with an exec ID).
Fire-and-forget (the `TtyControl` trait methods are sync, not async —
ADR-053; the bollard call is spawned).
- `signal()` — docker has no per-exec signal API in bollard's stable
surface. The best-effort path is `bollard::kill_container()` with
the mapped signal (SIGINT, SIGTERM, etc.); unknown signals fall back
to SIGKILL. For exec mode, the signal goes to the container's main
process, not just the exec's process — docker's exec isolation
limits per-exec signal targeting. This is within the
`TtyControl::signal` contract ("best-effort delivery to the
foreground process group") but is a docker-semantics limitation: a
user expecting `Ctrl-C` to kill only the exec process may see the
container killed. This is documented in the spec, not hidden.
The signal-name → bollard-signal mapping (`signal_name_to_bollard`)
covers the standard set (`HUP`, `INT`, `QUIT`, `TERM`, `KILL`, `USR1`,
`USR2`, `TSTP`, `CONT`) — the same set alknet-tty's control channel
supports (ADR-052 §Control Channel). Unknown names fall back to the
backend's default kill (`SIGKILL`).
### Cancel-Cleanup (ADR-056)
The `TtyHandle.exit_code` future's `Drop`-on-cancel (the adapter drops
the `TtyHandle` on connection drop, stream reset, or pump-task panic —
ADR-056) MUST kill the session target. The docker backend's
`exit_code` future wraps the kill in its `Drop`:
```rust
// The exit_code future holds a guard that kills on Drop.
struct DockerExitGuard {
docker: bollard::Docker,
container: String,
mode: String,
killed: bool,
}
impl Drop for DockerExitGuard {
fn drop(&mut self) {
if !self.killed {
let docker = self.docker.clone();
let container = self.container.clone();
tokio::spawn(async move {
let _ = docker.kill_container(&container,
Some(KillContainerOptions { signal: "SIGKILL".into() })).await;
});
}
}
}
```
- **Attach mode** — `kill_container` with `SIGKILL`. The container is
not removed (attach doesn't own the container's lifecycle — the
container may be a hosted service the operator wants to keep); the
kill terminates the process the terminal session was attached to.
- **Exec mode** — `kill_container` with `SIGKILL`, targeting the
container's main process (best-effort, per docker's exec isolation).
The exec instance's process is not separately killable in docker's
model; the container-level kill is the available mechanism.
The guard's `killed` flag prevents a double-kill if the exit_code
future resolved normally (the process exited) before the Drop. The
spawn-on-Drop pattern (tokio::spawn in the Drop impl) is the standard
way to run async cleanup from a sync `Drop`; the spawned task outlives
the dropping context.
This satisfies ADR-056's contract: dropping `exit_code` (cancel) kills
the session target. The "session target" for docker is the container
(attach) or the exec's process (exec, best-effort via the container).
A backend that returned an `exit_code` future without a kill-on-Drop
guard would violate the contract and could leave orphaned containers.
### Backend Registration
```rust
// assembly layer (behind `tty` feature)
let docker_backend = Arc::new(DockerTtyBackend::new(
docker.clone(),
"alknet".into(), // label prefix (ADR-060)
)) as Arc<dyn TtyBackend>;
let mut backends = HashMap::new();
backends.insert("docker".into(), docker_backend);
// (other backends: "local" from alknet-tty-local, "ssh" from alknet-ssh, etc.)
let tty_adapter = TtyAdapter::new(Arc::new(backends), ownership_provider);
```
The backend is registered under the `"docker"` key — the
`backend: "docker"` field in the negotiation frame selects it. A
deployment that doesn't want docker terminals doesn't register it; a
deployment that wants both docker and local terminals registers both.
## Constraints
- **The backend produces handles; the adapter owns the wire format.**
The backend does not write chunks to the bidi stream — the
`TtyAdapter` does (ADR-052). The backend's `allocate()` returns a
`TtyHandle`; the adapter pumps `stdin`, `stdout`, `exit_code`, and
dispatches `control` chunks. A backend that wrote to the wire
directly would break the wire-format invariants (the exit-chunk
ordering, ADR-055).
- **PTY mode merges stdout/stderr.** `TtyHandle.stderr` is `None` for
both attach and exec modes (both set `tty: true`). The adapter pumps
only stdout chunks (stream_type 1). This matches the PTY property
(one output stream from the slave) and the
[tty-backend.md](../tty/tty-backend.md) sketch.
- **`exit_code` future's `Drop` kills (ADR-056).** The `DockerExitGuard`
in the `exit_code` future issues `kill_container` on `Drop`-without-
resolve. The guard's `killed` flag prevents double-kill. The
spawned-on-Drop pattern runs the async kill from the sync `Drop`.
- **Signal delivery is best-effort.** docker has no per-exec signal
API in bollard's stable surface. `TtyControl::signal()` maps to
`kill_container` with the container's main process as the target.
Unknown signals fall back to `SIGKILL`. This is within the
`TtyControl::signal` contract (ADR-053) but is a docker-semantics
limitation, documented not hidden.
- **Attach doesn't own the container lifecycle.** The cancel-cleanup
kills the container (terminates the attached process) but does not
remove it. A container that was a hosted service (operator-created)
stays after the terminal session ends — the operator may want it
running. The kill (not remove) is the correct teardown for attach.
- **`resource_id()` delegates to the backend.** The adapter doesn't
parse docker JSON to extract the container ID; the backend's
`resource_id()` does. The adapter checks
`OwnershipProvider::owns(identity, "container", id, "tty")`. The
`"tty"` action is distinct from the call operation's `"exec"`
action — interactive terminal authorization is a separate scope.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| DockerTtyBackend in alknet-docker | [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) | Behind `tty` feature; attach/exec mode; POC `drive_attach_raw` as reference |
| TtyBackend trait and TtyHandle | [ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) | The trait this backend implements; backend params opaque; REQ-TTY-01 (backends need not be natively async — bollard is async, so no bridge needed) |
| Wire format | [ADR-052](../../decisions/052-alknet-tty-wire-format-and-two-carriage.md) | The chunk codec + control channel the adapter pumps to/from these handles |
| Exit code on a control chunk | [ADR-055](../../decisions/055-exit-code-on-control-chunk.md) | The adapter awaits `exit_code`, sends the exit chunk; the backend's `exit_code` resolves after output EOF + `inspect_exec` |
| Backend cleanup on session cancel | [ADR-056](../../decisions/056-backend-cleanup-on-session-cancel.md) | `exit_code` future's `Drop` kills the container (attach) or container's main process (exec, best-effort) |
| Container resource model | [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) | `resource_id()` delegation; `"tty"` action distinct from call op's `"exec"` |
| bollard 0.21 + features | [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) | No `websocket` feature — reliable `attach_container()` only |
## Open Questions
None. The backend's design is decided in ADR-061. The
signal-delivery path's precision in exec mode is a docker-semantics
limitation (docker has no per-exec signal API in bollard's stable
surface; `kill_container` targets the container's main process,
best-effort per the `TtyControl::signal` contract), not an open
question — the limitation is documented in §"DockerControl" and
acknowledged in ADR-061 §3.
## References
- [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md)
— the crate placement and attach/exec mode decision
- [ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) — the
trait this backend implements
- [ADR-052](../../decisions/052-alknet-tty-wire-format-and-two-carriage.md)
— the wire format the adapter pumps
- [ADR-055](../../decisions/055-exit-code-on-control-chunk.md) — the
exit-chunk ordering the `exit_code` future feeds into
- [ADR-056](../../decisions/056-backend-cleanup-on-session-cancel.md)
— the cancel-cleanup contract (`exit_code` future's `Drop` kills)
- [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md)
— the `resource_id()` delegation and the `"tty"` action
- `docs/research/alknet-docker/poc-summary.md` §"POC Target 1"
(interactive attach — the seed of this backend)
- `/workspace/alknet-docker-poc/src/ops.rs` — `drive_attach_raw` (the
reference for `allocate()`)
- `/workspace/bollard/src/container.rs` (`attach_container` :540,
`AttachContainerResults` :80, `LogOutput` :96,
`resize_container_tty` :687, `kill_container` :1059)
- `/workspace/bollard/src/exec.rs` (`create_exec` :172, `start_exec`
:225, `inspect_exec` :315, `resize_exec` :362, `StartExecResults` :99)
- [tty-backend.md](../tty/tty-backend.md) §"Backend implementations"
(the row this spec fills)
+334
View File
@@ -0,0 +1,334 @@
---
status: draft
last_updated: 2026-07-08
---
# alknet-docker — Overview
The docker operations crate: a thin, single-host bollard wrapper that
exposes docker container and image operations as call-protocol
operations on `alknet/call`, plus (behind a `tty` feature) a
`DockerTtyBackend` implementing `alknet-tty`'s `TtyBackend` trait. This
document covers the crate's purpose, the two-role design, its
dependency edges, the ALPN decision, the label namespace, the feature
gates, and the assembly-layer wiring. Component details are in the
sibling documents.
## What
alknet-docker is the docker integration point for the ALPN-as-service
architecture. It does two things:
1. **Registers docker operations on the shared `alknet/call`
`OperationRegistry`.** The operations — `docker/container/*` and
`docker/image/*` — are ordinary call-protocol operations with
`OperationSpec`s, `AccessControl`, and service discovery. The
existing `CallAdapter` dispatches them through `invoke()` /
`invoke_streaming()`. There is no `alknet/docker` ALPN and no
`DockerProtocolHandler`. Decided in
[ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md).
2. **Provides `DockerTtyBackend`** (behind the `tty` feature), an
`impl TtyBackend` (ADR-053) for interactive terminal sessions into
containers over `alknet/tty`. This is the backend the alknet-tty
spec listed as "future, out of scope here"
([tty-backend.md](../tty/tty-backend.md) §"Backend implementations").
Decided in [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md).
The two roles share the bollard `Docker` client, the container
identity model, and the label/ownership reasoning (ADR-060), but
diverge at the wire format: call operations use `EventEnvelope` on
`alknet/call`; interactive sessions use the raw chunk codec on
`alknet/tty` (ADR-052).
## Why
The overwhelming common use case (per the user's brief) is a hub that
runs as a call-protocol client, connects to worker nodes, and exposes
docker operations on those workers. alknet-docker is what runs on the
worker: it wraps the local docker daemon and registers the operations.
The hub composes them via `from_call` (ADR-017) + peer routing
(ADR-029) — the proxy pattern (ADR-050 §3). For interactive terminals
into the containers (an agent needing a shell), the hub opens an
`alknet/tty` session that the worker's `DockerTtyBackend` serves.
The two container use cases:
- **Disposable dev containers** (common by volume) — a coordinator
spawns a container for an implementation agent or an isolated env.
`docker/container/create` records ownership; the coordinator owns
the container; `docker/container/remove` revokes. Interactive
sessions go through `DockerTtyBackend`. The container is
short-lived; "burn it and start over" is the recovery model
(ADR-050 §4b).
- **Long-running hosted services** (less common, important) — the
production server (`/workspace/system/dev1`) hosts rarely-changing
services (reverse-proxy, postgres, redis, gitea) in docker. These
are created by an operator via `docker compose`, not via alknet.
alknet-docker's operations manage them (start/stop/inspect/logs);
the operator role reaches them via the static-resource fallback
(ADR-060 §3) — no ownership-store entry needed for pre-existing
containers an operator is statically authorized to manage.
Both use cases work through one `AccessControl` model (ADR-050) and
one operation surface. No special-casing, no "is this a managed
container?" branch.
## The Two Roles in Brief
### Role 1: Call operations on `alknet/call`
| Operation | Call op type | bollard method | Carriage |
|-----------|-------------|---------------|----------|
| `docker/container/list` | Query | `list_containers` | JSON |
| `docker/container/inspect` | Query | `inspect_container` | JSON |
| `docker/container/create` | Mutation | `create_container` | JSON |
| `docker/container/start` | Mutation | `start_container` | JSON |
| `docker/container/stop` | Mutation | `stop_container` | JSON |
| `docker/container/remove` | Mutation | `remove_container` | JSON |
| `docker/container/restart` | Mutation | `restart_container` | JSON |
| `docker/container/logs` | Subscription | `logs` | JSON (StreamingHandler) |
| `docker/container/exec` (tty:false) | Subscription | `create_exec` + `start_exec` + `inspect_exec` | JSON (StreamingHandler) |
| `docker/image/list` | Query | `list_images` | JSON |
| `docker/image/pull` | Subscription | `create_image` | JSON (StreamingHandler) |
| `docker/image/inspect` | Query | `inspect_image` | JSON |
Interactive exec (`tty: true`) and interactive attach are **not** call
operations — they are `alknet/tty` sessions via `DockerTtyBackend`.
A `docker/container/exec` call with `tty: true` in the input returns
`INVALID_INPUT` directing the caller to `alknet/tty`. The full surface,
access control, and the streaming shapes are in
[docker-operations.md](docker-operations.md).
### Role 2: `DockerTtyBackend` on `alknet/tty`
```rust
// behind the `tty` feature
pub struct DockerTtyBackend {
docker: Docker,
label_prefix: String, // "alknet" by default; configurable
}
#[async_trait]
impl TtyBackend for DockerTtyBackend {
async fn allocate(&self, params: &TtyParams) -> Result<TtyHandle, TtyError>;
fn resource_id(&self, params: &TtyParams) -> Option<(&'static str, String)>;
}
```
The backend wraps `bollard::attach_container()` (attach mode) or
`bollard::exec::start_exec` with `tty: true` (exec mode), producing a
`TtyHandle` the `TtyAdapter` pumps. The wire format, the session
lifecycle, and the exit-chunk ordering live in alknet-tty; the backend
produces bollard-backed handles. See
[docker-tty-backend.md](docker-tty-backend.md).
## Dependencies
```
alknet-docker
├── alknet-core (Identity, AccessControl, OwnershipProvider/Store — ADR-050)
├── alknet-call (OperationSpec, Handler, StreamingHandler, OperationRegistryBuilder)
├── bollard 0.21 (http + pipe + time features; the docker API — ADR-059)
├── alknet-tty (TtyBackend trait — only behind `tty` feature; ADR-061)
├── tokio (async runtime — bollard is tokio-based)
├── serde / serde_json (operation input/output, label values)
└── thiserror (DockerError enum)
```
The `alknet-call` edge is the protocol-foundation exception
(ADR-003 Am. 1): alknet-docker consumes `OperationSpec`, `Handler`,
`StreamingHandler`, and `OperationRegistryBuilder` from alknet-call
to register its operations. This is the same edge `alknet-http` has
(HTTP uses the call protocol types). alknet-docker does **not** depend
on `alknet-http`, `alknet-tty-local`, or any other handler crate — the
no-handler-depends-on-another-handler rule (ADR-003) is preserved.
The `alknet-tty` edge is **only** behind the `tty` feature. A
deployment using alknet-docker for call operations only (the common
case) does not pull in alknet-tty. See Feature Gates below.
### bollard version and features
bollard 0.21, verified as the latest published version on crates.io
(as of 2026-07-08). The POC's local checkout (0.21.0) matches the
published version — identical API surface. Features: `http` + `pipe`
(default, local daemon connect) + `time` (log timestamp typing). No
`ssl` (no remote daemon over TLS — fleet is call-protocol), no `ssh`
(the SSH-tunnel fleet model is what alknet replaces), no `websocket`
(reliable attach only, per the POC), no `buildkit` (deferred, OQ-049).
See [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md).
## ALPN
| ALPN | Role | Handler | Transport |
|------|------|---------|-----------|
| `alknet/call` | Call operations | `CallAdapter` (shared, not docker-specific) | QUIC bidi stream, `EventEnvelope` framing |
| `alknet/tty` | Interactive terminal sessions | `TtyAdapter` (shared, not docker-specific) | QUIC bidi stream, raw chunk codec (ADR-052) |
alknet-docker registers **no ALPN of its own**. The call operations
are dispatched by the shared `CallAdapter` on `alknet/call`; the
interactive sessions are dispatched by the shared `TtyAdapter` on
`alknet/tty`, using `DockerTtyBackend` as one of potentially several
registered backends. This is the ADR-058 decision: docker ops are
call-protocol operations, and the one operation that needed raw
carriage moved to alknet-tty.
## Label Namespace
alknet-docker applies two labels to containers it creates (ADR-060):
| Label | Value | Purpose |
|-------|-------|---------|
| `alknet.managed` | `"true"` | Marks the container as alknet-managed; the `list` owned_only filter and the ownership cross-check key on this. |
| `alknet.owner` | `<peer_id>` | The `Identity.id` of the spawner (the coordinator's identity, not the end user's — proxy pattern). |
The prefix (`alknet`) is configurable at assembly-layer wiring
(two-way-door). The label schema (two labels,
`<prefix>.managed` + `<prefix>.owner`) is the one-way commitment.
Containers created by `docker compose` or `docker run` (the
hosted-services case) have no alknet labels and no ownership-store
entry; they're reached via the static-resource operator fallback.
Full details: [docker-operations.md](docker-operations.md) §"Label
Namespace" and
[ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md).
## Feature Gates
```toml
# alknet-docker Cargo.toml
[features]
default = ["ops"]
ops = [] # docker/container/* and docker/image/* call operations
tty = ["dep:alknet-tty"] # DockerTtyBackend (impl TtyBackend)
```
- `default = ["ops"]` — the call operations. A hub or coordinator
wiring docker management over the call protocol uses the default
features. Pulls in bollard + alknet-call + alknet-core; no alknet-tty.
- `tty` — adds `DockerTtyBackend`, pulling in `alknet-tty` for the
`TtyBackend` trait. A deployment that wants interactive terminal
sessions into containers enables this and registers
`DockerTtyBackend` with the `TtyAdapter`. A call-operations-only
deployment leaves this off and doesn't pull in alknet-tty.
The `tty` feature mirrors `alknet-tty`'s `local` feature pattern
(ADR-054): the heavy backend edge is opt-in, so a deployment that
doesn't want it doesn't pay for it. See
[ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md).
## Assembly Layer Wiring
The assembly layer (the CLI binary or a hub/worker binary) constructs
the bollard client, the ownership store, and the label config, then
registers the docker operations and (optionally) the tty backend:
```rust
// 1. Construct the bollard client (local daemon, ADR-059 §2)
let docker = bollard::Docker::connect_with_local_defaults()?;
// 2. Construct the ownership store (in-memory default, ADR-050 §1)
let ownership_store = Arc::new(InMemoryOwnershipStore::new());
// 3. Register docker call operations on the shared registry
let mut builder = OperationRegistryBuilder::new();
// ... other operations (services/list, agent ops, etc.) ...
register_docker_ops(
&mut builder,
docker.clone(),
ownership_store.clone(),
&DockerLabels { prefix: "alknet".into() },
CompositionAuthority::new("docker-ops", ["container:list", "container:exec", /* ... */]),
);
// 4. (optional, behind `tty` feature) Register DockerTtyBackend
#[cfg(feature = "tty")]
{
let docker_backend = Arc::new(DockerTtyBackend::new(
docker.clone(),
"alknet".into(),
)) as Arc<dyn TtyBackend>;
// Insert into the TtyAdapter's backend map under "docker"
backends.insert("docker".into(), docker_backend);
}
// 5. Build the registry and start the endpoint
let registry = builder.build();
let call_adapter = CallAdapter::new(Arc::new(registry), identity_provider);
// ... register handlers on the endpoint, start ...
```
The `register_docker_ops` function (or a `DockerOps` struct the
builder consumes) adds each `docker/container/*` and `docker/image/*`
operation as a `HandlerRegistration` with the right `HandlerKind`
(`Once` for Query/Mutation, `Stream` for Subscription — ADR-049) and
`AccessControl` (per ADR-060). The assembly layer provides the
`CompositionAuthority` for the docker-ops handler (the scopes the
docker ops run under when composed) and the `Capabilities` (empty —
`Capabilities::new()` for local bollard; no outbound credentials
needed). The `Docker` client and `OwnershipStore` are
**closure-captured** into each handler at registration time (ADR-062)
— not read from `OperationContext` and not put in `Capabilities`
(`Capabilities` is for secret material only, per ADR-014).
The `OwnershipStore` is shared between the docker ops (which `record`
on create and `revoke` on remove) and the `AccessControl::check` path
(which consults the `OwnershipProvider` read side). The in-memory
default is sufficient for the single-host case; a persistence adapter
(ADR-035 shape) is built when a hub wants fleet ownership to survive
restarts.
## Architecture (component pointers)
- **[docker-operations.md](docker-operations.md)** — the operation
surface: lifecycle (Query/Mutation), logs/exec/pull (Subscription
via `StreamingHandler`), access control (ADR-050/050 application),
label namespace, teardown coupling, the exit-code-on-final-
`call.responded` pattern for exec.
- **[docker-tty-backend.md](docker-tty-backend.md)** — the
`DockerTtyBackend`: attach vs exec mode, `TtyHandle` field mapping
(PTY mode merges stderr), `TtyControl` → bollard resize/signal,
`exit_code` future `Drop`-kill (ADR-056), `resource_id()` delegation.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Docker ops on `alknet/call` (no separate ALPN) | [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md) | Call-protocol operations; raw-carriage dissolved by alknet-tty; no `carriage` field |
| bollard 0.21 + feature selection | [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) | Version verified current; `http`+`pipe`+`time`; no `ssl`/`ssh`/`websocket`/`buildkit` |
| Container resource model + label namespace | [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) | ADR-050 application; `alknet.managed`/`alknet.owner` labels; `list` owned_only; hosted-services static fallback |
| DockerTtyBackend in alknet-docker | [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) | Behind `tty` feature; attach/exec mode; POC `drive_attach_raw` as reference |
| Crate decomposition | [ADR-003](../../decisions/003-crate-decomposition.md) Am. 1 | alknet-docker depends on alknet-call (protocol-foundation exception) + alknet-tty (tty feature) |
| Call protocol stream model | [ADR-012](../../decisions/012-call-protocol-stream-model.md) | `EventEnvelope` framing, no `carriage` extension |
| Streaming handler for subscriptions | [ADR-049](../../decisions/049-streaming-handler-for-subscriptions.md) | `StreamingHandler` for logs/exec/pull; exit code on final `call.responded` |
| Dynamic resource ownership | [ADR-050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Containers as `AccessControl` resources; the model ADR-060 applies |
| Peer-graph routing | [ADR-029](../../decisions/029-peer-graph-routing-model.md) | Head→worker docker management via `PeerRef` / `invoke_peer` |
| Forwarded-for identity | [ADR-032](../../decisions/032-forwarded-for-identity.md) | End-user identity as metadata when proxying |
| TtyBackend trait | [ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) | The trait `DockerTtyBackend` implements |
| Exit code on control chunk | [ADR-055](../../decisions/055-exit-code-on-control-chunk.md) | The exit-chunk ordering `DockerTtyBackend`'s `exit_code` feeds into |
| Backend cleanup on cancel | [ADR-056](../../decisions/056-backend-cleanup-on-session-cancel.md) | `DockerTtyBackend`'s `exit_code` `Drop` kills the container/exec |
| Docker client + OwnershipStore injection | [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md) | Closure capture at registration; `Capabilities` for secrets only |
| Exit code on terminal `call.responded` | [ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md) | `{ "exitCode": N, "terminal": true }` on final `call.responded` for non-interactive exec |
## Open Questions
See [open-questions.md](../../open-questions.md) for full details.
- **OQ-048** (deferred(scope)): Network and volume operation surface.
- **OQ-049** (deferred(scope)): Image build (buildkit) scope.
- **OQ-050** (deferred(scope)): Docker system events subscription.
- **OQ-051** (deferred(scope)): Container create options surface.
## References
- `docs/research/alknet-docker/poc-summary.md` — the POC summary this
spec set builds on (validated targets, open unknowns)
- `/workspace/alknet-docker-poc/` — the POC source
- `/workspace/bollard/` — bollard 0.21.0 source (verified current)
- `/workspace/@alkdev/dispatch/` — the dispatch POC (prior art:
`dispatch.managed` labels, SSH-tunnel fleet model)
- `/workspace/system/dev1/docker.md` — the hosted-services use case
- `/workspace/@alkdev/reverse-proxy/deploy/docker-compose.yml` —
operator-created container example
+6 -5
View File
@@ -313,7 +313,7 @@ let mut backends = HashMap::new();
backends.insert("local".into(),
Arc::new(LocalTtyBackend::new()) as Arc<dyn TtyBackend>);
backends.insert("docker".into(),
Arc::new(DockerTtyBackend::new(docker_client)) as Arc<dyn TtyBackend>);
Arc::new(DockerTtyBackend::new(docker_client, "alknet".into())) as Arc<dyn TtyBackend>);
backends.insert("ssh".into(),
Arc::new(SshTtyBackend::new(ssh_session)) as Arc<dyn TtyBackend>);
let tty_adapter = TtyAdapter::new(Arc::new(backends));
@@ -329,12 +329,13 @@ layer chooses what's available.
| Backend | Crate | Status | Notes |
|---------|-------|--------|-------|
| `LocalTtyBackend` | `alknet-tty-local` (sibling, behind `local` feature) | in scope ([tty-local.md](tty-local.md)) | `portable_pty` (PTY) + `std::process` (pipe); the runner pattern |
| `DockerTtyBackend` | `alknet-docker` or sibling adapter | future, out of scope here | wraps `bollard::attach_container` / `exec` with `tty: true` |
| `DockerTtyBackend` | `alknet-docker` (behind `tty` feature) | specced ([docker-tty-backend.md](../docker/docker-tty-backend.md), [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md)) | wraps `bollard::attach_container` / `exec` with `tty: true`; attach vs exec mode |
| `SshTtyBackend` | `alknet-ssh` | future, out of scope here | wraps russh `pty_request` + `shell_request`/`exec_request`; dissolves alknet-ssh DP-5 PTY hedge |
The docker and SSH backend crates are future work; this spec commits the
trait shape they implement so they can be built against it without
re-spec'ing the seam. The `DockerTtyBackend` is the natural extension of
The SSH backend crate is future work; this spec commits the
trait shape it implements so it can be built against it without
re-spec'ing the seam. The `DockerTtyBackend` is now specced in
[alknet-docker](../docker/docker-tty-backend.md) (ADR-061) — the natural extension of
the alknet-docker POC's `drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`)
— with the trait, it becomes `impl TtyBackend for DockerTtyBackend`. The
`SshTtyBackend` dissolves the alknet-ssh research's PTY hedge (DP-5):
@@ -0,0 +1,267 @@
# ADR-058: alknet-docker Registers on `alknet/call` (No Separate ALPN)
## Status
Accepted
## Context
The alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`)
left two integration questions open (its "Open Unknowns" #1 and #2):
1. **Raw-carriage handoff in the dispatcher.** The POC's `drive_attach_raw`
reads the `call.requested` frame itself, then switches the bidi stream
to raw chunks for interactive attach. The POC noted this doesn't fit
the call dispatcher's `handle_stream` → `dispatch()` →
`DispatchResult::Stream(ResponseStream)` path, which pumps a stream
of `EventEnvelope`s (JSON carriage). The POC offered two options: (a)
branch `handle_stream` on a `carriage` field, handing raw streams to a
`RawHandler`; or (b) a separate ALPN (`alknet/docker-raw`) that owns
the whole stream.
2. **ALPN layout.** Should docker ops register on the shared
`alknet/call` ALPN (operations in the shared `OperationRegistry`) or
get their own `alknet/docker` ALPN (as a `ProtocolHandler`)? The POC
leaned shared but didn't decide.
Both questions were open **before** alknet-tty was extracted. The
alknet-tty crate spec (ADR-052, ADR-053) resolves both by moving the one
operation that needed raw carriage — interactive attach/exec with a PTY
— out of the call protocol entirely and into its own ALPN
(`alknet/tty`), backed by the `TtyBackend` trait.
### Why the raw-carriage problem dissolved
The POC's raw-carriage handoff was hard because it tried to do two
different things on one `alknet/call` stream:
- **Structured operations** (lifecycle, logs, inspect) — naturally
JSON-shaped, fit `call.responded`/`call.completed` exactly. This is
what the call protocol is for.
- **Interactive attach** — a bidirectional byte pump (stdin/stdout/stderr)
with a control sideband (resize, signal, exit). This is *not* a
request/response or subscription; it's a terminal session. Forcing it
through `EventEnvelope` framing is "wasteful and lossy" (POC §"Why not
JSON for everything?").
alknet-tty was created precisely to extract the second category into its
own protocol (`alknet/tty`) with its own wire format (ADR-052). The
docker POC's `drive_attach_raw` became the seed of alknet-tty's
`DockerTtyBackend` (ADR-061). What remains in alknet-docker's call
operations is the first category — structured operations that map cleanly
to `call.requested`/`call.responded`/`call.completed`.
The POC's concern ("the dispatcher would need a `carriage` field and a
`RawHandler` branch") is no longer operative: there is no raw-carriage
operation on `alknet/call`. The raw byte pump moved to `alknet/tty`; the
`DockerTtyBackend` implements `TtyBackend` and is reached through the
`TtyAdapter`, not the `CallAdapter`. See ADR-061 for the backend
placement.
### Why shared `alknet/call`, not a separate `alknet/docker` ALPN
The remaining docker operations (lifecycle, logs, exec-with-exit-code,
inspect, list, images) are ordinary call-protocol operations. They have:
- A structured input (JSON), a structured output (JSON), or a stream of
structured events (logs, image pull progress, exec output).
- Natural `OperationSpec`s with input/output JSON Schemas.
- `AccessControl` declarations against the ADR-050 container-as-resource
model (`resource_type: "container"`, `resource_id_path: "$.containerId"`).
- Service discovery through `services/list` / `services/schema`.
Putting them on a separate `alknet/docker` ALPN would mean a separate
`ProtocolHandler` that re-implements framing, dispatch, ACL, and service
discovery — or a thin wrapper that delegates to the call protocol's
machinery. Either way, it's a parallel dispatch surface for no benefit.
The shared registry is more composable: docker ops are callable from any
call client, including peer routing (`PeerRef` / `from_call` re-export),
which is the primary use case — a coordinator on the hub composing
`docker/container/exec` on a worker spoke.
A separate ALPN is warranted when the protocol's wire format is
incompatible with `EventEnvelope` framing (alknet-tty, alknet-ssh). The
docker operations remaining on `alknet/call` are all
`EventEnvelope`-shaped. The one that wasn't (raw attach) moved to
alknet-tty. See ADR-052 §"Why not JSON for everything?" for the boundary
criterion.
### Logs and exec streaming — still JSON carriage, still `alknet/call`
Two operations look like they might need raw carriage but don't:
- **`docker/container/logs`** — a `Subscription` operation. bollard's
`logs()` returns a `Stream<LogOutput>`; each `LogOutput` becomes a
`call.responded` carrying `{ "stream": "stdout"|"stderr", "text": "..." }`;
stream end → `call.completed`. This is the `StreamingHandler` shape
(ADR-049). The POC validated this path (`docker_logs_subscription_pumps_frames_and_completes`).
No raw carriage needed — each log line is naturally JSON-shaped.
- **`docker/container/exec`** (non-interactive, no TTY) — a `Subscription`
operation. bollard's `start_exec` returns `StartExecResults::Attached {
output, input }`; the output stream is pumped as `call.responded`
frames; after the stream ends, `inspect_exec()` gives the exit code,
which rides on a final `call.responded` with `{ "exitCode": N,
"terminal": true }` before `call.completed`. The POC validated this
(`docker_exec_streams_output_and_exit_code`). This is the exec path
for `tty: false` — a captured command with separate stdout/stderr and
an exit code, not an interactive terminal.
The interactive exec path (`tty: true`, bidirectional, resize/signal) is
**not** a call operation — it's a `DockerTtyBackend` session on
`alknet/tty`. ADR-061 covers that. The split is: non-interactive exec is
a call `Subscription`; interactive exec is a tty session. The `tty:
bool` field in the operation input selects the path; the caller picks.
## Decision
### 1. alknet-docker registers its operations on the shared `alknet/call` ALPN
alknet-docker does not register a `ProtocolHandler`. It constructs an
`OperationRegistry` (or a `DockerOps` registration bundle consumed by
the assembly layer's `OperationRegistryBuilder`) and the existing
`CallAdapter` dispatches docker operations through the shared
`OperationRegistry::invoke()` / `invoke_streaming()` paths. There is no
`alknet/docker` ALPN, no `DockerProtocolHandler`, and no parallel
dispatch surface.
This makes docker operations first-class citizens of the call protocol:
they appear in `services/list`, they have `OperationSpec`s with JSON
Schemas, they go through the standard `AccessControl::check`, and they
compose with `from_call` re-export and peer routing (ADR-029) like any
other operation.
### 2. There is no raw-carriage operation on `alknet/call`
The `carriage` field the POC proposed for `call.requested` is not added.
The call protocol's `EventEnvelope` framing is the only carriage on
`alknet/call`. Interactive attach/exec (the one operation that needed
raw carriage) is reached through `alknet/tty` via `DockerTtyBackend`
(ADR-061), not through a call operation. The POC's open question #1
(raw-carriage handoff in the dispatcher) is resolved by removing the
requirement, not by adding a branch.
### 3. The operation taxonomy
| Operation | Call op type | bollard method | Carriage |
|-----------|-------------|---------------|----------|
| `docker/container/list` | Query | `list_containers` | JSON |
| `docker/container/inspect` | Query | `inspect_container` | JSON |
| `docker/container/create` | Mutation | `create_container` | JSON |
| `docker/container/start` | Mutation | `start_container` | JSON |
| `docker/container/stop` | Mutation | `stop_container` | JSON |
| `docker/container/remove` | Mutation | `remove_container` | JSON |
| `docker/container/restart` | Mutation | `restart_container` | JSON |
| `docker/container/logs` | Subscription | `logs` | JSON (StreamingHandler) |
| `docker/container/exec` (tty:false) | Subscription | `create_exec` + `start_exec` + `inspect_exec` | JSON (StreamingHandler) |
| `docker/image/list` | Query | `list_images` | JSON |
| `docker/image/pull` | Subscription | `create_image` | JSON (StreamingHandler) |
| `docker/image/inspect` | Query | `inspect_image` | JSON |
Interactive exec (`tty: true`) and interactive attach are **not** in this
table — they are `alknet/tty` sessions via `DockerTtyBackend`
(ADR-061). A `docker/container/exec` call operation with `tty: true` in
the input returns an `call.error` with code `INVALID_INPUT` directing
the caller to use `alknet/tty`; the operation only accepts `tty: false`
(or absent). This keeps the one-operation-per-carriage invariant clean:
the call operation is captured output; the tty session is interactive.
The full operation surface (including which are in scope for v1 and
which are deferred) is in
[docker-operations.md](../crates/docker/docker-operations.md).
### 4. The assembly layer composes the registry
alknet-docker exports a `DockerOps` construct (a set of
`HandlerRegistration` bundles keyed by operation name) or a
`register_docker_ops(&mut builder, docker_client, ownership, labels)`
function. The assembly layer (the CLI binary or a hub binary) calls this
to add docker operations to the shared `OperationRegistryBuilder`
alongside other operations (services discovery, agent ops, etc.). The
`Docker` bollard client, the `OwnershipStore`, and the label namespace
config are injected by the assembly layer. See
[overview.md](../crates/docker/overview.md) §"Assembly Layer Wiring".
## Consequences
**Positive:**
- Docker operations inherit all call-protocol machinery for free:
service discovery, JSON Schema validation, `AccessControl`,
`from_call` re-export, peer routing, `forwarded_for` metadata, abort
cascade. No parallel dispatch surface.
- The POC's hardest open question (raw-carriage handoff) is resolved by
removal — the operation that needed it moved to alknet-tty. No
dispatcher change, no `carriage` field, no `RawHandler` trait.
- The operation taxonomy is clean: `Query`/`Mutation` for lifecycle,
`Subscription` (ADR-049 `StreamingHandler`) for logs/exec/pull. The
exit-code-on-final-`call.responded` pattern the POC validated (POC
target 3) works through the existing `StreamingHandler` → `pump_stream`
path with no protocol change.
- A hub that wants to manage docker on worker spokes composes
`docker/container/*` through `from_call` (ADR-017) + peer routing
(ADR-029) — the proxy pattern from ADR-050. This is the primary use
case and it works by construction.
**Negative:**
- The `docker/container/exec` operation has a `tty` field whose value
constrains the dispatch path: `tty: false` → call `Subscription`;
`tty: true` → reject, point to `alknet/tty`. This is a minor impedance
— a caller wanting interactive exec must use a different protocol
(`alknet/tty`) than a caller wanting captured exec. This is the
correct split (interactive terminal ≠ captured command), but it means
the "exec" concept spans two protocols. The `DockerTtyBackend`
(ADR-061) and the `docker/container/exec` call operation share the
underlying `bollard::exec` API but diverge at the wire format. This is
the same divergence as "SSH exec" (call op) vs "SSH PTY session"
(alknet-tty) — the boundary is principled, not accidental.
- A deployment that wants docker operations must wire the
`OperationRegistry` (the assembly layer's job). This is not a downside
— it's the same wiring every other operation set requires — but it
means alknet-docker is not a "drop-in `ProtocolHandler`" the way
alknet-tty is. The trade is composability (shared registry, peer
routing) for a slightly more involved assembly step.
## Door type
**One-way.** The decision to register on `alknet/call` rather than a
separate `alknet/docker` ALPN is a structural commitment: the operation
names (`docker/container/*`), the `OperationSpec`s, the
`AccessControl` shapes, and the `from_call` re-export surface all depend
on docker ops being call-protocol operations. Reversing this — moving
docker ops to their own ALPN — would be a rewrite of the operation
surface and a break for any client that calls `docker/container/*`
through the call protocol.
The decision *not* to add a `carriage` field to `call.requested` is
also one-way: the call protocol's wire format (ADR-012) stays
`EventEnvelope`-only on `alknet/call`. If a future operation needs raw
carriage on `alknet/call`, it would require a wire-format change (new
ALPN per ADR-006, or a `call.requested` field addition). The expected
path for raw-carriage operations is a separate ALPN (as alknet-tty did),
not a `call.requested` extension.
## References
- `docs/research/alknet-docker/poc-summary.md` §"Open Unknowns" #1 and #2
(the two questions this ADR resolves)
- [ADR-012](012-call-protocol-stream-model.md) — the call protocol wire
format this ADR keeps unchanged (no `carriage` field)
- [ADR-017](017-call-protocol-client-and-adapter-contract.md) — `from_call`
re-export, the proxy pattern's mechanism
- [ADR-024](024-operation-registry-layering.md) — the registry layering
docker ops register into
- [ADR-029](029-peer-graph-routing-model.md) — peer routing, the
head→worker docker management path
- [ADR-049](049-streaming-handler-for-subscriptions.md) —
`StreamingHandler`, the dispatch path for logs/exec/pull
- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md)
— containers as resources, the `AccessControl` shape docker ops declare
- [ADR-052](052-alknet-tty-wire-format-and-two-carriage.md) — alknet-tty's
wire format, where the raw-carriage operation moved
- [ADR-053](053-ttybackend-trait-and-ttyhandle.md) — `TtyBackend`, the
trait `DockerTtyBackend` implements (ADR-061)
- [ADR-061](061-docker-tty-backend-in-alknet-docker.md) —
`DockerTtyBackend` placement in alknet-docker
- Spec documents: [overview.md](../crates/docker/overview.md),
[docker-operations.md](../crates/docker/docker-operations.md)
@@ -0,0 +1,195 @@
# ADR-059: bollard 0.21 Dependency and Feature Selection
## Status
Accepted
## Context
The alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`)
depended on a local bollard checkout at `/workspace/bollard` (version
0.21.0). Its "Open Unknowns" #5 noted: "The real crate should depend on
published 0.21 from crates.io (the dispatch POC pinned 0.18 — a
3-version jump). The `websocket` feature is optional; the `http` and
`pipe` features are needed for socket/http connect. Confirm the
published 0.21 has the same API surface as the checkout."
Two concerns motivate an explicit version check:
1. **Agent training-data drift.** Implementation agents tend to default
to the bollard version represented in their training data unless
explicitly told to check. The dispatch POC used bollard 0.18 (3
versions behind); an agent that pattern-matched dispatch would pin
0.18 and silently use a 3-version-stale API. The cost of checking is
low (one crates.io lookup); the cost of using a stale version is
potentially high (API mismatches, missing methods, security regressions).
2. **API surface stability across the path-vs-published boundary.** The
POC used a path dependency; the crate uses a published version. A
version number match (0.21.0) should mean the API surface is
identical, but this is worth confirming rather than assuming.
### Version verification
A check against crates.io / docs.rs confirmed:
- **bollard 0.21.0 is the latest published version** on crates.io (as
of 2026-07-08). The local checkout at `/workspace/bollard` (0.21.0)
matches the published version. No version jump is needed; the POC's
dependency is current.
- The published 0.21.0 `Cargo.toml` (verified via docs.rs source) has
the same API surface the POC used: `attach_container` (container.rs:540),
`logs` (container.rs:928), `create_exec`/`start_exec`/`inspect_exec`
(exec.rs:172/225/315), `AttachContainerResults` (container.rs:80),
`StartExecResults` enum (exec.rs:99), `LogOutput` (container.rs:96),
`NewlineLogOutputDecoder` (read.rs:32). The POC's code-to-concept
mappings hold against the published version.
- bollard-stubs is pinned at `=1.53.1-rc.29.3.1` (the Docker API 1.53
schema), matching the Docker Engine 29.2.1 / API 1.53 the POC tested
against.
### Feature selection
bollard 0.21's feature surface (from the published `Cargo.toml`):
| Feature | What it pulls | Needed for alknet-docker? |
|---------|---------------|--------------------------|
| `default = ["http", "pipe"]` | hyper-util (http), hyperlocal + hyper-named-pipe (pipe) | **Yes** — the default. `pipe` connects to the unix socket (`/var/run/docker.sock`); `http` connects to a TCP/HTTP daemon. Both are the standard local-daemon paths. |
| `ssl` | rustls TLS for remote daemon connections | No — alknet-docker talks to a local daemon; remote daemons are reached over the call protocol (the fleet layer, not bollard's TLS). |
| `ssh` | openssh for SSH-tunneled daemons | No — the dispatch POC's SSH-tunnel model is exactly what alknet removes; the worker dials the hub over the call protocol instead of the hub SSHing into the worker. |
| `websocket` | tokio-tungstenite for the websocket attach endpoint | No — the POC deliberately used the reliable `attach_container()` (HTTP upgrade to TCP), not the websocket path (bollard's own docs warn of RFC 6455 compatibility issues). The websocket feature is for browser-attach scenarios alknet-docker doesn't have. |
| `buildkit` / `buildkit_providerless` | tonic + bollard-buildkit-proto for buildkit image builds | No — image build is deferred (POC "What the POC Does NOT Validate" #5; OQ-049 scopes this). |
| `chrono` / `time` | timestamp parsing for log timestamps | **Optional** — `docker/container/logs` sets `timestamps: true`; the `chrono` or `time` feature types the timestamp field. The choice between them is a two-way-door implementation detail (both are serde-compatible). `time` is the lighter dependency. |
| `json_data_content` | serde for body content types | No — alknet-docker serializes its own JSON shapes from bollard's types; it doesn't need bollard's content-type tagging. |
| `aws-lc-rs` / `webpki` / `test_*` | TLS provider variants and test-only features | No — not for production deps. |
The POC's `Cargo.toml` depended on bollard with default features (path
dependency). The crate does the same: `bollard = { version = "0.21",
default-features = true }`, with the `time` feature added for log
timestamp typing. No `ssl`, `ssh`, `websocket`, or `buildkit` features.
## Decision
### 1. Depend on published bollard 0.21 from crates.io
```toml
[dependencies]
bollard = { version = "0.21", features = ["time"] }
```
`default-features = true` (the default) pulls `http` + `pipe`, the two
local-daemon connect paths. The `time` feature adds timestamp typing
for `docker/container/logs`. No other features are enabled.
The local-checkout path dependency the POC used is retired. The
published 0.21.0 has the identical API surface (confirmed via docs.rs
source — same version number, same method signatures, same stubs pin).
### 2. Connect via `connect_with_local_defaults()` (unix socket)
The standard deployment talks to a local docker daemon over
`/var/run/docker.sock` (Unix) or the named pipe on Windows.
`Docker::connect_with_local_defaults()` handles both. The `Docker`
client is constructed once at assembly-layer startup and injected into
the docker ops registration. See
[overview.md](../crates/docker/overview.md) §"Assembly Layer Wiring".
A TCP daemon path (`connect_with_http_defaults`) is available if a
deployment runs the docker daemon on a different host — but this is the
exception, not the default. The fleet case (multiple docker daemons on
different machines) is handled at the call-protocol layer (a `CallClient`
per remote daemon, each running alknet-docker locally), not by pointing
one bollard client at a remote daemon over TCP. See ADR-058 §"Why shared
`alknet/call`" and the POC §6 (the normalization crate boundary).
### 3. Do not enable `websocket`, `ssl`, `ssh`, or `buildkit`
- **`websocket`** — the reliable `attach_container()` (HTTP upgrade to
TCP) is the attach path for `DockerTtyBackend` (ADR-061). The websocket
attach endpoint has documented reliability issues (bollard
`container.rs:577`) and is for browser-attach scenarios alknet-docker
doesn't have. The raw chunk format (ADR-052) is layered on top of the
reliable attach, not the websocket.
- **`ssl`** — remote daemons are reached over the call protocol (the
fleet layer), not bollard's TLS. alknet-docker is single-host by
construction (POC §6).
- **`ssh`** — the SSH-tunnel model is what the call-protocol fleet
layer replaces. The dispatch POC's `DockerProvider` SSHed into workers;
alknet-docker's workers dial the hub over `alknet/call` and expose
their local docker ops directly.
- **`buildkit`** — image build is deferred (OQ-049). The `build_image`
method is in bollard but the buildkit feature pulls tonic +
bollard-buildkit-proto, a heavy dependency for an out-of-scope feature.
### 4. `time` feature for log timestamps
`docker/container/logs` sets `timestamps: true` on the bollard `logs()`
query. Each `LogOutput` then carries a docker timestamp; the
`call.responded` frame separates `timestamp` and `text` into distinct
JSON fields (the POC noted this refinement over its single-`text`-field
POC shape). The `time` feature types the timestamp field as
`time::OffsetDateTime` (serde-compatible). The `chrono` feature is the
alternative; `time` is lighter and sufficient.
The timestamp-feature choice (`time` vs `chrono`) is a two-way-door
implementation detail within the one-way version-pin decision. Switching
features is a `Cargo.toml` change, not an API break.
## Consequences
**Positive:**
- bollard 0.21 is current (verified against crates.io); no stale-version
risk. The POC's API-surface mappings hold against the published version.
- The feature set is minimal: `http` + `pipe` (default) + `time`. No
`ssl`/`ssh`/`websocket`/`buildkit` weight. The compile-time and
dependency surface stays small.
- The single-host contract is clear: alknet-docker talks to one local
daemon over the unix socket. The fleet case is a call-protocol concern,
not a bollard-feature concern.
**Negative:**
- The `websocket` feature being disabled means the browser-attach
scenario (bollard's websocket attach for a browser that speaks docker's
websocket framing directly) is not available. This is intentional —
browsers reach docker through `alknet/tty` (via `DockerTtyBackend`) or
the call protocol, not bollard's websocket endpoint. If a future use
case forces the websocket path, enabling the feature is a `Cargo.toml`
addition (two-way door).
- The `ssh` feature being disabled means bollard's SSH-tunneled daemon
path is not available. This is intentional — the dispatch POC's
SSH-tunnel model is what alknet's call-protocol fleet layer replaces
(POC §6, "Prior art"). Enabling `ssh` would reintroduce the friction
(SSH key injection, port binding) the fleet layer exists to remove.
## Door type
**One-way (version pin) + two-way (feature set).** The version pin
(bollard 0.21, not 0.18 or a future 0.22) is a one-way commitment for
the implementation: the operation handlers are written against 0.21's
API surface, and a major-version bump is a migration. The feature set
(`http` + `pipe` + `time`, no `ssl`/`ssh`/`websocket`/`buildkit`) is
two-way-door: features can be added or removed in `Cargo.toml` without
an API break. The decision to *not* enable `ssl`/`ssh`/`buildkit` is a
scope decision (not needed for the current scope), not a capability
closure — the features can be added later if a use case forces them.
The `time`-vs-`chrono` choice is a two-way-door implementation detail.
## References
- `docs/research/alknet-docker/poc-summary.md` §"Open Unknowns" #5
(bollard version pinning)
- `docs/research/alknet-docker/poc-summary.md` §"What the POC Does NOT
Validate" #5 (image management / buildkit deferred)
- bollard 0.21.0 published `Cargo.toml` (verified via docs.rs source)
- bollard source: `src/container.rs` (`attach_container` :540, `logs`
:928), `src/exec.rs` (`create_exec` :172, `start_exec` :225,
`inspect_exec` :315), `src/read.rs` (`NewlineLogOutputDecoder` :32)
- [ADR-058](058-alknet-docker-on-alknet-call.md) — the single-host /
call-protocol-fleet contract this feature set reflects
- [ADR-061](061-docker-tty-backend-in-alknet-docker.md) — the
`DockerTtyBackend` that uses the reliable (non-websocket) attach path
- OQ-049 (build/image scope, deferred)
- Spec: [overview.md](../crates/docker/overview.md) §"Dependencies"
@@ -0,0 +1,338 @@
# ADR-060: Container Resource Model and Label Namespace
## Status
Accepted
## Context
ADR-050 established the dynamic resource ownership model for
runtime-spawned resources: a `OwnershipProvider` (read, sync) +
`OwnershipStore` (write, async) trait pair in alknet-core, with the
spawner owning the resource, proxy-to-share, teardown-revoke. ADR-050
explicitly named alknet-docker as the first consumer: containers as
`AccessControl` resources, `docker/container/exec` requiring
`resource: container/<id>:exec`, `docker/container/create` calling
`OwnershipStore::record` on success, `docker/container/remove` calling
`OwnershipStore::revoke` on teardown.
ADR-050 resolved the *model*. This ADR resolves the *application* of
that model to alknet-docker's concrete operation surface — the three
things ADR-050 left to the crate spec:
1. **The label namespace.** The dispatch POC (`/workspace/@alkdev/dispatch`)
used `dispatch.managed=true` to mark containers it created. The POC
summary §"What the POC Does NOT Validate" #6 noted the real crate
"needs a configurable label prefix and ownership mapping
(`alknet.owner=<peer-id>`) tied to the call protocol's identity model."
What is the label scheme, and how does it relate to the ownership store?
2. **The `list` result-filter.** ADR-050 specific #4a says the `list`
case (resource_type set, resource_id_path absent) is "scope-gate +
result-filter": the scope check gates the call, and the handler
filters the result via `OwnershipProvider::owned_resources`. How
does `docker/container/list` apply this against bollard's
`list_containers` (which returns all containers the daemon sees)?
3. **Teardown coupling.** ADR-050 specific #4b says teardown is
handler-driven: the docker handler calls `revoke` on container
removal. But docker containers can exit on their own (a process
that finishes, a `--rm` container). How does the ownership store
learn about autonomous container death, or does it rely solely on
explicit `docker/container/remove`?
### The two use cases and their ownership profiles
The user's brief named two container use cases:
- **Disposable dev containers** (the common case by volume) — a
coordinator spawns a container for an implementation agent or an
isolated env. The container is short-lived; the coordinator owns it
for its lifetime; it's removed when the agent is done. Ownership is
per-session, per-coordinator.
- **Long-running hosted services** (less common but important) — the
production server (`/workspace/system/dev1`) hosts rarely-changing
services (reverse-proxy, postgres, redis, gitea) in docker. These
containers are created by an operator (via `docker compose`), not by
a call-protocol coordinator. They're managed via alknet-docker's
operations (start/stop/restart/inspect/logs) but their *ownership* is
not in the alknet ownership store — they predate the connection.
These two cases have different ownership profiles, and the model must
handle both without forcing the hosted-services case through the
disposable-container path.
## Decision
### 1. Label namespace: `alknet.*` prefix, configurable, two labels
alknet-docker applies two labels to containers it creates:
| Label | Value | Purpose |
|-------|-------|---------|
| `alknet.managed` | `"true"` | Marks the container as alknet-managed. The `list` filter and the ownership-existence check key on this label. |
| `alknet.owner` | `<peer_id>` | The `Identity.id` (stable logical peer id, per ADR-030) of the spawner. For composing handlers, the *handler's* identity (the coordinator's `Identity`), not the end user's — the proxy pattern (ADR-050 §3). |
The label prefix (`alknet`) is **configurable** at assembly-layer wiring.
A deployment that runs alongside another alknet-managed fleet (or wants
to avoid collisions with its own `managed` labels) sets a different
prefix. The default is `alknet`. This is a two-way-door config choice;
the label *schema* (two labels, `<prefix>.managed` + `<prefix>.owner`)
is the one-way commitment.
The labels serve two purposes:
- **Ownership-store cross-check.** When `docker/container/exec` checks
`OwnershipProvider::owns(identity, "container", id, "exec")`, the
ownership store is the source of truth. The label is a *secondary*
signal: if the store says "yes" but the label is absent or names a
different owner, the container's ownership state is stale (the
container was removed and its ID reused, or the store and docker
diverged). The label is not authoritative — the store is — but it's a
debugging aid and a consistency check.
- **`list` filter.** `docker/container/list` with an `owned_only: true`
input flag filters to containers whose `alknet.owner` label matches
the caller's `Identity.id`. This is the result-filter path (ADR-050
#4a). Without the flag, the list returns all containers the daemon
sees (the hosted-services case — see §3 below).
**`alknet.*` label reservation.** The `<prefix>.*` label keys are
reserved: the `docker/container/create` handler overwrites
`<prefix>.managed` and `<prefix>.owner` on the container's labels
regardless of what the caller provided, preventing a caller from
spoofing ownership (setting `alknet.owner` to another peer's id). A
caller's own labels (non-`<prefix>.*` keys) are preserved. This is a
security property: ownership is set by the handler (from the
caller's `Identity.id`), not by the caller's input.
**Resource action vocabulary.** The `resource_action` field on
`AccessControl` for containers uses these actions:
| Action | Used by | Scope name |
|--------|---------|-----------|
| `exec` | `docker/container/exec` (call op, non-interactive) | `container:exec` |
| `tty` | `DockerTtyBackend` sessions (`alknet/tty`) | `container:tty` |
| `start` | `docker/container/start` | `container:start` |
| `stop` | `docker/container/stop` | `container:stop` |
| `remove` | `docker/container/remove` | `container:remove` |
| `restart` | `docker/container/restart` | `container:restart` |
| `manage` | (static, operator role) `Identity.resources["container"]` | `container:manage` |
| `list` | `docker/container/list` (scope-gate only, no resource_id) | `container:list` |
| `create` | `docker/container/create` (scope-gate only, no resource_id) | `container:create` |
The `tty` action is distinct from `exec` (ADR-061 §4): a caller
authorized for non-interactive exec (`container:exec`) is not
automatically authorized for an interactive terminal session
(`container:tty`). The `manage` action is the static operator-role
resource (§3) that subsumes the per-container actions for pre-existing
hosted-service containers. The `list` and `create` actions are
scope-gate-only (no `resource_id_path`); the rest target a specific
container via `resource_id_path: "$.containerId"`.
### 2. `docker/container/list`: scope-gate + optional result-filter
`docker/container/list` is a `Query` operation with
`resource_type: "container"`, no `resource_id_path` (the `list` case per
ADR-050 #4a). The input accepts an optional `owned_only: bool` flag
(default `false`):
- **`owned_only: false`** (default) — the handler calls
`bollard::list_containers()` and returns all containers the daemon
sees. The scope check gates the call (the caller needs
`container:list`); there is no result-filter. This is the
hosted-services case: an operator listing all containers on dev1,
including ones they didn't spawn through alknet.
- **`owned_only: true`** — the handler calls
`bollard::list_containers()` with a label filter
(`label: alknet.owner=<caller_peer_id>`), returning only the
containers the caller owns. This is the disposable-dev-container
case: a coordinator listing its own workspaces. The scope check still
gates the call; the label filter is the result-filter.
The two paths use the same bollard method (`list_containers` with
different `ListContainersOptions`), differing only in the label filter.
The `owned_only` flag is the caller's choice; the scope check is the
server's. This matches ADR-050 #4a's "allow if scoped, filter to owned"
default, with the filter opt-in (the hosted-services case needs the
unfiltered list).
### 3. Hosted services: ownership is not required for pre-existing containers
The hosted-services case (dev1's reverse-proxy, postgres, redis, gitea)
involves containers created by an operator via `docker compose`, not via
`docker/container/create`. These containers have no `alknet.owner`
label and no ownership-store entry. Operations on them (`start`,
`stop`, `inspect`, `logs`, `restart`) must still work — but
`AccessControl::check` with `resource_type: "container"` and a
`resource_id` would consult the ownership store, find no entry, and
deny.
The resolution: **the `AccessControl` for operations on a specific
container (`exec`, `stop`, `remove`, `inspect` with `resource_id_path`)
declares `resource_type: "container"` and `resource_action`, but the
ownership check is against the store, and the store's "no entry" result
falls through to the static `Identity.resources` path (ADR-050 §2
backward-compat).** A peer whose `Identity.resources["container"]`
includes `"manage"` (a deployment-configured scope for the operator
role) passes the check for any container, owned or not. A peer without
that resource but with `container:exec` scope passes only for containers
the ownership store says they own.
Concretely, the two roles:
| Role | Scope | Owns a specific container? | Can exec into container C? |
|------|-------|---------------------------|-----------------------------|
| Operator (manages hosted services) | `container:manage` in `Identity.resources["container"]` | No (no ownership entry) | Yes — the static-resource fallback passes (`manage` ⊇ `exec`) |
| Coordinator (spawns dev containers) | `container:exec` scope | Yes (ownership store: coordinator owns C) | Yes — the ownership check passes |
| Random peer | `container:exec` scope | No | No — ownership check fails, static fallback fails |
The operator role is configured statically (the dev1 operator's
`PeerEntry.resources` includes `container:manage`); the coordinator
role is dynamic (the ownership store records the spawn). Both paths
reach the same `AccessControl::check`; the difference is which branch
satisfies.
This means `docker/container/exec` and friends work on hosted-service
containers *for the operator role* without the operator having to
"claim" ownership of containers they didn't spawn. The
`docker/container/create` operation (the spawner path) always records
ownership; `docker/container/start` on a pre-existing container does
not (it's not a spawn). The model is: **spawn → own; manage (operator) →
static-resource-pass; neither → deny.**
### 4. Teardown coupling: handler-driven revoke + autonomous-death tolerance
ADR-050 #4b says teardown is handler-driven: `docker/container/remove`
calls `OwnershipStore::revoke("container", id)`. But containers can die
autonomously:
- A `--rm` container exits and is auto-removed by the daemon.
- A container crashes and an operator removes it via `docker rm` (not
through alknet-docker).
- The daemon restarts (all containers stop; container IDs are stale).
alknet-docker's teardown coupling is **handler-driven revoke on the
alknet-managed remove path, with autonomous-death tolerance**:
- **`docker/container/remove` (the alknet operation)** calls
`bollard::remove_container()`, and on success calls
`OwnershipStore::revoke("container", id)`. This is the ADR-050 #4b
contract: the handler that manages the lifecycle revokes.
- **Autonomous container death** (a `--rm` exit, an external `docker rm`,
a daemon restart) leaves a stale ownership-store entry. The entry is
not proactively cleaned up — there is no reaper, no event subscription
to the docker daemon's death events. Instead, the entry is cleaned up
lazily: the next operation against that container ID fails at the
bollard layer (the container doesn't exist), the error propagates as
a `call.error`, and a subsequent `docker/container/list` with
`owned_only: true` for the stale owner filters the container out
(it's not in the daemon's list). The ownership store's
`owned_resources` can be cross-checked against the live daemon list
on read, or left to be overwritten on the next spawn — the container ID
is unique per daemon run; a reused ID would be a new container.
The tolerance for stale entries is intentional: ownership is runtime
state (ADR-050 assumption 4), meaningless across restarts. The
in-memory default store loses all entries on restart; a persistence
adapter would cache and invalidate. A reaper that subscribes to docker
daemon death events would be more prompt but adds a subscription
surface for a marginal gain (the stale entry is inert — it can't grant
access to a container that doesn't exist, and a reused container ID
gets a fresh `record` on its next `create`).
A future "prompt stale-entry cleanup" feature is additive: a
`docker/system/events` subscription operation (deferred, OQ-050) that
the ownership store could subscribe to. The base model tolerates stale
entries; the prompt-cleanup path is a refinement.
### 5. `docker/container/create` records ownership; `start`/`stop`/`restart` do not
Only `docker/container/create` calls `OwnershipStore::record`. The
lifecycle operations (`start`, `stop`, `restart`, `remove`) do not
record — they act on an existing container whose ownership was recorded
at create time (or which pre-exists and is reached via the
static-resource operator path). `remove` calls `revoke` (the teardown
half); the others neither record nor revoke.
This matches ADR-050's "spawner owns" model: the *create* is the spawn
event; the lifecycle operations are management of an existing resource.
## Consequences
**Positive:**
- The two use cases (disposable dev containers, hosted services) both
work through one `AccessControl` model. The operator role reaches
hosted-service containers via the static-resource fallback; the
coordinator role reaches spawned containers via the ownership store.
No special-casing, no "is this a managed container?" branch in the
handlers.
- The label namespace is configurable and minimal (two labels). The
`managed` flag marks alknet-spawned containers; the `owner` label
carries the peer id for the `list` filter and the cross-check.
- Teardown is handler-driven on the alknet path and tolerant of
autonomous death on the non-alknet path. No reaper subscription
surface in the base model.
- ADR-050's model is applied without amendment. The `list` case
(specific #4a), the teardown coupling (#4b), and the
backward-compat static-resource fallback (#2) all map cleanly to
bollard's API (`list_containers` with a label filter, `remove_container`
+ `revoke`, the static `Identity.resources` path).
**Negative:**
- Stale ownership entries from autonomous container death are not
promptly cleaned up. An `owned_only: true` list could theoretically
return a container ID that no longer exists — but bollard's
`list_containers` returns *live* containers, so the list is always
accurate against the daemon; the stale entry only affects the
ownership store's internal map, not the list result. The cross-check
(label vs store) could flag divergence; this is a debugging aid, not
a correctness issue.
- The operator role requires static configuration
(`Identity.resources["container"] ⊇ "manage"`). A deployment that
wants the operator to manage hosted services must configure the
operator's `PeerEntry.resources`. This is not a downside — it's the
same static configuration every role requires — but it means the
hosted-services case isn't "zero-config"; the operator peer must be
declared.
- The `owned_only` flag on `list` is an opt-in. A coordinator that
forgets to set it gets all containers, not just its own. The default
(`false`) is correct for the hosted-services case (operator lists
all), but a coordinator expecting isolation must set the flag. This
is a caller-side convention, enforced by the scope check (a
non-operator peer without `container:manage` can't call `list` at all
— the scope gates it), but within the `container:list` scope, the
filter is the caller's choice.
## Door type
**One-way (label schema, ownership model application) + two-way (label
prefix, stale-entry policy).** The label schema (two labels,
`<prefix>.managed` + `<prefix>.owner`) and the
create-records/remove-revokes ownership coupling are one-way: clients
and the ownership store depend on the labels and the record/revoke
timing. The label prefix (default `alknet`) is two-way-door config. The
stale-entry policy (tolerate, no reaper) is two-way — a reaper
subscription is an additive refinement.
## References
- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md)
— the model this ADR applies (specifics #4a, #4b, #2)
- `docs/research/alknet-docker/poc-summary.md` §"What the POC Does NOT
Validate" #6 (label namespace / ownership mapping)
- `/workspace/@alkdev/dispatch/src/docker.rs` — the dispatch POC's
`dispatch.managed=true` label (prior art this ADR generalizes)
- [ADR-030](030-peerentry-and-identity-id-decoupling.md) — `Identity.id`
as the stable peer id used in the `alknet.owner` label
- [ADR-032](032-forwarded-for-identity.md) — `forwarded_for` for the
proxy pattern's end-user identity
- [ADR-015](015-privilege-model-and-authority-context.md) — the static
`Identity.resources` path (the operator-role fallback)
- `/workspace/system/dev1/docker.md` — the hosted-services use case
(reverse-proxy, postgres, redis, gitea on dev1)
- `/workspace/@alkdev/reverse-proxy/deploy/docker-compose.yml` — the
reverse-proxy's docker setup (operator-created, not alknet-spawned)
- Spec: [docker-operations.md](../crates/docker/docker-operations.md)
§"Access Control" and §"Label Namespace"
@@ -0,0 +1,272 @@
# ADR-061: DockerTtyBackend in alknet-docker
## Status
Accepted
## Context
The alknet-tty spec ([tty-backend.md](../crates/tty/tty-backend.md)
§"Backend implementations") lists three `TtyBackend` implementations:
`LocalTtyBackend` (in `alknet-tty-local`), `DockerTtyBackend` (location
"future, out of scope here"), and `SshTtyBackend` (in `alknet-ssh`).
The alknet-tty spec deliberately left the `DockerTtyBackend` location
open — it committed the trait shape but not the crate placement.
alknet-docker is now being specified. The `DockerTtyBackend` — the
`impl TtyBackend` that wraps `bollard::attach_container()` for
interactive attach and `bollard::exec::start_exec` with `tty: true` for
interactive exec — is the natural extension of the alknet-docker POC's
`drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`). The
question is which crate it lives in.
### The options
Three placements were considered:
1. **In `alknet-docker`** — the `DockerTtyBackend` is a type in
alknet-docker, behind a `tty` feature gate that pulls in
`alknet-tty` (for the `TtyBackend` trait) and `bollard` (for the
attach/exec API). alknet-docker already depends on bollard for its
call operations; the `tty` feature adds only the `alknet-tty` edge.
2. **In `alknet-tty-docker`** (a sibling adapter crate) — the
`DockerTtyBackend` in its own crate, depending on `alknet-tty` (the
trait) and `bollard` (the API). alknet-docker stays leaner (no
`alknet-tty` edge); the adapter is its own publishable unit.
3. **In `alknet-tty-local`** (extending the local sibling) — rejected
immediately. `alknet-tty-local` is for *local* backends (`portable_pty`,
`std::process`); docker is not local-process, and pulling bollard into
the local sibling would violate the feature-gate rationale (a
docker-only deployment pulling in PTY code, and a local-only
deployment pulling in bollard).
### Why option 1 (in alknet-docker)
The `DockerTtyBackend` is tightly coupled to alknet-docker's domain:
it uses the same bollard client, the same container identity model
(`resource_id` extraction from `backend_params.container`), and the
same label/ownership reasoning (ADR-060) as the call operations. A
separate `alknet-tty-docker` crate would duplicate the bollard client
construction, the container-id validation, and the ownership wiring —
or import them from alknet-docker, creating a cycle
(`alknet-tty-docker` → `alknet-docker` for the client, `alknet-docker`
→ `alknet-tty-docker` for the backend registration).
The dependency edge from alknet-docker to alknet-tty is clean:
alknet-docker depends on alknet-tty for the `TtyBackend` trait (and
`TtyHandle`, `TtyControl`, `TtyError`, `TtyParams`). This is the same
edge any backend crate has — `alknet-tty-local` depends on
`alknet-tty` for the trait. The trait is the inversion point
(ADR-053); the backend impls live where their transport deps live
(bollard, for `DockerTtyBackend`). bollard already lives in
alknet-docker; the `DockerTtyBackend` belongs with it.
The `tty` feature gate means a deployment that wants only the call
operations (no interactive terminal) doesn't pull in `alknet-tty`. The
feature is opt-in, matching `alknet-tty-local`'s `local` feature
pattern (ADR-054).
### The POC's `drive_attach_raw` as the reference
The POC's `drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`)
is the seed of `DockerTtyBackend::allocate()`. The mapping:
| POC `drive_attach_raw` | `DockerTtyBackend::allocate()` |
|---|---|
| reads `call.requested` frame, extracts container id | `TtyParams.backend_params` → `DockerBackendParams { container }` |
| `bollard::attach_container()` → `AttachContainerResults { output, input }` | same; `output` → `TtyHandle.stdout`, `input` → `TtyHandle.stdin` |
| `ChunkReader`/`ChunkWriter` on the bidi stream | the `TtyAdapter` owns the chunk codec (ADR-052); the backend produces handles, not wire bytes |
| zero-length stdin chunk → `container_input.shutdown()` | `TtyControl` / stdin EOF — the adapter signals; the backend's `AsyncWrite` close maps to `shutdown()` |
| completion: output stream ends → zero-length stdout sentinel → close | `TtyHandle.exit_code` future — for attach, the exit comes from `wait_container` or the output stream end; for exec with `tty: true`, `inspect_exec` after the output stream ends |
| `LogOutput` → `Chunk` stream_type mapping (StdOut→1, StdErr→2) | `TtyHandle.stdout`/`stderr` — PTY mode merges (stderr is `None`); the adapter pumps stdout only |
The POC's `drive_exec` (non-interactive, exit code via `inspect_exec`)
is the basis for the **call operation** `docker/container/exec`
(`tty: false`), not the `DockerTtyBackend` (which is the `tty: true`
interactive path). The split: non-interactive exec is a call
`Subscription` (ADR-058); interactive exec is a `DockerTtyBackend`
session. Both use `bollard::exec`; they diverge at the wire format.
## Decision
### 1. `DockerTtyBackend` is a type in alknet-docker, behind a `tty` feature
```toml
# alknet-docker Cargo.toml
[features]
default = ["ops"] # the call operations
ops = [] # docker/container/* and docker/image/* operations
tty = ["dep:alknet-tty"] # DockerTtyBackend (impl TtyBackend)
```
- `default = ["ops"]` — the call operations (lifecycle, logs, exec
non-interactive, images). This is what a hub or coordinator wires to
manage containers over the call protocol.
- `tty` — adds `DockerTtyBackend` (impl `TtyBackend`), pulling in
`alknet-tty` for the trait. A deployment that wants interactive
terminal sessions into docker containers enables this and registers
`DockerTtyBackend` with the `TtyAdapter` (alongside or instead of
`LocalTtyBackend`).
The feature gate means a deployment that uses alknet-docker only for
call operations (the common case — a hub managing dev containers)
doesn't pull in `alknet-tty`. A deployment that wants interactive
terminals into those containers enables `tty` and registers the backend.
### 2. `DockerTtyBackend::allocate()` wraps `attach_container` or exec-with-PTY
`allocate()` branches on `TtyParams.terminal` and a
`DockerBackendParams` field for the attach mode:
- **`attach` mode** — wraps `bollard::attach_container()` (the reliable
HTTP-upgrade path, per the POC and ADR-059 §3). Returns
`AttachContainerResults { output, input }`. `output` → `TtyHandle.stdout`
(as a `Stream<Bytes>`, mapping `LogOutput` → `Bytes` with the
stream_type flattened — PTY mode has no separate stderr, so
`TtyHandle.stderr` is `None`). `input` → `TtyHandle.stdin` (the
`AsyncWrite` bollard provides). Exit code: the output stream end +
`wait_container()` or `inspect_container()` for the exit status.
- **`exec` mode** — wraps `bollard::create_exec` + `start_exec` with
`tty: true`. `start_exec` returns `StartExecResults::Attached {
output, input }`; same mapping as attach. Exit code: after the output
stream ends, `inspect_exec()` for the exit code (the POC's pattern).
The mode is selected by a `DockerBackendParams.mode` field
(`"attach" | "exec"`, default `"exec"` for a new command, `"attach"`
for attaching to a running container's primary process). For `exec`
mode, `TtyParams.cmd` is the command vector; for `attach` mode, `cmd`
is ignored (the primary process is already running).
### 3. `TtyControl` maps to bollard's resize and signal
- `resize()` → `bollard::container::resize_container_tty()` (attach
mode) or `bollard::exec::resize_exec()` (exec mode).
- `signal()` → docker has no direct container signal API in bollard's
stable surface. The POC used `kill_container` for the attach case.
For exec mode, there is no per-exec signal; the signal is delivered
to the container's PID namespace via `kill_container` with the
container's main PID, or the exec's PID if exposed. This is a
best-effort mapping per the `TtyControl::signal` contract
("best-effort delivery to the foreground process group"); unknown
signals fall back to `kill_container` with `SIGKILL`. The exact
signal-delivery path (bollard's `kill_container` vs a future exec-PID
signal API) is a two-way-door implementation detail within the
one-way backend placement.
### 4. `resource_id()` returns `Some(("container", container_id))`
`DockerTtyBackend::resource_id(&params)` extracts the container ID
from `backend_params` and returns `Some(("container", id))`, per the
`tty-backend.md` sketch. The `TtyAdapter` calls
`OwnershipProvider::owns(identity, "container", id, "tty")` at
negotiation (ADR-050, ADR-060). The backend owns the extraction; the
adapter doesn't parse docker-specific JSON. This is the same shape
the `tty-backend.md` sketch anticipates.
### 5. `exit_code` future's Drop kills the container/exec (ADR-056)
The `TtyHandle.exit_code` future's `Drop`-on-cancel (ADR-056) issues:
- **Attach mode** — `bollard::kill_container()` (or the container's
natural exit if the stream end already triggered). The container is
not removed (attach doesn't own the container's lifecycle); it's
killed so the process doesn't outlive the session.
- **Exec mode** — the exec instance is not a separate killable
process in docker's model; the signal goes to the container. The
`Drop` issues `kill_container` with the container's main PID as a
best-effort. A container running an exec that the terminal session
cancels should have the exec's process terminated; docker's exec
isolation means this is best-effort, not guaranteed.
This satisfies ADR-056's contract: dropping `exit_code` (cancel) kills
the session target. The "session target" for docker is the container
(attach) or the exec's process (exec, best-effort via the container).
## Consequences
**Positive:**
- The `DockerTtyBackend` lives where its deps live (bollard in
alknet-docker), matching the ADR-053 inversion principle. No
`alknet-tty-docker` sibling crate, no cycle risk.
- The `tty` feature gate keeps the `alknet-tty` edge opt-in. A
call-operations-only deployment (the common case) doesn't pull in
alknet-tty.
- The POC's `drive_attach_raw` maps cleanly to
`DockerTtyBackend::allocate()` — the backend produces handles, the
`TtyAdapter` pumps the wire format. The POC's three-pump driver
(`session.rs`) is the reference; the backend extracts the bollard
interaction from it.
- The `resource_id()` delegation means the adapter's ownership check
(ADR-060) works without the adapter parsing docker JSON — the
backend extracts the container ID, the adapter checks the store.
**Negative:**
- alknet-docker gains an `alknet-tty` dependency edge (behind the
`tty` feature). This is the intended trade: the backend lives with
its transport dep, and the trait edge is feature-gated. A deployment
that wants interactive docker terminals accepts this; one that
doesn't is unaffected.
- Signal delivery in docker exec mode is best-effort (docker's exec
isolation limits per-exec signal targeting). The `TtyControl::signal`
contract is already best-effort (ADR-053), so this is within the
contract, but a user expecting `Ctrl-C` to kill only the exec
process (not the container) may see the container killed instead.
This is a docker-semantics limitation, not an alknet limitation;
the local backend (portable_pty) has precise process-group
targeting (REQ-TTY-02). The docker backend's signal semantics are
documented in the spec, not hidden.
- The `attach` vs `exec` mode split in `DockerBackendParams` is a
backend-specific detail the caller must know. The `TtyParams` is
opaque to the adapter (ADR-053), but the caller (the client opening
the tty session) must set `mode: "attach"` or `mode: "exec"` in the
negotiation frame's backend params. This is the same opacity as the
local backend's `cwd`/`env` — the caller knows their backend. No
adapter change is needed for a new mode; the backend parses its own
params.
## Door type
**One-way (crate placement) + two-way (feature gate, attach/exec mode
split, signal path).** The decision to put `DockerTtyBackend` in
alknet-docker (not a sibling crate) is one-way: the backend type, its
registration with the `TtyAdapter`, and its `bollard`/`alknet-tty`
edges are structural. Reversing to a sibling crate would move the type
and re-edge the dependencies.
The feature gate (`tty` opt-in), the attach/exec mode split in
`DockerBackendParams`, and the exact signal-delivery path
(`kill_container` vs a future exec-PID API) are two-way-door
implementation details within the one-way placement. The feature can
be renamed, the mode field can gain a third value, and the signal path
can be refined — all without a structural change.
## References
- [ADR-053](053-ttybackend-trait-and-ttyhandle.md) — the
`TtyBackend` trait this backend implements; the inversion principle
(backends live where their transport deps live)
- [ADR-052](052-alknet-tty-wire-format-and-two-carriage.md) — the wire
format the `TtyAdapter` pumps (the backend produces handles, not wire)
- [ADR-054](054-local-tty-backend-sibling-crate.md) — the sibling-crate
pattern for `LocalTtyBackend` (the parallel this ADR follows, with
the difference that docker's deps are already in alknet-docker)
- [ADR-056](056-backend-cleanup-on-session-cancel.md) — the
cancel-cleanup contract (`exit_code` future's `Drop` kills)
- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md)
— the ownership model `resource_id()` delegates to
- [ADR-058](058-alknet-docker-on-alknet-call.md) — why the
non-interactive exec is a call op, not this backend
- [ADR-059](059-bollard-021-dependency-and-features.md) — the bollard
dependency (no `websocket` feature — the reliable attach path)
- `docs/research/alknet-docker/poc-summary.md` §"POC Target 1"
(interactive attach — the seed of this backend)
- `/workspace/alknet-docker-poc/src/ops.rs` — `drive_attach_raw` (the
reference for `allocate()`)
- [tty-backend.md](../crates/tty/tty-backend.md) §"Backend
implementations" (the row this ADR fills: "DockerTtyBackend | alknet-docker
or sibling adapter | future, out of scope here")
- Spec: [docker-tty-backend.md](../crates/docker/docker-tty-backend.md)
@@ -0,0 +1,234 @@
# ADR-062: Docker Client and OwnershipStore Injection via Closure Capture
## Status
Accepted
## Context
The docker operation handlers need access to two non-secret, shared
runtime handles:
1. The bollard `Docker` client — a `Clone`-able handle (internally an
`Arc`) to the docker daemon connection, constructed once at
assembly-layer startup.
2. The `OwnershipStore` — the ADR-050 write-side trait, used by
`docker/container/create` (to `record`) and `docker/container/remove`
(to `revoke`).
An earlier draft of the docker spec
(`docker-operations.md`) showed handlers calling
`ctx.docker_client()` and speculated the handle lived on "a
`DockerOpsExt` extension trait on `OperationContext`, or a
`DockerClient` capability in `ctx.capabilities`." This was an inline
open question with a model conflict:
### Why `Capabilities` is the wrong channel
`Capabilities` (ADR-014, `operation-registry.md` §"Capability
Injection") is the type for **outbound secret material** — decrypted
API keys, signing keys, vault-derived credentials. Its contract
(`core-types.md`): non-`Serialize`, zeroized on drop, populated from
the vault at registration time. A bollard `Docker` handle to a local
unix socket is not secret material; it has no zeroizing semantics; it
must not be smuggled through the secret-injection channel. Putting a
`Docker` handle in `Capabilities` would be a category error against
the one-way `Capabilities` contract (ADR-014).
The same applies to the `OwnershipStore`: it's a shared state handle
(shared `Arc<dyn OwnershipStore>`), not secret material.
### Why `OperationContext` extension is the wrong channel
`OperationContext` (`operation-registry.md:230`) is a concrete struct
with a fixed field set, constructed by the dispatch path per call.
Adding a `docker_client` field (or an extension trait that reads one)
would either (a) make every non-docker handler pay for a field they
don't use, or (b) require a typed-map / downcast pattern that
violates the "context is concrete, not a bag" design. The
established pattern for per-handler-set shared state is not a
context field — it's closure capture.
### The established pattern: closure capture at registration
The `from_openapi` adapter (ADR-017, `client-and-adapters.md`,
`operation-registry.md:819`) captures its `reqwest::Client` in each
forwarding handler's closure at registration time:
```rust
// operation-registry.md:819 (the established pattern)
.with_local(vastai_listMachines_spec(), Arc::new(vastai_handler),
CompositionAuthority::new("vastai", ["vastai:query"]),
ScopedOperationEnv::new([...]),
Capabilities::new().with_api_key("vastai", vastai_token))
```
The `vastai_handler` closure captures its reqwest client (and its
auth token via `Capabilities`, which *is* secret material). The
non-secret shared handle (reqwest client) is captured in the closure;
the secret (the API key) goes through `Capabilities`. This is the
correct split: secret material through `Capabilities`, non-secret
shared state through closure capture.
The docker ops follow the same split. The `Docker` client and the
`OwnershipStore` are non-secret shared state — closure-captured. There
are no secret capabilities (local bollard needs no API key; the
docker daemon is local).
## Decision
### 1. `register_docker_ops` captures the `Docker` client and `OwnershipStore` in each handler closure
The `register_docker_ops` function (or the `DockerOps` builder)
takes `Arc<Docker>` and `Arc<dyn OwnershipStore>` (plus the label
config and the `CompositionAuthority`) as arguments, and constructs
each `Handler` / `StreamingHandler` closure with those handles
captured by reference (`Arc::clone` per handler):
```rust
pub fn register_docker_ops(
builder: &mut OperationRegistryBuilder,
docker: Arc<bollard::Docker>,
ownership_store: Arc<dyn OwnershipStore>,
labels: &DockerLabels,
authority: CompositionAuthority,
) {
let docker_clone = docker.clone();
let ownership_clone = ownership_store.clone();
let labels_clone = labels.clone();
builder.with_local(
container_inspect_spec(),
Arc::new(move |input, ctx| {
let docker = docker_clone.clone();
Box::pin(async move {
let container_id = input["containerId"].as_str()
.ok_or_else(|| /* INVALID_INPUT */)?;
match docker.inspect_container(container_id, None::<()>).await {
Ok(info) => ResponseEnvelope::ok(to_json_value(info)),
Err(e) => /* CONTAINER_NOT_FOUND / DOCKER_ERROR */,
}
})
}),
authority.clone(),
ScopedOperationEnv::empty(),
Capabilities::new(), // no secret caps — local bollard
);
// ... same pattern for each operation ...
}
```
Each handler closure captures the `Docker` client and (where needed)
the `OwnershipStore` by cloning the `Arc`. The handler does not read
these from `OperationContext`; they're baked into the closure at
registration time. The `Capabilities` passed to `with_local` is empty
(`Capabilities::new()`) — there are no secret capabilities for local
bollard operations.
### 2. The `Docker` client is shared, `Clone`-able, and constructed once
bollard's `Docker` is `Clone` (it holds an internal `Arc` to the
connection). The assembly layer constructs one
(`Docker::connect_with_local_defaults()`, ADR-059 §2), wraps it in an
`Arc`, and passes it to `register_docker_ops`. Each handler clones
the `Arc` (cheap — a refcount bump) into its closure. There is no
per-handler docker client; the one client is shared across all
docker operations.
### 3. The `OwnershipStore` is shared the same way
`Arc<dyn OwnershipStore>` is cloned into the `create` and `remove`
handler closures (the two operations that write to the store). The
other operations (inspect, start, stop, logs, exec, list, images)
don't capture the store — they don't write to it. The `create`
closure captures it for `record`; the `remove` closure captures it
for `revoke`.
### 4. No `docker_client()` accessor; no `OperationContext` extension
The handlers do not call `ctx.docker_client()` or any other accessor.
The `Docker` client and `OwnershipStore` are not on `OperationContext`
and are not accessible via an extension trait. They're closure-
captured, full stop. A handler that needs the docker client gets it
from its closure's captured `Arc<Docker>`, not from the context.
This means the handler signature is the standard
`Fn(Value, OperationContext) -> Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>`
(ADR-049) — no new parameter, no new context field. The docker-
specific state is in the closure, not the context.
## Consequences
**Positive:**
- The injection model matches the established `from_openapi` pattern
(non-secret shared state via closure capture, secret material via
`Capabilities`). No new mechanism; no `OperationContext` change; no
`Capabilities` conflict.
- The `Capabilities` contract (ADR-014) stays clean — it holds secret
material only. A local bollard handle is not smuggled through the
secret channel.
- The handlers are self-contained — the `Docker` client and store are
baked in at registration, not looked up at call time. This is
testable (construct a `Docker` client in a test, register the ops,
invoke the handler directly) and composable (the same handler works
whether the docker client points at a local socket or a test
daemon).
- No new trait, no extension mechanism, no downcast. The injection is
plain Rust closure capture — the simplest thing that works.
**Negative:**
- Each handler closure captures the `Docker` and (for create/remove)
the `OwnershipStore` by `Arc::clone`. This is a refcount bump per
handler construction (at startup, once) — negligible. The per-call
cost is zero (the closure already holds the `Arc`; calling the
handler doesn't re-clone).
- The `register_docker_ops` function signature is longer (takes the
`Docker` + store + labels + authority). This is the assembly layer's
wiring concern, not a handler concern — the handlers don't know
where their `Docker` came from.
- A test that wants to invoke a single docker handler must construct
the `Docker` client and (if the handler uses the store) the
`OwnershipStore` to register the op. This is the same as any
handler test that needs its dependencies — not new.
## Door type
**One-way.** The decision to use closure capture (not a context field,
not a `Capabilities` entry) is the injection model every docker
handler is written against. Reversing to a context field would be a
rewrite of every handler and an `OperationContext` change. The
specific captures (`Docker` + `OwnershipStore`) are the non-secret
state the handlers depend on; changing what's captured is a handler
rewrite.
The choice of `Arc<Docker>` vs a non-`Arc` shared reference is a
two-way-door implementation detail (bollard's `Docker` is already
`Clone`-with-`Arc`-internally; the outer `Arc` is for the trait-object
`OwnershipStore`'s sake, not the `Docker`'s). The capture pattern
(closure capture) is the one-way commitment.
## References
- [ADR-014](014-secret-material-flow-and-capability-injection.md) —
the `Capabilities` contract this ADR keeps clean (secret material
only; the `Docker` handle is not secret)
- [ADR-017](017-call-protocol-client-and-adapter-contract.md) — the
`from_openapi`/`from_call` adapter pattern this ADR follows
(non-secret shared state via closure capture)
- [ADR-022](022-handler-registration-provenance-and-composition-authority.md)
— `HandlerRegistration`, the bundle the `register_docker_ops`
builder populates
- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md)
— the `OwnershipStore` trait the `create`/`remove` handlers capture
- [ADR-058](058-alknet-docker-on-alknet-call.md) — the operation
surface this injection model serves
- `crates/call/operation-registry.md` §"Capability Injection"
(the `Capabilities` semantics this ADR respects) and
§"OperationContext" (the struct this ADR does *not* extend)
- `crates/call/client-and-adapters.md` §"from_openapi forwarding"
(the prior art: reqwest client closure-captured, API key in
`Capabilities`)
- Spec: [docker-operations.md](../crates/docker/docker-operations.md)
§"Handler injection" (the section this ADR adds, replacing the
earlier `ctx.docker_client()` sketches)
@@ -0,0 +1,185 @@
# ADR-063: Exit Code on a Terminal `call.responded` for Non-Interactive Exec
## Status
Accepted
## Context
The alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`
§"POC Target 3") validated a completion shape for streaming exec:
the exit code rides on a final `call.responded` frame before
`call.completed`. This keeps `call.completed`'s payload empty
(matching ADR-012's wire format — no core protocol change). The POC
used `{ "exitCode": N }` on the final `call.responded`.
The docker spec (`docker-operations.md` §"docker/container/exec")
carries this forward and adds a `"terminal": true` field to mark the
exit-code `call.responded` as the final value before completion. The
`terminal` flag is a docker-operation convention, not a call-protocol
wire-format change — it's a field in the `call.responded` payload
(the `Value`), not a new event type or a `call.completed` payload.
A reader asked: "why `terminal: true` and not a dedicated `call.exit`
event or an exit code on `call.completed`?" This ADR records the
decision.
### Why not an exit code on `call.completed`
ADR-012 defines `call.completed` with an empty payload (`{}`). The
POC's completion-shape decision (POC §"The completion-shape decision
this validates") chose to keep `call.completed` empty and put the
exit code on the preceding `call.responded`. Changing `call.completed`
to carry a payload would be a wire-format change (ADR-012 amendment,
new ALPN per ADR-006) — a one-way door that affects every
`call.completed` consumer, not just docker. The POC deliberately
avoided this.
### Why not a dedicated `call.exit` event
A new event type (`call.exit`) would be a wire-format addition (new
`EventEnvelope` variant, ADR-012 amendment). It would also duplicate
the "terminal value before completion" semantics that
`call.completed` already provides — `call.completed` is the
"stream end" signal; an exit code is a "final value" that rides
*before* the stream end. A `EventEnvelope` variant conflates these
two: `call.exit` would be both the final value and the stream-end
signal, but `call.completed` already handles stream end. Keeping
the exit code as a `call.responded` (a normal value) and letting
`call.completed` be the stream end preserves the
one-value-per-`call.responded` invariant: the exit code is just the
last value.
### Why `terminal: true`
A client consuming a `docker/container/exec` stream sees a sequence
of `call.responded` frames (stdout/stderr lines) and then a final
`call.responded` with `{ "exitCode": N, "terminal": true }`, then
`call.completed`. The `terminal: true` flag tells the client "this
is the last value; the next event is `call.completed`." Without it,
the client would have to infer "this `call.responded` has an
`exitCode` field, so it's the exit" — a content-sniffing heuristic
that breaks if a future stdout line happens to carry an `exitCode`
field.
The `terminal` flag is a docker-operation convention — a field in
the `call.responded` payload, not a protocol-level marker. It's
declared in the `docker/container/exec` `output_schema` (the exit
frame is a documented part of the output stream). Other
`Subscription` operations that produce a "terminal result before
completion" can adopt the same convention (a `terminal: true` field
on their final `call.responded`); it's not docker-specific, but
docker is the first operation to need it.
## Decision
### 1. The exec exit code rides on a final `call.responded` with `terminal: true`
`docker/container/exec` (non-interactive, `tty: false`) emits:
1. Zero or more `call.responded` frames with stdout/stderr output
(`{ "stream": "stdout"|"stderr", "text": "..." }`).
2. A final `call.responded` with `{ "exitCode": N, "terminal": true }`.
3. `call.completed` (empty payload, per ADR-012).
The `terminal: true` field marks the exit-code frame as the final
value before completion. The `exitCode` is `i32` (matching
`ExecInspectResponse.exit_code`; negative for signal-terminated,
though docker exec exit codes are typically non-negative).
### 2. `call.completed` stays empty
No `call.completed` payload change. ADR-012's wire format is
unchanged. The exit code is on the preceding `call.responded`, not
on `call.completed`.
### 3. `terminal: true` is an operation-level convention, not a protocol field
The `terminal` flag is a field in the `call.responded` payload (the
`Value`), declared in the operation's `output_schema`. It is not a
new `EventEnvelope` field, not a new event type, and not a
`call.completed` payload. The call protocol's wire format is
unchanged. Other `Subscription` operations that produce a terminal
result before completion may adopt the same convention; it's not
docker-specific, but docker is the first to use it.
The `OperationSpec.output_schema` for `docker/container/exec`
documents the two response shapes: the streaming output frames
(`{ "stream": ..., "text": ... }`) and the terminal exit frame
(`{ "exitCode": N, "terminal": true }`). A client reading the schema
knows to expect the terminal frame.
### 4. The exit-code frame is the last `call.responded` before `call.completed`
The handler ensures the exit-code `call.responded` is the last one
the stream produces. The `StreamingHandler` (ADR-049) pumps the
output stream, then emits the exit-code frame, then the stream ends
(the dispatcher writes `call.completed` on natural stream end). The
ordering is the handler's responsibility (the handler controls the
stream's item order); the dispatcher's `pump_stream` writes them in
order. This mirrors the POC's validated pattern.
## Consequences
**Positive:**
- The exit code propagates through the streaming completion path
without a wire-format change. `call.completed` stays empty
(ADR-012 unchanged); the exit code is a normal `call.responded`
value that happens to be the last one.
- The `terminal: true` flag gives clients a deterministic "this is
the exit" signal without content-sniffing the `exitCode` field. A
client can stop reading output frames when it sees `terminal: true`
and read the exit code.
- The convention is reusable: any `Subscription` operation with a
terminal result before completion can use `terminal: true` on its
final `call.responded`. No protocol change; just an
operation-level field.
**Negative:**
- The `terminal` flag is a convention, not enforced by the wire
format. A `Subscription` operation that emits a terminal result
*without* `terminal: true` would be non-conformant but not
protocol-violating. The `output_schema` documents the convention;
clients that read the schema know to expect it.
- A client that doesn't check `terminal: true` and instead
content-sniffs `exitCode` would work for docker exec but break on
a hypothetical stdout line that carries an `exitCode` field. The
flag is the robust path; content-sniffing is the fragile path. The
spec recommends the flag; the schema documents it.
## Door type
**One-way.** The exit-code-on-`call.responded`-before-`call.completed`
shape is the completion contract `docker/container/exec` commits to.
Clients consuming the exec stream depend on this ordering; changing
it (moving the exit code to `call.completed`, or adding a `call.exit`
event) would be a break for those clients.
The `terminal: true` field name is two-way-door within the one-way
shape — it could be renamed (e.g., `final: true`) without a wire-
format change, since it's a payload field, not a protocol marker.
But the *presence* of a terminal marker on the final `call.responded`
is the one-way commitment: clients depend on "the last
`call.responded` before `call.completed` is marked as terminal and
carries the exit code."
## References
- `docs/research/alknet-docker/poc-summary.md` §"POC Target 3"
(exec with exit code — the validated pattern this ADR formalizes)
- `docs/research/alknet-docker/poc-summary.md` §"The completion-shape
decision this validates" (the POC's reasoning for exit-on-
`call.responded`, not on `call.completed`)
- [ADR-012](012-call-protocol-stream-model.md) — the call protocol
wire format this ADR keeps unchanged (`call.completed` stays empty)
- [ADR-049](049-streaming-handler-for-subscriptions.md) —
`StreamingHandler`, the dispatch path that pumps the
`call.responded` stream and writes `call.completed` on stream end
- [ADR-023](023-operation-error-schemas.md) — `output_schema`
documents the terminal frame shape
- [ADR-058](058-alknet-docker-on-alknet-call.md) — the operation
taxonomy (exec is a `Subscription`)
- Spec: [docker-operations.md](../crates/docker/docker-operations.md)
§"docker/container/exec" (the operation this ADR's shape serves)
+34 -1
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-07-07
last_updated: 2026-07-08
---
# Open Questions
@@ -149,6 +149,15 @@ Door type is separate from whether a decision is made. A two-way door is a decis
| [OQ-46](questions/046-runner-api-surface.md) | Runner API Surface | deferred(scope) | two | low |
| [OQ-47](questions/047-stdin-closure-canonical-signal.md) | Stdin Closure Canonical Signal | resolved | two | low |
### alknet-docker
| OQ | Title | Status | Door | Pri |
|----|-------|--------|------|-----|
| [OQ-48](questions/048-network-and-volume-operation-surface.md) | Network and Volume Operation Surface | deferred(scope) | two | low |
| [OQ-49](questions/049-image-build-buildkit-scope.md) | Image Build (buildkit) Scope | deferred(scope) | two | low |
| [OQ-50](questions/050-docker-system-events-subscription.md) | Docker System Events Subscription | deferred(scope) | two | low |
| [OQ-51](questions/051-container-create-options-surface.md) | Container Create Options Surface | deferred(scope) | two | med |
## Deferred / Blocked
The safe-exit visibility surface. These questions are parked because the
@@ -194,3 +203,27 @@ filtering the tables above.
- **Priority**: low
- **Full file**: [OQ-46](questions/046-runner-api-surface.md)
### OQ-48: Network and Volume Operation Surface
- **Blocked on**: a concrete use case for network or volume management over the call protocol. Dev containers use the default bridge network; hosted services declare networks/volumes in `docker compose`.
- **Priority**: low
- **Full file**: [OQ-48](questions/048-network-and-volume-operation-surface.md)
### OQ-49: Image Build (buildkit) Scope
- **Blocked on**: a concrete use case for building images over the call protocol. The current use cases pull pre-built images, not build them.
- **Priority**: low
- **Full file**: [OQ-49](questions/049-image-build-buildkit-scope.md)
### OQ-50: Docker System Events Subscription
- **Blocked on**: a concrete use case for prompt stale-ownership cleanup. The base model (ADR-060 §4) tolerates inert stale entries; the events subscription is a refinement.
- **Priority**: low
- **Full file**: [OQ-50](questions/050-docker-system-events-subscription.md)
### OQ-51: Container Create Options Surface
- **Blocked on**: v1 implementation — the `create` input JSON Schema is finalized when `register_docker_ops` is written and tested against bollard's `Config` struct. An architectural decision (ADR-060 §5), not a deferral past implementation.
- **Priority**: medium
- **Full file**: [OQ-51](questions/051-container-create-options-surface.md)
@@ -0,0 +1,30 @@
# OQ-48: Network and Volume Operation Surface
- **Origin**:
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
§"Operation Surface (v1 scope)" (out-of-scope list).
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Blocked on**: a concrete use case for network or volume management
over the call protocol. The two container use cases (disposable dev
containers, hosted services) don't currently require
network/volume CRUD over the call protocol — dev containers use
the default bridge network; hosted services are configured via
`docker compose` with networks/volumes declared in the compose file.
- **Resolution**: Not yet decidable. bollard has the API surface
(`network.rs`: `create_network`, `remove_network`, `inspect_network`,
`list_networks`, `connect_network`, `disconnect_network`;
`volume.rs`: `list_volumes`, `create_volume`, `inspect_volume`,
`remove_volume`, `prune_volumes`) — the mapping is mechanical, the
same shape as the container lifecycle ops (`Query`/`Mutation`,
single `call.responded`). The deferral is scope, not feasibility:
v1 is containers + images; networks and volumes are added when a
use case forces them (e.g., a coordinator that needs to create
isolated networks for dev containers, or a fleet layer that manages
volumes across hosts). Adding them is additive (new operations in
the registry) and does not break the existing surface.
- **Cross-references**:
[ADR-058](decisions/058-alknet-docker-on-alknet-call.md) (the shared
`alknet/call` registration model new ops would follow),
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
@@ -0,0 +1,30 @@
# OQ-49: Image Build (buildkit) Scope
- **Origin**:
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
§"Out of scope for v1"; [ADR-059](decisions/059-bollard-021-dependency-and-features.md)
§3 (no `buildkit` feature).
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Blocked on**: a concrete use case for building images over the
call protocol. The two container use cases (disposable dev
containers, hosted services) pull pre-built images (`docker/image/pull`)
rather than building them. The reverse-proxy and other hosted
services build via `docker compose build` (operator-side), not via
alknet.
- **Resolution**: Not yet decidable. bollard's `build_image` (image.rs:655)
and the `buildkit` feature (which pulls tonic +
bollard-buildkit-proto) are available but deferred. Build is a large
feature (build context upload, layer caching, multi-stage, buildkit
progress streaming) and is not needed for the current scope. When a
use case forces it, the operation is a `Subscription` (progress
events → `call.responded`, build complete → `call.completed`) and
the `buildkit` feature is enabled in `Cargo.toml` (two-way-door
feature addition, per ADR-059). The v1 surface has `image/pull` +
`image/list` + `image/inspect`; `image/build` is added when needed.
- **Cross-references**:
[ADR-059](decisions/059-bollard-021-dependency-and-features.md)
(feature set decision — `buildkit` not enabled),
[crates/docker/overview.md](crates/docker/overview.md) §"bollard
version and features"
@@ -0,0 +1,34 @@
# OQ-50: Docker System Events Subscription
- **Origin**:
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
§"Out of scope for v1";
[ADR-060](decisions/060-container-resource-model-and-label-namespace.md)
§4 (autonomous-death tolerance).
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: low
- **Blocked on**: a concrete use case for prompt stale-ownership
cleanup. The base model (ADR-060 §4) tolerates stale ownership
entries from autonomous container death (a `--rm` exit, external
`docker rm`, daemon restart) — they're inert (a reused container ID
gets a fresh `record` on its next `create`). A reaper that
subscribes to docker daemon events would clean them promptly, but
the promptness gain is marginal for the current use cases.
- **Resolution**: Not yet decidable. bollard's `events()`
(system.rs:128) returns a `Stream<SystemEventsMessage>` of daemon
events (container start/stop/die/destroy, image pull, etc.). A
`docker/system/events` `Subscription` operation would surface these
as `call.responded` frames; the ownership store could subscribe
internally to revoke on `destroy` events. The deferral is scope:
the base model works without it; the events subscription is a
refinement for when prompt cleanup matters (e.g., a high-churn
coordinator that spawns/removes many containers and wants the
ownership store to stay tight). Adding it is additive (a new
operation + an internal store subscription) and does not break the
existing surface.
- **Cross-references**:
[ADR-060](decisions/060-container-resource-model-and-label-namespace.md)
§4 (teardown coupling — stale-entry tolerance is the base model
this subscription would refine),
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
@@ -0,0 +1,37 @@
# OQ-51: Container Create Options Surface
- **Origin**:
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
§"Out of scope for v1";
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
§"`docker/container/create`".
- **Status**: deferred(scope)
- **Door type**: Two-way
- **Priority**: medium
- **Blocked on**: v1 implementation. The v1 `create` input schema
accepts the common fields (image, command, env, labels, name) and
the full `CreateContainerOptions` surface (mounts, port bindings,
networks, volumes, capabilities, etc.) is deferred to the
implementation pass, where the input JSON Schema can be designed
against bollard's `Config` struct and tested against real create
calls. This is not an architectural decision (the `create` operation
is already decided — ADR-060 §5); it's a schema-detail decision
best made with the bollard types in hand.
- **Resolution**: Not yet decidable. bollard's `create_container`
takes a `Config` struct (`container.rs:296`) with ~40 fields
(`Image`, `Cmd`, `Env`, `Labels`, `HostConfig` with mounts/ports/
networks/etc.). The v1 input schema accepts the high-frequency
fields and omits the long tail; the full surface is a JSON Schema
design task (which fields are required, which are optional, how
`HostConfig` is nested, how the `alknet.*` labels merge with
caller labels). The deferral is to the implementation pass, not
past it — the schema is finalized when `register_docker_ops` is
written and tested. The `OperationSpec.input_schema` is the
one-way surface; its exact field set is a two-way-door refinement
within it.
- **Cross-references**:
[ADR-060](decisions/060-container-resource-model-and-label-namespace.md)
§5 (`create` records ownership — the architectural decision this
OQ defers the schema details of),
[crates/docker/docker-operations.md](crates/docker/docker-operations.md)
§"`docker/container/create`"