From 1ce11717d83bfb3292063db3cfb139e6a337a252 Mon Sep 17 00:00:00 2001 From: "glm-5.2" Date: Wed, 8 Jul 2026 09:29:22 +0000 Subject: [PATCH] docs(arch): draft alknet-docker architecture specs (ADRs 058-063, OQs 048-051) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Greenfield architecture spec set for the alknet-docker crate — a thin, single-host bollard wrapper exposing docker container/image operations as call-protocol ops on the shared alknet/call ALPN, plus a DockerTtyBackend (impl TtyBackend) behind a tty feature for interactive terminal sessions into containers over alknet/tty. Six ADRs: - 058: docker ops on alknet/call (no separate ALPN; raw-carriage handoff dissolved by alknet-tty extraction — interactive attach moved to alknet/tty via DockerTtyBackend, no carriage field on call.requested) - 059: bollard 0.21 (verified current on crates.io) + feature selection (http+pipe+time; no ssl/ssh/websocket/buildkit) - 060: container resource model (ADR-050 application) — alknet.managed/ alknet.owner labels, list owned_only flag, hosted-services operator role via static-resource fallback, handler-driven revoke with autonomous-death tolerance, resource-action vocabulary - 061: DockerTtyBackend in alknet-docker behind tty feature (attach vs exec mode; POC drive_attach_raw as reference) - 062: Docker client + OwnershipStore injection via closure capture (not Capabilities, not OperationContext — matches from_openapi pattern) - 063: exit code on terminal call.responded for non-interactive exec (call.completed stays empty, ADR-012 unchanged) Four spec docs (crates/docker/): README, overview, docker-operations, docker-tty-backend. Four deferred-scope OQs (048-051): network/volume ops, buildkit, system events subscription, create options surface. Updates the tty-backend.md Backend implementations table (DockerTtyBackend row now specced) and the architecture README (doc table, ADR table, current-state paragraph, OQ count). Grounded in the alknet-docker POC (docs/research/alknet-docker/poc-summary.md) which validated the hard parts; the remaining lifecycle ops are mechanical bollard wrapping. Reviewed by architecture-reviewer subagent; criticals (the docker_client injection model conflicting with the Capabilities contract) resolved via ADR-062/063 before commit. --- docs/architecture/README.md | 18 +- docs/architecture/crates/docker/README.md | 166 ++++++ .../crates/docker/docker-operations.md | 537 ++++++++++++++++++ .../crates/docker/docker-tty-backend.md | 482 ++++++++++++++++ docs/architecture/crates/docker/overview.md | 334 +++++++++++ docs/architecture/crates/tty/tty-backend.md | 11 +- .../058-alknet-docker-on-alknet-call.md | 267 +++++++++ ...059-bollard-021-dependency-and-features.md | 195 +++++++ ...iner-resource-model-and-label-namespace.md | 338 +++++++++++ ...061-docker-tty-backend-in-alknet-docker.md | 272 +++++++++ ...er-client-injection-via-closure-capture.md | 234 ++++++++ ...63-exit-code-on-terminal-call-responded.md | 185 ++++++ docs/architecture/open-questions.md | 35 +- ...48-network-and-volume-operation-surface.md | 30 + .../049-image-build-buildkit-scope.md | 30 + .../050-docker-system-events-subscription.md | 34 ++ .../051-container-create-options-surface.md | 37 ++ 17 files changed, 3196 insertions(+), 9 deletions(-) create mode 100644 docs/architecture/crates/docker/README.md create mode 100644 docs/architecture/crates/docker/docker-operations.md create mode 100644 docs/architecture/crates/docker/docker-tty-backend.md create mode 100644 docs/architecture/crates/docker/overview.md create mode 100644 docs/architecture/decisions/058-alknet-docker-on-alknet-call.md create mode 100644 docs/architecture/decisions/059-bollard-021-dependency-and-features.md create mode 100644 docs/architecture/decisions/060-container-resource-model-and-label-namespace.md create mode 100644 docs/architecture/decisions/061-docker-tty-backend-in-alknet-docker.md create mode 100644 docs/architecture/decisions/062-docker-client-injection-via-closure-capture.md create mode 100644 docs/architecture/decisions/063-exit-code-on-terminal-call-responded.md create mode 100644 docs/architecture/questions/048-network-and-volume-operation-surface.md create mode 100644 docs/architecture/questions/049-image-build-buildkit-scope.md create mode 100644 docs/architecture/questions/050-docker-system-events-subscription.md create mode 100644 docs/architecture/questions/051-container-create-options-surface.md diff --git a/docs/architecture/README.md b/docs/architecture/README.md index e8d7204..38d6cfd 100644 --- a/docs/architecture/README.md +++ b/docs/architecture/README.md @@ -1,6 +1,6 @@ --- status: draft -last_updated: 2026-07-06 +last_updated: 2026-07-08 --- # Alknet Architecture @@ -24,12 +24,14 @@ The alknet-call crate is **implemented and reviewed** — both the server-side c **alknet-tty specs drafted.** The alknet-tty crate (terminal session protocol handler — `ProtocolHandler` on `alknet/tty`, two-carriage wire format with a raw chunk codec + JSON control channel, backend-agnostic via a `TtyBackend` trait) now has architecture specs: [crates/tty/](crates/tty/) (overview, tty-wire, tty-backend, tty-adapter, tty-local) and six ADRs — [ADR-052](decisions/052-alknet-tty-wire-format-and-two-carriage.md) (wire format: `alknet/tty` ALPN, JSON negotiation frame then raw chunks, fixed channel set 0-3, control as JSON), [ADR-053](decisions/053-ttybackend-trait-and-ttyhandle.md) (`TtyBackend` trait + `TtyHandle`; `exit_code` as a `Future`; backends need not be natively async — REQ-TTY-01 from the local-PTY POC; `TtyControlHandle` newtype for `Clone`-ability — `Clone` is not object-safe), [ADR-054](decisions/054-local-tty-backend-sibling-crate.md) (`alknet-tty-local` sibling crate behind a `local` feature re-export; PTY vs pipe per-session; the runner pattern preserved), [ADR-055](decisions/055-exit-code-on-control-chunk.md) (exit code on a stream_type 3 control chunk; "exit chunk is last" invariant; adapter owns the ordering), [ADR-056](decisions/056-backend-cleanup-on-session-cancel.md) (backend cleanup contract: dropping the `exit_code` future on session cancel MUST kill the session target — closes the orphaned-process gap the local-PTY POC surfaced for the waiter thread), [ADR-057](decisions/057-alknet-tty-no-alknet-call-dep.md) (alknet-tty does not depend on alknet-call — the negotiation framing is self-contained; the earlier "reuse `FrameFramedReader`" claim was unsound because the utility is welded to `EventEnvelope` deserialization). The specs are grounded in the alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`) and the alknet-tty POC (`/workspace/alknet-tty-poc/`, built 2026-07-05), which validated the wire format, the control channel, the local-PTY bridge, and the signal-delivery contract (REQ-TTY-02). The docker and SSH backends are future crates that implement the `TtyBackend` trait — out of scope for this spec set, but the trait shape is committed so they can be built against it. +**alknet-docker specs drafted.** The alknet-docker crate (docker operations on the shared `alknet/call` ALPN + `DockerTtyBackend` behind a `tty` feature) now has architecture specs: [crates/docker/](crates/docker/) (overview, docker-operations, docker-tty-backend) and six ADRs — [ADR-058](decisions/058-alknet-docker-on-alknet-call.md) (docker ops register on `alknet/call`, not a separate `alknet/docker` ALPN; the raw-carriage handoff the POC struggled with is dissolved by the alknet-tty extraction — interactive attach moved to `alknet/tty` via `DockerTtyBackend`, no `carriage` field on `call.requested`), [ADR-059](decisions/059-bollard-021-dependency-and-features.md) (bollard 0.21, verified current on crates.io; features `http`+`pipe`+`time`, no `ssl`/`ssh`/`websocket`/`buildkit` — single-host by construction, fleet is a call-protocol concern), [ADR-060](decisions/060-container-resource-model-and-label-namespace.md) (ADR-050 application to bollard: `alknet.managed`/`alknet.owner` labels; `list` `owned_only` flag; hosted-services operator role via the static-resource fallback; handler-driven `revoke` on `remove` with autonomous-death tolerance), [ADR-061](decisions/061-docker-tty-backend-in-alknet-docker.md) (`DockerTtyBackend` in alknet-docker behind a `tty` feature, not a sibling crate; attach vs exec mode; the POC's `drive_attach_raw` as the reference), [ADR-062](decisions/062-docker-client-injection-via-closure-capture.md) (the `Docker` client + `OwnershipStore` are closure-captured at registration time, not read from `OperationContext` and not smuggled through `Capabilities` — `Capabilities` is for secret material only per ADR-014; matches the `from_openapi` pattern), [ADR-063](decisions/063-exit-code-on-terminal-call-responded.md) (non-interactive exec puts `{ "exitCode": N, "terminal": true }` on a final `call.responded` before `call.completed` — `call.completed` stays empty, ADR-012 unchanged). The specs are grounded in the alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`, `/workspace/alknet-docker-poc/`), which validated the hard parts (interactive attach, logs subscription, exec with exit code); the remaining lifecycle operations are mechanical bollard wrapping. The two use cases — disposable dev containers (coordinator-spawned, ownership-recorded) and long-running hosted services (operator-managed, static-resource fallback, per `/workspace/system/dev1/docker.md`) — both work through one `AccessControl` model (ADR-050/060). The `DockerTtyBackend` fills the `TtyBackend` row the alknet-tty spec left open. Four OQs (048–051) track deferred scope: network/volume ops, buildkit, system events subscription, and the full `CreateContainerOptions` surface (deferred to v1 implementation). + ## Architecture Documents | Document | Status | Description | |----------|--------|-------------| | [overview.md](overview.md) | draft | Workspace-level overview, crate graph, shared types, design principles | -| [open-questions.md](open-questions.md) | draft | OQ index — theme-grouped tables + Deferred/Blocked section; per-OQ files in [`questions/`](open-questions/questions/) | +| [open-questions.md](open-questions.md) | draft | OQ index — theme-grouped tables + Deferred/Blocked section; per-OQ files in [`questions/`](questions/) | | [crates/core/README.md](crates/core/README.md) | draft | alknet-core crate index | | [crates/core/core-types.md](crates/core/core-types.md) | draft | ProtocolHandler, HandlerError, Connection, BiStream, StreamError | | [crates/core/endpoint.md](crates/core/endpoint.md) | draft | ALPN router, HandlerRegistry, accept loop, shutdown | @@ -52,6 +54,10 @@ The alknet-call crate is **implemented and reviewed** — both the server-side c | [crates/tty/tty-backend.md](crates/tty/tty-backend.md) | draft | `TtyBackend` trait, `TtyParams`, `TtyHandle`, `TtyControl` — the backend inversion point. Carries REQ-TTY-01 (backends need not be natively async) | | [crates/tty/tty-adapter.md](crates/tty/tty-adapter.md) | draft | `TtyAdapter` (`ProtocolHandler` on `alknet/tty`): session lifecycle, three-pump bidirectional driver, negotiation errors, exit-chunk ordering (ADR-055), access control | | [crates/tty/tty-local.md](crates/tty/tty-local.md) | draft | `alknet-tty-local` sibling crate: `LocalTtyBackend` via `portable_pty` (PTY) and `std::process::Command` (pipe/runner). Carries REQ-TTY-02 (signal forwarding to the process group) | +| [crates/docker/README.md](crates/docker/README.md) | draft | alknet-docker crate index | +| [crates/docker/overview.md](crates/docker/overview.md) | draft | Crate purpose, two-role design (call ops + DockerTtyBackend), dependencies, ALPN, label namespace, feature gates, assembly-layer wiring | +| [crates/docker/docker-operations.md](crates/docker/docker-operations.md) | draft | Operation surface: lifecycle (Query/Mutation), logs/exec/pull (Subscription via StreamingHandler), access control (ADR-050/060), label namespace, teardown coupling | +| [crates/docker/docker-tty-backend.md](crates/docker/docker-tty-backend.md) | draft | `DockerTtyBackend` (impl `TtyBackend`): attach vs exec mode, `TtyHandle` field mapping, `TtyControl` → bollard resize/signal, `exit_code` Drop-kill (ADR-056) | | [crates/vault/README.md](crates/vault/README.md) | stable | alknet-vault crate index | | [crates/vault/mnemonic-derivation.md](crates/vault/mnemonic-derivation.md) | stable | BIP39, SLIP-0010, BIP-0032, derivation paths, key types | | [crates/vault/encryption.md](crates/vault/encryption.md) | stable | AES-256-GCM, EncryptedData, key versioning, salt (Phase B reserved) | @@ -119,10 +125,16 @@ The alknet-call crate is **implemented and reviewed** — both the server-side c | [055](decisions/055-exit-code-on-control-chunk.md) | Exit Code on a Control Chunk (the Last Chunk Before Stream Close) | Accepted | | [056](decisions/056-backend-cleanup-on-session-cancel.md) | Backend Cleanup on Session Cancel (Drop of `exit_code` Kills) | Accepted | | [057](decisions/057-alknet-tty-no-alknet-call-dep.md) | alknet-tty Does Not Depend on alknet-call (Self-Contained Negotiation Framing) | Accepted | +| [058](decisions/058-alknet-docker-on-alknet-call.md) | alknet-docker Registers on `alknet/call` (No Separate ALPN) | Accepted | +| [059](decisions/059-bollard-021-dependency-and-features.md) | bollard 0.21 Dependency and Feature Selection | Accepted | +| [060](decisions/060-container-resource-model-and-label-namespace.md) | Container Resource Model and Label Namespace | Accepted | +| [061](decisions/061-docker-tty-backend-in-alknet-docker.md) | DockerTtyBackend in alknet-docker | Accepted | +| [062](decisions/062-docker-client-injection-via-closure-capture.md) | Docker Client and OwnershipStore Injection via Closure Capture | Accepted | +| [063](decisions/063-exit-code-on-terminal-call-responded.md) | Exit Code on a Terminal `call.responded` for Non-Interactive Exec | Accepted | ## Open Questions -Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (47 OQs across 15 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](open-questions/questions/) (`NNN-slug.md`, mirroring the ADR convention). +Open questions are tracked in [open-questions.md](open-questions.md) — an index of theme-grouped tables (51 OQs across 16 themes) with a cross-theme [Deferred / Blocked](open-questions.md#deferred--blocked) section surfacing the safe-exit deferrals. Each OQ lives in its own file under [`questions/`](questions/) (`NNN-slug.md`, mirroring the ADR convention). ## Document Lifecycle diff --git a/docs/architecture/crates/docker/README.md b/docs/architecture/crates/docker/README.md new file mode 100644 index 0000000..c9c585a --- /dev/null +++ b/docs/architecture/crates/docker/README.md @@ -0,0 +1,166 @@ +--- +status: draft +last_updated: 2026-07-08 +--- + +# alknet-docker + +The docker operations crate for the ALPN-as-service architecture: a +thin, single-host bollard wrapper that exposes docker container and +image operations as call-protocol operations on the shared +`alknet/call` ALPN, plus (behind a `tty` feature) a `DockerTtyBackend` +implementing `alknet-tty`'s `TtyBackend` trait for interactive terminal +sessions into containers. + +## What + +alknet-docker does two things: + +1. **Call operations.** A set of `OperationSpec`-registered operations + (`docker/container/*`, `docker/image/*`) on the shared + `alknet/call` ALPN, mapping bollard's docker API to the call + protocol's `Query` / `Mutation` / `Subscription` dispatch paths. + Lifecycle (create/start/stop/remove/list/inspect) is `Query`/ + `Mutation`; logs, non-interactive exec, and image pull are + `Subscription` (streaming via `StreamingHandler`, ADR-049). The + operations declare `AccessControl` against the ADR-050 container- + as-resource model. Decided in + [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md). + +2. **`DockerTtyBackend`** (behind the `tty` feature). An + `impl TtyBackend` (ADR-053) wrapping `bollard::attach_container()` + and `bollard::exec::start_exec` with `tty: true`, for interactive + terminal sessions into containers over `alknet/tty`. This is the + `TtyBackend` row the alknet-tty spec left open ("future, out of + scope here"). Decided in + [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md). + +The two use cases the crate serves (per the user's brief and the system +docs): + +- **Disposable dev containers** (the common case by volume) — a + coordinator spawns a container for an implementation agent or an + isolated env. alknet-docker's `create` records ownership + (`OwnershipStore::record`); the coordinator owns the container for + its lifetime; `remove` revokes. Interactive sessions into the + container go through `DockerTtyBackend` on `alknet/tty`. +- **Long-running hosted services** (the production server case) — + rarely-changing services (reverse-proxy, postgres, redis, gitea on + dev1, per `/workspace/system/dev1/docker.md`) created by an operator + via `docker compose`, not via alknet. alknet-docker's operations + (start/stop/inspect/logs) manage them; the operator role reaches them + via the static-resource fallback (ADR-060 §3). + +## Documents + +| Document | Status | Description | +|----------|--------|-------------| +| [overview.md](overview.md) | draft | Crate purpose, two-role design (call ops + DockerTtyBackend), dependencies, ALPN, label namespace, feature gates, assembly-layer wiring | +| [docker-operations.md](docker-operations.md) | draft | The operation surface: lifecycle (Query/Mutation), logs/exec/pull (Subscription via StreamingHandler), access control, label namespace, teardown coupling | +| [docker-tty-backend.md](docker-tty-backend.md) | draft | `DockerTtyBackend` (impl `TtyBackend`): attach vs exec mode, `TtyHandle` field mapping, `TtyControl` → bollard resize/signal, `exit_code` Drop-kill (ADR-056) | + +## Applicable ADRs + +| ADR | Title | Relevance | +|-----|-------|-----------| +| [058](../../decisions/058-alknet-docker-on-alknet-call.md) | alknet-docker Registers on `alknet/call` | Docker ops are call-protocol operations, not a separate ALPN; raw-carriage dissolved by alknet-tty extraction | +| [059](../../decisions/059-bollard-021-dependency-and-features.md) | bollard 0.21 Dependency and Feature Selection | Version pin (0.21, verified current); features `http`+`pipe`+`time`, no `ssl`/`ssh`/`websocket`/`buildkit` | +| [060](../../decisions/060-container-resource-model-and-label-namespace.md) | Container Resource Model and Label Namespace | ADR-050 application: `alknet.managed`/`alknet.owner` labels; `list` owned_only flag; hosted-services static-resource fallback; handler-driven revoke + autonomous-death tolerance | +| [061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) | DockerTtyBackend in alknet-docker | Backend in alknet-docker behind `tty` feature; attach/exec mode split; POC `drive_attach_raw` as reference | +| [003](../../decisions/003-crate-decomposition.md) | Crate Decomposition | alknet-docker depends on alknet-core + alknet-call (ops) and alknet-tty (tty feature); no handler-depends-on-handler violation | +| [012](../../decisions/012-call-protocol-stream-model.md) | Call Protocol Stream Model | The wire format docker ops use (`EventEnvelope`, no `carriage` field) | +| [017](../../decisions/017-call-protocol-client-and-adapter-contract.md) | Call Protocol Client and Adapter Contract | `from_call` re-export — the proxy pattern's mechanism for hub→worker docker management | +| [023](../../decisions/023-operation-error-schemas.md) | Operation Error Schemas | Docker ops declare domain error codes (CONTAINER_NOT_FOUND, IMAGE_NOT_FOUND, etc.) | +| [024](../../decisions/024-operation-registry-layering.md) | Operation Registry Layering | Docker ops register in the curated Layer 0 at startup | +| [029](../../decisions/029-peer-graph-routing-model.md) | Peer-Graph Routing Model | Head→worker docker management via `PeerRef` / `invoke_peer` | +| [032](../../decisions/032-forwarded-for-identity.md) | Forwarded-For Identity | End-user identity as metadata when a coordinator proxies docker ops | +| [049](../../decisions/049-streaming-handler-for-subscriptions.md) | Streaming Handler for Subscriptions | `StreamingHandler` for logs/exec/pull; exit code on final `call.responded` | +| [050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Dynamic Resource Ownership | Containers as `AccessControl` resources; the model ADR-060 applies | +| [052](../../decisions/052-alknet-tty-wire-format-and-two-carriage.md) | alknet-tty Wire Format | The `alknet/tty` wire format `DockerTtyBackend` sessions use | +| [053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) | TtyBackend Trait and TtyHandle | The trait `DockerTtyBackend` implements | +| [055](../../decisions/055-exit-code-on-control-chunk.md) | Exit Code on a Control Chunk | The exit-chunk ordering `DockerTtyBackend`'s `exit_code` feeds into | +| [056](../../decisions/056-backend-cleanup-on-session-cancel.md) | Backend Cleanup on Session Cancel | `DockerTtyBackend`'s `exit_code` future `Drop` kills the container/exec | +| [014](../../decisions/014-secret-material-flow-and-capability-injection.md) | Secret Material Flow and Capability Injection | `Capabilities` is for secret material only — the `Docker` handle is not secret (ADR-062) | +| [062](../../decisions/062-docker-client-injection-via-closure-capture.md) | Docker Client and OwnershipStore Injection via Closure Capture | Closure capture at registration; `Capabilities` for secrets only; matches `from_openapi` pattern | +| [063](../../decisions/063-exit-code-on-terminal-call-responded.md) | Exit Code on a Terminal `call.responded` for Non-Interactive Exec | `{ "exitCode": N, "terminal": true }` on final `call.responded`; `call.completed` stays empty | + +## Relevant Open Questions + +| OQ | Title | Status | Relevance | +|----|-------|--------|-----------| +| OQ-048 | Network and volume operation surface | deferred(scope) | Network/volume CRUD deferred; v1 is containers + images | +| OQ-049 | Image build (buildkit) scope | deferred(scope) | `buildkit` feature deferred; v1 has `image/pull` + `image/list` + `image/inspect` | +| OQ-050 | Docker system events subscription | deferred(scope) | `docker/system/events` subscription for stale-ownership cleanup deferred | +| OQ-051 | Container create options surface | deferred(scope) | Full `CreateContainerOptions` (mounts, port bindings, networks) surface deferred to v1 implementation | + +## Key Design Principles + +1. **Single-host, bollard-specific.** alknet-docker talks to one local + docker daemon over `/var/run/docker.sock`. The fleet case (multiple + daemons on different machines) is a call-protocol concern — a + `CallClient` per remote daemon, each running alknet-docker locally + — not a bollard-feature concern. No `ssl`/`ssh` features, no remote + daemon over TLS. See [overview.md](overview.md) and + [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md). + +2. **Call operations on `alknet/call`, not a separate ALPN.** Docker + ops are ordinary call-protocol operations with `OperationSpec`s, + `AccessControl`, and service discovery. They inherit `from_call` + re-export and peer routing. The one operation that didn't fit + (interactive attach) moved to `alknet/tty` via `DockerTtyBackend`. + No `carriage` field on `call.requested`, no parallel dispatch + surface. See [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md). + +3. **Containers are runtime-spawned resources (ADR-050).** `create` + records ownership; `remove` revokes; `exec`/`stop`/`inspect` check + ownership via `OperationSpec.resource_id_path`. The + hosted-services case (operator-managed, pre-existing containers) + works through the static-resource fallback. Two labels + (`alknet.managed`, `alknet.owner`) mark alknet-spawned containers + for the `list` filter and the cross-check. See + [docker-operations.md](docker-operations.md) and + [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md). + +4. **Interactive terminal is a tty concern, not a docker concern.** + `DockerTtyBackend` implements `alknet-tty`'s `TtyBackend` trait, + behind a `tty` feature. The wire format, the session lifecycle, and + the exit-chunk ordering live in alknet-tty; the backend produces + bollard-backed handles. This is the same inversion as + `LocalTtyBackend` (local process) and `SshTtyBackend` (SSH). See + [docker-tty-backend.md](docker-tty-backend.md) and + [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md). + +5. **bollard 0.21, verified current.** The POC used a local 0.21 + checkout; the crate depends on published 0.21 from crates.io + (verified as latest). Features: `http` + `pipe` (default, local + daemon) + `time` (log timestamps). No `websocket` (reliable attach + only), no `ssl`/`ssh` (fleet is call-protocol), no `buildkit` + (deferred). See [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md). + +6. **Mechanical mapping, no feasibility risk.** The POC validated the + hard parts (raw carriage attach, logs subscription, exec with exit + code). The remaining lifecycle operations (create/start/stop/remove/ + list/inspect) are mechanical bollard wrapping — `Query`/`Mutation` + with single `call.responded` responses, "the boring case" (POC §"What + the POC Does NOT Validate" #4). See + [docker-operations.md](docker-operations.md). + +## References + +- `docs/research/alknet-docker/poc-summary.md` — the POC that validated + the hard parts (interactive attach, logs subscription, exec with exit + code) and surfaced the open unknowns this spec set resolves +- `/workspace/alknet-docker-poc/` — the POC source (`src/ops.rs` + `DockerOps`, `src/raw.rs` chunk codec, `src/frame.rs` EventEnvelope + mirror, `tests/integration.rs` 6 tests against a live daemon) +- `/workspace/bollard/` — bollard 0.21.0 source (the local checkout the + POC used; identical API surface to the published 0.21) +- `/workspace/@alkdev/dispatch/` — the dispatch POC (bollard 0.18, + `dispatch.managed=true` labels, SSH-tunnel fleet model — the prior + art this crate generalizes and the friction it removes) +- `/workspace/system/dev1/docker.md` — the production hosted-services + use case (reverse-proxy, postgres, redis, gitea on dev1) +- `/workspace/@alkdev/reverse-proxy/deploy/docker-compose.yml` — the + reverse-proxy's docker setup (operator-created, not alknet-spawned) +- `docs/architecture/crates/tty/` — the alknet-tty spec + (`DockerTtyBackend` is the row `tty-backend.md` left open) \ No newline at end of file diff --git a/docs/architecture/crates/docker/docker-operations.md b/docs/architecture/crates/docker/docker-operations.md new file mode 100644 index 0000000..fbf47cf --- /dev/null +++ b/docs/architecture/crates/docker/docker-operations.md @@ -0,0 +1,537 @@ +--- +status: draft +last_updated: 2026-07-08 +--- + +# alknet-docker — Operations + +The operation surface alknet-docker registers on the shared +`alknet/call` `OperationRegistry`: lifecycle operations (Query/ +Mutation), streaming operations (Subscription via `StreamingHandler`), +access control (ADR-050/060 application), the label namespace, and +teardown coupling. This document specifies what an implementer builds +against. The ALPN decision (shared `alknet/call`, no separate +`alknet/docker`) is in [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md). + +## What + +alknet-docker exports a `register_docker_ops()` function (or a +`DockerOps` registration bundle) that the assembly layer calls to add +docker operations to the shared `OperationRegistryBuilder`. Each +operation is a `HandlerRegistration` with: + +- An `OperationSpec` (name, namespace, op_type, schemas, + `access_control`, `resource_id_path` per ADR-050). +- A `HandlerKind` (`Once` for Query/Mutation, `Stream` for + Subscription — ADR-049). +- `provenance: Local` (assembly-registered, can compose). +- A `CompositionAuthority` (the scopes the docker ops run under when + composed — e.g., `container:list`, `container:exec`). +- Empty `Capabilities` (local bollard, no outbound credentials). + +The operations are dispatched by the existing `CallAdapter` through +`invoke()` / `invoke_streaming()`. There is no docker-specific +dispatch code in alknet-call. + +## Why + +The POC (`docs/research/alknet-docker/poc-summary.md`) validated the +hard parts (interactive attach, logs subscription, exec with exit +code) and noted the remaining lifecycle operations are "mechanical +bollard wrapping, no feasibility risk... `Query`/`Mutation` operations +with single `call.responded` responses, the boring case" (POC §"What +the POC Does NOT Validate" #4). This document specifies the boring +case + the validated streaming cases as a concrete operation surface. + +The operations declare `AccessControl` against the ADR-050 +container-as-resource model, applied to bollard's API via ADR-060 +(label namespace, `list` filter, hosted-services fallback, teardown +coupling). No new auth model is invented; the existing model is +applied. + +## Architecture + +### Operation Surface (v1 scope) + +| Operation | Type | bollard method | `resource_id_path` | Notes | +|-----------|------|---------------|-------------------|-------| +| `docker/container/list` | Query | `list_containers` | (none — list case) | `owned_only` input flag (ADR-060 §2) | +| `docker/container/inspect` | Query | `inspect_container` | `$.containerId` | | +| `docker/container/create` | Mutation | `create_container` | (none — creates the resource) | Records ownership (ADR-060 §5) | +| `docker/container/start` | Mutation | `start_container` | `$.containerId` | | +| `docker/container/stop` | Mutation | `stop_container` | `$.containerId` | | +| `docker/container/remove` | Mutation | `remove_container` | `$.containerId` | Revokes ownership (ADR-060 §4) | +| `docker/container/restart` | Mutation | `restart_container` | `$.containerId` | | +| `docker/container/logs` | Subscription | `logs` | `$.containerId` | StreamingHandler; `follow` input flag | +| `docker/container/exec` | Subscription | `create_exec` + `start_exec` + `inspect_exec` | `$.containerId` | StreamingHandler; `tty:false` only; exit code on final `call.responded` | +| `docker/image/list` | Query | `list_images` | (none) | | +| `docker/image/pull` | Subscription | `create_image` | (none) | StreamingHandler; progress events | +| `docker/image/inspect` | Query | `inspect_image` | (none) | | + +**Out of scope for v1 (deferred):** + +- Network operations (`docker/network/*`) — OQ-048. +- Volume operations (`docker/volume/*`) — OQ-048. +- Image build (`docker/image/build` with buildkit) — OQ-049. +- System events (`docker/system/events`) — OQ-050. +- Full `CreateContainerOptions` surface (mounts, port bindings, + networks) — OQ-051. v1 `create` accepts the image, command, env, + labels, and name; the full options surface is a v1 implementation + refinement. +- Interactive exec / attach (`tty: true`) — not a call operation; + `DockerTtyBackend` on `alknet/tty` (ADR-061). + +### Lifecycle Operations (Query / Mutation) + +The lifecycle operations are the "boring case": a single bollard +async method call, a single `call.responded` with the JSON result +(or `call.error` on failure). The handler is a `Handler` (not +`StreamingHandler`), registered as `HandlerKind::Once`. + +The `start`, `stop`, and `restart` operations share the `inspect` +shape (a `Mutation` with `resource_id_path: "$.containerId"`, +`container:` scope, a single bollard method call, a single +`call.responded`). The representative example below is `inspect`; +the others differ only in the bollard method and the declared +scope. + +```rust +// docker/container/inspect — representative lifecycle op. +// The handler closure captures the bollard Docker client by +// Arc::clone at registration time (ADR-062); it does not read the +// client from OperationContext. +let docker_clone = docker.clone(); // Arc +let handler: Handler = Arc::new(move |input, ctx| { + let docker = docker_clone.clone(); + Box::pin(async move { + let container_id = input["containerId"].as_str() + .ok_or_else(|| /* INVALID_INPUT */)?; + match docker.inspect_container(container_id, None::<()>).await { + Ok(info) => ResponseEnvelope::ok(to_json_value(info)), + Err(bollard::Error::DockerResponseServerError { status_code: 404, .. }) => + ResponseEnvelope::error(CallError { + code: "CONTAINER_NOT_FOUND".into(), + message: format!("container {} not found", container_id), + retryable: false, details: None, + }), + Err(e) => ResponseEnvelope::error(CallError { + code: "DOCKER_ERROR".into(), + message: e.to_string(), + retryable: false, details: None, + }), + } + }) +}); +``` + +The handler extracts the container ID from the input, calls the +bollard method, and maps the result to a `ResponseEnvelope`. The +bollard error is mapped to a declared domain error code +(`CONTAINER_NOT_FOUND`, `DOCKER_ERROR`) per ADR-023. The +`OperationSpec.error_schemas` declare these codes so clients get +typed error enums. + +### Handler Injection (ADR-062) + +The `Docker` client and `OwnershipStore` are **closure-captured** at +registration time, not read from `OperationContext` or smuggled +through `Capabilities`. This matches the established `from_openapi` +pattern (ADR-017): non-secret shared state via closure capture, +secret material via `Capabilities`. A bollard `Docker` handle to a +local unix socket is not secret material — it must not go through +`Capabilities` (ADR-014's contract: `Capabilities` is for outbound +secret material only). The `register_docker_ops` function takes +`Arc` + `Arc` + the label config + the +`CompositionAuthority`, and constructs each handler closure with the +handles it needs captured by `Arc::clone`. The `Capabilities` passed +to each registration is empty (`Capabilities::new()`) — local bollard +needs no API key. See [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md) +for the full decision and the rationale for why `Capabilities` and +`OperationContext` extension are the wrong channels. + +### `docker/container/create` — ownership recording + +`create` is the spawn event (ADR-060 §5). The handler: + +1. Parses the input (image, command, env, labels, name, and v1-limited + `CreateContainerOptions` — OQ-051 for the full surface). +2. Injects the alknet labels: `alknet.managed=true`, + `alknet.owner=`. The caller's peer ID comes from + `ctx.identity` (`Identity.id`, ADR-030). For a composed call (the + coordinator composing `create`), this is the coordinator's + identity, not the end user's — the proxy pattern (ADR-050 §3). +3. Calls `bollard::create_container()`. +4. On success, calls `OwnershipStore::record(identity, "container", + container_id)` (ADR-050 §1, ADR-060 §5). +5. Returns the `ContainerCreateResponse` (with the container ID) as + `call.responded`. + +The label injection is the handler's job, not bollard's — bollard's +`CreateContainerOptions.labels` accepts a map; the handler merges the +alknet labels with any caller-provided labels (caller labels win on +conflict for non-`alknet.*` keys; `alknet.*` keys are reserved and +overwritten by the handler to prevent a caller spoofing ownership). + +### `docker/container/remove` — ownership revocation + +`remove` is the teardown event (ADR-060 §4). The handler: + +1. Calls `bollard::remove_container()`. +2. On success, calls `OwnershipStore::revoke("container", container_id)`. +3. Returns `call.responded` with an empty/ok result. + +Autonomous container death (a `--rm` exit, external `docker rm`, +daemon restart) is tolerated — stale ownership entries are not +promptly cleaned up (no reaper); they're inert (a reused container ID +gets a fresh `record` on its next `create`). See ADR-060 §4. + +### `docker/container/list` — scope-gate + optional result-filter + +`list` is the ADR-050 #4a "list case": `resource_type: "container"`, +no `resource_id_path`. The input accepts an `owned_only: bool` flag +(default `false`): + +- `owned_only: false` — `bollard::list_containers()` with no label + filter; returns all containers the daemon sees. The scope check + gates the call (caller needs `container:list`). This is the + hosted-services case (operator lists all). +- `owned_only: true` — `bollard::list_containers()` with a label + filter (`label: alknet.owner=`); returns only the + caller's containers. This is the disposable-dev-container case + (coordinator lists its own). + +The result-filter is bollard-side (the label filter is in the +`ListContainersOptions`), not handler-side — bollard filters at the +daemon. The handler doesn't call `OwnershipProvider::owned_resources` +and filter in Rust; it pushes the filter to bollard via the label. +This is more efficient (the daemon filters before sending) and +correct (the label and the ownership store agree for alknet-spawned +containers; the cross-check is ADR-060 §1's secondary signal). + +### Streaming Operations (Subscription) + +The streaming operations (logs, exec, image pull) are `Subscription` +operations using `StreamingHandler` (ADR-049). The handler returns a +`Stream`; the dispatcher's `pump_stream` writes +each `Ok(value)` as `call.responded`, an `Err` as `call.error` +(terminal), and natural stream end as `call.completed`. + +#### `docker/container/logs` + +Maps `bollard::container::logs()` (returns +`Stream>`) to a stream of +`call.responded` frames. The POC validated this path (POC target 2: +`docker_logs_subscription_pumps_frames_and_completes`). + +```rust +// The StreamingHandler closure captures the bollard Docker client +// by Arc::clone at registration time (ADR-062). ResponseStream is +// the type alias Pin + Send>> +// (ADR-049, operation-registry.md §"Handler"). +let docker_clone = docker.clone(); +let handler: StreamingHandler = Arc::new(move |input, ctx| { + let docker = docker_clone.clone(); + Box::pin(async_stream::stream! { + let container_id = input["containerId"].as_str().unwrap(); + let follow = input["follow"].as_bool().unwrap_or(false); + let options = LogsOptionsBuilder::default() + .follow(follow) + .stdout(true).stderr(true) + .timestamps(true) // ADR-059 §4: time feature + .build(); + let mut stream = docker.logs(container_id, Some(options)); + while let Some(result) = stream.next().await { + match result { + Ok(LogOutput::StdOut { message }) | + Ok(LogOutput::Console { message }) => yield ResponseEnvelope::ok(json!({ + "stream": "stdout", + "timestamp": /* from LogOutput */, + "text": /* message bytes as UTF-8 */ + })), + Ok(LogOutput::StdErr { message }) => yield ResponseEnvelope::ok(json!({ + "stream": "stderr", + "timestamp": /* ... */, + "text": /* ... */ + })), + Ok(_) => yield ResponseEnvelope::ok(json!({})), // other variants + Err(e) => yield ResponseEnvelope::error(CallError { + code: "DOCKER_ERROR".into(), + message: e.to_string(), + retryable: false, details: None, + }), + } + } + // stream end → call.completed (the dispatcher's pump_stream + // writes call.completed on natural stream end — ADR-049) + }) as ResponseStream +}); +``` + +Each `LogOutput` becomes a `call.responded` with `stream` +(stdout/stderr), `timestamp` (from the `time` feature, ADR-059 §4), +and `text` (the log line, as UTF-8 — the POC's single-`text`-field +refinement separates timestamp and text). Stream end (the container +exits for `follow: true`, or historical logs drain for `follow: false`) +→ `call.completed` (the dispatcher's `pump_stream` handles this on +natural stream end). + +#### `docker/container/exec` (non-interactive, `tty: false`) + +Maps `bollard::create_exec` + `start_exec` + `inspect_exec` to a +stream of `call.responded` frames, with the exit code on a final +`call.responded` before `call.completed`. The POC validated this +path (POC target 3: `docker_exec_streams_output_and_exit_code`). + +The `tty` field in the input MUST be `false` (or absent). If `tty: +true`, the handler returns a single `ResponseEnvelope::error` with +code `INVALID_INPUT` and a message directing the caller to +`alknet/tty` for interactive exec (ADR-058 §3). This is the +one-operation-per-carriage invariant: the call operation is captured +output; the tty session is interactive. + +The handler: +1. `create_exec(container_id, CreateExecOptions { cmd, env, tty: false, ... })` + → exec ID. +2. `start_exec(exec_id, None)` → `StartExecResults::Attached { output, input }`. +3. Pump `output` (a `Stream`) as `call.responded` frames + (stdout/stderr separated, since `tty: false`). +4. After the output stream ends, `inspect_exec(exec_id)` → + `ExecInspectResponse { exit_code, ... }`. +5. Emit a final `call.responded` with `{ "exitCode": N, "terminal": true }`. +6. Stream end → `call.completed` (the exit-code `call.responded` is + the last one before `completed`). + +The exit code rides on a `call.responded`, not on `call.completed`. +The `terminal: true` flag marks this `call.responded` as the final +value before `call.completed`. The full shape decision (why +`terminal: true` on `call.responded` and not an exit code on +`call.completed` or a `call.exit` event) is in +[ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md). + +#### `docker/image/pull` + +Maps `bollard::image::create_image()` (returns +`Stream>`) to a stream of +`call.responded` frames with progress events. Same shape as logs: +each progress event → `call.responded`, stream end → `call.completed`. +No exit code (image pull has no exit code; success is stream end +without error). + +### Access Control + +The operations declare `AccessControl` against the ADR-050 model, +applied via ADR-060. Three patterns: + +#### Specific-container operations (`exec`, `inspect`, `start`, `stop`, `remove`, `restart`, `logs`) + +```rust +OperationSpec { + name: "docker/container/exec", + access_control: AccessControl { + required_scopes: vec!["container:exec".into()], + resource_type: Some("container".into()), + resource_action: Some("exec".into()), + .. + }, + resource_id_path: Some("$.containerId".into()), + .. +} +``` + +The dispatcher extracts `containerId` from the input via +`resource_id_path` and passes it to `AccessControl::check`, which +consults the `OwnershipProvider` (ADR-050 §2). The check passes if: +- the caller owns the container (the ownership store says so — the + `create` path recorded it), OR +- the caller's static `Identity.resources["container"]` includes an + action that subsumes the required action (the operator-role + fallback — ADR-060 §3, e.g., `container:manage` ⊇ `container:exec`). + +#### The `list` case (`container/list`) + +```rust +OperationSpec { + name: "docker/container/list", + access_control: AccessControl { + required_scopes: vec!["container:list".into()], + resource_type: Some("container".into()), + // no resource_action — the list case + .. + }, + // no resource_id_path — the list case (ADR-050 #4a) + .. +} +``` + +The scope check gates the call. The `owned_only` input flag selects +the result-filter (label-based, bollard-side). See ADR-060 §2. + +#### The `create` case (`container/create`) + +```rust +OperationSpec { + name: "docker/container/create", + access_control: AccessControl { + required_scopes: vec!["container:create".into()], + // no resource_type — create spawns the resource; the + // ownership is recorded after the bollard call succeeds. + // The scope check gates; the ownership is a side effect. + .. + }, + // no resource_id_path — the container doesn't exist yet + .. +} +``` + +`create` has no `resource_type` (the resource doesn't exist yet); +the scope check gates. The ownership recording is a handler side +effect (ADR-060 §5), not an ACL check. + +#### Image operations (`image/list`, `image/pull`, `image/inspect`) + +```rust +OperationSpec { + name: "docker/image/pull", + access_control: AccessControl { + required_scopes: vec!["image:pull".into()], + // no resource_type — images are not runtime-spawned resources + // in the ADR-050 sense (they're pulled, not spawned with + // per-caller ownership). Scope-gate only. + .. + }, + .. +} +``` + +Images are not runtime-spawned resources with per-caller ownership +— they're shared daemon state. The scope check gates; no ownership +provider consultation. + +### Error Schemas (ADR-023) + +Each operation declares its domain error codes in +`OperationSpec.error_schemas`: + +| Code | Operations | HTTP status (for to_openapi projection) | Description | +|------|-----------|----------------------------------------|-------------| +| `CONTAINER_NOT_FOUND` | inspect, start, stop, remove, restart, logs, exec | 404 | The container ID doesn't exist on the daemon | +| `IMAGE_NOT_FOUND` | image/inspect | 404 | The image doesn't exist locally | +| `IMAGE_PULL_FAILED` | image/pull | 502 | The pull failed (registry error, network) | +| `DOCKER_ERROR` | all | 500 | A bollard error not covered by a specific code | +| `INVALID_INPUT` | exec (tty:true) | 400 | `tty: true` on `docker/container/exec` — use `alknet/tty` | + +The `INVALID_INPUT` for `tty: true` on exec carries a message +directing the caller to `alknet/tty` for interactive exec. This is +the one-operation-per-carriage invariant's user-visible surface +(ADR-058 §3). + +### Composition and Peer Routing + +Docker operations compose through the standard call-protocol +mechanisms: + +- **`from_call` re-export (ADR-017).** A hub that wants to manage + docker on a worker imports the worker's `docker/container/*` + operations via `from_call` and re-registers them locally. The hub's + clients call the re-exported operations; the hub is the direct + caller to the worker; the worker's ownership store sees the hub as + the owner (the proxy pattern, ADR-050 §3). The end user's identity + rides as `forwarded_for` (ADR-032). +- **Peer routing (ADR-029).** A hub with multiple worker connections + composes `docker/container/exec` via `invoke_peer(PeerRef::Specific( + worker_peer_id), ...)`. The `PeerCompositeEnv` routes to the named + worker's sub-overlay. This is the head→worker fan-out primitive. +- **Local composition.** A coordinator handler that spawns a + container and then execs into it composes `docker/container/create` + then `docker/container/exec` through `OperationEnv::invoke()`. The + coordinator's `CompositionAuthority` has `container:create` + + `container:exec` scopes (static, ADR-022); the ownership check + passes because `create` recorded the coordinator as the owner. + +The `docker/container/exec` with `tty: true` rejection is a +composition boundary: a coordinator composing exec must use `tty: +false` (captured output); interactive exec requires an `alknet/tty` +session, which is a different protocol (the coordinator opens a tty +connection, not a call operation). This is the same boundary as +"SSH exec" (call op) vs "SSH PTY session" (alknet-tty). + +## Constraints + +- **No `carriage` field on `call.requested`.** The call protocol's + wire format (ADR-012) is `EventEnvelope`-only on `alknet/call`. + The one operation that needed raw carriage (interactive attach) + moved to `alknet/tty` via `DockerTtyBackend` (ADR-058 §2, ADR-061). + A `docker/container/exec` call with `tty: true` is rejected with + `INVALID_INPUT`. +- **`resource_id_path` is required for specific-container ops.** The + dispatcher extracts the container ID from the input via the JSON + pointer; `AccessControl::check` receives it. Operations without + `resource_id_path` (`list`, `create`, image ops) don't reference a + specific container and don't consult the ownership provider for a + specific resource. +- **`create` records ownership; `remove` revokes; lifecycle ops do + neither.** Only the spawn (`create`) and the teardown (`remove`) + touch the ownership store. `start`/`stop`/`restart` act on an + existing container whose ownership was recorded at create time + (or which pre-exists and is reached via the static fallback). +- **Stale ownership entries are tolerated.** Autonomous container + death leaves a stale entry; no reaper cleans it promptly. The + entry is inert (a reused container ID gets a fresh `record`). A + `docker/system/events` subscription for prompt cleanup is deferred + (OQ-050). +- **`owned_only` on `list` is the caller's choice.** Default + `false` (returns all); `true` filters to the caller's containers. + The default and the two cases are decided in ADR-060 §2; the scope + check gates regardless of the flag. + +## Design Decisions + +| Decision | ADR | Summary | +|----------|-----|---------| +| Docker ops on `alknet/call` | [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md) | No separate ALPN; raw-carriage dissolved; `tty:true` exec rejected | +| Container resource model + labels | [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) | `alknet.managed`/`alknet.owner` labels; `list` `owned_only`; hosted-services fallback; teardown coupling | +| Streaming handler for subscriptions | [ADR-049](../../decisions/049-streaming-handler-for-subscriptions.md) | `StreamingHandler` for logs/exec/pull; exit code on final `call.responded` | +| Dynamic resource ownership | [ADR-050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Containers as `AccessControl` resources; `resource_id_path` | +| Operation error schemas | [ADR-023](../../decisions/023-operation-error-schemas.md) | `CONTAINER_NOT_FOUND`, `IMAGE_NOT_FOUND`, `DOCKER_ERROR`, `INVALID_INPUT` | +| Peer-graph routing | [ADR-029](../../decisions/029-peer-graph-routing-model.md) | Head→worker docker management | +| Forwarded-for identity | [ADR-032](../../decisions/032-forwarded-for-identity.md) | End-user identity as metadata in the proxy pattern | +| bollard 0.21 + features | [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) | `time` feature for log timestamps | +| Docker client + OwnershipStore injection | [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md) | Closure capture at registration (not `Capabilities`, not `OperationContext`); matches `from_openapi` pattern | +| Exit code on terminal `call.responded` | [ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md) | `{ "exitCode": N, "terminal": true }` on final `call.responded`; `call.completed` stays empty | + +## Open Questions + +See [open-questions.md](../../open-questions.md) for full details. + +- **OQ-048** (deferred(scope)): Network and volume operation surface. +- **OQ-049** (deferred(scope)): Image build (buildkit) scope. +- **OQ-050** (deferred(scope)): Docker system events subscription. +- **OQ-051** (deferred(scope)): Container create options surface. + +## References + +- [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md) — the + ALPN decision (shared `alknet/call`, no `carriage` field) +- [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) + — the resource model and label namespace this surface applies +- [ADR-049](../../decisions/049-streaming-handler-for-subscriptions.md) + — `StreamingHandler` for logs/exec/pull +- [ADR-050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) + — the ownership model +- [ADR-023](../../decisions/023-operation-error-schemas.md) — declared + domain error codes +- [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md) + — the Docker client + OwnershipStore injection model (closure + capture, not `Capabilities`) +- [ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md) + — the exec exit-code frame shape (`terminal: true`) +- `docs/research/alknet-docker/poc-summary.md` — the POC (validated + logs subscription, exec with exit code; lifecycle ops are "the + boring case") +- `/workspace/alknet-docker-poc/src/ops.rs` — `DockerOps`: + `drive_logs`, `drive_exec` (the streaming-handler reference) +- `/workspace/bollard/src/container.rs` (`list_containers` :245, + `create_container` :296, `inspect_container` :777, `logs` :928), + `/workspace/bollard/src/exec.rs` (`create_exec` :172, `start_exec` + :225, `inspect_exec` :315), `/workspace/bollard/src/image.rs` + (`create_image` :120, `list_images` :66, `inspect_image` :177) \ No newline at end of file diff --git a/docs/architecture/crates/docker/docker-tty-backend.md b/docs/architecture/crates/docker/docker-tty-backend.md new file mode 100644 index 0000000..96d86ee --- /dev/null +++ b/docs/architecture/crates/docker/docker-tty-backend.md @@ -0,0 +1,482 @@ +--- +status: draft +last_updated: 2026-07-08 +--- + +# alknet-docker — DockerTtyBackend + +The `DockerTtyBackend` is alknet-docker's `impl TtyBackend` (ADR-053) +for interactive terminal sessions into docker containers over +`alknet/tty`. It wraps `bollard::attach_container()` (attach mode) and +`bollard::exec::start_exec` with `tty: true` (exec mode), producing a +`TtyHandle` the `TtyAdapter` pumps. This document specifies the +backend's allocation, control, and cancel-cleanup. The crate +placement is decided in +[ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md). + +## What + +`DockerTtyBackend` is the `TtyBackend` implementation the +`TtyAdapter` dispatches to when a client opens an `alknet/tty` session +with `backend: "docker"` in the negotiation frame. The adapter selects +the backend by the `backend` field, calls `allocate()`, and pumps the +resulting `TtyHandle`'s fields bidirectionally using the alknet-tty +wire format (ADR-052). The backend produces handles; the adapter owns +the wire format. + +```rust +// behind the `tty` feature in alknet-docker +pub struct DockerTtyBackend { + docker: bollard::Docker, + label_prefix: String, // "alknet" by default; for the ownership cross-check +} + +#[async_trait] +impl TtyBackend for DockerTtyBackend { + async fn allocate(&self, params: &TtyParams) -> Result; + fn resource_id(&self, params: &TtyParams) -> Option<(&'static str, String)>; +} +``` + +The backend is constructed by the assembly layer with the bollard +client (shared with the call operations) and the label prefix (shared +with ADR-060's label namespace), then registered in the +`TtyAdapter`'s backend map under the `"docker"` key. + +## Why + +The alknet-tty spec ([tty-backend.md](../tty/tty-backend.md) §"Backend +implementations") listed the `DockerTtyBackend` location as "future, +out of scope here." This spec fills that row: the backend lives in +alknet-docker (ADR-061), behind the `tty` feature, and wraps the +bollard attach/exec API the POC validated. + +The POC's `drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`) +is the seed of `allocate()`. The POC proved the mechanism (raw chunks +on a bidi stream after a JSON request); the `TtyBackend` trait +extracts the bollard interaction into a backend that produces handles, +while the `TtyAdapter` (in alknet-tty) owns the wire format and +session lifecycle. The backend is the inversion point +([ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md)). + +## Architecture + +### Backend Params + +The `DockerTtyBackend` deserializes its params from +`TtyParams.backend_params` (the opaque `serde_json::Map` the adapter +passes through verbatim, per ADR-053 §"Backend params are opaque"): + +```rust +#[derive(Deserialize)] +struct DockerBackendParams { + /// The container to attach to or exec in. + container: String, + /// "attach" (attach to the running primary process) or "exec" + /// (start a new process with a PTY). Default: "exec". + #[serde(default = "default_mode")] + mode: String, +} + +fn default_mode() -> String { "exec".into() } +``` + +- `container` — the container ID or name. Required. The backend + validates it exists (the bollard call will fail if not; the error + maps to `TtyError::Backend`). +- `mode` — `"attach"` or `"exec"`. `"attach"` attaches to the + container's primary process (the running process, pid 1 or the + `ENTRYPOINT`); `"exec"` starts a new process in the container with a + PTY. Default `"exec"` (the common case — a new shell). For `exec` + mode, `TtyParams.cmd` is the command vector; for `attach` mode, + `cmd` is ignored (the primary process is already running). + +The adapter does not interpret these fields — it passes the +negotiation frame's backend-specific JSON through as +`backend_params`, and the backend deserializes its own +strongly-typed struct (ADR-053). + +### `allocate()` — attach mode + +Wraps `bollard::container::attach_container()` (the reliable +HTTP-upgrade-to-TCP path, per the POC and +[ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) +§3 — not the websocket path). + +```rust +async fn allocate_attach(&self, params: &TtyParams, container: &str) -> Result { + let options = AttachContainerOptionsBuilder::default() + .stdout(true).stderr(true).stdin(true) + .stream(true).logs(false) + .build(); + let AttachContainerResults { output, input } = self.docker + .attach_container(container, Some(options)) + .await + .map_err(|e| TtyError::Backend { message: e.to_string() })?; + + // output: Stream → TtyHandle.stdout (Stream) + // PTY mode (tty: true on the container) merges stdout/stderr; + // bollard's LogOutput on a TTY returns StdOut only, so stderr + // is None. (tty-backend.md §DockerTtyBackend notes this.) + let stdout = Box::pin(output.map(|r| match r { + Ok(LogOutput::StdOut { message }) | Ok(LogOutput::Console { message }) => + message, // Bytes + Ok(_) => bytes::Bytes::new(), + Err(_) => bytes::Bytes::new(), + })); + + // input: AsyncWrite → TtyHandle.stdin (Box) + let stdin: Box = Box::new(input); + + // exit_code: wait for the container's process to exit, then read + // the exit status from inspect_container. This future resolves + // independently of the output stream — the adapter (ADR-055) + // awaits BOTH the stdout pump's EOF AND this future's resolve + // before sending the exit chunk, so the backend does not need to + // drain stdout itself. For a container that exits with a non-zero + // status, the State.exit_code carries it. The Drop of this future + // (cancel) kills the container (ADR-056) — see "Cancel-Cleanup". + let docker = self.docker.clone(); + let container_for_exit = container.to_string(); + let exit_code: BoxFuture<'static, Result> = Box::pin(async move { + // Poll the container until it exits. wait_container returns a + // stream of WaitContainerResponse; the first (and usually + // only) one carries the exit code. Alternatively, poll + // inspect_container until State.running is false. The + // adapter's stdout pump drains the output stream in + // parallel; this future resolves when the process exits. + let mut wait_stream = docker.wait_container(&container_for_exit, + None::>); + if let Some(result) = wait_stream.next().await { + match result { + Ok(resp) => Ok(resp.status_code.unwrap_or(-1) as i32), + Err(e) => Err(TtyError::WaitFailed { message: e.to_string() }), + } + } else { + Ok(-1) + } + }); + + Ok(TtyHandle { + stdin, + stdout, + stderr: None, // PTY mode merges + exit_code, + control: Some(TtyControlHandle::new(Arc::new(DockerControl { + docker: self.docker.clone(), + container: container.to_string(), + mode: "attach".into(), + }))), + }) +} +``` + +The output stream maps `LogOutput` → `Bytes` (the adapter pumps these +as stdout chunks, stream_type 1). PTY mode (`tty: true` on the +container) merges stdout/stderr into `StdOut`, so `TtyHandle.stderr` is +`None` — the adapter pumps only stdout chunks. This matches the +`tty-backend.md` sketch: "stdout/stderr are merged when `tty: true` +(bollard's `LogOutput` on a TTY exec returns `StdOut` only), so +`TtyHandle.stderr` is `None` for the PTY case." + +### `allocate()` — exec mode + +Wraps `bollard::exec::create_exec` + `start_exec` with `tty: true`: + +```rust +async fn allocate_exec(&self, params: &TtyParams, container: &str) -> Result { + let cmd = ¶ms.cmd; // argv[0] + args + let config = CreateExecOptions { + cmd: Some(cmd.clone()), + env: Some(params.env.clone()), + attach_stdout: Some(true), + attach_stderr: Some(true), + attach_stdin: Some(true), + tty: Some(true), // PTY mode + ..Default::default() + }; + let exec = self.docker.create_exec(container, config).await + .map_err(|e| TtyError::Backend { message: e.to_string() })?; + let exec_id = exec.id; + + let StartExecResults::Attached { output, input } = self.docker + .start_exec(&exec_id, None).await + .map_err(|e| TtyError::Backend { message: e.to_string() })?; + + // Same mapping as attach mode: output → stdout, input → stdin, + // stderr None (PTY mode merges). + let stdout = /* output → Stream */; + let stdin: Box = Box::new(input); + + // exit_code: after output stream ends, inspect_exec for the exit code + let docker = self.docker.clone(); + let exec_id = exec_id.clone(); + let exit_code: BoxFuture<'static, Result> = Box::pin(async move { + // Wait for output stream end (the future resolves after the + // pump drains output — the adapter's exit-chunk ordering + // awaits both stdout EOF and exit_code resolve, ADR-055). + // Then inspect_exec: + let inspect = docker.inspect_exec(&exec_id).await + .map_err(|e| TtyError::WaitFailed { message: e.to_string() })?; + Ok(inspect.exit_code.unwrap_or(-1)) + }); + + Ok(TtyHandle { stdin, stdout, stderr: None, exit_code, + control: Some(TtyControlHandle::new(Arc::new(DockerControl { + docker: self.docker.clone(), + container: container.to_string(), + mode: "exec".into(), + exec_id: Some(exec_id), + }))) }) +} +``` + +The exit-code future for exec mode follows the POC's pattern (POC +target 3): pump the output stream, then `inspect_exec` for the exit +code. The adapter's exit-chunk ordering (ADR-055) awaits both the +stdout pump's EOF and the `exit_code` future's resolve before sending +the exit chunk — so the `exit_code` future must not resolve until the +output stream has ended. For exec mode, the output stream end is the +signal that the exec process exited; `inspect_exec` after that returns +the exit code. + +### `resource_id()` — ownership check delegation + +```rust +fn resource_id(&self, params: &TtyParams) -> Option<(&'static str, String)> { + let p: DockerBackendParams = serde_json::from_value( + serde_json::Value::Object(params.backend_params.clone()) + ).ok()?; + Some(("container", p.container)) +} +``` + +The adapter calls this at negotiation and checks +`OwnershipProvider::owns(identity, "container", container_id, "tty")` +if an ownership provider is wired (ADR-050, ADR-060). The backend owns +the extraction; the adapter doesn't parse docker-specific JSON. The +`"tty"` action is a distinct resource action from `"exec"` (the call +operation's action) — a caller authorized to `docker/container/exec` +(captured, non-interactive) is not automatically authorized for an +interactive tty session; the `tty` action is a separate scope. This +matches the "interactive terminal ≠ captured command" boundary +(ADR-058 §3). + +### `DockerControl` — `TtyControl` impl + +```rust +struct DockerControl { + docker: bollard::Docker, + container: String, + mode: String, // "attach" or "exec" + exec_id: Option, // Some for exec mode (for resize_exec) +} + +impl TtyControl for DockerControl { + fn resize(&self, cols: u16, rows: u16, pixel_width: u16, pixel_height: u16) { + let docker = self.docker.clone(); + let container = self.container.clone(); + let exec_id = self.exec_id.clone(); + let mode = self.mode.clone(); + tokio::spawn(async move { + let _ = match mode.as_str() { + "attach" => docker.resize_container_tty(&container, + ResizeContainerTtyOptions { width: cols, height: rows, .. }).await, + "exec" if exec_id.is_some() => docker.resize_exec( + exec_id.as_ref().unwrap(), + ResizeExecOptions { height: rows, width: cols, .. }).await, + _ => Ok(()), + }; + }); + } + + fn signal(&self, name: &str) { + // docker has no per-exec signal API in bollard's stable surface. + // Best-effort: kill_container with the mapped signal. + // Unknown signals fall back to SIGKILL (TtyControl::signal contract). + let docker = self.docker.clone(); + let container = self.container.clone(); + let sig = signal_name_to_bollard(name); // SIGINT, SIGTERM, etc. + tokio::spawn(async move { + let _ = docker.kill_container(&container, + Some(KillContainerOptions { signal: sig })).await; + }); + } +} +``` + +- `resize()` — `bollard::container::resize_container_tty()` (attach + mode) or `bollard::exec::resize_exec()` (exec mode with an exec ID). + Fire-and-forget (the `TtyControl` trait methods are sync, not async — + ADR-053; the bollard call is spawned). +- `signal()` — docker has no per-exec signal API in bollard's stable + surface. The best-effort path is `bollard::kill_container()` with + the mapped signal (SIGINT, SIGTERM, etc.); unknown signals fall back + to SIGKILL. For exec mode, the signal goes to the container's main + process, not just the exec's process — docker's exec isolation + limits per-exec signal targeting. This is within the + `TtyControl::signal` contract ("best-effort delivery to the + foreground process group") but is a docker-semantics limitation: a + user expecting `Ctrl-C` to kill only the exec process may see the + container killed. This is documented in the spec, not hidden. + +The signal-name → bollard-signal mapping (`signal_name_to_bollard`) +covers the standard set (`HUP`, `INT`, `QUIT`, `TERM`, `KILL`, `USR1`, +`USR2`, `TSTP`, `CONT`) — the same set alknet-tty's control channel +supports (ADR-052 §Control Channel). Unknown names fall back to the +backend's default kill (`SIGKILL`). + +### Cancel-Cleanup (ADR-056) + +The `TtyHandle.exit_code` future's `Drop`-on-cancel (the adapter drops +the `TtyHandle` on connection drop, stream reset, or pump-task panic — +ADR-056) MUST kill the session target. The docker backend's +`exit_code` future wraps the kill in its `Drop`: + +```rust +// The exit_code future holds a guard that kills on Drop. +struct DockerExitGuard { + docker: bollard::Docker, + container: String, + mode: String, + killed: bool, +} + +impl Drop for DockerExitGuard { + fn drop(&mut self) { + if !self.killed { + let docker = self.docker.clone(); + let container = self.container.clone(); + tokio::spawn(async move { + let _ = docker.kill_container(&container, + Some(KillContainerOptions { signal: "SIGKILL".into() })).await; + }); + } + } +} +``` + +- **Attach mode** — `kill_container` with `SIGKILL`. The container is + not removed (attach doesn't own the container's lifecycle — the + container may be a hosted service the operator wants to keep); the + kill terminates the process the terminal session was attached to. +- **Exec mode** — `kill_container` with `SIGKILL`, targeting the + container's main process (best-effort, per docker's exec isolation). + The exec instance's process is not separately killable in docker's + model; the container-level kill is the available mechanism. + +The guard's `killed` flag prevents a double-kill if the exit_code +future resolved normally (the process exited) before the Drop. The +spawn-on-Drop pattern (tokio::spawn in the Drop impl) is the standard +way to run async cleanup from a sync `Drop`; the spawned task outlives +the dropping context. + +This satisfies ADR-056's contract: dropping `exit_code` (cancel) kills +the session target. The "session target" for docker is the container +(attach) or the exec's process (exec, best-effort via the container). +A backend that returned an `exit_code` future without a kill-on-Drop +guard would violate the contract and could leave orphaned containers. + +### Backend Registration + +```rust +// assembly layer (behind `tty` feature) +let docker_backend = Arc::new(DockerTtyBackend::new( + docker.clone(), + "alknet".into(), // label prefix (ADR-060) +)) as Arc; + +let mut backends = HashMap::new(); +backends.insert("docker".into(), docker_backend); +// (other backends: "local" from alknet-tty-local, "ssh" from alknet-ssh, etc.) +let tty_adapter = TtyAdapter::new(Arc::new(backends), ownership_provider); +``` + +The backend is registered under the `"docker"` key — the +`backend: "docker"` field in the negotiation frame selects it. A +deployment that doesn't want docker terminals doesn't register it; a +deployment that wants both docker and local terminals registers both. + +## Constraints + +- **The backend produces handles; the adapter owns the wire format.** + The backend does not write chunks to the bidi stream — the + `TtyAdapter` does (ADR-052). The backend's `allocate()` returns a + `TtyHandle`; the adapter pumps `stdin`, `stdout`, `exit_code`, and + dispatches `control` chunks. A backend that wrote to the wire + directly would break the wire-format invariants (the exit-chunk + ordering, ADR-055). +- **PTY mode merges stdout/stderr.** `TtyHandle.stderr` is `None` for + both attach and exec modes (both set `tty: true`). The adapter pumps + only stdout chunks (stream_type 1). This matches the PTY property + (one output stream from the slave) and the + [tty-backend.md](../tty/tty-backend.md) sketch. +- **`exit_code` future's `Drop` kills (ADR-056).** The `DockerExitGuard` + in the `exit_code` future issues `kill_container` on `Drop`-without- + resolve. The guard's `killed` flag prevents double-kill. The + spawned-on-Drop pattern runs the async kill from the sync `Drop`. +- **Signal delivery is best-effort.** docker has no per-exec signal + API in bollard's stable surface. `TtyControl::signal()` maps to + `kill_container` with the container's main process as the target. + Unknown signals fall back to `SIGKILL`. This is within the + `TtyControl::signal` contract (ADR-053) but is a docker-semantics + limitation, documented not hidden. +- **Attach doesn't own the container lifecycle.** The cancel-cleanup + kills the container (terminates the attached process) but does not + remove it. A container that was a hosted service (operator-created) + stays after the terminal session ends — the operator may want it + running. The kill (not remove) is the correct teardown for attach. +- **`resource_id()` delegates to the backend.** The adapter doesn't + parse docker JSON to extract the container ID; the backend's + `resource_id()` does. The adapter checks + `OwnershipProvider::owns(identity, "container", id, "tty")`. The + `"tty"` action is distinct from the call operation's `"exec"` + action — interactive terminal authorization is a separate scope. + +## Design Decisions + +| Decision | ADR | Summary | +|----------|-----|---------| +| DockerTtyBackend in alknet-docker | [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) | Behind `tty` feature; attach/exec mode; POC `drive_attach_raw` as reference | +| TtyBackend trait and TtyHandle | [ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) | The trait this backend implements; backend params opaque; REQ-TTY-01 (backends need not be natively async — bollard is async, so no bridge needed) | +| Wire format | [ADR-052](../../decisions/052-alknet-tty-wire-format-and-two-carriage.md) | The chunk codec + control channel the adapter pumps to/from these handles | +| Exit code on a control chunk | [ADR-055](../../decisions/055-exit-code-on-control-chunk.md) | The adapter awaits `exit_code`, sends the exit chunk; the backend's `exit_code` resolves after output EOF + `inspect_exec` | +| Backend cleanup on session cancel | [ADR-056](../../decisions/056-backend-cleanup-on-session-cancel.md) | `exit_code` future's `Drop` kills the container (attach) or container's main process (exec, best-effort) | +| Container resource model | [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) | `resource_id()` delegation; `"tty"` action distinct from call op's `"exec"` | +| bollard 0.21 + features | [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) | No `websocket` feature — reliable `attach_container()` only | + +## Open Questions + +None. The backend's design is decided in ADR-061. The +signal-delivery path's precision in exec mode is a docker-semantics +limitation (docker has no per-exec signal API in bollard's stable +surface; `kill_container` targets the container's main process, +best-effort per the `TtyControl::signal` contract), not an open +question — the limitation is documented in §"DockerControl" and +acknowledged in ADR-061 §3. + +## References + +- [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) + — the crate placement and attach/exec mode decision +- [ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) — the + trait this backend implements +- [ADR-052](../../decisions/052-alknet-tty-wire-format-and-two-carriage.md) + — the wire format the adapter pumps +- [ADR-055](../../decisions/055-exit-code-on-control-chunk.md) — the + exit-chunk ordering the `exit_code` future feeds into +- [ADR-056](../../decisions/056-backend-cleanup-on-session-cancel.md) + — the cancel-cleanup contract (`exit_code` future's `Drop` kills) +- [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) + — the `resource_id()` delegation and the `"tty"` action +- `docs/research/alknet-docker/poc-summary.md` §"POC Target 1" + (interactive attach — the seed of this backend) +- `/workspace/alknet-docker-poc/src/ops.rs` — `drive_attach_raw` (the + reference for `allocate()`) +- `/workspace/bollard/src/container.rs` (`attach_container` :540, + `AttachContainerResults` :80, `LogOutput` :96, + `resize_container_tty` :687, `kill_container` :1059) +- `/workspace/bollard/src/exec.rs` (`create_exec` :172, `start_exec` + :225, `inspect_exec` :315, `resize_exec` :362, `StartExecResults` :99) +- [tty-backend.md](../tty/tty-backend.md) §"Backend implementations" + (the row this spec fills) \ No newline at end of file diff --git a/docs/architecture/crates/docker/overview.md b/docs/architecture/crates/docker/overview.md new file mode 100644 index 0000000..2ff89a5 --- /dev/null +++ b/docs/architecture/crates/docker/overview.md @@ -0,0 +1,334 @@ +--- +status: draft +last_updated: 2026-07-08 +--- + +# alknet-docker — Overview + +The docker operations crate: a thin, single-host bollard wrapper that +exposes docker container and image operations as call-protocol +operations on `alknet/call`, plus (behind a `tty` feature) a +`DockerTtyBackend` implementing `alknet-tty`'s `TtyBackend` trait. This +document covers the crate's purpose, the two-role design, its +dependency edges, the ALPN decision, the label namespace, the feature +gates, and the assembly-layer wiring. Component details are in the +sibling documents. + +## What + +alknet-docker is the docker integration point for the ALPN-as-service +architecture. It does two things: + +1. **Registers docker operations on the shared `alknet/call` + `OperationRegistry`.** The operations — `docker/container/*` and + `docker/image/*` — are ordinary call-protocol operations with + `OperationSpec`s, `AccessControl`, and service discovery. The + existing `CallAdapter` dispatches them through `invoke()` / + `invoke_streaming()`. There is no `alknet/docker` ALPN and no + `DockerProtocolHandler`. Decided in + [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md). + +2. **Provides `DockerTtyBackend`** (behind the `tty` feature), an + `impl TtyBackend` (ADR-053) for interactive terminal sessions into + containers over `alknet/tty`. This is the backend the alknet-tty + spec listed as "future, out of scope here" + ([tty-backend.md](../tty/tty-backend.md) §"Backend implementations"). + Decided in [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md). + +The two roles share the bollard `Docker` client, the container +identity model, and the label/ownership reasoning (ADR-060), but +diverge at the wire format: call operations use `EventEnvelope` on +`alknet/call`; interactive sessions use the raw chunk codec on +`alknet/tty` (ADR-052). + +## Why + +The overwhelming common use case (per the user's brief) is a hub that +runs as a call-protocol client, connects to worker nodes, and exposes +docker operations on those workers. alknet-docker is what runs on the +worker: it wraps the local docker daemon and registers the operations. +The hub composes them via `from_call` (ADR-017) + peer routing +(ADR-029) — the proxy pattern (ADR-050 §3). For interactive terminals +into the containers (an agent needing a shell), the hub opens an +`alknet/tty` session that the worker's `DockerTtyBackend` serves. + +The two container use cases: + +- **Disposable dev containers** (common by volume) — a coordinator + spawns a container for an implementation agent or an isolated env. + `docker/container/create` records ownership; the coordinator owns + the container; `docker/container/remove` revokes. Interactive + sessions go through `DockerTtyBackend`. The container is + short-lived; "burn it and start over" is the recovery model + (ADR-050 §4b). + +- **Long-running hosted services** (less common, important) — the + production server (`/workspace/system/dev1`) hosts rarely-changing + services (reverse-proxy, postgres, redis, gitea) in docker. These + are created by an operator via `docker compose`, not via alknet. + alknet-docker's operations manage them (start/stop/inspect/logs); + the operator role reaches them via the static-resource fallback + (ADR-060 §3) — no ownership-store entry needed for pre-existing + containers an operator is statically authorized to manage. + +Both use cases work through one `AccessControl` model (ADR-050) and +one operation surface. No special-casing, no "is this a managed +container?" branch. + +## The Two Roles in Brief + +### Role 1: Call operations on `alknet/call` + +| Operation | Call op type | bollard method | Carriage | +|-----------|-------------|---------------|----------| +| `docker/container/list` | Query | `list_containers` | JSON | +| `docker/container/inspect` | Query | `inspect_container` | JSON | +| `docker/container/create` | Mutation | `create_container` | JSON | +| `docker/container/start` | Mutation | `start_container` | JSON | +| `docker/container/stop` | Mutation | `stop_container` | JSON | +| `docker/container/remove` | Mutation | `remove_container` | JSON | +| `docker/container/restart` | Mutation | `restart_container` | JSON | +| `docker/container/logs` | Subscription | `logs` | JSON (StreamingHandler) | +| `docker/container/exec` (tty:false) | Subscription | `create_exec` + `start_exec` + `inspect_exec` | JSON (StreamingHandler) | +| `docker/image/list` | Query | `list_images` | JSON | +| `docker/image/pull` | Subscription | `create_image` | JSON (StreamingHandler) | +| `docker/image/inspect` | Query | `inspect_image` | JSON | + +Interactive exec (`tty: true`) and interactive attach are **not** call +operations — they are `alknet/tty` sessions via `DockerTtyBackend`. +A `docker/container/exec` call with `tty: true` in the input returns +`INVALID_INPUT` directing the caller to `alknet/tty`. The full surface, +access control, and the streaming shapes are in +[docker-operations.md](docker-operations.md). + +### Role 2: `DockerTtyBackend` on `alknet/tty` + +```rust +// behind the `tty` feature +pub struct DockerTtyBackend { + docker: Docker, + label_prefix: String, // "alknet" by default; configurable +} + +#[async_trait] +impl TtyBackend for DockerTtyBackend { + async fn allocate(&self, params: &TtyParams) -> Result; + fn resource_id(&self, params: &TtyParams) -> Option<(&'static str, String)>; +} +``` + +The backend wraps `bollard::attach_container()` (attach mode) or +`bollard::exec::start_exec` with `tty: true` (exec mode), producing a +`TtyHandle` the `TtyAdapter` pumps. The wire format, the session +lifecycle, and the exit-chunk ordering live in alknet-tty; the backend +produces bollard-backed handles. See +[docker-tty-backend.md](docker-tty-backend.md). + +## Dependencies + +``` +alknet-docker +├── alknet-core (Identity, AccessControl, OwnershipProvider/Store — ADR-050) +├── alknet-call (OperationSpec, Handler, StreamingHandler, OperationRegistryBuilder) +├── bollard 0.21 (http + pipe + time features; the docker API — ADR-059) +├── alknet-tty (TtyBackend trait — only behind `tty` feature; ADR-061) +├── tokio (async runtime — bollard is tokio-based) +├── serde / serde_json (operation input/output, label values) +└── thiserror (DockerError enum) +``` + +The `alknet-call` edge is the protocol-foundation exception +(ADR-003 Am. 1): alknet-docker consumes `OperationSpec`, `Handler`, +`StreamingHandler`, and `OperationRegistryBuilder` from alknet-call +to register its operations. This is the same edge `alknet-http` has +(HTTP uses the call protocol types). alknet-docker does **not** depend +on `alknet-http`, `alknet-tty-local`, or any other handler crate — the +no-handler-depends-on-another-handler rule (ADR-003) is preserved. + +The `alknet-tty` edge is **only** behind the `tty` feature. A +deployment using alknet-docker for call operations only (the common +case) does not pull in alknet-tty. See Feature Gates below. + +### bollard version and features + +bollard 0.21, verified as the latest published version on crates.io +(as of 2026-07-08). The POC's local checkout (0.21.0) matches the +published version — identical API surface. Features: `http` + `pipe` +(default, local daemon connect) + `time` (log timestamp typing). No +`ssl` (no remote daemon over TLS — fleet is call-protocol), no `ssh` +(the SSH-tunnel fleet model is what alknet replaces), no `websocket` +(reliable attach only, per the POC), no `buildkit` (deferred, OQ-049). +See [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md). + +## ALPN + +| ALPN | Role | Handler | Transport | +|------|------|---------|-----------| +| `alknet/call` | Call operations | `CallAdapter` (shared, not docker-specific) | QUIC bidi stream, `EventEnvelope` framing | +| `alknet/tty` | Interactive terminal sessions | `TtyAdapter` (shared, not docker-specific) | QUIC bidi stream, raw chunk codec (ADR-052) | + +alknet-docker registers **no ALPN of its own**. The call operations +are dispatched by the shared `CallAdapter` on `alknet/call`; the +interactive sessions are dispatched by the shared `TtyAdapter` on +`alknet/tty`, using `DockerTtyBackend` as one of potentially several +registered backends. This is the ADR-058 decision: docker ops are +call-protocol operations, and the one operation that needed raw +carriage moved to alknet-tty. + +## Label Namespace + +alknet-docker applies two labels to containers it creates (ADR-060): + +| Label | Value | Purpose | +|-------|-------|---------| +| `alknet.managed` | `"true"` | Marks the container as alknet-managed; the `list` owned_only filter and the ownership cross-check key on this. | +| `alknet.owner` | `` | The `Identity.id` of the spawner (the coordinator's identity, not the end user's — proxy pattern). | + +The prefix (`alknet`) is configurable at assembly-layer wiring +(two-way-door). The label schema (two labels, +`.managed` + `.owner`) is the one-way commitment. +Containers created by `docker compose` or `docker run` (the +hosted-services case) have no alknet labels and no ownership-store +entry; they're reached via the static-resource operator fallback. + +Full details: [docker-operations.md](docker-operations.md) §"Label +Namespace" and +[ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md). + +## Feature Gates + +```toml +# alknet-docker Cargo.toml +[features] +default = ["ops"] +ops = [] # docker/container/* and docker/image/* call operations +tty = ["dep:alknet-tty"] # DockerTtyBackend (impl TtyBackend) +``` + +- `default = ["ops"]` — the call operations. A hub or coordinator + wiring docker management over the call protocol uses the default + features. Pulls in bollard + alknet-call + alknet-core; no alknet-tty. +- `tty` — adds `DockerTtyBackend`, pulling in `alknet-tty` for the + `TtyBackend` trait. A deployment that wants interactive terminal + sessions into containers enables this and registers + `DockerTtyBackend` with the `TtyAdapter`. A call-operations-only + deployment leaves this off and doesn't pull in alknet-tty. + +The `tty` feature mirrors `alknet-tty`'s `local` feature pattern +(ADR-054): the heavy backend edge is opt-in, so a deployment that +doesn't want it doesn't pay for it. See +[ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md). + +## Assembly Layer Wiring + +The assembly layer (the CLI binary or a hub/worker binary) constructs +the bollard client, the ownership store, and the label config, then +registers the docker operations and (optionally) the tty backend: + +```rust +// 1. Construct the bollard client (local daemon, ADR-059 §2) +let docker = bollard::Docker::connect_with_local_defaults()?; + +// 2. Construct the ownership store (in-memory default, ADR-050 §1) +let ownership_store = Arc::new(InMemoryOwnershipStore::new()); + +// 3. Register docker call operations on the shared registry +let mut builder = OperationRegistryBuilder::new(); +// ... other operations (services/list, agent ops, etc.) ... +register_docker_ops( + &mut builder, + docker.clone(), + ownership_store.clone(), + &DockerLabels { prefix: "alknet".into() }, + CompositionAuthority::new("docker-ops", ["container:list", "container:exec", /* ... */]), +); + +// 4. (optional, behind `tty` feature) Register DockerTtyBackend +#[cfg(feature = "tty")] +{ + let docker_backend = Arc::new(DockerTtyBackend::new( + docker.clone(), + "alknet".into(), + )) as Arc; + // Insert into the TtyAdapter's backend map under "docker" + backends.insert("docker".into(), docker_backend); +} + +// 5. Build the registry and start the endpoint +let registry = builder.build(); +let call_adapter = CallAdapter::new(Arc::new(registry), identity_provider); +// ... register handlers on the endpoint, start ... +``` + +The `register_docker_ops` function (or a `DockerOps` struct the +builder consumes) adds each `docker/container/*` and `docker/image/*` +operation as a `HandlerRegistration` with the right `HandlerKind` +(`Once` for Query/Mutation, `Stream` for Subscription — ADR-049) and +`AccessControl` (per ADR-060). The assembly layer provides the +`CompositionAuthority` for the docker-ops handler (the scopes the +docker ops run under when composed) and the `Capabilities` (empty — +`Capabilities::new()` for local bollard; no outbound credentials +needed). The `Docker` client and `OwnershipStore` are +**closure-captured** into each handler at registration time (ADR-062) +— not read from `OperationContext` and not put in `Capabilities` +(`Capabilities` is for secret material only, per ADR-014). + +The `OwnershipStore` is shared between the docker ops (which `record` +on create and `revoke` on remove) and the `AccessControl::check` path +(which consults the `OwnershipProvider` read side). The in-memory +default is sufficient for the single-host case; a persistence adapter +(ADR-035 shape) is built when a hub wants fleet ownership to survive +restarts. + +## Architecture (component pointers) + +- **[docker-operations.md](docker-operations.md)** — the operation + surface: lifecycle (Query/Mutation), logs/exec/pull (Subscription + via `StreamingHandler`), access control (ADR-050/050 application), + label namespace, teardown coupling, the exit-code-on-final- + `call.responded` pattern for exec. +- **[docker-tty-backend.md](docker-tty-backend.md)** — the + `DockerTtyBackend`: attach vs exec mode, `TtyHandle` field mapping + (PTY mode merges stderr), `TtyControl` → bollard resize/signal, + `exit_code` future `Drop`-kill (ADR-056), `resource_id()` delegation. + +## Design Decisions + +| Decision | ADR | Summary | +|----------|-----|---------| +| Docker ops on `alknet/call` (no separate ALPN) | [ADR-058](../../decisions/058-alknet-docker-on-alknet-call.md) | Call-protocol operations; raw-carriage dissolved by alknet-tty; no `carriage` field | +| bollard 0.21 + feature selection | [ADR-059](../../decisions/059-bollard-021-dependency-and-features.md) | Version verified current; `http`+`pipe`+`time`; no `ssl`/`ssh`/`websocket`/`buildkit` | +| Container resource model + label namespace | [ADR-060](../../decisions/060-container-resource-model-and-label-namespace.md) | ADR-050 application; `alknet.managed`/`alknet.owner` labels; `list` owned_only; hosted-services static fallback | +| DockerTtyBackend in alknet-docker | [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md) | Behind `tty` feature; attach/exec mode; POC `drive_attach_raw` as reference | +| Crate decomposition | [ADR-003](../../decisions/003-crate-decomposition.md) Am. 1 | alknet-docker depends on alknet-call (protocol-foundation exception) + alknet-tty (tty feature) | +| Call protocol stream model | [ADR-012](../../decisions/012-call-protocol-stream-model.md) | `EventEnvelope` framing, no `carriage` extension | +| Streaming handler for subscriptions | [ADR-049](../../decisions/049-streaming-handler-for-subscriptions.md) | `StreamingHandler` for logs/exec/pull; exit code on final `call.responded` | +| Dynamic resource ownership | [ADR-050](../../decisions/050-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Containers as `AccessControl` resources; the model ADR-060 applies | +| Peer-graph routing | [ADR-029](../../decisions/029-peer-graph-routing-model.md) | Head→worker docker management via `PeerRef` / `invoke_peer` | +| Forwarded-for identity | [ADR-032](../../decisions/032-forwarded-for-identity.md) | End-user identity as metadata when proxying | +| TtyBackend trait | [ADR-053](../../decisions/053-ttybackend-trait-and-ttyhandle.md) | The trait `DockerTtyBackend` implements | +| Exit code on control chunk | [ADR-055](../../decisions/055-exit-code-on-control-chunk.md) | The exit-chunk ordering `DockerTtyBackend`'s `exit_code` feeds into | +| Backend cleanup on cancel | [ADR-056](../../decisions/056-backend-cleanup-on-session-cancel.md) | `DockerTtyBackend`'s `exit_code` `Drop` kills the container/exec | +| Docker client + OwnershipStore injection | [ADR-062](../../decisions/062-docker-client-injection-via-closure-capture.md) | Closure capture at registration; `Capabilities` for secrets only | +| Exit code on terminal `call.responded` | [ADR-063](../../decisions/063-exit-code-on-terminal-call-responded.md) | `{ "exitCode": N, "terminal": true }` on final `call.responded` for non-interactive exec | + +## Open Questions + +See [open-questions.md](../../open-questions.md) for full details. + +- **OQ-048** (deferred(scope)): Network and volume operation surface. +- **OQ-049** (deferred(scope)): Image build (buildkit) scope. +- **OQ-050** (deferred(scope)): Docker system events subscription. +- **OQ-051** (deferred(scope)): Container create options surface. + +## References + +- `docs/research/alknet-docker/poc-summary.md` — the POC summary this + spec set builds on (validated targets, open unknowns) +- `/workspace/alknet-docker-poc/` — the POC source +- `/workspace/bollard/` — bollard 0.21.0 source (verified current) +- `/workspace/@alkdev/dispatch/` — the dispatch POC (prior art: + `dispatch.managed` labels, SSH-tunnel fleet model) +- `/workspace/system/dev1/docker.md` — the hosted-services use case +- `/workspace/@alkdev/reverse-proxy/deploy/docker-compose.yml` — + operator-created container example \ No newline at end of file diff --git a/docs/architecture/crates/tty/tty-backend.md b/docs/architecture/crates/tty/tty-backend.md index 5f09205..cf87a5d 100644 --- a/docs/architecture/crates/tty/tty-backend.md +++ b/docs/architecture/crates/tty/tty-backend.md @@ -313,7 +313,7 @@ let mut backends = HashMap::new(); backends.insert("local".into(), Arc::new(LocalTtyBackend::new()) as Arc); backends.insert("docker".into(), - Arc::new(DockerTtyBackend::new(docker_client)) as Arc); + Arc::new(DockerTtyBackend::new(docker_client, "alknet".into())) as Arc); backends.insert("ssh".into(), Arc::new(SshTtyBackend::new(ssh_session)) as Arc); let tty_adapter = TtyAdapter::new(Arc::new(backends)); @@ -329,12 +329,13 @@ layer chooses what's available. | Backend | Crate | Status | Notes | |---------|-------|--------|-------| | `LocalTtyBackend` | `alknet-tty-local` (sibling, behind `local` feature) | in scope ([tty-local.md](tty-local.md)) | `portable_pty` (PTY) + `std::process` (pipe); the runner pattern | -| `DockerTtyBackend` | `alknet-docker` or sibling adapter | future, out of scope here | wraps `bollard::attach_container` / `exec` with `tty: true` | +| `DockerTtyBackend` | `alknet-docker` (behind `tty` feature) | specced ([docker-tty-backend.md](../docker/docker-tty-backend.md), [ADR-061](../../decisions/061-docker-tty-backend-in-alknet-docker.md)) | wraps `bollard::attach_container` / `exec` with `tty: true`; attach vs exec mode | | `SshTtyBackend` | `alknet-ssh` | future, out of scope here | wraps russh `pty_request` + `shell_request`/`exec_request`; dissolves alknet-ssh DP-5 PTY hedge | -The docker and SSH backend crates are future work; this spec commits the -trait shape they implement so they can be built against it without -re-spec'ing the seam. The `DockerTtyBackend` is the natural extension of +The SSH backend crate is future work; this spec commits the +trait shape it implements so it can be built against it without +re-spec'ing the seam. The `DockerTtyBackend` is now specced in +[alknet-docker](../docker/docker-tty-backend.md) (ADR-061) — the natural extension of the alknet-docker POC's `drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`) — with the trait, it becomes `impl TtyBackend for DockerTtyBackend`. The `SshTtyBackend` dissolves the alknet-ssh research's PTY hedge (DP-5): diff --git a/docs/architecture/decisions/058-alknet-docker-on-alknet-call.md b/docs/architecture/decisions/058-alknet-docker-on-alknet-call.md new file mode 100644 index 0000000..9890cab --- /dev/null +++ b/docs/architecture/decisions/058-alknet-docker-on-alknet-call.md @@ -0,0 +1,267 @@ +# ADR-058: alknet-docker Registers on `alknet/call` (No Separate ALPN) + +## Status + +Accepted + +## Context + +The alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`) +left two integration questions open (its "Open Unknowns" #1 and #2): + +1. **Raw-carriage handoff in the dispatcher.** The POC's `drive_attach_raw` + reads the `call.requested` frame itself, then switches the bidi stream + to raw chunks for interactive attach. The POC noted this doesn't fit + the call dispatcher's `handle_stream` → `dispatch()` → + `DispatchResult::Stream(ResponseStream)` path, which pumps a stream + of `EventEnvelope`s (JSON carriage). The POC offered two options: (a) + branch `handle_stream` on a `carriage` field, handing raw streams to a + `RawHandler`; or (b) a separate ALPN (`alknet/docker-raw`) that owns + the whole stream. + +2. **ALPN layout.** Should docker ops register on the shared + `alknet/call` ALPN (operations in the shared `OperationRegistry`) or + get their own `alknet/docker` ALPN (as a `ProtocolHandler`)? The POC + leaned shared but didn't decide. + +Both questions were open **before** alknet-tty was extracted. The +alknet-tty crate spec (ADR-052, ADR-053) resolves both by moving the one +operation that needed raw carriage — interactive attach/exec with a PTY +— out of the call protocol entirely and into its own ALPN +(`alknet/tty`), backed by the `TtyBackend` trait. + +### Why the raw-carriage problem dissolved + +The POC's raw-carriage handoff was hard because it tried to do two +different things on one `alknet/call` stream: + +- **Structured operations** (lifecycle, logs, inspect) — naturally + JSON-shaped, fit `call.responded`/`call.completed` exactly. This is + what the call protocol is for. +- **Interactive attach** — a bidirectional byte pump (stdin/stdout/stderr) + with a control sideband (resize, signal, exit). This is *not* a + request/response or subscription; it's a terminal session. Forcing it + through `EventEnvelope` framing is "wasteful and lossy" (POC §"Why not + JSON for everything?"). + +alknet-tty was created precisely to extract the second category into its +own protocol (`alknet/tty`) with its own wire format (ADR-052). The +docker POC's `drive_attach_raw` became the seed of alknet-tty's +`DockerTtyBackend` (ADR-061). What remains in alknet-docker's call +operations is the first category — structured operations that map cleanly +to `call.requested`/`call.responded`/`call.completed`. + +The POC's concern ("the dispatcher would need a `carriage` field and a +`RawHandler` branch") is no longer operative: there is no raw-carriage +operation on `alknet/call`. The raw byte pump moved to `alknet/tty`; the +`DockerTtyBackend` implements `TtyBackend` and is reached through the +`TtyAdapter`, not the `CallAdapter`. See ADR-061 for the backend +placement. + +### Why shared `alknet/call`, not a separate `alknet/docker` ALPN + +The remaining docker operations (lifecycle, logs, exec-with-exit-code, +inspect, list, images) are ordinary call-protocol operations. They have: + +- A structured input (JSON), a structured output (JSON), or a stream of + structured events (logs, image pull progress, exec output). +- Natural `OperationSpec`s with input/output JSON Schemas. +- `AccessControl` declarations against the ADR-050 container-as-resource + model (`resource_type: "container"`, `resource_id_path: "$.containerId"`). +- Service discovery through `services/list` / `services/schema`. + +Putting them on a separate `alknet/docker` ALPN would mean a separate +`ProtocolHandler` that re-implements framing, dispatch, ACL, and service +discovery — or a thin wrapper that delegates to the call protocol's +machinery. Either way, it's a parallel dispatch surface for no benefit. +The shared registry is more composable: docker ops are callable from any +call client, including peer routing (`PeerRef` / `from_call` re-export), +which is the primary use case — a coordinator on the hub composing +`docker/container/exec` on a worker spoke. + +A separate ALPN is warranted when the protocol's wire format is +incompatible with `EventEnvelope` framing (alknet-tty, alknet-ssh). The +docker operations remaining on `alknet/call` are all +`EventEnvelope`-shaped. The one that wasn't (raw attach) moved to +alknet-tty. See ADR-052 §"Why not JSON for everything?" for the boundary +criterion. + +### Logs and exec streaming — still JSON carriage, still `alknet/call` + +Two operations look like they might need raw carriage but don't: + +- **`docker/container/logs`** — a `Subscription` operation. bollard's + `logs()` returns a `Stream`; each `LogOutput` becomes a + `call.responded` carrying `{ "stream": "stdout"|"stderr", "text": "..." }`; + stream end → `call.completed`. This is the `StreamingHandler` shape + (ADR-049). The POC validated this path (`docker_logs_subscription_pumps_frames_and_completes`). + No raw carriage needed — each log line is naturally JSON-shaped. +- **`docker/container/exec`** (non-interactive, no TTY) — a `Subscription` + operation. bollard's `start_exec` returns `StartExecResults::Attached { + output, input }`; the output stream is pumped as `call.responded` + frames; after the stream ends, `inspect_exec()` gives the exit code, + which rides on a final `call.responded` with `{ "exitCode": N, + "terminal": true }` before `call.completed`. The POC validated this + (`docker_exec_streams_output_and_exit_code`). This is the exec path + for `tty: false` — a captured command with separate stdout/stderr and + an exit code, not an interactive terminal. + +The interactive exec path (`tty: true`, bidirectional, resize/signal) is +**not** a call operation — it's a `DockerTtyBackend` session on +`alknet/tty`. ADR-061 covers that. The split is: non-interactive exec is +a call `Subscription`; interactive exec is a tty session. The `tty: +bool` field in the operation input selects the path; the caller picks. + +## Decision + +### 1. alknet-docker registers its operations on the shared `alknet/call` ALPN + +alknet-docker does not register a `ProtocolHandler`. It constructs an +`OperationRegistry` (or a `DockerOps` registration bundle consumed by +the assembly layer's `OperationRegistryBuilder`) and the existing +`CallAdapter` dispatches docker operations through the shared +`OperationRegistry::invoke()` / `invoke_streaming()` paths. There is no +`alknet/docker` ALPN, no `DockerProtocolHandler`, and no parallel +dispatch surface. + +This makes docker operations first-class citizens of the call protocol: +they appear in `services/list`, they have `OperationSpec`s with JSON +Schemas, they go through the standard `AccessControl::check`, and they +compose with `from_call` re-export and peer routing (ADR-029) like any +other operation. + +### 2. There is no raw-carriage operation on `alknet/call` + +The `carriage` field the POC proposed for `call.requested` is not added. +The call protocol's `EventEnvelope` framing is the only carriage on +`alknet/call`. Interactive attach/exec (the one operation that needed +raw carriage) is reached through `alknet/tty` via `DockerTtyBackend` +(ADR-061), not through a call operation. The POC's open question #1 +(raw-carriage handoff in the dispatcher) is resolved by removing the +requirement, not by adding a branch. + +### 3. The operation taxonomy + +| Operation | Call op type | bollard method | Carriage | +|-----------|-------------|---------------|----------| +| `docker/container/list` | Query | `list_containers` | JSON | +| `docker/container/inspect` | Query | `inspect_container` | JSON | +| `docker/container/create` | Mutation | `create_container` | JSON | +| `docker/container/start` | Mutation | `start_container` | JSON | +| `docker/container/stop` | Mutation | `stop_container` | JSON | +| `docker/container/remove` | Mutation | `remove_container` | JSON | +| `docker/container/restart` | Mutation | `restart_container` | JSON | +| `docker/container/logs` | Subscription | `logs` | JSON (StreamingHandler) | +| `docker/container/exec` (tty:false) | Subscription | `create_exec` + `start_exec` + `inspect_exec` | JSON (StreamingHandler) | +| `docker/image/list` | Query | `list_images` | JSON | +| `docker/image/pull` | Subscription | `create_image` | JSON (StreamingHandler) | +| `docker/image/inspect` | Query | `inspect_image` | JSON | + +Interactive exec (`tty: true`) and interactive attach are **not** in this +table — they are `alknet/tty` sessions via `DockerTtyBackend` +(ADR-061). A `docker/container/exec` call operation with `tty: true` in +the input returns an `call.error` with code `INVALID_INPUT` directing +the caller to use `alknet/tty`; the operation only accepts `tty: false` +(or absent). This keeps the one-operation-per-carriage invariant clean: +the call operation is captured output; the tty session is interactive. + +The full operation surface (including which are in scope for v1 and +which are deferred) is in +[docker-operations.md](../crates/docker/docker-operations.md). + +### 4. The assembly layer composes the registry + +alknet-docker exports a `DockerOps` construct (a set of +`HandlerRegistration` bundles keyed by operation name) or a +`register_docker_ops(&mut builder, docker_client, ownership, labels)` +function. The assembly layer (the CLI binary or a hub binary) calls this +to add docker operations to the shared `OperationRegistryBuilder` +alongside other operations (services discovery, agent ops, etc.). The +`Docker` bollard client, the `OwnershipStore`, and the label namespace +config are injected by the assembly layer. See +[overview.md](../crates/docker/overview.md) §"Assembly Layer Wiring". + +## Consequences + +**Positive:** + +- Docker operations inherit all call-protocol machinery for free: + service discovery, JSON Schema validation, `AccessControl`, + `from_call` re-export, peer routing, `forwarded_for` metadata, abort + cascade. No parallel dispatch surface. +- The POC's hardest open question (raw-carriage handoff) is resolved by + removal — the operation that needed it moved to alknet-tty. No + dispatcher change, no `carriage` field, no `RawHandler` trait. +- The operation taxonomy is clean: `Query`/`Mutation` for lifecycle, + `Subscription` (ADR-049 `StreamingHandler`) for logs/exec/pull. The + exit-code-on-final-`call.responded` pattern the POC validated (POC + target 3) works through the existing `StreamingHandler` → `pump_stream` + path with no protocol change. +- A hub that wants to manage docker on worker spokes composes + `docker/container/*` through `from_call` (ADR-017) + peer routing + (ADR-029) — the proxy pattern from ADR-050. This is the primary use + case and it works by construction. + +**Negative:** + +- The `docker/container/exec` operation has a `tty` field whose value + constrains the dispatch path: `tty: false` → call `Subscription`; + `tty: true` → reject, point to `alknet/tty`. This is a minor impedance + — a caller wanting interactive exec must use a different protocol + (`alknet/tty`) than a caller wanting captured exec. This is the + correct split (interactive terminal ≠ captured command), but it means + the "exec" concept spans two protocols. The `DockerTtyBackend` + (ADR-061) and the `docker/container/exec` call operation share the + underlying `bollard::exec` API but diverge at the wire format. This is + the same divergence as "SSH exec" (call op) vs "SSH PTY session" + (alknet-tty) — the boundary is principled, not accidental. +- A deployment that wants docker operations must wire the + `OperationRegistry` (the assembly layer's job). This is not a downside + — it's the same wiring every other operation set requires — but it + means alknet-docker is not a "drop-in `ProtocolHandler`" the way + alknet-tty is. The trade is composability (shared registry, peer + routing) for a slightly more involved assembly step. + +## Door type + +**One-way.** The decision to register on `alknet/call` rather than a +separate `alknet/docker` ALPN is a structural commitment: the operation +names (`docker/container/*`), the `OperationSpec`s, the +`AccessControl` shapes, and the `from_call` re-export surface all depend +on docker ops being call-protocol operations. Reversing this — moving +docker ops to their own ALPN — would be a rewrite of the operation +surface and a break for any client that calls `docker/container/*` +through the call protocol. + +The decision *not* to add a `carriage` field to `call.requested` is +also one-way: the call protocol's wire format (ADR-012) stays +`EventEnvelope`-only on `alknet/call`. If a future operation needs raw +carriage on `alknet/call`, it would require a wire-format change (new +ALPN per ADR-006, or a `call.requested` field addition). The expected +path for raw-carriage operations is a separate ALPN (as alknet-tty did), +not a `call.requested` extension. + +## References + +- `docs/research/alknet-docker/poc-summary.md` §"Open Unknowns" #1 and #2 + (the two questions this ADR resolves) +- [ADR-012](012-call-protocol-stream-model.md) — the call protocol wire + format this ADR keeps unchanged (no `carriage` field) +- [ADR-017](017-call-protocol-client-and-adapter-contract.md) — `from_call` + re-export, the proxy pattern's mechanism +- [ADR-024](024-operation-registry-layering.md) — the registry layering + docker ops register into +- [ADR-029](029-peer-graph-routing-model.md) — peer routing, the + head→worker docker management path +- [ADR-049](049-streaming-handler-for-subscriptions.md) — + `StreamingHandler`, the dispatch path for logs/exec/pull +- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md) + — containers as resources, the `AccessControl` shape docker ops declare +- [ADR-052](052-alknet-tty-wire-format-and-two-carriage.md) — alknet-tty's + wire format, where the raw-carriage operation moved +- [ADR-053](053-ttybackend-trait-and-ttyhandle.md) — `TtyBackend`, the + trait `DockerTtyBackend` implements (ADR-061) +- [ADR-061](061-docker-tty-backend-in-alknet-docker.md) — + `DockerTtyBackend` placement in alknet-docker +- Spec documents: [overview.md](../crates/docker/overview.md), + [docker-operations.md](../crates/docker/docker-operations.md) \ No newline at end of file diff --git a/docs/architecture/decisions/059-bollard-021-dependency-and-features.md b/docs/architecture/decisions/059-bollard-021-dependency-and-features.md new file mode 100644 index 0000000..dd8fb65 --- /dev/null +++ b/docs/architecture/decisions/059-bollard-021-dependency-and-features.md @@ -0,0 +1,195 @@ +# ADR-059: bollard 0.21 Dependency and Feature Selection + +## Status + +Accepted + +## Context + +The alknet-docker POC (`docs/research/alknet-docker/poc-summary.md`) +depended on a local bollard checkout at `/workspace/bollard` (version +0.21.0). Its "Open Unknowns" #5 noted: "The real crate should depend on +published 0.21 from crates.io (the dispatch POC pinned 0.18 — a +3-version jump). The `websocket` feature is optional; the `http` and +`pipe` features are needed for socket/http connect. Confirm the +published 0.21 has the same API surface as the checkout." + +Two concerns motivate an explicit version check: + +1. **Agent training-data drift.** Implementation agents tend to default + to the bollard version represented in their training data unless + explicitly told to check. The dispatch POC used bollard 0.18 (3 + versions behind); an agent that pattern-matched dispatch would pin + 0.18 and silently use a 3-version-stale API. The cost of checking is + low (one crates.io lookup); the cost of using a stale version is + potentially high (API mismatches, missing methods, security regressions). + +2. **API surface stability across the path-vs-published boundary.** The + POC used a path dependency; the crate uses a published version. A + version number match (0.21.0) should mean the API surface is + identical, but this is worth confirming rather than assuming. + +### Version verification + +A check against crates.io / docs.rs confirmed: + +- **bollard 0.21.0 is the latest published version** on crates.io (as + of 2026-07-08). The local checkout at `/workspace/bollard` (0.21.0) + matches the published version. No version jump is needed; the POC's + dependency is current. +- The published 0.21.0 `Cargo.toml` (verified via docs.rs source) has + the same API surface the POC used: `attach_container` (container.rs:540), + `logs` (container.rs:928), `create_exec`/`start_exec`/`inspect_exec` + (exec.rs:172/225/315), `AttachContainerResults` (container.rs:80), + `StartExecResults` enum (exec.rs:99), `LogOutput` (container.rs:96), + `NewlineLogOutputDecoder` (read.rs:32). The POC's code-to-concept + mappings hold against the published version. +- bollard-stubs is pinned at `=1.53.1-rc.29.3.1` (the Docker API 1.53 + schema), matching the Docker Engine 29.2.1 / API 1.53 the POC tested + against. + +### Feature selection + +bollard 0.21's feature surface (from the published `Cargo.toml`): + +| Feature | What it pulls | Needed for alknet-docker? | +|---------|---------------|--------------------------| +| `default = ["http", "pipe"]` | hyper-util (http), hyperlocal + hyper-named-pipe (pipe) | **Yes** — the default. `pipe` connects to the unix socket (`/var/run/docker.sock`); `http` connects to a TCP/HTTP daemon. Both are the standard local-daemon paths. | +| `ssl` | rustls TLS for remote daemon connections | No — alknet-docker talks to a local daemon; remote daemons are reached over the call protocol (the fleet layer, not bollard's TLS). | +| `ssh` | openssh for SSH-tunneled daemons | No — the dispatch POC's SSH-tunnel model is exactly what alknet removes; the worker dials the hub over the call protocol instead of the hub SSHing into the worker. | +| `websocket` | tokio-tungstenite for the websocket attach endpoint | No — the POC deliberately used the reliable `attach_container()` (HTTP upgrade to TCP), not the websocket path (bollard's own docs warn of RFC 6455 compatibility issues). The websocket feature is for browser-attach scenarios alknet-docker doesn't have. | +| `buildkit` / `buildkit_providerless` | tonic + bollard-buildkit-proto for buildkit image builds | No — image build is deferred (POC "What the POC Does NOT Validate" #5; OQ-049 scopes this). | +| `chrono` / `time` | timestamp parsing for log timestamps | **Optional** — `docker/container/logs` sets `timestamps: true`; the `chrono` or `time` feature types the timestamp field. The choice between them is a two-way-door implementation detail (both are serde-compatible). `time` is the lighter dependency. | +| `json_data_content` | serde for body content types | No — alknet-docker serializes its own JSON shapes from bollard's types; it doesn't need bollard's content-type tagging. | +| `aws-lc-rs` / `webpki` / `test_*` | TLS provider variants and test-only features | No — not for production deps. | + +The POC's `Cargo.toml` depended on bollard with default features (path +dependency). The crate does the same: `bollard = { version = "0.21", +default-features = true }`, with the `time` feature added for log +timestamp typing. No `ssl`, `ssh`, `websocket`, or `buildkit` features. + +## Decision + +### 1. Depend on published bollard 0.21 from crates.io + +```toml +[dependencies] +bollard = { version = "0.21", features = ["time"] } +``` + +`default-features = true` (the default) pulls `http` + `pipe`, the two +local-daemon connect paths. The `time` feature adds timestamp typing +for `docker/container/logs`. No other features are enabled. + +The local-checkout path dependency the POC used is retired. The +published 0.21.0 has the identical API surface (confirmed via docs.rs +source — same version number, same method signatures, same stubs pin). + +### 2. Connect via `connect_with_local_defaults()` (unix socket) + +The standard deployment talks to a local docker daemon over +`/var/run/docker.sock` (Unix) or the named pipe on Windows. +`Docker::connect_with_local_defaults()` handles both. The `Docker` +client is constructed once at assembly-layer startup and injected into +the docker ops registration. See +[overview.md](../crates/docker/overview.md) §"Assembly Layer Wiring". + +A TCP daemon path (`connect_with_http_defaults`) is available if a +deployment runs the docker daemon on a different host — but this is the +exception, not the default. The fleet case (multiple docker daemons on +different machines) is handled at the call-protocol layer (a `CallClient` +per remote daemon, each running alknet-docker locally), not by pointing +one bollard client at a remote daemon over TCP. See ADR-058 §"Why shared +`alknet/call`" and the POC §6 (the normalization crate boundary). + +### 3. Do not enable `websocket`, `ssl`, `ssh`, or `buildkit` + +- **`websocket`** — the reliable `attach_container()` (HTTP upgrade to + TCP) is the attach path for `DockerTtyBackend` (ADR-061). The websocket + attach endpoint has documented reliability issues (bollard + `container.rs:577`) and is for browser-attach scenarios alknet-docker + doesn't have. The raw chunk format (ADR-052) is layered on top of the + reliable attach, not the websocket. +- **`ssl`** — remote daemons are reached over the call protocol (the + fleet layer), not bollard's TLS. alknet-docker is single-host by + construction (POC §6). +- **`ssh`** — the SSH-tunnel model is what the call-protocol fleet + layer replaces. The dispatch POC's `DockerProvider` SSHed into workers; + alknet-docker's workers dial the hub over `alknet/call` and expose + their local docker ops directly. +- **`buildkit`** — image build is deferred (OQ-049). The `build_image` + method is in bollard but the buildkit feature pulls tonic + + bollard-buildkit-proto, a heavy dependency for an out-of-scope feature. + +### 4. `time` feature for log timestamps + +`docker/container/logs` sets `timestamps: true` on the bollard `logs()` +query. Each `LogOutput` then carries a docker timestamp; the +`call.responded` frame separates `timestamp` and `text` into distinct +JSON fields (the POC noted this refinement over its single-`text`-field +POC shape). The `time` feature types the timestamp field as +`time::OffsetDateTime` (serde-compatible). The `chrono` feature is the +alternative; `time` is lighter and sufficient. + +The timestamp-feature choice (`time` vs `chrono`) is a two-way-door +implementation detail within the one-way version-pin decision. Switching +features is a `Cargo.toml` change, not an API break. + +## Consequences + +**Positive:** + +- bollard 0.21 is current (verified against crates.io); no stale-version + risk. The POC's API-surface mappings hold against the published version. +- The feature set is minimal: `http` + `pipe` (default) + `time`. No + `ssl`/`ssh`/`websocket`/`buildkit` weight. The compile-time and + dependency surface stays small. +- The single-host contract is clear: alknet-docker talks to one local + daemon over the unix socket. The fleet case is a call-protocol concern, + not a bollard-feature concern. + +**Negative:** + +- The `websocket` feature being disabled means the browser-attach + scenario (bollard's websocket attach for a browser that speaks docker's + websocket framing directly) is not available. This is intentional — + browsers reach docker through `alknet/tty` (via `DockerTtyBackend`) or + the call protocol, not bollard's websocket endpoint. If a future use + case forces the websocket path, enabling the feature is a `Cargo.toml` + addition (two-way door). +- The `ssh` feature being disabled means bollard's SSH-tunneled daemon + path is not available. This is intentional — the dispatch POC's + SSH-tunnel model is what alknet's call-protocol fleet layer replaces + (POC §6, "Prior art"). Enabling `ssh` would reintroduce the friction + (SSH key injection, port binding) the fleet layer exists to remove. + +## Door type + +**One-way (version pin) + two-way (feature set).** The version pin +(bollard 0.21, not 0.18 or a future 0.22) is a one-way commitment for +the implementation: the operation handlers are written against 0.21's +API surface, and a major-version bump is a migration. The feature set +(`http` + `pipe` + `time`, no `ssl`/`ssh`/`websocket`/`buildkit`) is +two-way-door: features can be added or removed in `Cargo.toml` without +an API break. The decision to *not* enable `ssl`/`ssh`/`buildkit` is a +scope decision (not needed for the current scope), not a capability +closure — the features can be added later if a use case forces them. + +The `time`-vs-`chrono` choice is a two-way-door implementation detail. + +## References + +- `docs/research/alknet-docker/poc-summary.md` §"Open Unknowns" #5 + (bollard version pinning) +- `docs/research/alknet-docker/poc-summary.md` §"What the POC Does NOT + Validate" #5 (image management / buildkit deferred) +- bollard 0.21.0 published `Cargo.toml` (verified via docs.rs source) +- bollard source: `src/container.rs` (`attach_container` :540, `logs` + :928), `src/exec.rs` (`create_exec` :172, `start_exec` :225, + `inspect_exec` :315), `src/read.rs` (`NewlineLogOutputDecoder` :32) +- [ADR-058](058-alknet-docker-on-alknet-call.md) — the single-host / + call-protocol-fleet contract this feature set reflects +- [ADR-061](061-docker-tty-backend-in-alknet-docker.md) — the + `DockerTtyBackend` that uses the reliable (non-websocket) attach path +- OQ-049 (build/image scope, deferred) +- Spec: [overview.md](../crates/docker/overview.md) §"Dependencies" \ No newline at end of file diff --git a/docs/architecture/decisions/060-container-resource-model-and-label-namespace.md b/docs/architecture/decisions/060-container-resource-model-and-label-namespace.md new file mode 100644 index 0000000..e446a45 --- /dev/null +++ b/docs/architecture/decisions/060-container-resource-model-and-label-namespace.md @@ -0,0 +1,338 @@ +# ADR-060: Container Resource Model and Label Namespace + +## Status + +Accepted + +## Context + +ADR-050 established the dynamic resource ownership model for +runtime-spawned resources: a `OwnershipProvider` (read, sync) + +`OwnershipStore` (write, async) trait pair in alknet-core, with the +spawner owning the resource, proxy-to-share, teardown-revoke. ADR-050 +explicitly named alknet-docker as the first consumer: containers as +`AccessControl` resources, `docker/container/exec` requiring +`resource: container/:exec`, `docker/container/create` calling +`OwnershipStore::record` on success, `docker/container/remove` calling +`OwnershipStore::revoke` on teardown. + +ADR-050 resolved the *model*. This ADR resolves the *application* of +that model to alknet-docker's concrete operation surface — the three +things ADR-050 left to the crate spec: + +1. **The label namespace.** The dispatch POC (`/workspace/@alkdev/dispatch`) + used `dispatch.managed=true` to mark containers it created. The POC + summary §"What the POC Does NOT Validate" #6 noted the real crate + "needs a configurable label prefix and ownership mapping + (`alknet.owner=`) tied to the call protocol's identity model." + What is the label scheme, and how does it relate to the ownership store? + +2. **The `list` result-filter.** ADR-050 specific #4a says the `list` + case (resource_type set, resource_id_path absent) is "scope-gate + + result-filter": the scope check gates the call, and the handler + filters the result via `OwnershipProvider::owned_resources`. How + does `docker/container/list` apply this against bollard's + `list_containers` (which returns all containers the daemon sees)? + +3. **Teardown coupling.** ADR-050 specific #4b says teardown is + handler-driven: the docker handler calls `revoke` on container + removal. But docker containers can exit on their own (a process + that finishes, a `--rm` container). How does the ownership store + learn about autonomous container death, or does it rely solely on + explicit `docker/container/remove`? + +### The two use cases and their ownership profiles + +The user's brief named two container use cases: + +- **Disposable dev containers** (the common case by volume) — a + coordinator spawns a container for an implementation agent or an + isolated env. The container is short-lived; the coordinator owns it + for its lifetime; it's removed when the agent is done. Ownership is + per-session, per-coordinator. +- **Long-running hosted services** (less common but important) — the + production server (`/workspace/system/dev1`) hosts rarely-changing + services (reverse-proxy, postgres, redis, gitea) in docker. These + containers are created by an operator (via `docker compose`), not by + a call-protocol coordinator. They're managed via alknet-docker's + operations (start/stop/restart/inspect/logs) but their *ownership* is + not in the alknet ownership store — they predate the connection. + +These two cases have different ownership profiles, and the model must +handle both without forcing the hosted-services case through the +disposable-container path. + +## Decision + +### 1. Label namespace: `alknet.*` prefix, configurable, two labels + +alknet-docker applies two labels to containers it creates: + +| Label | Value | Purpose | +|-------|-------|---------| +| `alknet.managed` | `"true"` | Marks the container as alknet-managed. The `list` filter and the ownership-existence check key on this label. | +| `alknet.owner` | `` | The `Identity.id` (stable logical peer id, per ADR-030) of the spawner. For composing handlers, the *handler's* identity (the coordinator's `Identity`), not the end user's — the proxy pattern (ADR-050 §3). | + +The label prefix (`alknet`) is **configurable** at assembly-layer wiring. +A deployment that runs alongside another alknet-managed fleet (or wants +to avoid collisions with its own `managed` labels) sets a different +prefix. The default is `alknet`. This is a two-way-door config choice; +the label *schema* (two labels, `.managed` + `.owner`) +is the one-way commitment. + +The labels serve two purposes: + +- **Ownership-store cross-check.** When `docker/container/exec` checks + `OwnershipProvider::owns(identity, "container", id, "exec")`, the + ownership store is the source of truth. The label is a *secondary* + signal: if the store says "yes" but the label is absent or names a + different owner, the container's ownership state is stale (the + container was removed and its ID reused, or the store and docker + diverged). The label is not authoritative — the store is — but it's a + debugging aid and a consistency check. +- **`list` filter.** `docker/container/list` with an `owned_only: true` + input flag filters to containers whose `alknet.owner` label matches + the caller's `Identity.id`. This is the result-filter path (ADR-050 + #4a). Without the flag, the list returns all containers the daemon + sees (the hosted-services case — see §3 below). + +**`alknet.*` label reservation.** The `.*` label keys are +reserved: the `docker/container/create` handler overwrites +`.managed` and `.owner` on the container's labels +regardless of what the caller provided, preventing a caller from +spoofing ownership (setting `alknet.owner` to another peer's id). A +caller's own labels (non-`.*` keys) are preserved. This is a +security property: ownership is set by the handler (from the +caller's `Identity.id`), not by the caller's input. + +**Resource action vocabulary.** The `resource_action` field on +`AccessControl` for containers uses these actions: + +| Action | Used by | Scope name | +|--------|---------|-----------| +| `exec` | `docker/container/exec` (call op, non-interactive) | `container:exec` | +| `tty` | `DockerTtyBackend` sessions (`alknet/tty`) | `container:tty` | +| `start` | `docker/container/start` | `container:start` | +| `stop` | `docker/container/stop` | `container:stop` | +| `remove` | `docker/container/remove` | `container:remove` | +| `restart` | `docker/container/restart` | `container:restart` | +| `manage` | (static, operator role) `Identity.resources["container"]` | `container:manage` | +| `list` | `docker/container/list` (scope-gate only, no resource_id) | `container:list` | +| `create` | `docker/container/create` (scope-gate only, no resource_id) | `container:create` | + +The `tty` action is distinct from `exec` (ADR-061 §4): a caller +authorized for non-interactive exec (`container:exec`) is not +automatically authorized for an interactive terminal session +(`container:tty`). The `manage` action is the static operator-role +resource (§3) that subsumes the per-container actions for pre-existing +hosted-service containers. The `list` and `create` actions are +scope-gate-only (no `resource_id_path`); the rest target a specific +container via `resource_id_path: "$.containerId"`. + +### 2. `docker/container/list`: scope-gate + optional result-filter + +`docker/container/list` is a `Query` operation with +`resource_type: "container"`, no `resource_id_path` (the `list` case per +ADR-050 #4a). The input accepts an optional `owned_only: bool` flag +(default `false`): + +- **`owned_only: false`** (default) — the handler calls + `bollard::list_containers()` and returns all containers the daemon + sees. The scope check gates the call (the caller needs + `container:list`); there is no result-filter. This is the + hosted-services case: an operator listing all containers on dev1, + including ones they didn't spawn through alknet. +- **`owned_only: true`** — the handler calls + `bollard::list_containers()` with a label filter + (`label: alknet.owner=`), returning only the + containers the caller owns. This is the disposable-dev-container + case: a coordinator listing its own workspaces. The scope check still + gates the call; the label filter is the result-filter. + +The two paths use the same bollard method (`list_containers` with +different `ListContainersOptions`), differing only in the label filter. +The `owned_only` flag is the caller's choice; the scope check is the +server's. This matches ADR-050 #4a's "allow if scoped, filter to owned" +default, with the filter opt-in (the hosted-services case needs the +unfiltered list). + +### 3. Hosted services: ownership is not required for pre-existing containers + +The hosted-services case (dev1's reverse-proxy, postgres, redis, gitea) +involves containers created by an operator via `docker compose`, not via +`docker/container/create`. These containers have no `alknet.owner` +label and no ownership-store entry. Operations on them (`start`, +`stop`, `inspect`, `logs`, `restart`) must still work — but +`AccessControl::check` with `resource_type: "container"` and a +`resource_id` would consult the ownership store, find no entry, and +deny. + +The resolution: **the `AccessControl` for operations on a specific +container (`exec`, `stop`, `remove`, `inspect` with `resource_id_path`) +declares `resource_type: "container"` and `resource_action`, but the +ownership check is against the store, and the store's "no entry" result +falls through to the static `Identity.resources` path (ADR-050 §2 +backward-compat).** A peer whose `Identity.resources["container"]` +includes `"manage"` (a deployment-configured scope for the operator +role) passes the check for any container, owned or not. A peer without +that resource but with `container:exec` scope passes only for containers +the ownership store says they own. + +Concretely, the two roles: + +| Role | Scope | Owns a specific container? | Can exec into container C? | +|------|-------|---------------------------|-----------------------------| +| Operator (manages hosted services) | `container:manage` in `Identity.resources["container"]` | No (no ownership entry) | Yes — the static-resource fallback passes (`manage` ⊇ `exec`) | +| Coordinator (spawns dev containers) | `container:exec` scope | Yes (ownership store: coordinator owns C) | Yes — the ownership check passes | +| Random peer | `container:exec` scope | No | No — ownership check fails, static fallback fails | + +The operator role is configured statically (the dev1 operator's +`PeerEntry.resources` includes `container:manage`); the coordinator +role is dynamic (the ownership store records the spawn). Both paths +reach the same `AccessControl::check`; the difference is which branch +satisfies. + +This means `docker/container/exec` and friends work on hosted-service +containers *for the operator role* without the operator having to +"claim" ownership of containers they didn't spawn. The +`docker/container/create` operation (the spawner path) always records +ownership; `docker/container/start` on a pre-existing container does +not (it's not a spawn). The model is: **spawn → own; manage (operator) → +static-resource-pass; neither → deny.** + +### 4. Teardown coupling: handler-driven revoke + autonomous-death tolerance + +ADR-050 #4b says teardown is handler-driven: `docker/container/remove` +calls `OwnershipStore::revoke("container", id)`. But containers can die +autonomously: + +- A `--rm` container exits and is auto-removed by the daemon. +- A container crashes and an operator removes it via `docker rm` (not + through alknet-docker). +- The daemon restarts (all containers stop; container IDs are stale). + +alknet-docker's teardown coupling is **handler-driven revoke on the +alknet-managed remove path, with autonomous-death tolerance**: + +- **`docker/container/remove` (the alknet operation)** calls + `bollard::remove_container()`, and on success calls + `OwnershipStore::revoke("container", id)`. This is the ADR-050 #4b + contract: the handler that manages the lifecycle revokes. +- **Autonomous container death** (a `--rm` exit, an external `docker rm`, + a daemon restart) leaves a stale ownership-store entry. The entry is + not proactively cleaned up — there is no reaper, no event subscription + to the docker daemon's death events. Instead, the entry is cleaned up + lazily: the next operation against that container ID fails at the + bollard layer (the container doesn't exist), the error propagates as + a `call.error`, and a subsequent `docker/container/list` with + `owned_only: true` for the stale owner filters the container out + (it's not in the daemon's list). The ownership store's + `owned_resources` can be cross-checked against the live daemon list + on read, or left to be overwritten on the next spawn — the container ID + is unique per daemon run; a reused ID would be a new container. + +The tolerance for stale entries is intentional: ownership is runtime +state (ADR-050 assumption 4), meaningless across restarts. The +in-memory default store loses all entries on restart; a persistence +adapter would cache and invalidate. A reaper that subscribes to docker +daemon death events would be more prompt but adds a subscription +surface for a marginal gain (the stale entry is inert — it can't grant +access to a container that doesn't exist, and a reused container ID +gets a fresh `record` on its next `create`). + +A future "prompt stale-entry cleanup" feature is additive: a +`docker/system/events` subscription operation (deferred, OQ-050) that +the ownership store could subscribe to. The base model tolerates stale +entries; the prompt-cleanup path is a refinement. + +### 5. `docker/container/create` records ownership; `start`/`stop`/`restart` do not + +Only `docker/container/create` calls `OwnershipStore::record`. The +lifecycle operations (`start`, `stop`, `restart`, `remove`) do not +record — they act on an existing container whose ownership was recorded +at create time (or which pre-exists and is reached via the +static-resource operator path). `remove` calls `revoke` (the teardown +half); the others neither record nor revoke. + +This matches ADR-050's "spawner owns" model: the *create* is the spawn +event; the lifecycle operations are management of an existing resource. + +## Consequences + +**Positive:** + +- The two use cases (disposable dev containers, hosted services) both + work through one `AccessControl` model. The operator role reaches + hosted-service containers via the static-resource fallback; the + coordinator role reaches spawned containers via the ownership store. + No special-casing, no "is this a managed container?" branch in the + handlers. +- The label namespace is configurable and minimal (two labels). The + `managed` flag marks alknet-spawned containers; the `owner` label + carries the peer id for the `list` filter and the cross-check. +- Teardown is handler-driven on the alknet path and tolerant of + autonomous death on the non-alknet path. No reaper subscription + surface in the base model. +- ADR-050's model is applied without amendment. The `list` case + (specific #4a), the teardown coupling (#4b), and the + backward-compat static-resource fallback (#2) all map cleanly to + bollard's API (`list_containers` with a label filter, `remove_container` + + `revoke`, the static `Identity.resources` path). + +**Negative:** + +- Stale ownership entries from autonomous container death are not + promptly cleaned up. An `owned_only: true` list could theoretically + return a container ID that no longer exists — but bollard's + `list_containers` returns *live* containers, so the list is always + accurate against the daemon; the stale entry only affects the + ownership store's internal map, not the list result. The cross-check + (label vs store) could flag divergence; this is a debugging aid, not + a correctness issue. +- The operator role requires static configuration + (`Identity.resources["container"] ⊇ "manage"`). A deployment that + wants the operator to manage hosted services must configure the + operator's `PeerEntry.resources`. This is not a downside — it's the + same static configuration every role requires — but it means the + hosted-services case isn't "zero-config"; the operator peer must be + declared. +- The `owned_only` flag on `list` is an opt-in. A coordinator that + forgets to set it gets all containers, not just its own. The default + (`false`) is correct for the hosted-services case (operator lists + all), but a coordinator expecting isolation must set the flag. This + is a caller-side convention, enforced by the scope check (a + non-operator peer without `container:manage` can't call `list` at all + — the scope gates it), but within the `container:list` scope, the + filter is the caller's choice. + +## Door type + +**One-way (label schema, ownership model application) + two-way (label +prefix, stale-entry policy).** The label schema (two labels, +`.managed` + `.owner`) and the +create-records/remove-revokes ownership coupling are one-way: clients +and the ownership store depend on the labels and the record/revoke +timing. The label prefix (default `alknet`) is two-way-door config. The +stale-entry policy (tolerate, no reaper) is two-way — a reaper +subscription is an additive refinement. + +## References + +- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md) + — the model this ADR applies (specifics #4a, #4b, #2) +- `docs/research/alknet-docker/poc-summary.md` §"What the POC Does NOT + Validate" #6 (label namespace / ownership mapping) +- `/workspace/@alkdev/dispatch/src/docker.rs` — the dispatch POC's + `dispatch.managed=true` label (prior art this ADR generalizes) +- [ADR-030](030-peerentry-and-identity-id-decoupling.md) — `Identity.id` + as the stable peer id used in the `alknet.owner` label +- [ADR-032](032-forwarded-for-identity.md) — `forwarded_for` for the + proxy pattern's end-user identity +- [ADR-015](015-privilege-model-and-authority-context.md) — the static + `Identity.resources` path (the operator-role fallback) +- `/workspace/system/dev1/docker.md` — the hosted-services use case + (reverse-proxy, postgres, redis, gitea on dev1) +- `/workspace/@alkdev/reverse-proxy/deploy/docker-compose.yml` — the + reverse-proxy's docker setup (operator-created, not alknet-spawned) +- Spec: [docker-operations.md](../crates/docker/docker-operations.md) + §"Access Control" and §"Label Namespace" \ No newline at end of file diff --git a/docs/architecture/decisions/061-docker-tty-backend-in-alknet-docker.md b/docs/architecture/decisions/061-docker-tty-backend-in-alknet-docker.md new file mode 100644 index 0000000..4b6ebe8 --- /dev/null +++ b/docs/architecture/decisions/061-docker-tty-backend-in-alknet-docker.md @@ -0,0 +1,272 @@ +# ADR-061: DockerTtyBackend in alknet-docker + +## Status + +Accepted + +## Context + +The alknet-tty spec ([tty-backend.md](../crates/tty/tty-backend.md) +§"Backend implementations") lists three `TtyBackend` implementations: +`LocalTtyBackend` (in `alknet-tty-local`), `DockerTtyBackend` (location +"future, out of scope here"), and `SshTtyBackend` (in `alknet-ssh`). +The alknet-tty spec deliberately left the `DockerTtyBackend` location +open — it committed the trait shape but not the crate placement. + +alknet-docker is now being specified. The `DockerTtyBackend` — the +`impl TtyBackend` that wraps `bollard::attach_container()` for +interactive attach and `bollard::exec::start_exec` with `tty: true` for +interactive exec — is the natural extension of the alknet-docker POC's +`drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`). The +question is which crate it lives in. + +### The options + +Three placements were considered: + +1. **In `alknet-docker`** — the `DockerTtyBackend` is a type in + alknet-docker, behind a `tty` feature gate that pulls in + `alknet-tty` (for the `TtyBackend` trait) and `bollard` (for the + attach/exec API). alknet-docker already depends on bollard for its + call operations; the `tty` feature adds only the `alknet-tty` edge. + +2. **In `alknet-tty-docker`** (a sibling adapter crate) — the + `DockerTtyBackend` in its own crate, depending on `alknet-tty` (the + trait) and `bollard` (the API). alknet-docker stays leaner (no + `alknet-tty` edge); the adapter is its own publishable unit. + +3. **In `alknet-tty-local`** (extending the local sibling) — rejected + immediately. `alknet-tty-local` is for *local* backends (`portable_pty`, + `std::process`); docker is not local-process, and pulling bollard into + the local sibling would violate the feature-gate rationale (a + docker-only deployment pulling in PTY code, and a local-only + deployment pulling in bollard). + +### Why option 1 (in alknet-docker) + +The `DockerTtyBackend` is tightly coupled to alknet-docker's domain: +it uses the same bollard client, the same container identity model +(`resource_id` extraction from `backend_params.container`), and the +same label/ownership reasoning (ADR-060) as the call operations. A +separate `alknet-tty-docker` crate would duplicate the bollard client +construction, the container-id validation, and the ownership wiring — +or import them from alknet-docker, creating a cycle +(`alknet-tty-docker` → `alknet-docker` for the client, `alknet-docker` +→ `alknet-tty-docker` for the backend registration). + +The dependency edge from alknet-docker to alknet-tty is clean: +alknet-docker depends on alknet-tty for the `TtyBackend` trait (and +`TtyHandle`, `TtyControl`, `TtyError`, `TtyParams`). This is the same +edge any backend crate has — `alknet-tty-local` depends on +`alknet-tty` for the trait. The trait is the inversion point +(ADR-053); the backend impls live where their transport deps live +(bollard, for `DockerTtyBackend`). bollard already lives in +alknet-docker; the `DockerTtyBackend` belongs with it. + +The `tty` feature gate means a deployment that wants only the call +operations (no interactive terminal) doesn't pull in `alknet-tty`. The +feature is opt-in, matching `alknet-tty-local`'s `local` feature +pattern (ADR-054). + +### The POC's `drive_attach_raw` as the reference + +The POC's `drive_attach_raw` (`/workspace/alknet-docker-poc/src/ops.rs`) +is the seed of `DockerTtyBackend::allocate()`. The mapping: + +| POC `drive_attach_raw` | `DockerTtyBackend::allocate()` | +|---|---| +| reads `call.requested` frame, extracts container id | `TtyParams.backend_params` → `DockerBackendParams { container }` | +| `bollard::attach_container()` → `AttachContainerResults { output, input }` | same; `output` → `TtyHandle.stdout`, `input` → `TtyHandle.stdin` | +| `ChunkReader`/`ChunkWriter` on the bidi stream | the `TtyAdapter` owns the chunk codec (ADR-052); the backend produces handles, not wire bytes | +| zero-length stdin chunk → `container_input.shutdown()` | `TtyControl` / stdin EOF — the adapter signals; the backend's `AsyncWrite` close maps to `shutdown()` | +| completion: output stream ends → zero-length stdout sentinel → close | `TtyHandle.exit_code` future — for attach, the exit comes from `wait_container` or the output stream end; for exec with `tty: true`, `inspect_exec` after the output stream ends | +| `LogOutput` → `Chunk` stream_type mapping (StdOut→1, StdErr→2) | `TtyHandle.stdout`/`stderr` — PTY mode merges (stderr is `None`); the adapter pumps stdout only | + +The POC's `drive_exec` (non-interactive, exit code via `inspect_exec`) +is the basis for the **call operation** `docker/container/exec` +(`tty: false`), not the `DockerTtyBackend` (which is the `tty: true` +interactive path). The split: non-interactive exec is a call +`Subscription` (ADR-058); interactive exec is a `DockerTtyBackend` +session. Both use `bollard::exec`; they diverge at the wire format. + +## Decision + +### 1. `DockerTtyBackend` is a type in alknet-docker, behind a `tty` feature + +```toml +# alknet-docker Cargo.toml +[features] +default = ["ops"] # the call operations +ops = [] # docker/container/* and docker/image/* operations +tty = ["dep:alknet-tty"] # DockerTtyBackend (impl TtyBackend) +``` + +- `default = ["ops"]` — the call operations (lifecycle, logs, exec + non-interactive, images). This is what a hub or coordinator wires to + manage containers over the call protocol. +- `tty` — adds `DockerTtyBackend` (impl `TtyBackend`), pulling in + `alknet-tty` for the trait. A deployment that wants interactive + terminal sessions into docker containers enables this and registers + `DockerTtyBackend` with the `TtyAdapter` (alongside or instead of + `LocalTtyBackend`). + +The feature gate means a deployment that uses alknet-docker only for +call operations (the common case — a hub managing dev containers) +doesn't pull in `alknet-tty`. A deployment that wants interactive +terminals into those containers enables `tty` and registers the backend. + +### 2. `DockerTtyBackend::allocate()` wraps `attach_container` or exec-with-PTY + +`allocate()` branches on `TtyParams.terminal` and a +`DockerBackendParams` field for the attach mode: + +- **`attach` mode** — wraps `bollard::attach_container()` (the reliable + HTTP-upgrade path, per the POC and ADR-059 §3). Returns + `AttachContainerResults { output, input }`. `output` → `TtyHandle.stdout` + (as a `Stream`, mapping `LogOutput` → `Bytes` with the + stream_type flattened — PTY mode has no separate stderr, so + `TtyHandle.stderr` is `None`). `input` → `TtyHandle.stdin` (the + `AsyncWrite` bollard provides). Exit code: the output stream end + + `wait_container()` or `inspect_container()` for the exit status. +- **`exec` mode** — wraps `bollard::create_exec` + `start_exec` with + `tty: true`. `start_exec` returns `StartExecResults::Attached { + output, input }`; same mapping as attach. Exit code: after the output + stream ends, `inspect_exec()` for the exit code (the POC's pattern). + +The mode is selected by a `DockerBackendParams.mode` field +(`"attach" | "exec"`, default `"exec"` for a new command, `"attach"` +for attaching to a running container's primary process). For `exec` +mode, `TtyParams.cmd` is the command vector; for `attach` mode, `cmd` +is ignored (the primary process is already running). + +### 3. `TtyControl` maps to bollard's resize and signal + +- `resize()` → `bollard::container::resize_container_tty()` (attach + mode) or `bollard::exec::resize_exec()` (exec mode). +- `signal()` → docker has no direct container signal API in bollard's + stable surface. The POC used `kill_container` for the attach case. + For exec mode, there is no per-exec signal; the signal is delivered + to the container's PID namespace via `kill_container` with the + container's main PID, or the exec's PID if exposed. This is a + best-effort mapping per the `TtyControl::signal` contract + ("best-effort delivery to the foreground process group"); unknown + signals fall back to `kill_container` with `SIGKILL`. The exact + signal-delivery path (bollard's `kill_container` vs a future exec-PID + signal API) is a two-way-door implementation detail within the + one-way backend placement. + +### 4. `resource_id()` returns `Some(("container", container_id))` + +`DockerTtyBackend::resource_id(¶ms)` extracts the container ID +from `backend_params` and returns `Some(("container", id))`, per the +`tty-backend.md` sketch. The `TtyAdapter` calls +`OwnershipProvider::owns(identity, "container", id, "tty")` at +negotiation (ADR-050, ADR-060). The backend owns the extraction; the +adapter doesn't parse docker-specific JSON. This is the same shape +the `tty-backend.md` sketch anticipates. + +### 5. `exit_code` future's Drop kills the container/exec (ADR-056) + +The `TtyHandle.exit_code` future's `Drop`-on-cancel (ADR-056) issues: + +- **Attach mode** — `bollard::kill_container()` (or the container's + natural exit if the stream end already triggered). The container is + not removed (attach doesn't own the container's lifecycle); it's + killed so the process doesn't outlive the session. +- **Exec mode** — the exec instance is not a separate killable + process in docker's model; the signal goes to the container. The + `Drop` issues `kill_container` with the container's main PID as a + best-effort. A container running an exec that the terminal session + cancels should have the exec's process terminated; docker's exec + isolation means this is best-effort, not guaranteed. + +This satisfies ADR-056's contract: dropping `exit_code` (cancel) kills +the session target. The "session target" for docker is the container +(attach) or the exec's process (exec, best-effort via the container). + +## Consequences + +**Positive:** + +- The `DockerTtyBackend` lives where its deps live (bollard in + alknet-docker), matching the ADR-053 inversion principle. No + `alknet-tty-docker` sibling crate, no cycle risk. +- The `tty` feature gate keeps the `alknet-tty` edge opt-in. A + call-operations-only deployment (the common case) doesn't pull in + alknet-tty. +- The POC's `drive_attach_raw` maps cleanly to + `DockerTtyBackend::allocate()` — the backend produces handles, the + `TtyAdapter` pumps the wire format. The POC's three-pump driver + (`session.rs`) is the reference; the backend extracts the bollard + interaction from it. +- The `resource_id()` delegation means the adapter's ownership check + (ADR-060) works without the adapter parsing docker JSON — the + backend extracts the container ID, the adapter checks the store. + +**Negative:** + +- alknet-docker gains an `alknet-tty` dependency edge (behind the + `tty` feature). This is the intended trade: the backend lives with + its transport dep, and the trait edge is feature-gated. A deployment + that wants interactive docker terminals accepts this; one that + doesn't is unaffected. +- Signal delivery in docker exec mode is best-effort (docker's exec + isolation limits per-exec signal targeting). The `TtyControl::signal` + contract is already best-effort (ADR-053), so this is within the + contract, but a user expecting `Ctrl-C` to kill only the exec + process (not the container) may see the container killed instead. + This is a docker-semantics limitation, not an alknet limitation; + the local backend (portable_pty) has precise process-group + targeting (REQ-TTY-02). The docker backend's signal semantics are + documented in the spec, not hidden. +- The `attach` vs `exec` mode split in `DockerBackendParams` is a + backend-specific detail the caller must know. The `TtyParams` is + opaque to the adapter (ADR-053), but the caller (the client opening + the tty session) must set `mode: "attach"` or `mode: "exec"` in the + negotiation frame's backend params. This is the same opacity as the + local backend's `cwd`/`env` — the caller knows their backend. No + adapter change is needed for a new mode; the backend parses its own + params. + +## Door type + +**One-way (crate placement) + two-way (feature gate, attach/exec mode +split, signal path).** The decision to put `DockerTtyBackend` in +alknet-docker (not a sibling crate) is one-way: the backend type, its +registration with the `TtyAdapter`, and its `bollard`/`alknet-tty` +edges are structural. Reversing to a sibling crate would move the type +and re-edge the dependencies. + +The feature gate (`tty` opt-in), the attach/exec mode split in +`DockerBackendParams`, and the exact signal-delivery path +(`kill_container` vs a future exec-PID API) are two-way-door +implementation details within the one-way placement. The feature can +be renamed, the mode field can gain a third value, and the signal path +can be refined — all without a structural change. + +## References + +- [ADR-053](053-ttybackend-trait-and-ttyhandle.md) — the + `TtyBackend` trait this backend implements; the inversion principle + (backends live where their transport deps live) +- [ADR-052](052-alknet-tty-wire-format-and-two-carriage.md) — the wire + format the `TtyAdapter` pumps (the backend produces handles, not wire) +- [ADR-054](054-local-tty-backend-sibling-crate.md) — the sibling-crate + pattern for `LocalTtyBackend` (the parallel this ADR follows, with + the difference that docker's deps are already in alknet-docker) +- [ADR-056](056-backend-cleanup-on-session-cancel.md) — the + cancel-cleanup contract (`exit_code` future's `Drop` kills) +- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md) + — the ownership model `resource_id()` delegates to +- [ADR-058](058-alknet-docker-on-alknet-call.md) — why the + non-interactive exec is a call op, not this backend +- [ADR-059](059-bollard-021-dependency-and-features.md) — the bollard + dependency (no `websocket` feature — the reliable attach path) +- `docs/research/alknet-docker/poc-summary.md` §"POC Target 1" + (interactive attach — the seed of this backend) +- `/workspace/alknet-docker-poc/src/ops.rs` — `drive_attach_raw` (the + reference for `allocate()`) +- [tty-backend.md](../crates/tty/tty-backend.md) §"Backend + implementations" (the row this ADR fills: "DockerTtyBackend | alknet-docker + or sibling adapter | future, out of scope here") +- Spec: [docker-tty-backend.md](../crates/docker/docker-tty-backend.md) \ No newline at end of file diff --git a/docs/architecture/decisions/062-docker-client-injection-via-closure-capture.md b/docs/architecture/decisions/062-docker-client-injection-via-closure-capture.md new file mode 100644 index 0000000..f97ffc3 --- /dev/null +++ b/docs/architecture/decisions/062-docker-client-injection-via-closure-capture.md @@ -0,0 +1,234 @@ +# ADR-062: Docker Client and OwnershipStore Injection via Closure Capture + +## Status + +Accepted + +## Context + +The docker operation handlers need access to two non-secret, shared +runtime handles: + +1. The bollard `Docker` client — a `Clone`-able handle (internally an + `Arc`) to the docker daemon connection, constructed once at + assembly-layer startup. +2. The `OwnershipStore` — the ADR-050 write-side trait, used by + `docker/container/create` (to `record`) and `docker/container/remove` + (to `revoke`). + +An earlier draft of the docker spec +(`docker-operations.md`) showed handlers calling +`ctx.docker_client()` and speculated the handle lived on "a +`DockerOpsExt` extension trait on `OperationContext`, or a +`DockerClient` capability in `ctx.capabilities`." This was an inline +open question with a model conflict: + +### Why `Capabilities` is the wrong channel + +`Capabilities` (ADR-014, `operation-registry.md` §"Capability +Injection") is the type for **outbound secret material** — decrypted +API keys, signing keys, vault-derived credentials. Its contract +(`core-types.md`): non-`Serialize`, zeroized on drop, populated from +the vault at registration time. A bollard `Docker` handle to a local +unix socket is not secret material; it has no zeroizing semantics; it +must not be smuggled through the secret-injection channel. Putting a +`Docker` handle in `Capabilities` would be a category error against +the one-way `Capabilities` contract (ADR-014). + +The same applies to the `OwnershipStore`: it's a shared state handle +(shared `Arc`), not secret material. + +### Why `OperationContext` extension is the wrong channel + +`OperationContext` (`operation-registry.md:230`) is a concrete struct +with a fixed field set, constructed by the dispatch path per call. +Adding a `docker_client` field (or an extension trait that reads one) +would either (a) make every non-docker handler pay for a field they +don't use, or (b) require a typed-map / downcast pattern that +violates the "context is concrete, not a bag" design. The +established pattern for per-handler-set shared state is not a +context field — it's closure capture. + +### The established pattern: closure capture at registration + +The `from_openapi` adapter (ADR-017, `client-and-adapters.md`, +`operation-registry.md:819`) captures its `reqwest::Client` in each +forwarding handler's closure at registration time: + +```rust +// operation-registry.md:819 (the established pattern) +.with_local(vastai_listMachines_spec(), Arc::new(vastai_handler), + CompositionAuthority::new("vastai", ["vastai:query"]), + ScopedOperationEnv::new([...]), + Capabilities::new().with_api_key("vastai", vastai_token)) +``` + +The `vastai_handler` closure captures its reqwest client (and its +auth token via `Capabilities`, which *is* secret material). The +non-secret shared handle (reqwest client) is captured in the closure; +the secret (the API key) goes through `Capabilities`. This is the +correct split: secret material through `Capabilities`, non-secret +shared state through closure capture. + +The docker ops follow the same split. The `Docker` client and the +`OwnershipStore` are non-secret shared state — closure-captured. There +are no secret capabilities (local bollard needs no API key; the +docker daemon is local). + +## Decision + +### 1. `register_docker_ops` captures the `Docker` client and `OwnershipStore` in each handler closure + +The `register_docker_ops` function (or the `DockerOps` builder) +takes `Arc` and `Arc` (plus the label +config and the `CompositionAuthority`) as arguments, and constructs +each `Handler` / `StreamingHandler` closure with those handles +captured by reference (`Arc::clone` per handler): + +```rust +pub fn register_docker_ops( + builder: &mut OperationRegistryBuilder, + docker: Arc, + ownership_store: Arc, + labels: &DockerLabels, + authority: CompositionAuthority, +) { + let docker_clone = docker.clone(); + let ownership_clone = ownership_store.clone(); + let labels_clone = labels.clone(); + builder.with_local( + container_inspect_spec(), + Arc::new(move |input, ctx| { + let docker = docker_clone.clone(); + Box::pin(async move { + let container_id = input["containerId"].as_str() + .ok_or_else(|| /* INVALID_INPUT */)?; + match docker.inspect_container(container_id, None::<()>).await { + Ok(info) => ResponseEnvelope::ok(to_json_value(info)), + Err(e) => /* CONTAINER_NOT_FOUND / DOCKER_ERROR */, + } + }) + }), + authority.clone(), + ScopedOperationEnv::empty(), + Capabilities::new(), // no secret caps — local bollard + ); + // ... same pattern for each operation ... +} +``` + +Each handler closure captures the `Docker` client and (where needed) +the `OwnershipStore` by cloning the `Arc`. The handler does not read +these from `OperationContext`; they're baked into the closure at +registration time. The `Capabilities` passed to `with_local` is empty +(`Capabilities::new()`) — there are no secret capabilities for local +bollard operations. + +### 2. The `Docker` client is shared, `Clone`-able, and constructed once + +bollard's `Docker` is `Clone` (it holds an internal `Arc` to the +connection). The assembly layer constructs one +(`Docker::connect_with_local_defaults()`, ADR-059 §2), wraps it in an +`Arc`, and passes it to `register_docker_ops`. Each handler clones +the `Arc` (cheap — a refcount bump) into its closure. There is no +per-handler docker client; the one client is shared across all +docker operations. + +### 3. The `OwnershipStore` is shared the same way + +`Arc` is cloned into the `create` and `remove` +handler closures (the two operations that write to the store). The +other operations (inspect, start, stop, logs, exec, list, images) +don't capture the store — they don't write to it. The `create` +closure captures it for `record`; the `remove` closure captures it +for `revoke`. + +### 4. No `docker_client()` accessor; no `OperationContext` extension + +The handlers do not call `ctx.docker_client()` or any other accessor. +The `Docker` client and `OwnershipStore` are not on `OperationContext` +and are not accessible via an extension trait. They're closure- +captured, full stop. A handler that needs the docker client gets it +from its closure's captured `Arc`, not from the context. + +This means the handler signature is the standard +`Fn(Value, OperationContext) -> Pin + Send>>` +(ADR-049) — no new parameter, no new context field. The docker- +specific state is in the closure, not the context. + +## Consequences + +**Positive:** + +- The injection model matches the established `from_openapi` pattern + (non-secret shared state via closure capture, secret material via + `Capabilities`). No new mechanism; no `OperationContext` change; no + `Capabilities` conflict. +- The `Capabilities` contract (ADR-014) stays clean — it holds secret + material only. A local bollard handle is not smuggled through the + secret channel. +- The handlers are self-contained — the `Docker` client and store are + baked in at registration, not looked up at call time. This is + testable (construct a `Docker` client in a test, register the ops, + invoke the handler directly) and composable (the same handler works + whether the docker client points at a local socket or a test + daemon). +- No new trait, no extension mechanism, no downcast. The injection is + plain Rust closure capture — the simplest thing that works. + +**Negative:** + +- Each handler closure captures the `Docker` and (for create/remove) + the `OwnershipStore` by `Arc::clone`. This is a refcount bump per + handler construction (at startup, once) — negligible. The per-call + cost is zero (the closure already holds the `Arc`; calling the + handler doesn't re-clone). +- The `register_docker_ops` function signature is longer (takes the + `Docker` + store + labels + authority). This is the assembly layer's + wiring concern, not a handler concern — the handlers don't know + where their `Docker` came from. +- A test that wants to invoke a single docker handler must construct + the `Docker` client and (if the handler uses the store) the + `OwnershipStore` to register the op. This is the same as any + handler test that needs its dependencies — not new. + +## Door type + +**One-way.** The decision to use closure capture (not a context field, +not a `Capabilities` entry) is the injection model every docker +handler is written against. Reversing to a context field would be a +rewrite of every handler and an `OperationContext` change. The +specific captures (`Docker` + `OwnershipStore`) are the non-secret +state the handlers depend on; changing what's captured is a handler +rewrite. + +The choice of `Arc` vs a non-`Arc` shared reference is a +two-way-door implementation detail (bollard's `Docker` is already +`Clone`-with-`Arc`-internally; the outer `Arc` is for the trait-object +`OwnershipStore`'s sake, not the `Docker`'s). The capture pattern +(closure capture) is the one-way commitment. + +## References + +- [ADR-014](014-secret-material-flow-and-capability-injection.md) — + the `Capabilities` contract this ADR keeps clean (secret material + only; the `Docker` handle is not secret) +- [ADR-017](017-call-protocol-client-and-adapter-contract.md) — the + `from_openapi`/`from_call` adapter pattern this ADR follows + (non-secret shared state via closure capture) +- [ADR-022](022-handler-registration-provenance-and-composition-authority.md) + — `HandlerRegistration`, the bundle the `register_docker_ops` + builder populates +- [ADR-050](050-dynamic-resource-ownership-for-runtime-spawned-resources.md) + — the `OwnershipStore` trait the `create`/`remove` handlers capture +- [ADR-058](058-alknet-docker-on-alknet-call.md) — the operation + surface this injection model serves +- `crates/call/operation-registry.md` §"Capability Injection" + (the `Capabilities` semantics this ADR respects) and + §"OperationContext" (the struct this ADR does *not* extend) +- `crates/call/client-and-adapters.md` §"from_openapi forwarding" + (the prior art: reqwest client closure-captured, API key in + `Capabilities`) +- Spec: [docker-operations.md](../crates/docker/docker-operations.md) + §"Handler injection" (the section this ADR adds, replacing the + earlier `ctx.docker_client()` sketches) \ No newline at end of file diff --git a/docs/architecture/decisions/063-exit-code-on-terminal-call-responded.md b/docs/architecture/decisions/063-exit-code-on-terminal-call-responded.md new file mode 100644 index 0000000..08f93f5 --- /dev/null +++ b/docs/architecture/decisions/063-exit-code-on-terminal-call-responded.md @@ -0,0 +1,185 @@ +# ADR-063: Exit Code on a Terminal `call.responded` for Non-Interactive Exec + +## Status + +Accepted + +## Context + +The alknet-docker POC (`docs/research/alknet-docker/poc-summary.md` +§"POC Target 3") validated a completion shape for streaming exec: +the exit code rides on a final `call.responded` frame before +`call.completed`. This keeps `call.completed`'s payload empty +(matching ADR-012's wire format — no core protocol change). The POC +used `{ "exitCode": N }` on the final `call.responded`. + +The docker spec (`docker-operations.md` §"docker/container/exec") +carries this forward and adds a `"terminal": true` field to mark the +exit-code `call.responded` as the final value before completion. The +`terminal` flag is a docker-operation convention, not a call-protocol +wire-format change — it's a field in the `call.responded` payload +(the `Value`), not a new event type or a `call.completed` payload. + +A reader asked: "why `terminal: true` and not a dedicated `call.exit` +event or an exit code on `call.completed`?" This ADR records the +decision. + +### Why not an exit code on `call.completed` + +ADR-012 defines `call.completed` with an empty payload (`{}`). The +POC's completion-shape decision (POC §"The completion-shape decision +this validates") chose to keep `call.completed` empty and put the +exit code on the preceding `call.responded`. Changing `call.completed` +to carry a payload would be a wire-format change (ADR-012 amendment, +new ALPN per ADR-006) — a one-way door that affects every +`call.completed` consumer, not just docker. The POC deliberately +avoided this. + +### Why not a dedicated `call.exit` event + +A new event type (`call.exit`) would be a wire-format addition (new +`EventEnvelope` variant, ADR-012 amendment). It would also duplicate +the "terminal value before completion" semantics that +`call.completed` already provides — `call.completed` is the +"stream end" signal; an exit code is a "final value" that rides +*before* the stream end. A `EventEnvelope` variant conflates these +two: `call.exit` would be both the final value and the stream-end +signal, but `call.completed` already handles stream end. Keeping +the exit code as a `call.responded` (a normal value) and letting +`call.completed` be the stream end preserves the +one-value-per-`call.responded` invariant: the exit code is just the +last value. + +### Why `terminal: true` + +A client consuming a `docker/container/exec` stream sees a sequence +of `call.responded` frames (stdout/stderr lines) and then a final +`call.responded` with `{ "exitCode": N, "terminal": true }`, then +`call.completed`. The `terminal: true` flag tells the client "this +is the last value; the next event is `call.completed`." Without it, +the client would have to infer "this `call.responded` has an +`exitCode` field, so it's the exit" — a content-sniffing heuristic +that breaks if a future stdout line happens to carry an `exitCode` +field. + +The `terminal` flag is a docker-operation convention — a field in +the `call.responded` payload, not a protocol-level marker. It's +declared in the `docker/container/exec` `output_schema` (the exit +frame is a documented part of the output stream). Other +`Subscription` operations that produce a "terminal result before +completion" can adopt the same convention (a `terminal: true` field +on their final `call.responded`); it's not docker-specific, but +docker is the first operation to need it. + +## Decision + +### 1. The exec exit code rides on a final `call.responded` with `terminal: true` + +`docker/container/exec` (non-interactive, `tty: false`) emits: + +1. Zero or more `call.responded` frames with stdout/stderr output + (`{ "stream": "stdout"|"stderr", "text": "..." }`). +2. A final `call.responded` with `{ "exitCode": N, "terminal": true }`. +3. `call.completed` (empty payload, per ADR-012). + +The `terminal: true` field marks the exit-code frame as the final +value before completion. The `exitCode` is `i32` (matching +`ExecInspectResponse.exit_code`; negative for signal-terminated, +though docker exec exit codes are typically non-negative). + +### 2. `call.completed` stays empty + +No `call.completed` payload change. ADR-012's wire format is +unchanged. The exit code is on the preceding `call.responded`, not +on `call.completed`. + +### 3. `terminal: true` is an operation-level convention, not a protocol field + +The `terminal` flag is a field in the `call.responded` payload (the +`Value`), declared in the operation's `output_schema`. It is not a +new `EventEnvelope` field, not a new event type, and not a +`call.completed` payload. The call protocol's wire format is +unchanged. Other `Subscription` operations that produce a terminal +result before completion may adopt the same convention; it's not +docker-specific, but docker is the first to use it. + +The `OperationSpec.output_schema` for `docker/container/exec` +documents the two response shapes: the streaming output frames +(`{ "stream": ..., "text": ... }`) and the terminal exit frame +(`{ "exitCode": N, "terminal": true }`). A client reading the schema +knows to expect the terminal frame. + +### 4. The exit-code frame is the last `call.responded` before `call.completed` + +The handler ensures the exit-code `call.responded` is the last one +the stream produces. The `StreamingHandler` (ADR-049) pumps the +output stream, then emits the exit-code frame, then the stream ends +(the dispatcher writes `call.completed` on natural stream end). The +ordering is the handler's responsibility (the handler controls the +stream's item order); the dispatcher's `pump_stream` writes them in +order. This mirrors the POC's validated pattern. + +## Consequences + +**Positive:** + +- The exit code propagates through the streaming completion path + without a wire-format change. `call.completed` stays empty + (ADR-012 unchanged); the exit code is a normal `call.responded` + value that happens to be the last one. +- The `terminal: true` flag gives clients a deterministic "this is + the exit" signal without content-sniffing the `exitCode` field. A + client can stop reading output frames when it sees `terminal: true` + and read the exit code. +- The convention is reusable: any `Subscription` operation with a + terminal result before completion can use `terminal: true` on its + final `call.responded`. No protocol change; just an + operation-level field. + +**Negative:** + +- The `terminal` flag is a convention, not enforced by the wire + format. A `Subscription` operation that emits a terminal result + *without* `terminal: true` would be non-conformant but not + protocol-violating. The `output_schema` documents the convention; + clients that read the schema know to expect it. +- A client that doesn't check `terminal: true` and instead + content-sniffs `exitCode` would work for docker exec but break on + a hypothetical stdout line that carries an `exitCode` field. The + flag is the robust path; content-sniffing is the fragile path. The + spec recommends the flag; the schema documents it. + +## Door type + +**One-way.** The exit-code-on-`call.responded`-before-`call.completed` +shape is the completion contract `docker/container/exec` commits to. +Clients consuming the exec stream depend on this ordering; changing +it (moving the exit code to `call.completed`, or adding a `call.exit` +event) would be a break for those clients. + +The `terminal: true` field name is two-way-door within the one-way +shape — it could be renamed (e.g., `final: true`) without a wire- +format change, since it's a payload field, not a protocol marker. +But the *presence* of a terminal marker on the final `call.responded` +is the one-way commitment: clients depend on "the last +`call.responded` before `call.completed` is marked as terminal and +carries the exit code." + +## References + +- `docs/research/alknet-docker/poc-summary.md` §"POC Target 3" + (exec with exit code — the validated pattern this ADR formalizes) +- `docs/research/alknet-docker/poc-summary.md` §"The completion-shape + decision this validates" (the POC's reasoning for exit-on- + `call.responded`, not on `call.completed`) +- [ADR-012](012-call-protocol-stream-model.md) — the call protocol + wire format this ADR keeps unchanged (`call.completed` stays empty) +- [ADR-049](049-streaming-handler-for-subscriptions.md) — + `StreamingHandler`, the dispatch path that pumps the + `call.responded` stream and writes `call.completed` on stream end +- [ADR-023](023-operation-error-schemas.md) — `output_schema` + documents the terminal frame shape +- [ADR-058](058-alknet-docker-on-alknet-call.md) — the operation + taxonomy (exec is a `Subscription`) +- Spec: [docker-operations.md](../crates/docker/docker-operations.md) + §"docker/container/exec" (the operation this ADR's shape serves) \ No newline at end of file diff --git a/docs/architecture/open-questions.md b/docs/architecture/open-questions.md index ffc2b63..665a349 100644 --- a/docs/architecture/open-questions.md +++ b/docs/architecture/open-questions.md @@ -1,6 +1,6 @@ --- status: draft -last_updated: 2026-07-07 +last_updated: 2026-07-08 --- # Open Questions @@ -149,6 +149,15 @@ Door type is separate from whether a decision is made. A two-way door is a decis | [OQ-46](questions/046-runner-api-surface.md) | Runner API Surface | deferred(scope) | two | low | | [OQ-47](questions/047-stdin-closure-canonical-signal.md) | Stdin Closure Canonical Signal | resolved | two | low | +### alknet-docker + +| OQ | Title | Status | Door | Pri | +|----|-------|--------|------|-----| +| [OQ-48](questions/048-network-and-volume-operation-surface.md) | Network and Volume Operation Surface | deferred(scope) | two | low | +| [OQ-49](questions/049-image-build-buildkit-scope.md) | Image Build (buildkit) Scope | deferred(scope) | two | low | +| [OQ-50](questions/050-docker-system-events-subscription.md) | Docker System Events Subscription | deferred(scope) | two | low | +| [OQ-51](questions/051-container-create-options-surface.md) | Container Create Options Surface | deferred(scope) | two | med | + ## Deferred / Blocked The safe-exit visibility surface. These questions are parked because the @@ -194,3 +203,27 @@ filtering the tables above. - **Priority**: low - **Full file**: [OQ-46](questions/046-runner-api-surface.md) +### OQ-48: Network and Volume Operation Surface + +- **Blocked on**: a concrete use case for network or volume management over the call protocol. Dev containers use the default bridge network; hosted services declare networks/volumes in `docker compose`. +- **Priority**: low +- **Full file**: [OQ-48](questions/048-network-and-volume-operation-surface.md) + +### OQ-49: Image Build (buildkit) Scope + +- **Blocked on**: a concrete use case for building images over the call protocol. The current use cases pull pre-built images, not build them. +- **Priority**: low +- **Full file**: [OQ-49](questions/049-image-build-buildkit-scope.md) + +### OQ-50: Docker System Events Subscription + +- **Blocked on**: a concrete use case for prompt stale-ownership cleanup. The base model (ADR-060 §4) tolerates inert stale entries; the events subscription is a refinement. +- **Priority**: low +- **Full file**: [OQ-50](questions/050-docker-system-events-subscription.md) + +### OQ-51: Container Create Options Surface + +- **Blocked on**: v1 implementation — the `create` input JSON Schema is finalized when `register_docker_ops` is written and tested against bollard's `Config` struct. An architectural decision (ADR-060 §5), not a deferral past implementation. +- **Priority**: medium +- **Full file**: [OQ-51](questions/051-container-create-options-surface.md) + diff --git a/docs/architecture/questions/048-network-and-volume-operation-surface.md b/docs/architecture/questions/048-network-and-volume-operation-surface.md new file mode 100644 index 0000000..ffe685f --- /dev/null +++ b/docs/architecture/questions/048-network-and-volume-operation-surface.md @@ -0,0 +1,30 @@ +# OQ-48: Network and Volume Operation Surface + +- **Origin**: + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) + §"Operation Surface (v1 scope)" (out-of-scope list). +- **Status**: deferred(scope) +- **Door type**: Two-way +- **Priority**: low +- **Blocked on**: a concrete use case for network or volume management + over the call protocol. The two container use cases (disposable dev + containers, hosted services) don't currently require + network/volume CRUD over the call protocol — dev containers use + the default bridge network; hosted services are configured via + `docker compose` with networks/volumes declared in the compose file. +- **Resolution**: Not yet decidable. bollard has the API surface + (`network.rs`: `create_network`, `remove_network`, `inspect_network`, + `list_networks`, `connect_network`, `disconnect_network`; + `volume.rs`: `list_volumes`, `create_volume`, `inspect_volume`, + `remove_volume`, `prune_volumes`) — the mapping is mechanical, the + same shape as the container lifecycle ops (`Query`/`Mutation`, + single `call.responded`). The deferral is scope, not feasibility: + v1 is containers + images; networks and volumes are added when a + use case forces them (e.g., a coordinator that needs to create + isolated networks for dev containers, or a fleet layer that manages + volumes across hosts). Adding them is additive (new operations in + the registry) and does not break the existing surface. +- **Cross-references**: + [ADR-058](decisions/058-alknet-docker-on-alknet-call.md) (the shared + `alknet/call` registration model new ops would follow), + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) \ No newline at end of file diff --git a/docs/architecture/questions/049-image-build-buildkit-scope.md b/docs/architecture/questions/049-image-build-buildkit-scope.md new file mode 100644 index 0000000..0506f9b --- /dev/null +++ b/docs/architecture/questions/049-image-build-buildkit-scope.md @@ -0,0 +1,30 @@ +# OQ-49: Image Build (buildkit) Scope + +- **Origin**: + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) + §"Out of scope for v1"; [ADR-059](decisions/059-bollard-021-dependency-and-features.md) + §3 (no `buildkit` feature). +- **Status**: deferred(scope) +- **Door type**: Two-way +- **Priority**: low +- **Blocked on**: a concrete use case for building images over the + call protocol. The two container use cases (disposable dev + containers, hosted services) pull pre-built images (`docker/image/pull`) + rather than building them. The reverse-proxy and other hosted + services build via `docker compose build` (operator-side), not via + alknet. +- **Resolution**: Not yet decidable. bollard's `build_image` (image.rs:655) + and the `buildkit` feature (which pulls tonic + + bollard-buildkit-proto) are available but deferred. Build is a large + feature (build context upload, layer caching, multi-stage, buildkit + progress streaming) and is not needed for the current scope. When a + use case forces it, the operation is a `Subscription` (progress + events → `call.responded`, build complete → `call.completed`) and + the `buildkit` feature is enabled in `Cargo.toml` (two-way-door + feature addition, per ADR-059). The v1 surface has `image/pull` + + `image/list` + `image/inspect`; `image/build` is added when needed. +- **Cross-references**: + [ADR-059](decisions/059-bollard-021-dependency-and-features.md) + (feature set decision — `buildkit` not enabled), + [crates/docker/overview.md](crates/docker/overview.md) §"bollard + version and features" \ No newline at end of file diff --git a/docs/architecture/questions/050-docker-system-events-subscription.md b/docs/architecture/questions/050-docker-system-events-subscription.md new file mode 100644 index 0000000..384a450 --- /dev/null +++ b/docs/architecture/questions/050-docker-system-events-subscription.md @@ -0,0 +1,34 @@ +# OQ-50: Docker System Events Subscription + +- **Origin**: + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) + §"Out of scope for v1"; + [ADR-060](decisions/060-container-resource-model-and-label-namespace.md) + §4 (autonomous-death tolerance). +- **Status**: deferred(scope) +- **Door type**: Two-way +- **Priority**: low +- **Blocked on**: a concrete use case for prompt stale-ownership + cleanup. The base model (ADR-060 §4) tolerates stale ownership + entries from autonomous container death (a `--rm` exit, external + `docker rm`, daemon restart) — they're inert (a reused container ID + gets a fresh `record` on its next `create`). A reaper that + subscribes to docker daemon events would clean them promptly, but + the promptness gain is marginal for the current use cases. +- **Resolution**: Not yet decidable. bollard's `events()` + (system.rs:128) returns a `Stream` of daemon + events (container start/stop/die/destroy, image pull, etc.). A + `docker/system/events` `Subscription` operation would surface these + as `call.responded` frames; the ownership store could subscribe + internally to revoke on `destroy` events. The deferral is scope: + the base model works without it; the events subscription is a + refinement for when prompt cleanup matters (e.g., a high-churn + coordinator that spawns/removes many containers and wants the + ownership store to stay tight). Adding it is additive (a new + operation + an internal store subscription) and does not break the + existing surface. +- **Cross-references**: + [ADR-060](decisions/060-container-resource-model-and-label-namespace.md) + §4 (teardown coupling — stale-entry tolerance is the base model + this subscription would refine), + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) \ No newline at end of file diff --git a/docs/architecture/questions/051-container-create-options-surface.md b/docs/architecture/questions/051-container-create-options-surface.md new file mode 100644 index 0000000..a03a65a --- /dev/null +++ b/docs/architecture/questions/051-container-create-options-surface.md @@ -0,0 +1,37 @@ +# OQ-51: Container Create Options Surface + +- **Origin**: + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) + §"Out of scope for v1"; + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) + §"`docker/container/create`". +- **Status**: deferred(scope) +- **Door type**: Two-way +- **Priority**: medium +- **Blocked on**: v1 implementation. The v1 `create` input schema + accepts the common fields (image, command, env, labels, name) and + the full `CreateContainerOptions` surface (mounts, port bindings, + networks, volumes, capabilities, etc.) is deferred to the + implementation pass, where the input JSON Schema can be designed + against bollard's `Config` struct and tested against real create + calls. This is not an architectural decision (the `create` operation + is already decided — ADR-060 §5); it's a schema-detail decision + best made with the bollard types in hand. +- **Resolution**: Not yet decidable. bollard's `create_container` + takes a `Config` struct (`container.rs:296`) with ~40 fields + (`Image`, `Cmd`, `Env`, `Labels`, `HostConfig` with mounts/ports/ + networks/etc.). The v1 input schema accepts the high-frequency + fields and omits the long tail; the full surface is a JSON Schema + design task (which fields are required, which are optional, how + `HostConfig` is nested, how the `alknet.*` labels merge with + caller labels). The deferral is to the implementation pass, not + past it — the schema is finalized when `register_docker_ops` is + written and tested. The `OperationSpec.input_schema` is the + one-way surface; its exact field set is a two-way-door refinement + within it. +- **Cross-references**: + [ADR-060](decisions/060-container-resource-model-and-label-namespace.md) + §5 (`create` records ownership — the architectural decision this + OQ defers the schema details of), + [crates/docker/docker-operations.md](crates/docker/docker-operations.md) + §"`docker/container/create`" \ No newline at end of file