31 Commits
Author SHA1 Message Date
glm-5.3-flash d22cecd317 docs: ADR-049/050 in indexes; changelog link refs; ported-ADR range 001..048
- architecture README: index rows for ADR-049 (establishment phase) and
  ADR-050 (pump_bidi); stale ADR-001..045 range labels corrected
- channel-operations.md: design-decision table + references list the
  two new ADRs
- CHANGELOG: missing link references for 0.4.0/0.4.1/0.5.0/0.6.0
- AGENTS.md: ADR range 001..050
2026-09-07 09:03:10 +00:00
glm-5.3-flash d25deb5eaf docs(review 007): mark R-01/R-02/R-03 resolved; record the two sketch deviations
Status → Resolved with the three unit commits; both deviations from
the review's sketches (typed-opaque ChannelPlan over Option<Value>;
(u64, u64) over io::Result) recorded up top with their rationale, as
the ADRs carry them.
2026-09-07 08:48:06 +00:00
glm-5.3-flash 9c6fec17ca feat(review 007 Unit 3): pump_bidi two-pump helper (R-03, ADR-050)
Extracts the two-pump data-plane helper alknet ADR-078 deferred until
the shapes converged (they have: alktunnels POC pump_halves + alktty's
channels session). Purely additive.

- channels::pump::pump_bidi(channel, peer_read, peer_write) -> (u64, u64):
  two joined pumps, shutdown-on-completion per direction; copy counts
  for observability. The channel side is a single AsyncRead +
  AsyncWrite value (the accept_bi BiStream); the peer side takes split
  halves — the establisher's natural dial result (into_split).
- Return (u64, u64), not the review sketch's io::Result<(u64, u64)>:
  both pumps swallow copy errors by contract (mid-stream error =
  abrupt close, no error channel mid-stream per ADR-049 §6), so an
  Err state would be dead code. Deviation recorded in ADR-050.
- alktty's three-pump session does not fit (exit future as a third
  signal) and stays as-is, per the review's scope.
- Tests reproduce the POC's two-pump semantics through the helper:
  bidirectional flow with exact counts, EOF-from-one-side completes
  the other's shutdown (clean EOF at the far end), dead-source =
  EOF-shaped teardown.
- ADR-050 records the decision, deviations, and two-way door type.

Verification: cargo test (625 passed, +2), clippy -D warnings, fmt
--check, doc clean, test --all-features clean, wasm32-unknown-unknown
check clean.
2026-09-07 08:47:46 +00:00
glm-5.3-flash 8f122b0a38 feat(review 007 Unit 2): birth-teardown telemetry — unaccepted-stream debug hint (R-02)
The optional hardening half of R-02 (the doc notes landed with Unit 1,
dc4ad2b). The wrapper's teardown task now observes whether the pump
handler ever accepted the channel's stream:

- ChannelBidiStreamSource gains a shared acceptance flag
  (with_accepted_flag / accepted()); channel_source_with_accepted_flag
  threads it from run_open_wrapper.
- On handler exit without accept, the teardown task logs a debug!
  naming the contract: "the returned JoinHandle must track the
  data-plane lifetime ... await pumps inline" — the telemetry hint for
  the teardown-at-birth shape the alktunnels POC hit empirically.
- Telemetry only: no behavior change; the benign no-data close still
  tears down identically.

Verification: cargo test (623 passed, +2: accepted-flag flip on the
yield-once accept; accepting vs non-accepting handler teardown
behavior), clippy -D warnings, fmt.
2026-09-07 08:45:09 +00:00
glm-5.3-flash dc4ad2bc6d feat(review 007 Unit 1): Establishment carries the channel plan (R-01) + lifetime doc (R-02)
Implements ADR-049 amendment 2 — the reserved Establishment payload is
filled, and the OpenHandler lifetime contract is documented.

- Establishment { plan: Option<ChannelPlan> } with ChannelPlan =
  Arc<dyn Any + Send + Sync>: typed-opaque, because the payload an
  establisher hands the pump handler is a live handle (dialed socket,
  TTY handle), not JSON — the review's Option<Value> sketch could not
  satisfy its own verification gate. #[non_exhaustive] keeps a future
  carrier change from being another break. Construction:
  Establishment::new(plan) / Establishment::default().
- OpenHandler gains the plan parameter:
  Fn(Value, Option<ChannelPlan>, Connection, AuthContext) ->
  JoinHandle<()>. Separate parameter (not merged into input) — a
  typed payload cannot ride the JSON input; no schema collision.
  Wire surface unchanged: the plan is process-local (establisher ->
  wrapper -> handler).
- run_open_wrapper threads establishment.plan to the handler; None
  when no establisher is registered. Kills the alktunnels-POC
  side-channel handoff (resource-keyed slot + poll loop) whose
  concurrent same-resource race is now unreachable — each open's
  establisher result flows to its own handler.
- Lifetime contract documented (R-02, doc-only half): the returned
  JoinHandle must track the data-plane lifetime — the wrapper awaits
  it and its completion triggers teardown; early return = teardown
  at birth. Noted on the OpenHandler type docs and both registration
  entry points.
- Breaking at 0.6.0 (the point of landing it before alktunnels
  Phase 1): Ok(Establishment {}) sites become
  Ok(Establishment::default()) mechanically.

Verification: cargo test (621 passed, +4: plan-flows-to-handler,
concurrent same-resource opens get distinct plans, no-establisher
None plan, Establishment construction), clippy -D warnings, fmt
--check, doc clean, test --all-features clean.
2026-09-07 08:43:17 +00:00
glm-5.3-flash 6590ab005f docs: review 007 — establishment follow-ups (from the alktunnels UDP POC)
Findings filed to prevent a second fix->publish->update-dependents
cycle; each was reached by building working code against 0.5.0:

- R-01 [major] — Establishment is payloadless but the channel plan
  is exactly what establishers need to hand to the pump handler
  (ADR-049's own 'reserved for a channel plan' note). Costs verified
  in two consumers: alktunnels POC side-channel handoff (same-resource
  opens race the slot), alktty forced to keep backend allocate
  post-open in-band (allocate_failed stays an in-band frame — the
  shape ADR-049 eliminates, alive one layer down). Ask: fill the
  reserved field (plan: Option<Value>, process-local, wire unchanged)
  in a 0.6.0 sweep; the break is mechanical (Establishment::default).
- R-02 [minor] — the OpenHandler JoinHandle lifetime contract is
  undocumented and load-bearing: the wrapper's await of the returned
  handle IS the teardown trigger; a handler that returns before its
  pumps finish tears the channel down at birth (the POC found this
  empirically — every tunnel EOF'd instantly). Doc note + ADR-049
  amendment; optional debug warning.
- R-03 [minor, optional] — the ADR-078 two-pump helper's convergence
  test is satisfied (POC gives both shapes); extracting pump_bidi
  now is additive (no break) and pins the contract upstream. Decide
  in the same 0.6 sweep.

Non-findings: reverse-flow (-R) needs no upstream mechanism
(from_connection_with_serving + serving-side allocation, verified by
trace); establisher receives registry-validated input; typed
establishment errors complete on the wire; early-arrival cap
unchanged; EstablishmentError reason set sufficient.

Goal stated in the review: 0.6 is the last breaking sweep forced by
known work. alkcall tests: 617 passed (docs-only change; baseline
check).
2026-09-07 08:18:11 +00:00
glm-5.3-flash 36e74cda11 feat(review 006 Unit 3): teardown-race log, early-arrival bound docs, count accessor (E-03, E-04, N-2)
- E-03: the open wrapper's handler-exit teardown no longer discards
  UnknownChannel silently — debug log + benign-race pinning comment
  (ledger take is the atomic gate; no double-decrement)
- E-04: consumer-facing doc note on the 64-parked-chunks observable
  bound (channel-client.md + EARLY_ARRIVAL_CAP const doc) for
  tunnel-style push-first producers
- N-2: ChannelManager::early_arrival_count() accessor (the
  observability choice over removing the write-only counter),
  documented monotonic, with a park/adopt-drain monotonicity test
- review 006: Unit 3 marked implemented in Status and remediation plan

Verification: 617 tests pass; clippy -D warnings clean (host +
wasm32 check); fmt clean; doc clean
2026-09-06 19:35:18 +00:00
glm-5.3-flash f8dad9dbc8 feat(review 006 Unit 2): additive OperationSpec.description disclosed via discovery (E-02)
- OperationSpec gains description: Option<String> (builder
  with_description, defaults None; no struct-literal construction
  sites exist, so additive by construction)
- spec_to_json_pub emits description when set; rebuild_spec_for
  parses it back — the field survives from_call discovery and
  op/register announcement (same round-trip pattern as
  resource_id_path / publish_schema)
- services/list and the local-ops half of services/list-peers emit
  description when set; output-schema docs on both listing specs and
  operation_spec_schema advertise the field
- Tests: builder/default, emit/omit, listing emission, schema
  disclosure, schema-doc presence, round-trip + absent-stays-absent
  (8 new; 616 total)
- Docs: review 006 Unit 2 marked IMPLEMENTED; ADR-047 §6 amendment
  records the E-02 discovery decision (listing enrichment lands, the
  channel/resources/subscribe half stays deferred); OQ-40 gains the
  load-bearing note; operation-registry.md struct + listing docs;
  CHANGELOG

Verification: cargo test (616 pass), clippy -D warnings (host +
wasm32), fmt --check, cargo doc --no-deps, wasm32 check — all clean
2026-09-06 19:26:51 +00:00
glm-5.3-flash 2586c3b217 feat(review 006 Unit 1): channel-open establishment phase + typed client error (E-01, N-1)
Implements ADR-049 Unit 1 — the open-op wrapper gains an awaited,
bounded establishment phase, and the client stops erasing the error.

- OpenEstablisher hook + Establishment/EstablishmentError types:
  register_openable_with_establisher awaits the establisher bounded
  (earlier of dispatch deadline and per-registration timeout, else
  ESTABLISHMENT_TIMEOUT = 10s) after allocation, before the reply and
  before the pump handler is spawned (ADR-049 §1/§2). Implementation
  note: the establisher takes (input, auth) only — the channel's
  yield-once BiStream belongs exclusively to the pump handler
  (amendment recorded in ADR-049).
- Establishment failure: teardown_channel + opener-ledger take +
  policy.on_close un-increment (allocation and teardown balance;
  the ledger take is the atomic gate, ADR-047 §7), reply
  channel:open_failed with details {reason, message} — reason ∈
  dial_failed / unknown_resource / resource_shortage / handler_error
  / timeout (ADR-049 §3). SSH contract consumer-visible: a failed
  open never returns a channel_id.
- register_openable unchanged (no establisher = always-OK; existing
  registrations compile and behave identically — compat gate test).
- ChannelClient::open_channel returns ChannelOpenError (breaking at
  0.5.0): CallFailed { error: CallError } carries the wire error
  verbatim (establishment_reason() branches on details.reason);
  MissingChannelId / AdoptFailed cover the local-only shapes
  (ADR-049 §4, review 006 N-1).
- Tests cover all four verification gates from the review: e2e
  establisher failure through a real channels connection (typed
  reason + no-channel + ledger un-increment), bounded timeout,
  no-establisher compat, establisher-success pump round-trip; plus
  reason-vocabulary mapping and bound arithmetic.
- Bump to 0.5.0 (open_channel error-type change is semver-relevant).

Verification: cargo test (608 passed), clippy --all-targets -D
warnings, fmt --check, doc --no-deps, wasm32 check — all clean.
2026-09-06 19:00:04 +00:00
glm-5.3-flash 48ceeba55c docs: ADR-049 channel-open establishment phase; verify review 006
Verify review 006's findings against source at 88e3f5e (E-01..E-04
all confirmed; E-02 cost corrected — OperationSpec has no description
field, four touchpoints) and file three additional findings from the
same sweep (N-1 client error-type gap, N-2 write-only early-arrival
counter, N-3 pump-panic posture).

ADR-049 resolves E-01 + N-1: split-hook OpenEstablisher awaited
bounded by the open-op wrapper (restoring ADR-047 §3's "channel
plan" shape), teardown + typed channel:open_failed reply on
establishment failure, ChannelClient::open_channel typed error.
Review 006 gains the post-verification remediation plan and verdict
appendix.

Verification: cargo test (597 passed), cargo doc --no-deps clean.
2026-09-06 10:57:20 +00:00
glm-5.3-flash 88e3f5e9c3 docs: review 006 — channel-open establishment gap (from alktunnels phase 0)
Design review from the alktunnels Phase 0 research pass, verified
against tree a22b2b8 (0.4.1). Findings numbered E-01..E-04:

- E-01 [major] — the open op cannot fail after allocation: the
  wrapper replies {channel_id} the moment the OpenHandler is spawned;
  establishment failures (params-valid-but-rejected, backend lookup
  failure, target dial failure) present to the consumer as a
  successful open followed by an instant, indistinguishable clean
  EOF (implicit-EOF mux path + unified poll_read EOF arms). SSH
  semantics (RFC 4254 §5.1 open-failure reply with reason codes;
  channel never exists opener-side), SOCKS5 reply codes, and
  udpgw's opaque ERR bit (counterexample) surveyed in
  alktunnels/docs/research/ssh-socks5-survey.md. alktty's in-band
  error-frame mechanism (send_negotiation_error, 0x00-peek) is the
  per-crate workaround this upstream establisher obsoletes for the
  channels path. Proposed shape: an awaited establishment hook
  (OpenEstablisher) or await-and-inspect OpenHandler, tearing down on
  failure and replying channel:open_failed with SSH-four reason codes
  in ADR-016 details. Remediation sketch + verification gates
  included.
- E-02 [minor] — services/list discloses no per-op metadata; OQ-40
  (channel/resources/subscribe) becomes load-bearing for the first
  time via the alktunnels discovery resolution (OQ-TN-08).
- E-03 [minor] — OpenHandler-exit vs channel/close teardown race is
  benign (ledger take is the gate) but the let _ = discard at
  operations.rs:509 is silent; recommend log-or-comment.
- E-04 [minor] — early-arrival park cap (64) is an observable bound
  for push-first producers under slow adopters; no change requested,
  filed so the constraint is visible to the next consumer.

Non-findings recorded: open-op ACL path complete across all three
dispatch entry points; input_schema enforcement covers open params;
EOF arms unified; channel_open marker + resource_id_path wire
round-trip intact; opener-ledger decrement atomic at every call site.

alkcall tests: 597 passed (docs-only change; baseline check).
2026-09-06 09:54:13 +00:00
glm-5.3-flash a22b2b84c9 fix: park early-arrival chunks for un-adopted channels (open/first-data race)
The connect side adopts a channel (installs local routing state) only
after the open-op response arrives, but the accept side's OpenHandler
can start pumping data the moment the channel opens — the two race and
the demux's lenient unknown-channel drop (REQ-CH-04) silently lost the
producer's first chunks (a TTY backend's banner, a sub protocol's
greeting).

route_payload now parks up to 64 payloads per unknown channel_id in a
bounded early-arrival buffer; adopt_channel drains them into the new
receiver in order. Beyond the cap the chunk drops with the existing
debug log + dropped_unknown_chunks counter (which now also counts
overflow). clear_all drops parked buffers with the connection.

Surfaced by alktty's consumer end-to-end test (review #001 L3): the
session never resolved because the producer's first chunks (stdout
sentinel + exit chunk for an immediately-resolving backend) arrived
before the adopt and were dropped. REQ-CH-04 wording updated by this
behavior; ADR-039 §demux loop describes the lenient drop for genuinely
unknown channels, which remains the case past the cap.

Verification: cargo test 597 (2 rewritten for the new semantics +
route_payload_to_unknown_channel_parks_until_adopt gate);
--all-features 614; clippy -D warnings clean; fmt clean; doc 0
warnings; publish --dry-run ok; standalone probe (handler-writes-first
e2e over one connection) shows 0 dropped chunks with the fix vs 1
without.
2026-09-05 07:04:30 +00:00
glm-5.3-flash 574f58442a feat: enforce input_schema at call time (ADR-016 INVALID_INPUT leg)
- OperationSpec.input_schema was advertise-only: services/schema
  disclosed it but no dispatch entry point consulted it (the only
  enforced schema was publish_schema per-chunk on Pub ops, P-03).
- compile input_schema once at registration, same fail-closed rule as
  publish_schema/CF-003: an un-compilable schema is a registration
  error, never a silently-skipped contract. Validator cache mirrored
  on fork and in OperationRegistryBuilder like the publish validators.
- check after the ACL gate in all three dispatch entry points:
  invoke, invoke_streaming, invoke_sink (via resolve_sink_handler,
  preserving the P-08 single-source-of-truth property). Violations
  return INVALID_INPUT with the input echoed in details.
- raw-JSON-Schema semantics (permissive on unknown keys); adapters
  wanting closed-by-default keep their own hardening (alkhttp's
  CompiledInputSchema composes unchanged).
- motivated by alktty review #001 L1: the channels open-op wrapper
  hands the registry-checked input to the OpenHandler as the
  authoritative params, which requires the registry to validate it.

Verification: cargo test 596 lib (6 new: invoke/streaming/sink
enforcement, fail-closed registration, permissive-{} compile,
fork-carries-validator); --all-features 613; clippy -D warnings
clean; fmt clean; doc 0 warnings; publish --dry-run ok. alkhttp
438+16 tests pass against 0.3.1 (registry) — re-verify against the
published 0.4.0 after upload.
2026-09-05 06:33:00 +00:00
glm-5.3-flash ba94c70eb2 chore: bump to 0.3.1, changelog for the UP-03 list-peers fix
Verification: 590 default / 607 all-features tests, clippy
(all-targets, all-features, wasm32) clean, fmt clean, publish
dry-run OK.
2026-09-04 16:01:56 +00:00
glm-5.3-flash fd212307e2 fix(up-03): PeerCompositeEnv::peer_operations override — list-peers sees peer-announced ops
Surfaced by alkhttp review 006 (UP-03): services/list-peers showed
every peer with an empty operations array. PeerCompositeEnv overrode
peer_ids only, so peer_operations fell to the trait default
(Vec::new()) and the ADR-022 amendment's "announced op is discoverable
via services/list-peers" promise never resolved on the wire. ADR-030
prescribed the fix but it had never been ported into alkcall. The
existing list-peers unit tests passed because they mock
peer_operations with hand-rolled envs.

Implements ADR-030 as specified:
- OperationEnv gains list_operation_names (default Vec::new(),
  back-compat for all existing implementors)
- OverlayOperationEnv overrides it with its overlay's registered names
- PeerCompositeEnv::peer_operations delegates to the peer overlay's
  list_operation_names; PeerCompositeEnv::list_operation_names
  aggregates session + connections + base (mirrors its contains())
- LocalOperationEnv enumerates its registry; ChannelsSessionEnv
  delegates to base

Gate: announced_op_is_discoverable_via_services_list_peers in
src/registry/op_register.rs — announces an op through op/register,
then asserts both the direct peer_operations probe and the
services/list-peers wire shape attribute the announced op to the peer,
over the exact compose_root_env shape (PeerCompositeEnv + attached
connection overlay). Verified load-bearing: reverting the
peer_operations override fails the gate.

ADR-030 status Proposed -> Accepted with the UP-03 provenance note.

Verification: 590 default / 607 all-features tests, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean,
semver-checks 196 pass against v0.3.0 (defaulted trait method is
non-breaking).
2026-09-04 16:00:23 +00:00
glm-5.3-flash 1e20bb77d0 chore: prepublish review — bump to 0.3.0, changelog, doc alignment
0.2.0 is already on crates.io (2026-08-31, c16b069); the review
004/005 remediation work is unreleased on top of it and lands as
0.3.0.

- Bump version 0.2.0 -> 0.3.0. Two OperationRegistry methods changed
  borrowed returns to owned (registration, list_operations) —
  source-breaking for annotated call sites, minor bump per 0.x
  semver rules. cargo semver-checks passes (196 checks) against the
  published baseline; the return-type changes were caught by manual
  diff review.
- CHANGELOG 0.3.0: connect-side serving (from_connection_with_serving
  + ServingConfig), OperationRegistry::fork + builder from_registry,
  registry::op_register (bootstrap op, collision policy,
  ALREADY_EXISTS), install_bootstrap_discovery, spec_to_json_pub +
  resource_id_path round-trip, overlay accessors, concurrent serving
  loops, &self registration.
- README: serving-as-consumer section, consumer role table update,
  drop the stale `mut` on the registry example.
- AGENTS.md: ADR range 001..047 -> 001..048.
- Fix rustdoc private-intra-doc-link warning on StartedDispatch.

Verification: 589 default / 606 all-features tests, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean,
publish dry-run OK.
2026-09-04 14:14:22 +00:00
glm-5.3-flash d5b2661b38 fix(review 005 Unit 3): resource_id_path wire round-trip + bootstrap-list doc alignment (G-04, G-05)
- resource_id_path rides both halves of the spec wire round-trip:
  spec_to_json_pub serializes it (optional string key), rebuild_spec_for
  parses it. Additive optional field - absent stays absent. Previously
  an announced (or from_call-imported) op declaring ownership-scoped
  resource extraction silently rebuilt with resource_id: None, so ACL
  checks ran without the resource ID.
- Gates: spec_round_trips_resource_id_path (serialize -> parse ->
  field intact) + spec_without_resource_id_path_stays_absent (additive
  field breaks no consumer).
- ADR-022 amendment: bootstrap-op set gains services/list-peers with a
  dated G-05 note (the installer has registered it since the amendment
  landed; the doc lagged the code). Set remains closed at four.

Verification: cargo test 589 / --all-features 606, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean.

Refs docs/reviews/005-...md (G-04, G-05; all findings closed).
2026-09-04 09:41:43 +00:00
glm-5.3-flash 23c9b28c6b fix(review 005 Unit 2): op/register serving-registry collision gate (G-03)
- op_register_handler takes the serving registry alongside the
  connection and rejects announced names that collide with the serving
  side's own registrations (ALREADY_EXISTS regardless of replace).
  Peer-announced ops may collide with peer-announced ops (replace
  governs, the reconnect path) but never shadow the deployment's own
  ops: the connection overlay resolves before base in PeerCompositeEnv,
  so an unscreened same-name announce would silently rewrite what a
  wire-dispatched handler's ctx.env.invoke resolves. Composition
  authority (ADR-018) stays with the deployer.
- ADR-022 amendment (2026-09-04): collision policy recorded in the
  2026-09-03 amendment's op/register section (rationale + visibility
  irrelevance); status line notes the sub-amendment.
- Gates: base-External collision rejected even with replace (overlay
  stays clean, serving registration untouched); Internal base op
  equally protected; overlay/overlay collisions still follow replace;
  nested composition of a base op resolves the serving side's own op
  after an unrelated announce (real compose_root_env env shape).
- PeerCompositeEnv resolution order deliberately unchanged.

Verification: cargo test 587 / --all-features 604, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean.

Refs docs/reviews/005-...md (G-03; Unit 3 open).
2026-09-04 09:40:11 +00:00
glm-5.3-flash 1cbb7c6536 fix(review 005 Unit 1): concurrent serving loops + stub-exercising gates (G-01, G-02)
- Split dispatch() into dispatch_start() (sync prefix) + spawned
  invocation: both single-stream loops (serve_single_stream and the
  accept-side run_loop_single_stream) spawn Once invocations, Sub
  pumps, and sink response writers; only the Pub sink start stays
  inline (chunk_tx must register before the next call.published).
  Inline dispatch deadlocked same-connection nested composition: the
  read loop awaited the parent handler, which awaited a nested call
  whose response only the same read loop could resolve (resolved only
  via the 30s sweeper). Spawned handles tracked + aborted at loop exit;
  in_flight_sinks behind an Arc<parking_lot::Mutex> with guards dropped
  before awaits.
- run_loop_single_stream gains the pending-resolution arms
  (RESPONDED/COMPLETED/ERROR): the accept side previously served only
  and had no loop resolving its own outbound pendings in single-stream
  mode — the latent accept-side imported-op composition hazard is
  mechanized shut.
- Write-failure in the spawned Once path warns instead of closing the
  loop (matches the Sink arm; dying transport still surfaces via
  ConnectionClosed on the next read).
- G-02 gate: hub_handler_composes_peer_announced_op_via_nested_composition
  — announce -> consumer calls hub/compose -> hub's serving loop
  wire-dispatches it -> handler composes via ctx.env -> forwarding
  stub's nested call crosses back to the consumer. The F-05 gate
  bypassed this path entirely.
- Interleaved-directions gate: outbound_call_resolves_while_inbound_
  subscription_is_being_served — consumer serves a live Sub while a
  wire-dispatched hub handler issues an outbound call on the same
  connection.
- Both gates verified load-bearing: run against the pre-fix loop each
  reproduces the G-01 hang (no progress, bounded-timeout failure);
  post-fix both resolve in <0.2s, no sweeper evictions.

Verification: cargo test 583 / --all-features 600, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean.

Refs docs/reviews/005-...md (G-01, G-02; Units 2-3 open).
2026-09-04 09:36:08 +00:00
glm-5.3-flash 435ae9da2f docs(review 005): serving-loop concurrency + op/register composition findings
Post-remediation review of f84d214 (review 004 Units 1-3). Five
findings, verified in source and (for G-01) empirically via a probe
test that was added, run, and removed:

- G-01 [major]: serve_single_stream awaits dispatch inline; a
  wire-dispatched handler composing a peer-announced op (or a
  from_call import) over the same connection deadlocks — the nested
  call resolves only via the 30s sweeper (probe: TIMEOUT at 30.0007s).
- G-02 [major]: the F-05 e2e gate calls the announced op directly,
  bypassing the forwarding stub — the one path G-01 breaks.
- G-03 [major]: op/register's collision gate is overlay-only;
  PeerCompositeEnv resolves connections before base, so an announced
  op can shadow the serving side's own ops in nested composition.
- G-04 [minor]: resource_id_path does not survive the spec wire
  round-trip (pre-existing shape, load-bearing for op/register).
- G-05 [minor]: install_bootstrap_discovery registers
  services/list-peers; ADR-022's bootstrap set doesn't name it.

Non-findings bound the re-review: fork surface lock discipline,
bootstrap discovery closure, frame-arm equivalence of the composed
loop, unchanged pure-consumer default, alkhttp cross-repo claims,
CJK sweep (none), all gates reproduce (581/598, clippy, fmt, wasm,
doc).

Remediation plan: Unit 1 (concurrent serving loop + stub-exercising
gate) gates Unit 4 downstream; Unit 2 (collision policy); Unit 3
(round-trip completeness + doc alignment).

Verification: cargo doc --no-deps clean; tree unchanged apart from
this review doc.
2026-09-04 07:25:41 +00:00
glm-5.3-flash f84d214173 feat: per-session fork registry, connect-side serving loop, op/register (review 004 Units 1-3)
Remediates all six findings of review 004 (per-connection dispatch
resolution and client-side op serving). All claims re-verified in
source before remediation; F-02's member list gains ScopedPeerEnv
(also Clone — fork surface simpler than estimated).

- OperationRegistry: interior mutability (parking_lot RwLock on both
  maps); register takes &self; registration/list_operations return
  owned clones; fork() deep-copies registrations + cached publish-schema
  validators (F-02/F-03); OperationRegistryBuilder::from_registry.
- install_bootstrap_discovery: services/list, services/list-peers,
  services/schema registered closed over the fork itself, so
  per-session openables are discoverable and services/schema answers
  from the fork (F-06).
- Dispatcher::serve_single_stream: full-duplex single-stream loop —
  call.requested dispatches inbound; responded/completed/error resolve
  outbound pendings; aborted tries both tables (in-flight sink aborts
  + pending cascade); published routes inbound sinks (F-04).
- ChannelClient::from_connection_with_serving(connection,
  Option<ServingConfig>): opt-in serving; from_connection keeps the
  pure-consumer default.
- registry::op_register: OpRegisterRequest wire DTO (spec in
  services/schema JSON + replace flag), op_register_spec,
  op_register_handler (rebuild -> forwarding stub -> register_imported,
  forced Internal/FromCall), announce_op; CallError::already_exists;
  spec_to_json_pub; from_call's rebuild_spec_for + forwarding-handler
  constructors crate-shared (F-05).
- ADR-047 §4 amendment #2: per-session fork is the dispatch-registry
  mechanism; overlay stays nested-invocation/peer-announced landing
  zone (F-01/F-02).
- ADR-022 amendment 2026-09-03: bootstrap-op set (services/list,
  services/schema, op/register), opt-in connect-side serving, op/register
  wire shape (F-04/F-05).
- alkhttp ADR-048 reconciliation note + OQ-05 re-pointed at the alkcall
  ADRs (Unit 1b).
- Review 004 status -> remediated; remediation log with gates.

Verification:
- cargo test: 581 passed, 0 failed (565 baseline + 16 new)
- cargo test --all-features: 598 passed, 0 failed
- cargo clippy --all-targets -- -D warnings: clean
- cargo clippy --all-features --all-targets -- -D warnings: clean
- cargo clippy --target wasm32-unknown-unknown -- -D warnings: clean
- cargo fmt --check: clean
- cargo doc --no-deps: clean

Gates: fork_registry_open_op_resolves_and_is_discoverable (open op via
fork + services/list shows openable + services/schema validates),
serving_loop_hub_to_consumer_call_resolves (hub->consumer call through
consumer's serving loop, consumer->hub still resolves),
op_register_announce_then_hub_call_routes_back_to_consumer (announce ->
overlay -> hub call -> forwarding stub -> consumer serves).
2026-09-03 17:32:45 +00:00
glm-5.3-flash c0dbf82518 docs(review 004): per-connection dispatch resolution + client-side op serving
Focused design-mismatch review found via alkhttp's WS data-channel
drill-down (alkhttp review 003 WS-24/WS-25). The channel machinery is
done and proven; the gaps are the dispatch-resolution mechanism and
the connect-side serving half — both upstream of any transport.

Findings:
- F-01 [major]: top-level dispatch consults only the dispatcher's
  base registry — ops registered per the ADR-047 §4 amendment's
  overlay mechanism resolve NOT_FOUND on the wire; the only proven
  shape (per-session registry as the dispatcher's base) differs from
  the ADR's wording
- F-02 [major]: no Clone/fork surface on OperationRegistry or
  HandlerRegistration — every inner payload type IS Clone-able
  (verified type-by-type, incl. jsonschema::Validator and
  Capabilities), so the fork is a small addition
- F-03 [minor]: fork must carry handlers + validators, not just specs
- F-04 [major]: connect-side channel-0 read pump resolves responses
  only; inbound call.requested frames are silently dropped — no
  serving half on the single-stream shape (ADR-022/AGENTS §8
  bidirectionality unreachable from the connect side)
- F-05 [major]: no wire mechanism announces client-side ops; the
  six call.* kinds are closed. Resolution candidate: bootstrap op
  (op/register) served per-session, handler writes into the
  connection-local overlay; discovery rides services/list-peers
- F-06 [minor]: per-session fork must carry bootstrap discovery ops
  for per-session openables to be discoverable

Includes a non-findings section (e2e reference shape,
register_openable completeness, channel-id split, envelope-kind
closure, wasm-cleanliness) and a 4-unit plan: ADR decisions (Unit 1)
-> fork surface (Unit 2) -> client serving + bootstrap op (Unit 3)
-> alkhttp wiring downstream (Unit 4, tracked in alkhttp review 003).

Verification: cargo test (565), clippy --all-targets -D warnings,
fmt, doc --no-deps — all clean at c16b069. No source changes.
2026-09-03 14:42:57 +00:00
glm-5.3-flash c16b0697e3 chore: drop dead Cargo.lock from package exclude list
cargo always includes a git-tracked lockfile in the package regardless
of the exclude entry; keeping it listed implied a behavior that does not
exist.
2026-08-31 10:01:48 +00:00
glm-5.3-flash c3d4fa1b30 chore: bump to 0.2.0
First release carrying the consumer-findings remediation (CF-001..004),
the feature-gated gateway dispatch spine (ADR-048), and the
registration-time publish_schema validation behavior change (CF-003).
Gate is version-only for existing consumers: the public 0.1.1 API
surface is unchanged (probe-verified).
2026-08-31 09:56:15 +00:00
glm-5.3-flash ae372c0c7e test+docs: prepublish hardening from review (failure paths, doc hygiene)
Coverage:
- CF-002: demux skipped-bytes budget teardown test (268 MiB skip in-memory;
  a budgetless demux wedges in the 17th skip, the real one tears down) and
  the skip-hits-EOF arm (truncated oversized payload ends the loop).
- CF-001: retryable CONNECTION_CLOSED pinned for subscribe write failures
  in both stream modes and single-stream publish request-frame failures;
  non-retryable INTERNAL pinned for mid-publish failures in both modes
  (deterministic FailOnFlushN write half).
- Gateway: invoke_sink with with_deadline(None) completes a slow sink.

Docs:
- Fix broken intra-doc link on lib.rs's feature-gated gateway mention
  (rustdoc warned on default-feature builds).
- Re-point 40 src/ references from the old alknet mono-repo ADR numbering
  (049/050/052/065/070/074/092) to this crate's numbering
  (021/011/034/007/008/009/005); drop into_sub_streams references
  removed by ADR-035.

Verification: 565 default / 582 all-features (8 new), clippy -D warnings
on default/gateway/all-features/wasm32, fmt clean, rustdoc warning-free
on default and all-features, publish dry-run clean. Consumer-facing API
continuity 0.1.1 -> 0.2.0 verified by compiling an API-surface probe
against both versions.
2026-08-31 09:56:07 +00:00
glm-5.3-flash d5fd548b8d feat: promote dispatch spine to gateway module (ADR-048, feature-gated)
Promote alkhttp's transport-neutral dispatch spine into alkcall as
alkcall::gateway behind the opt-in gateway cargo feature (default off;
adds no dependencies):

- GatewayDispatch: deadline-bounded invoke spine over OperationRegistry
  (invoke / invoke_streaming / invoke_sink) with the root-context
  discipline (internal: false, forwarded_for: None) hubs and spokes
  relaying calls (ADR-042 translate path) need identically to alkhttp's
  HTTP gateway. The 30 s deadline becomes a constructor knob
  (with_deadline).
- schema_disclosure_denial: the shared is-internal + ACL check for
  services/schema inner-name disclosure; ACL denial returns FORBIDDEN
  (identity-aware refinement), Internal visibility returns spec-404.
  One implementation so transports cannot drift (CF-004).
- MAX_BATCH_OPERATIONS / CallRequest / HTTP error mapping stay in
  alkhttp (projection + transport concerns); alkhttp migrates to this
  module in a follow-up session and drops its local copy.

Docs: ADR-048 (decision + divergence rationale), ADR index entry,
CHANGELOG.

Verification: 574 tests pass with --features gateway (16 new), 558 pass
default, clippy -D warnings clean both feature sets, --all-features
clean, fmt clean, wasm32 target clean, rustdoc warning-free.
2026-08-31 08:45:57 +00:00
glm-5.3-flash 8cb2a6eb6d fix: remediate consumer findings CF-001..004 (alkhttp ledger)
- CF-004: services_schema_handler now applies the same visibility +
  AccessControl gates as invoke() (identity resolution mirrors invoke:
  handler_identity under internal). Restricted ops return spec-404 NOT_FOUND
  — matches "restricted ops don't exist" and leaks nothing about the
  restricted surface. Closes the unauthenticated /call-path disclosure.
- CF-003: publish_schema compiled at registration time (both
  OperationRegistry::register and OperationRegistryBuilder::store);
  un-compilable schemas are a registration error — an unvalidated ingest
  path can no longer be constructed. Compiled validator cached per-op
  (publish_validator) and consumed by dispatch; per-request compile gone.
  BEHAVIOR CHANGE: register/builder reject un-compilable publish_schema.
- CF-002: demux TooLarge skip streams through a fixed 64 KiB buffer
  instead of allocating the peer-declared length (u32, up to ~4 GiB);
  cumulative 256 MiB skipped-bytes budget tears down dribbling peers.
  Existing resync test passes unchanged.
- CF-001: new retryable CallError::connection_closed (CONNECTION_CLOSED)
  applied only where the call is provably undelivered — request-frame
  write failures on all consumer paths (call/subscribe/publish, both
  stream modes; publish pump tags write stages). Mid-publish failures and
  producer-side fail_all stay non-retryable INTERNAL (delivery ambiguous).
  New code string is additive; retryable flag is the machine-readable
  signal.

Verification: cargo test (558 pass, 15 new), clippy --all-targets -D
warnings, fmt --check, wasm32-unknown-unknown check.
2026-08-31 08:22:55 +00:00
glm-5.3-flash a2d72f9737 docs(ledger): file CF-004 — services_schema_handler discloses Internal/ACL-restricted op specs
Found via alkhttp Review 002 (PRJ-16): the services/schema handler
does a bare registry.registration(name) with no Visibility and no
AccessControl check, so POST /call (and the MCP call tool) can fetch
any Internal op's complete spec unauthenticated. The GET /schema route
and MCP schema tool enforce the pre-checks; the /call path is the
hole.
2026-08-30 10:50:28 +00:00
glm-5.3-flash 1f08f1e385 docs(ledger): file CF-003 — wire-path publish_schema compile failure is fail-open
Found while fixing alkhttp's HTTP-side instance
(review-001-publish-schema-validation-robust): the identical
warn-and-skip pattern exists at src/protocol/dispatch.rs:352-366 — a
compile failure of a Pub op's publish_schema proceeds with
validator: None, so arbitrary unvalidated JSON reaches the sink
handler over the wire. Suggested direction: fail-closed + per-
registration validator cache (worked shape in alkhttp 1572a9d,
src/gateway/schema_cache.rs) + registration-time schema compilation.
Single compile site verified (both pump arms consume the one
InFlightSink.publish_validator).
2026-08-30 08:33:29 +00:00
glm-5.3-flash 84fe94c4d7 docs(ledger): file CF-002 — demux TooLarge skip allocates peer-declared length
Cross-crate finding from alkhttp Review 001 (WS-12), filed here per the
ledger's purpose (consumer-surfaced alkcall findings).

src/channels/adapter.rs:144: the ChunkError::TooLarge arm allocates
vec![0u8; length] from the peer's untrusted 8-byte header before
reading; length is u32, so ~4 GiB can be pinned per connection and held
indefinitely by a dribbling peer. Reachable via the alkhttp WS path by
any authenticated browser. Skip/resync logic is correct; the memory
shape is wrong — stream-skip with a bounded buffer instead.

The normal payload arm (:159) is safe (parse_header bounds it at
MAX_CHUNK_LEN); only the TooLarge arm is unbounded.
2026-08-30 06:16:30 +00:00
glm-5.3-flash 4c99b877e3 docs(reviews): consumer-findings ledger for alkhttp-as-consumer findings (CF-001 write-failure retryability) 2026-08-29 09:43:17 +00:00
49 changed files with 10520 additions and 394 deletions

No files matched your search

+4 -2
View File
@@ -215,13 +215,15 @@ cargo semver-checks check-release # before a release
## Architecture Context
- `docs/architecture/` — the authoritative spec. Read it before
non-trivial changes. ADRs are numbered 001..047; OQs (open questions)
non-trivial changes. ADRs are numbered 001..050; OQs (open questions)
track resolved/deferred decisions.
- This crate unifies `alknet-call` and `alknet-channels` from the
alknet mono-repo (`/workspace/@alkdev/alknet`). The source
architecture docs were ported from
`/workspace/@alkdev/alknet/docs/architecture/` and renumbered as
alkcall ADRs (001..047). The ALPN strings (`alk/call`,
alkcall ADRs (001..048; ADR-049 and later were authored in this
crate). The ALPN
strings (`alk/call`,
`alk/channels`) are wire-stable going forward (renamed from
`alknet/` to `alk/` in v0.1.1, before the first published consumer).
- Key ADRs that inform this crate's design:
+377
View File
@@ -4,6 +4,376 @@ All notable changes to this crate are documented here. The format is
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this crate adheres to [Semantic Versioning](https://semver.org/).
## [0.6.0] - 2026-09-07
The establishment follow-ups sweep (review 007): the `Establishment`
plan payload lands (R-01), the `OpenHandler` lifetime contract is
documented (R-02), and the two-pump helper is extracted (R-03). One
breaking change (below); the wire surface is unchanged.
### Added
- **`pump_bidi` (review 007 R-03 / ADR-050).** The two-pump helper —
`channels::pump_bidi(channel, peer_read, peer_write)` — pumps a
channel stream (`BiStream`-shaped) against a peer's split
read/write halves: one pump per direction, each shutting down the
opposite sink on completion (alknet ADR-078's shutdown-on-
completion), joined, returning `(u64, u64)` copy counts. Errors
are EOF-shaped by design (the POC + ADR-078 semantics — no `Err`
state; the review sketch's `io::Result` return was dead code).
The helper pins the contract in one place instead of three
(alktunnels POC `pump_halves`, alktty's channels session, the next
consumer). Purely additive.
### Changed
- **`Establishment` carries the channel plan (review 007 R-01 —
breaking at 0.6.0).** The reserved field is filled:
`Establishment { plan: Option<ChannelPlan> }` with
`ChannelPlan = Arc<dyn Any + Send + Sync>` — typed-opaque, because
the payload an establisher hands the pump handler is a live handle
(a dialed socket, a TTY handle), not JSON. The wrapper threads
`establishment.plan` to the `OpenHandler`'s new second parameter
(`Fn(Value, Option<ChannelPlan>, Connection, AuthContext) ->
JoinHandle<()>`); `None` when no establisher is registered or it
returned `Establishment::default()`. Process-local: establisher →
wrapper → handler; nothing new crosses the transport. This kills
the side-channel handoff the alktunnels POC shipped (resource-keyed
slot + poll loop) with its concurrent same-resource race — each
open's establisher result flows to its own handler. Migration:
`Ok(Establishment {})` → `Ok(Establishment::default())` (or
`Establishment::new(handle)` to deliver a handle); handler closures
gain a `_plan` (or `plan`) parameter.
- **`OpenHandler` lifetime contract documented (review 007 R-02).**
Doc-only semantics note on the `OpenHandler` type and the
registration entry points: the returned `JoinHandle` must track the
data-plane lifetime — the wrapper awaits it and its completion
triggers channel teardown (drop of the demux sender = EOF to the
handler's read half); a handler that returns before its pumps
finish tears the channel down at birth (await pumps inline, never
spawn-and-forget). Plus a `debug!` telemetry line in
`run_open_wrapper` when a handler exits without having accepted the
channel's `BiStream` (the birth-teardown hint).
## [0.5.0] - 2026-09-06
The channel-open establishment phase (ADR-049 — review 006 E-01 +
N-1). Additive establishment machinery; one breaking change (the
`ChannelClient::open_channel` error type, below).
### Added
- **`OperationSpec.description` (review 006 E-02).** Additive
`Option<String>` op description (builder: `with_description`),
disclosed by `services/list` and `services/list-peers` (local
listings) when set, carried in the `services/schema` wire shape
(`spec_to_json_pub` emit / `rebuild_spec_for` parse — the field
survives `from_call` discovery and `op/register` announcement).
Describes the op, not the produced resource set (ADR-047 §6
amendment: the live resource-enumeration half stays deferred,
OQ-40). Absent on the wire when unset — additive for all consumers.
- **Channel-open establishment phase (ADR-049 — review 006 E-01).**
`ChannelCore::register_openable_with_establisher` registers a
per-ALPN open op with an establisher hook (`OpenEstablisher`): an
awaited establishment phase — validate params semantically, dial the
backend — bounded by the dispatch deadline or the registration's
timeout override, else `ESTABLISHMENT_TIMEOUT` (10s; the earlier of
the two). On establisher failure (error or deadline) the wrapper
tears down the just-allocated channel (demux sender, opener-ledger
take, `policy.on_close` un-increment — the allocation and teardown
balance) and replies `channel:open_failed` with
`details: { reason, message }`, reason ∈ `dial_failed` /
`unknown_resource` / `resource_shortage` / `handler_error` /
`timeout`. The SSH contract holds consumer-visibly: a failed open
never returns a `channel_id`. `register_openable` is unchanged
(no establisher = always-OK, existing registrations compile and
behave identically); the establisher takes `(input, auth)` — the
channel's yield-once `BiStream` belongs exclusively to the pump
handler. Additive wire surface (new error-code string + `details`
shape).
### Changed
- **`ChannelClient::open_channel` returns a typed error (ADR-049 §4 —
review 006 N-1, breaking at 0.5.0).** The open op's `CallError` is
carried verbatim in `ChannelOpenError::CallFailed` instead of being
flattened into a debug-formatted string, so consumers branch on
`channel:open_failed`'s typed reason (`establishment_reason()`).
Other variants: `MissingChannelId` (malformed success reply),
`AdoptFailed` (local adoption failure). Mechanical for consumers —
the `String` was a debug-formatting wrapper.
## [0.4.1] - 2026-09-05
Bug-fix release: chunks arriving for a not-yet-adopted channel are
parked instead of dropped (the open-op response / first-data race).
No API changes.
### Fixed
- **Early-arrival chunks for un-adopted channels are parked, not
dropped.** The connect side adopts a channel (installs local routing
state) only after the open-op response arrives, but the accept side's
`OpenHandler` can start pumping data the moment the channel opens —
the two race. Previously the connect side's demux dropped those
chunks (REQ-CH-04's lenient unknown-channel drop), silently losing
the first chunks of any push-first producer (a TTY backend's banner
or greeting, a sub protocol's initial frame). `route_payload` now
parks up to 64 payloads per unknown `channel_id` in a bounded
early-arrival buffer and `adopt_channel` drains them into the new
receiver in order; chunks beyond the cap drop with the pre-existing
debug log and `dropped_unknown_chunks` counter (the counter now means
"early-arrival overflow or genuinely unknown channel", not just the
latter). `clear_all` drops the parked buffers with the connection.
Surfaced by alktty's consumer end-to-end test (review #001 L3): the
session never resolved because the producer's first chunks (the
stdout sentinel + exit chunk for an immediately-resolving backend)
arrived before the adopt and were dropped.
## [0.4.0] - 2026-09-05
`input_schema` is now enforced at call time (the advertise-vs-enforce
leg ADR-016 promised: `INVALID_INPUT` is the registry's
schema-mismatch error code). Minor bump — behavioral break for any
caller that was passing schema-violating inputs to ops declaring a
non-trivial `input_schema` and relying on the validation being
documentation-only.
### Changed
- **`OperationSpec.input_schema` is enforced by the registry on every
dispatch.** Until now the schema was advertise-only: `services/schema`
disclosed it, the wire and gateway paths disclosed it, but no dispatch
entry point consulted it (the only schema alkcall actually enforced
was `publish_schema` per-chunk on Pub ops, P-03). The schema now
compiles once at registration (same fail-closed rule as
`publish_schema`/CF-003: an un-compilable schema is a registration
error, never a silently-skipped contract) and is checked after the
ACL gate in all three dispatch entry points — `invoke`,
`invoke_streaming`, and `invoke_sink` (via `resolve_sink_handler`,
preserving the P-08 single-source-of-truth property). Violations
return `CallError::invalid_input("input failed input_schema
validation")` with the input echoed in `details`. Raw-JSON-Schema
semantics (permissive on unknown keys) — adapters that want
closed-by-default enforcement keep doing their own hardening, as
alkhttp's `CompiledInputSchema` already does; the two checks compose.
Built-in ops are unaffected: alkcall's own specs declare the empty
object schema (accepts anything) and `services/schema` already
returned `INVALID_INPUT` for a missing `name` before its handler ran.
Motivated by alktty's code review #001 (L1): the channels open-op
wrapper hands the registry-checked `input` to the `OpenHandler` as
the authoritative params, which requires the registry to actually
validate it.
- `OperationRegistry` gained an `input_validator` cache (compiled at
registration, mirrored on `fork` and in `OperationRegistryBuilder`
like the publish validators); `input_validator()` is `pub(crate)` —
the public surface is unchanged.
## [0.3.1] - 2026-09-04
Bug-fix release: `services/list-peers` can now list peer-announced
ops. No API breaks — `OperationEnv` gains one defaulted trait method
(non-breaking for all implementors; `cargo semver-checks` 196 checks
pass against the 0.3.0 baseline).
### Fixed
- **`services/list-peers` shows each peer's operations**
(UP-03, surfaced by alkhttp's review 006). `PeerCompositeEnv`
overrode `peer_ids` only, so `peer_operations` fell to the trait
default (`Vec::new()`) and every peer listed with an empty
operations array — the ADR-022 amendment's "announced op is
discoverable via `services/list-peers`" promise never resolved on
the wire. ADR-030 prescribed the fix but it had never been ported
into alkcall; the existing `list-peers` unit tests passed because
they mock `peer_operations` with hand-rolled envs. Implemented per
ADR-030: `OperationEnv::list_operation_names` (default `Vec::new()`)
+ overrides on `OverlayOperationEnv` (overlay map keys),
`PeerCompositeEnv` (`peer_operations` delegates to the peer's
overlay; its own aggregate mirrors `contains()`: session +
connections + base), `LocalOperationEnv` (registry names), and
`ChannelsSessionEnv` (delegates to base). Gate:
`announced_op_is_discoverable_via_services_list_peers` exercises the
exact `compose_root_env` shape (announce via `op/register`, then
assert both the `peer_operations` probe and the `services/list-peers`
wire shape attribute the op to the peer) — verified load-bearing
(reverting the override fails the gate). ADR-030 status is now
Accepted with the UP-03 provenance note.
## [0.3.0] - 2026-09-04
The connect side can serve (two-way ops over one channels connection),
per-session fork registries, and the `op/register` bootstrap op — the
remediation of reviews 004 and 005
(`docs/reviews/004-*.md`, `docs/reviews/005-*.md`). Two
`OperationRegistry` methods changed their return types from borrowed to
owned (source-breaking for annotated call sites; minor bump per 0.x
semver rules) — verified with `cargo semver-checks` (196 checks pass
against the published 0.2.0 baseline; the return-type changes were
caught by manual diff review, as the tool has no lint for that
pattern).
### Added
- **`OperationRegistry::fork`** — deep-copies a registry (handlers,
provenance, composition authority, capabilities, cached
publish-schema validators) into an independently-mutable copy. The
per-session shape: fork the deployment's base registry, register the
session's ops on the fork, dispatch the session over the fork
(ADR-047 §4 amendment, 2026-09-03). Internally mutable, so a fork
shared as an `Arc` can receive registrations after the dispatcher
was built — `install_bootstrap_discovery` relies on this.
- **`OperationRegistryBuilder::from_registry`** — seeds a builder from
an existing registry's registrations (the fork surface expressed
through the builder; registration order is not preserved — HashMap
iteration).
- **Connect-side serving (two-way ops over one connection).** The
channels connect side's channel-0 read pump previously resolved
outbound responses only and silently dropped inbound
`call.requested` frames. New `ChannelClient::from_connection_with_serving`
takes `Option<ServingConfig>` (`registry` + `identity_provider`);
with `Some(..)` the read pump becomes the full-duplex serving loop
(`Dispatcher::serve_single_stream`), so a connected peer can call the
consumer's ops. `from_connection` keeps the resolution-only pump
(pure-consumer default). Serving is opt-in: the protocol is
symmetric, the API is explicit (ADR-022 amendment, 2026-09-03).
- **`Dispatcher::serve_single_stream` and `Dispatcher::dispatch_start`
+ `StartedDispatch`** — the shared serving-loop machinery. Both
single-stream loops (accept side and connect side) now run the same
concurrency model: Once invocations, Sub pumps, and sink response
writers are spawned, so same-connection nested composition resolves
concurrently instead of deadlocking until the 30 s sweeper (review
005 G-01). Handles are tracked and aborted at loop exit; no lock is
held across an await.
- **`registry::op_register` module — the `op/register` bootstrap op**
(review 004 F-05). A connected peer announces the ops it serves over
the wire: `OP_REGISTER_NAME` (`op/register`), `OpRegisterRequest`
(wire round-trip via `to_json`/`from_json`), `op_register_spec` /
`op_register_handler`. Announcements land in the connection overlay
and are served through forwarding handlers, so hub→consumer import
(`from_call`) and consumer→hub announce are symmetric. **Collision
policy** (review 005 G-03): a peer-announced op may replace other
peer-announced ops (when `replace` is set) but never the serving
side's own registrations — a name present on the serving registry
rejects with the new `CallError::already_exists` (`ALREADY_EXISTS`,
non-retryable), regardless of `replace`. The collision gate checks
the serving registry, the session fork, and the connection overlay
itself.
- **`install_bootstrap_discovery`** (review 004 F-06) — registers
`services/list`, `services/list-peers`, and `services/schema`
against a registry with handlers closed over that same `Arc`, so
per-session forks are the discovery source for their own openables
(handlers see every op the registry serves at call time). ACL
filtering stays per-caller. The bootstrap-op set on channel 0 is
closed and now includes `services/list-peers` (review 005 G-05 —
doc alignment; the code had installed it since the amendment landed)
plus `op/register` (ADR-022 amendment).
- **`spec_to_json_pub`** — public serialization of an `OperationSpec`
into the `services/schema` wire shape (the shape `op/register`
announces with and `services/schema` serves). **`resource_id_path`
now survives the spec wire round-trip** (review 005 G-04): both the
serializer and the `op/register` parser carry the field.
- **`CallConnection::overlay_contains` /
`CallConnection::overlay_registration`** — read access to the
connection overlay, for `op/register` replace semantics and
composition reachability checks.
### Changed
- **`OperationRegistry::registration` returns
`Option<HandlerRegistration>`** (was `Option<&HandlerRegistration>`)
and **`list_operations` returns `Vec<OperationSpec>`** (was
`Vec<&OperationSpec>`). Source-breaking for call sites that annotate
the borrowed types — clone-at-the-boundary instead. Minor bump per
0.x semver rules.
- **Registration is `&self` throughout** — `OperationRegistry::register`,
`ChannelOperations::register_on`, and `ChannelCore::register_openable`
take `&OperationRegistry` (was `&mut`). Source-compatible for
callers (reborrow); this is what makes post-construction
registration on a forked, `Arc`-shared registry possible.
- **`from_call` composition reachability** — the forwarding-handler
path declares reachable namespaces on the composing handler's
`scoped_env` (empty is deny-by-default) and the connection overlay
is attached only for identity-carrying connections (ADR-030 §5).
## [0.2.0] - 2026-08-31
Consumer-findings remediation (CF-001..004 from
`docs/reviews/consumer-findings-ledger.md` — alkhttp as the first real
consumer), the promoted `gateway` dispatch spine (ADR-048), and a
docs-hygiene pass. One behavior change noted below; otherwise additive.
### Added
- **`gateway` feature: the transport-neutral dispatch spine**
(ADR-048). New `alkcall::gateway` module behind the opt-in `gateway`
cargo feature (default off; adds no dependencies). `GatewayDispatch`
is the deadline-bounded, re-rooted-context invoke spine over
`OperationRegistry` (`invoke` / `invoke_streaming` / `invoke_sink`)
promoted from alkhttp's gateway after it proved transport-agnostic —
hubs and spokes relaying calls (ADR-042 translate path) need the
identical root-context discipline (`internal: false`,
`forwarded_for: None`) without any HTTP. `schema_disclosure_denial`
is the shared is-internal + ACL check for `services/schema`
inner-op-name disclosure (one implementation so transports cannot
drift; ACL denial returns `FORBIDDEN`, Internal visibility returns
spec-404 — see ADR-048 for the split from the wire handler's
conservative spec-404). The handler deadline is a constructor knob
(`with_deadline`); the default remains 30 s. alkhttp migrates to
this module in a follow-up and drops its local copy.
### Changed
- **Un-compilable `publish_schema` values are rejected at registration
time** (CF-003). `OperationRegistry::register` and every
`OperationRegistryBuilder` method now return `Err` when a Pub op's
`publish_schema` fails to compile. Previously the bad schema was
accepted and the dispatch path failed **open** (warn log, chunks
unvalidated). The compiled validator is now cached per-op
(`OperationRegistry::publish_validator`) and consumed by the dispatch
path — the per-request compile is gone. Code that registered
un-compilable schemas will now get a registration error instead of a
silently-unvalidated op.
### Fixed
- **`services/schema` no longer discloses Internal or ACL-restricted op
specs** (CF-004). The schema handler now applies the same gates as
`invoke()` — Internal-visibility rejection and `AccessControl::check`
with matching identity resolution. Restricted ops return spec-404
(`NOT_FOUND`), consistent with "restricted ops don't exist" elsewhere;
no information leaks about the restricted surface. The gate is inside
the handler, so every transport is covered.
- **Demux `TooLarge` skip no longer allocates from the peer's header**
(CF-002). The skip now streams through a fixed 64 KiB buffer, and a
cumulative 256 MiB skipped-bytes budget tears down connections that
loop oversized headers. Previously the buffer was sized from the
untrusted `u32` length (up to ~4 GiB pinned per dribbling peer).
- **Write failures before request delivery are retryable** (CF-001).
New `CallError::connection_closed` (`CONNECTION_CLOSED`, `retryable:
true`) is returned when the `call.requested` frame write fails on any
consumer path (`call`/`subscribe`/`publish`, both stream modes) — the
call provably never reached the producer, so reconnect/retry is safe.
Mid-publish write failures and producer-side connection-teardown
failures remain non-retryable `INTERNAL` (delivery ambiguous). The
new code string is additive; the `retryable` flag is the
machine-readable signal for consumers.
### Fixed (continued)
- **Doc hygiene.** Resolved the broken intra-doc link on `lib.rs`'s
feature-gated `gateway` mention (rustdoc warned on default-feature
builds), and re-pointed 40 `src/` doc references from the old alknet
mono-repo ADR numbering (ADR-049/050/052/065/070/074/092) to this
crate's numbering (ADR-021/011/034/007/008/009/005) so published docs
reference ADRs this crate actually carries. Failure-path coverage was
added alongside: the CF-002 skipped-bytes teardown and skip-EOF arms,
CF-001's subscribe write-failure sites (both stream modes), the
publish request-frame failure in single-stream mode, and
non-retryability pinning for mid-publish failures.
## [0.1.1] - 2026-08-17
A minor release that renames the ALPN prefix from `alknet/` to `alk/`
@@ -61,5 +431,12 @@ Vendored core types (`Connection`, `ProtocolHandler`, `BiStream`,
(ADR-046), the channels protocol with openable-ALPNs-as-operations
(ADR-047), and the `ChannelClient` transport-agnostic client.
[0.6.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.6.0
[0.5.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.5.0
[0.4.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.4.1
[0.4.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.4.0
[0.3.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.3.1
[0.3.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.3.0
[0.2.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.2.0
[0.1.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.1.1
[0.1.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.1.0
Generated
+1 -1
View File
@@ -27,7 +27,7 @@ dependencies = [
[[package]]
name = "alkcall"
version = "0.1.1"
version = "0.6.0"
dependencies = [
"async-trait",
"bytes",
+3 -2
View File
@@ -1,6 +1,6 @@
[package]
name = "alkcall"
version = "0.1.1"
version = "0.6.0"
edition = "2021"
rust-version = "1.85"
license = "MIT OR Apache-2.0"
@@ -9,13 +9,14 @@ readme = "README.md"
repository = "https://git.alk.dev/alkdev/alkcall"
keywords = ["rpc", "json-rpc", "multiplexing", "wire-format", "alpn"]
categories = ["network-programming", "asynchronous", "encoding"]
exclude = [".opencode/", "AGENTS.md", "docs/reviews/", "docs/sdd_process.md", "Cargo.lock"]
exclude = [".opencode/", "AGENTS.md", "docs/reviews/", "docs/sdd_process.md"]
[lib]
name = "alkcall"
[features]
default = []
gateway = []
[dependencies]
jsonschema = { version = "0.46", default-features = false }
+30 -2
View File
@@ -21,7 +21,7 @@ use alkcall::registry::{
spec::{OperationSpec, OperationType, Visibility, AccessControl},
};
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry.register(HandlerRegistration::new(
OperationSpec::new(
"echo/run",
@@ -98,6 +98,34 @@ for reg in registrations {
let response = conn.call("remote/status", serde_json::json!({})).await;
```
### Serving your own ops as a connected consumer
The call protocol is symmetric — both sides of a connection can serve
ops. A `ChannelClient` built with `from_connection` is a pure consumer
(inbound `call.requested` frames are dropped); pass a `ServingConfig`
to also serve your registry to the peer, and use `op/register` to
announce which ops you serve:
```rust
use std::sync::Arc;
use alkcall::channels::client::{ChannelClient, ServingConfig};
use alkcall::registry::discovery::install_bootstrap_discovery;
let registry = Arc::new(OperationRegistry::new());
// ... register your ops on the registry, then:
install_bootstrap_discovery(&registry)?;
let client = ChannelClient::from_connection_with_serving(
connection,
Some(ServingConfig {
registry: Arc::clone(&registry),
identity_provider: provider,
}),
).await?;
// peer-callable ops resolve against `registry` on channel 0;
// `client.call_open_op` still works — both directions share the pump
```
## Architecture
alkcall is a pure protocol crate — no networking, no transport
@@ -107,7 +135,7 @@ Downstream crates compose on top of it in a layered dependency chain.
| Role | Call protocol | Channels protocol |
|------|---------------|-------------------|
| **Producer** | Registers ops on an `OperationRegistry`, runs a `Dispatcher` | Runs a `ChannelsAdapter`, registers openable ALPNs via `ChannelCore::register_openable` |
| **Consumer** | Uses `CallConnection` to call ops, uses `from_call` to discover/import remote ops | Uses `ChannelClient` to open channels via `call_open_op` + `open_channel` |
| **Consumer** | Uses `CallConnection` to call ops, uses `from_call` to discover/import remote ops; may also serve its own ops (`from_connection_with_serving`) | Uses `ChannelClient` to open channels via `call_open_op` + `open_channel` |
| **Hub** | Both: runs a `Dispatcher` for ops it produces, holds `CallConnection`s to spokes for ops it consumes | Both: runs a `ChannelsAdapter` for inbound connections, holds `ChannelClient`s to spokes |
| **Spoke / Worker** | Both: produces ops (its own services), consumes hub ops | Both: produces channels (TTY, tunnel), may consume hub channels |
+6 -3
View File
@@ -13,7 +13,7 @@ This crate unifies `alknet-call` and `alknet-channels` from the alknet
mono-repo, plus the vendored core types formerly in `alknet-core`. The
source architecture docs were ported from
`/workspace/@alkdev/alknet/docs/architecture/` and renumbered as alkcall
ADRs (ADR-001..045). The ALPN strings (`alk/call`, `alk/channels`)
ADRs (ADR-001..050). The ALPN strings (`alk/call`, `alk/channels`)
are wire-stable and unchanged — see ADR-004.
## Documents
@@ -100,6 +100,9 @@ are wire-stable and unchanged — see ADR-004.
| [045](decisions/045-alknetclient-native-dial-seam.md) | AlknetClient Dial Seam | spawn_dispatch / from_connection take-over; dial in consumer |
| [046](decisions/046-publish-operation-type-and-handler-kind-sink.md) | Publish Operation Type and HandlerKind::Sink | `OperationType::Pub` (producer→consumer streaming); `SinkHandler` + `HandlerKind::Sink`; `call.published` wire event; `invoke_sink()` dispatch; `Subscription` renamed to `Sub` |
| [047](decisions/047-openable-alpns-are-operations.md) | Openable ALPNs Are Operations | `channel/open` dissolves into per-ALPN ops `channels/<alpn>/sub`/`pub`; `channel_open` marker on `OperationSpec`; `ChannelCore` wrapper; extension-trait `ChannelOperationEnv`; connection-owner allocates `channel_id`; opener ledger (Gap 2 fix); ALPNs are call apps |
| [048](decisions/048-dispatch-spine-gateway-module.md) | Dispatch Spine (feature-gated `gateway` module) | `alkcall::gateway` behind the `gateway` feature; `GatewayDispatch` invoke spine (deadline knob, re-rooted context) + `schema_disclosure_denial` (FORBIDDEN for ACL deny, spec-404 for Internal); promoted from alkhttp for hub/spoke reuse |
| [049](decisions/049-channel-open-establishment-phase.md) | Channel-Open Establishment Phase | `OpenEstablisher` + `register_openable_with_establisher` (awaited, bounded); typed `channel:open_failed` with `details.reason`; `Establishment.plan` (`ChannelPlan`) threaded to the `OpenHandler` (amendment 2); the `JoinHandle` data-plane lifetime contract |
| [050](decisions/050-pump-bidi-two-pump-helper.md) | `pump_bidi` Two-Pump Helper | `channels::pump_bidi` — shutdown-on-completion two-pump data plane pinned in one place (alknet ADR-078) |
## Relevant Open Questions
@@ -271,7 +274,7 @@ feature flags) or in the downstream alknet crate.
- `@alkdev/alknet: docs/architecture/` — the source architecture docs
these were ported from (renumbered from alknet ADR-001..094 to alkcall
ADR-001..045)
ADR-001..048; ADR-049 and later were authored in this crate)
- `@alkdev/alktype` — the binary struct engine; compiles BAST documents
(e.g. `chunk-header.bast.json`) into readers/writers/validators
- `@alkdev/pubsub` — the TypeScript EventEnvelope prior art the call
@@ -287,5 +290,5 @@ feature flags) or in the downstream alknet crate.
> TTY's wire format, ADR-082 for alknet-tls, ADR-086 for endpoint types).
> These are ADRs for sibling crates that are not part of alkcall. They
> retain their alknet numbering (052, 082, 086, etc.) — any ADR number
> outside the alkcall range 001..045 is an alknet source ADR, found at
> outside the alkcall range 001..048 is an alknet source ADR, found at
> `/workspace/@alkdev/alknet/docs/architecture/decisions/`.
+42 -2
View File
@@ -68,12 +68,19 @@ impl ChannelClient {
/// the `MpscRecvStream` (read half). The caller can build a
/// `Connection` from these via `channel_source` and
/// `Connection::from_source`.
///
/// The error is typed (ADR-049 §4 — review 006 N-1):
/// `ChannelOpenError::CallFailed` carries the wire `CallError`
/// verbatim (branch on `establishment_reason()` for
/// `channel:open_failed`'s `details.reason`); `MissingChannelId`
/// covers a malformed success reply; `AdoptFailed` covers local
/// adoption failure.
pub async fn open_channel(
&self,
operation_id: &str,
input: Value,
alpn: &str,
) -> Result<(u32, MpscSendStream, MpscRecvStream), String>;
) -> Result<(u32, MpscSendStream, MpscRecvStream), ChannelOpenError>;
/// Take the `CallConnection` — used by the consumer to register
/// imported ops (`from_call`) on the connection's overlay. After
@@ -88,7 +95,11 @@ impl ChannelClient {
ADR-047 dissolved the generic `channel/open` operation into per-ALPN
open ops (`channels/<alpn>/sub`, `channels/<alpn>/pub`). Each ALPN
crate registers its own open op via `ChannelCore::register_openable`
(ADR-047 §3, as amended 2026-08-13 — per-connection registration).
(ADR-047 §3, as amended 2026-08-13 — per-connection registration),
optionally with an establisher via
`register_openable_with_establisher` (ADR-049 — the awaited,
bounded establishment phase; establishment failure resolves
`Err` on the client with `channel:open_failed` + `details.reason`).
The `ChannelClient` calls these ops by name on channel 0 via
`call_open_op`; `open_channel` wraps `call_open_op` + `adopt_channel`.
@@ -97,6 +108,35 @@ and `subscribe_resources` method are deferred (OQ-40, OQ-41). The
`ResourceEntry.access` preview was dropped by ADR-047 §6 — it is
available via `services/schema` on the op spec.
## The early-arrival park bound (push-first producers)
The open-op response carrying `channel_id` races the producer's first
data-plane writes: the producer's handler can start pumping before the
consumer's `adopt_channel` runs. The demux parks those first chunks
per-channel (FIFO) instead of dropping them, and `adopt_channel` drains
the parked chunks into the new receiver — a push-first producer (a TTY
backend's banner, a sub protocol's greeting) must not lose its first
chunks.
The park is bounded: **up to 64 chunks per channel** are parked
(`EARLY_ARRIVAL_CAP`, ADR-040's memory bounds); chunks arriving past
the cap are dropped silently — at the consumer this presents as a
truncated stream with clean framing everywhere else, not as an error.
Consumer-facing consequences (review 006 E-04, noted for tunnel-style
consumers):
- Size any pre-adopt buffering math (e.g. a UDP POC's MTU-vs-buffer
sizing) against **64 parked chunks as the observable bound**, and
adopt promptly — the open reply resolves *before* the first data
arrives by design, so the adopt is the consumer's next step, not a
slow path.
- The producer side observes both halves of the race via
`ChannelManager::early_arrival_count()` (chunks parked) and
`ChannelManager::dropped_unknown_chunks()` (chunks lost to the cap) —
a non-zero dropped counter under a push-first producer means the
adopter was too slow for the producer's burst, and the affected
streams were truncated at the park boundary.
## Transport-agnostic by construction
`ChannelClient` is the client side of the channels protocol. The channels
+32
View File
@@ -77,12 +77,39 @@ round-trip, so the open round-trip is not additive latency.
| `channel:forbidden` | `AccessControl::check` denied the open | false |
| `channel:allocation_failed` | Handler allocate failed (e.g., backend couldn't start) | true (often transient) |
| `channel:too_many_channels` | Per-connection or per-identity channel limit hit (ADR-040, ADR-041) | false |
| `channel:open_failed` | The establishment phase failed (ADR-049) — `details: { reason, message }`, reason ∈ `dial_failed` / `unknown_resource` / `resource_shortage` / `handler_error` / `timeout` | false |
| `channel:no_channels_session` | The op was invoked outside a channels session (no `ChannelManager` — ADR-047 §2) | false |
`channel:unknown_alpn` and `channel:invalid_params` (ADR-037) are gone —
an unregistered ALPN is an ordinary `NOT_FOUND` (op not registered); a
bad `input` is ordinary schema rejection.
### The establishment phase (ADR-049)
An openable ALPN may register an **establisher** alongside its open
handler (`ChannelCore::register_openable_with_establisher`). The
establisher is the awaited preparation step ADR-047 §3 described
("validate params, consult ownership, prepare the backend, return a
channel plan"): the wrapper awaits it **bounded** (the dispatch
deadline when the op carries one, else the registration's override or
the 10s `ESTABLISHMENT_TIMEOUT` — the earlier of the two), **before
replying and before spawning the pump handler**. On establisher
failure (error or deadline) the wrapper tears down the just-allocated
channel (demux sender, opener-ledger entry, `policy.on_close`
un-increment — the same atomic ledger-take gate every teardown path
runs) and replies `channel:open_failed` with
`details: { reason, message }` (the table row above). The SSH
contract holds consumer-visibly: a failed open never returns a
`channel_id`.
Registrations without an establisher behave exactly as before (the
open op cannot fail post-allocation; establishment work inside the
handler is invisible to the open reply). The bound applies only to
the establisher — the spawned pump handler's lifetime is governed by
the existing teardown machinery, unchanged. Head-of-line safety: the
serving loop spawns Once invocations as independent tasks, so a slow
establisher on one open op does not block other calls on channel 0.
### `channels/<alpn>/pub` — publish a binary stream
`OperationType::Pub` (ADR-046). The initiator publishes a stream of
@@ -443,6 +470,8 @@ All design decisions are documented as ADRs in [decisions/](decisions/).
| [035](decisions/035-channels-pure-channel-multiplexing.md) | Pure Channel Multiplexing | No stream_types; handler owns sub-mux |
| [021](decisions/021-streaming-handler-for-subscriptions.md) | StreamingHandler | The machinery `channel/resources/subscribe` uses |
| [046](decisions/046-publish-operation-type-and-handler-kind-sink.md) | Pub Operation Type | The `Pub`/`Sub` primitives the per-ALPN open ops build on |
| [049](decisions/049-channel-open-establishment-phase.md) | Channel-Open Establishment Phase | The establisher hook; typed `channel:open_failed`; `Establishment.plan` (amendment 2); the `JoinHandle` lifetime contract |
| [050](decisions/050-pump-bidi-two-pump-helper.md) | `pump_bidi` Two-Pump Helper | The data-plane helper openable-ALPN handlers await inline |
| [026](decisions/026-forwarded-for-identity.md) | Forwarded-For Identity | The auth chain for hub-relayed opens (and why the cap is per direct-caller, not per `forwarded_for`) |
| [011](decisions/011-dynamic-resource-ownership-for-runtime-spawned-resources.md) | Dynamic Resource Ownership | The parallel — a channel slot is a resource, the cap is a quota check; `resource_id_path` works again under per-ALPN ops |
@@ -450,6 +479,9 @@ All design decisions are documented as ADRs in [decisions/](decisions/).
- ADR-047: openable ALPNs are operations (the unifying ADR — per-ALPN
open ops, `channel_open` marker, `ChannelCore` wrapper, opener ledger)
- ADR-049: channel-open establishment phase (the establisher hook,
`channel:open_failed`, the plan payload)
- ADR-050: `pump_bidi` (the two-pump helper)
- ADR-037: channel lifecycle operations (amended by ADR-047)
- ADR-041: per-identity channel cap (amended by ADR-047 §7 — opener
ledger, every teardown path)
@@ -2,7 +2,7 @@
## Status
Accepted (amended 2026-06-26, 2026-07-13, and 2026-07-16 — see "Amendments" below; the 2026-07-16 amendment per ADR-045 §5 removes `CallClient::connect`)
Accepted (amended 2026-06-26, 2026-07-13, and 2026-07-16 — see "Amendments" below; the 2026-07-16 amendment per ADR-045 §5 removes `CallClient::connect`; amendment 2026-09-03 — the bootstrap-op set and the connect-side serving loop, see "Amendment (2026-09-03)" below; amendment 2026-09-04 — the `op/register` collision policy, in that amendment's "Collision policy" paragraph)
## Context
@@ -359,6 +359,123 @@ same as `from_openapi` receives HTTP credentials.
prior art
- POC at `/workspace/@alkdev/dispatch` — head/worker dispatch over SSH+axum
## Amendment (2026-09-03): bootstrap-op set and the connect-side serving loop
Review 004 (F-04/F-05) verified two gaps between this ADR's
bidirectionality promise (§2 "connection direction is independent of
call direction") and the single-stream channel-0 implementation:
1. **The connect side's read pump resolved responses only** — inbound
`call.requested` frames were silently dropped. Serving existed only
on the accept side (`run_loop_single_stream`).
2. **No wire mechanism announced client-side ops** — `from_call`
imports hub ops into the consumer, but a connected peer had no way
to announce "here are the ops I serve" over the wire.
The amendment (implemented in alkcall; the e2e gates are
`serving_loop_hub_to_consumer_call_resolves` and
`op_register_announce_then_hub_call_routes_back_to_consumer` in
`src/channels/client.rs`):
### The bootstrap-op set (one-way)
Each side of a channels connection **may serve** the bootstrap ops on
channel 0. The set is closed:
- `services/list` — discovery (the op `from_call` dials on every
import; each side is expected to serve it)
- `services/schema` — per-op schema disclosure
- `services/list-peers` — peer-keyed discovery (the op that makes
peer-announced ops discoverable; `from_call`-import discovery
relies on it). Added to this list 2026-09-04 (review 005 G-05) —
`install_bootstrap_discovery` has installed it since the amendment
landed; the set was declared closed and the doc lagged the code.
- `op/register` — peer op announcement (below)
These are the ops a peer may assume are reachable (subject to each
op's `AccessControl`); a peer that does not serve one answers
`NOT_FOUND` and the caller treats it accordingly (`from_call` already
surfaces discovery failure as `AdapterError::DiscoveryFailed`).
### Connect-side serving is opt-in (two-way door at the API level)
`ChannelClient::from_connection_with_serving(connection,
Option<ServingConfig>)` — with `None` (the pure-consumer default) the
read pump resolves outbound pendings only (previous behavior). With
`Some(ServingConfig { registry, identity_provider })` the read pump
becomes the full-duplex serving loop (`Dispatcher::serve_single_stream`):
inbound `call.requested` frames dispatch against the configured
registry and resolve back to the peer; outbound pendings still resolve
in the same loop. Serving is opt-in because a pure consumer has no
registry to serve; the *protocol* is symmetric, the *API* is explicit.
Direction disambiguation in the loop is by table membership, not
framing: an id that is one of *our* outbound pendings resolves there;
an id that is an inbound request dispatches; `call.aborted` tries both
tables (in-flight sink aborts and the pending map's cascade). IDs are
UUID-generated per side, so cross-correlation is not a hazard.
### `op/register` (one-way in wire shape)
The peer→hub direction of `from_call`'s bundle flow: a peer sends
`call.requested` for `op/register` carrying the `OperationSpec` in the
`services/schema` wire shape (`spec_to_json`) plus a `replace` flag.
The serving-side handler (`registry::op_register::op_register_handler`):
1. Rebuilds the spec (the same parser `from_call` uses — one spec
serialization on the wire).
2. Wraps a **call-forwarding handler** that issues a nested
`call.requested` back over channel 0 to the announcing peer (the
same shape `from_call`'s imported bundles use — the in-process twin
at `protocol/adapter.rs`).
3. Writes the bundle into **that connection's overlay** via
`register_imported` with `FromCall` provenance and
`Visibility::Internal` (composition material, ADR-017 — never
directly callable from the serving side's own wire).
The announced op is discoverable via `services/list-peers` (the
overlay is peer-keyed, `compose_root_env` attaches it) and invocable
via nested composition (`env.invoke`). Announced Sub/Pub ops register
as stubs that answer `INVALID_OPERATION_TYPE` — nested composition is
request/response-only (`OverlayOperationEnv`'s contract); the
streaming/sink forwarding shapes ride on the `from_call` import path.
Access control: `op/register` itself carries an `AccessControl` (an
unprivileged peer cannot reach the handler — the registry's normal
invoke path enforces it). Replace semantics: a collision with an
existing overlay registration is rejected with `ALREADY_EXISTS`
unless `replace: true` (the reconnect path re-announces). The
overlay dies with the connection (Layer 2), so reconnect re-announce
is naturally scoped.
Collision policy (amended 2026-09-04, review 005 G-03): a
peer-announced op may collide with other *peer-announced* ops on the
same connection (`replace` governs) but **never** with the serving
side's own registrations — a name present on the serving registry
rejects with `ALREADY_EXISTS` regardless of `replace`. The connection
overlay shadows the base registry in `PeerCompositeEnv` (connections
resolve before base, ADR-024 §1), so an unscreened same-name announce
would silently rewrite what a wire-dispatched handler's
`ctx.env.invoke` resolves for any name the deployment registered:
composition authority (ADR-018) belongs to the composing handler's
deployer, not the connected peer. Visibility is irrelevant to this
gate (`Internal` ops are as shadowable as `External` —
`OverlayOperationEnv` gates on `AccessControl`, not visibility; the
composed child is `internal: true` by design).
The envelope kind set stays closed at six — bootstrap ops over channel
0 are the door (AGENTS.md §7 allows adding kinds; none is needed).
### Cross-references
- Review 004 (`docs/reviews/004-per-connection-dispatch-and-client-serving-review.md`)
F-04/F-05 — the verification and the design decision
- ADR-047 §4 amendment #2 (2026-09-03) — the per-session fork the
bootstrap ops compose on
- alkhttp OQ-05 / ADR-048 — the browser data-channel wiring this
unblocks (the alkhttp-side gap was wiring; these two mechanisms are
what it wires to)
## Amendments (2026-06-26)
This ADR left four decisions as two-way doors (§1 Consequences flagged DC-1's
@@ -2,7 +2,15 @@
## Status
Proposed
Accepted (implemented 2026-09-04 — surfaced as UP-03 in alkhttp's
review 006 `docs/…/006-alkcall-0.3.0-consequence-review.md`: the
`op/register` amendment's "announced op is discoverable via
`services/list-peers`" promise did not resolve on the wire because
this override had never been ported into alkcall; the gate is
`announced_op_is_discoverable_via_services_list_peers` in
`src/registry/op_register.rs`. The `services/list-peers` unit tests
did not catch it because they mock `peer_operations` with hand-rolled
envs.)
## Context
@@ -5,7 +5,132 @@
Accepted (amends ADR-037; refines ADR-044, ADR-046; §4 amended
2026-08-13 — open ops are registered per-connection, not resolved via
`context.env` downcast — see "Amendment (§4 per-connection
registration, 2026-08-13)" below)
registration, 2026-08-13)"; amendment #2 (2026-09-03) — the
per-connection registration mechanism is the **per-session fork of the
base registry installed as the session's dispatch registry**, not the
connection overlay — see "Amendment (§4 mechanism, 2026-09-03)" below;
§6 amended 2026-09-06 — the static half of discovery gains an additive
per-op `description` on the listing, the dynamic half stays deferred
— see "Amendment (§6 listing enrichment, 2026-09-06)" below)
## Amendment (§6 listing enrichment, 2026-09-06)
Review 006 E-02 (from the alktunnels Phase 0 sweep) made §6's dynamic
half load-bearing for the first time and asked for a decision on the
static half. Decision: **both halves of the "static per-op" split gain
what is cheap today; the dynamic half stays deferred** (OQ-40, now
with a "load-bearing for alktunnels discovery UI" note in
`open-questions.md`; alktunnels v1 uses config-known op names):
- `OperationSpec` gains `description: Option<String>` — a human-
readable op description, set via `with_description` at registration.
Additive (defaults `None`; no struct-literal construction sites
exist — all sites use `OperationSpec::new`).
- `services/list` (and the local-ops half of `services/list-peers`)
emit `description` when set — one round-trip answers "which ops
exist and what are they for" without the N+1 `services/schema`
sweep. The output-schema docs on both listing specs and on
`services/schema`'s `operation_spec_schema` advertise the field.
- `spec_to_json_pub` emits it when set; `rebuild_spec_for` parses it
back — the description survives discovery and peer announcement
(`from_call`, `op/register`) like every other additive spec field
(`resource_id_path`, `publish_schema` round-trip the same way).
Scope note from E-02 stands: the listing field describes **the op**,
not **the produced resource set** (a tunnel producer registers one op
and N resources). "Which tunnel resources may I open, live" remains
`channel/resources/subscribe`'s job (§6's dynamic half, ADR-037 §
`channel/resources/subscribe`) — deferred until a consumer needs live
resource discovery.
## Amendment (§4 mechanism, 2026-09-03)
The 2026-08-13 amendment named the registration target as "the
connection overlay registry (Layer 2 per ADR-019)". Review 004 (F-01)
verified that this shape cannot dispatch: the top-level dispatch path
(`Dispatcher::dispatch` / `run_loop_single_stream`) resolves and
invokes against the dispatcher's **base registry only**; the
connection overlay is reachable solely as a layer of `context.env`
(for nested invocations — a handler calling `env.invoke(...)), never
for resolving the incoming `call.requested` itself. An open op
registered on the overlay resolves `NOT_FOUND` on the wire.
The one shape proven end-to-end (alkcall's own e2e gate) is different:
the `install_channel_zero` hook builds a **fresh per-connection
registry containing the open op and passes it as the dispatcher's base
registry**. This amendment makes that the operative mechanism.
**The decision: per-connection registration happens on a fork of the
deployment's base registry, installed as the session's dispatch
registry.** The `install_channel_zero` hook (and any future
session-establishment seam):
1. **Forks** the deployment's base registry
(`OperationRegistry::fork` — a deep copy carrying handlers,
provenance, composition authority, capabilities, and the cached
publish-schema validators; review 004 F-02/F-03).
2. **Registers the per-session ops on the fork** — the generic channel
ops (`ChannelOperations::register_on`), the openables
(`ChannelCore::register_openable`), and the bootstrap discovery ops
(`install_bootstrap_discovery`, closed over the fork itself so
`services/list` sees the fork's per-session ops — review 004 F-06).
3. **Dispatches channel 0 over the fork** (`Dispatcher::new(fork,
...)`).
The fork is possible because `OperationRegistry` is internally
mutable (`parking_lot::RwLock` around both maps) — a fork shared as an
`Arc<OperationRegistry>` can receive bootstrap ops after the
dispatcher was built, and the self-referential discovery closure sees
every post-install registration.
The **connection overlay (Layer 2) remains what ADR-019/ADR-024
describe**: the landing zone for peer-announced ops (`op/register`,
review 004 F-05 — ADR-022 amendment) and the nested-invocation target
for imported ops. It is not the dispatch-resolution path for the
session's own ops.
Rationale for the fork shape over an overlay-aware dispatch fallback
(F-02 option (b)): the fork is the only shape with an end-to-end
proof, it needs no change to the shared dispatch loop, and it keeps
the overlay's `invoke_with_policy` shape (namespace-scoped,
parent-context-driven — built for nested composition) out of the
top-level call path, where it does not match the frame-handling
contract.
This preserves every invariant the 2026-08-13 amendment protected:
- **Layering (ADR-044):** unchanged — the open-op wrapper is in
`channels-call`; the call crate's `OperationRegistry` gains only
`fork` (and interior mutability), no channels types.
- **Per-connection resolution:** the open op gets the *right*
`ChannelManager` because the fork is built per-connection and its
openable closes over that connection's `ChannelCore`.
- **"Marked ops invoked outside a channels session" (ADR-047 §2):**
unchanged in effect — a `channels/<alpn>/sub` op registered only on
a session fork is not reachable on a bare `alk/call` connection (the
fork isn't that session's dispatch registry) — the dispatch path
returns `NOT_FOUND`.
### Door type
**Two-way (implementation detail), as before.** The registration
*target* mechanism (fork as base registry) sits within the same
wrapper-shape detail the 2026-08-13 amendment already marked two-way.
The one-way decisions (per-ALPN op names, the `channel_open` marker,
removal of `channel/open`/`direction`) are unchanged.
### References
- Review 004 F-01/F-02/F-03/F-06
(`docs/reviews/004-per-connection-dispatch-and-client-serving-review.md`)
— the verification and the mechanism decision
- ADR-022 amendment (2026-09-03) — the bootstrap-op set (`services/list`,
`services/schema`, `op/register`) and the connect-side serving loop
- ADR-019: operation registry layering (the overlay stays the nested
invocation / peer-announced-ops landing zone)
- The e2e gate: `fork_registry_open_op_resolves_and_is_discoverable`
(`src/channels/client.rs`) — open op resolves through the fork,
per-session openable in `services/list`, `services/schema` validates
## Amendment (§4 per-connection registration, 2026-08-13)
@@ -0,0 +1,151 @@
# ADR-048: Dispatch Spine (feature-gated `gateway` module)
## Status
Accepted
## Context
The first real downstream consumer of this crate — alkhttp — built its
HTTP gateway (`POST /call`, `/subscribe`, `/publish`, the MCP `call`
tool, the to_openapi projections) on a small internal component it
calls the **dispatch spine**: a thin struct over
`Arc<OperationRegistry>` that owns the *non-HTTP* half of gateway
dispatch. The HTTP layer resolves bearer tokens to an `Identity`,
frames NDJSON/SSE, maps `CallError` to HTTP statuses, and wraps axum
handlers; the spine does everything that happens after that:
- constructs the root `OperationContext` identically for every
transport (`internal: false` — ACL runs against the caller's
identity, not a handler's composition authority; `forwarded_for:
None` — wire-ingress only; identity supplied per-call),
- resolves the registration's `composition_authority` /
`capabilities` / `scoped_env` into that context,
- bounds Once-ops and sink dispatch with a deadline while leaving
streaming subscriptions unbounded (ADR-021: subscriptions are
long-lived),
- and applies the `services/schema` disclosure guard when the
dispatched operation's *input* names another operation.
This is transport-neutral work. It contains no HTTP concepts: no
statuses, no headers, no body framing. And it is not HTTP-shaped by
accident — the same shape is exactly what a **hub** needs when it
terminates channel 0 on both legs and relays calls to spokes
(ADR-042's "translate, not forward" rule): the hub must re-root the
context at itself (its own identity, `internal: false`, no
`forwarded_for`), re-resolve its capabilities for the outbound leg, and
bound the relay so a hung spoke does not wedge the browser-facing
connection. alkhttp needed it; the hub relay needs the same thing;
any protocol crate that exposes a call surface to a less-trusted
in-transport caller (a WS-native relay, a CLI bridge, a test harness)
will need it again.
Duplicating it per consumer is the failure mode this crate already
paid for once: the spine's `services/schema` guard existed because the
registry handler and the HTTP route were written against different
disclosure rules (CF-004, filed from alkhttp's consumer review). One
shared implementation is the fix; a second copy in a second crate
would re-open the drift.
alkhttp is not yet published. This is the cheapest moment to move the
component into alkcall (behind a feature, so the base crate stays lean
and the surface is opt-in) and have alkhttp consume it rather than
carry its own copy.
## Decision
Promote the dispatch spine into alkcall as a new `gateway` module,
**feature-gated**:
- `alkcall/gateway` behind the `gateway` cargo feature (default off —
the base crate stays lean; the module adds no dependencies, the gate
exists to keep the audit surface explicit and opt-in).
- `GatewayDispatch` — the spine struct: `invoke()`,
`invoke_streaming()`, `invoke_sink()`, registry access, and root
context construction. The 30 s deadline alkhttp hardcodes becomes a
constructor knob (`with_deadline`) so consumers keep their own
policy; the default remains 30 s for drop-in equivalence.
- `schema_disclosure_denial()` — the shared check that a spec the
caller could not invoke is not disclosed: `NOT_FOUND` for
Internal-visibility ops, **`FORBIDDEN` for ACL-denied ops**.
- `CallRequest` stays in alkhttp (payload framing is transport
business); `MAX_BATCH_OPERATIONS` stays in alkhttp (batch is a
projection concern).
### ACL denial returns FORBIDDEN, not spec-404
The wire-path `services/schema` handler (CF-004, discovery.rs) returns
spec-404 for both Internal visibility and ACL denial — the right
answer for an unauthenticated wire caller, where even acknowledging
the op's existence is a leak vector. The spine's guard is invoked by a
transport that has *already* resolved the caller's identity and often
already admitted the op elsewhere (the alkhttp GET `/schema` route
answers `403` for ACL-denied ops, deliberately). The spine therefore
keeps alkhttp's split:
- Internal visibility → `NOT_FOUND` (never acknowledged, at any
authority).
- ACL denial → `FORBIDDEN` (informative for a caller whose identity
the transport resolved; alkcall's own registry `invoke()` produces
the same code for the same caller state, so the guard's answer is
never *more* restrictive or *less* restrictive than the invoke that
would follow it).
The two layers are consistent by construction: the wire handler's
spec-404 is the conservative outer bound, the spine's FORBIDDEN is the
identity-aware refinement. When an unauthenticated caller hits both,
`FORBIDDEN` and `NOT_FOUND` differ only in which exists — and
`AccessControl::check` with no identity returns
`"authentication required"`, which HTTP-side mappers translate to 401.
alkhttp consumes this module in a later session and drops its local
copy; until then the two implementations coexist (byte-identical in
behavior, one in each crate).
### What the spine deliberately does NOT include
- **No HTTP mapping.** `CallError` → status/body is the consumer's
(alkhttp `gateway::error`). The neutral wire error is the module's
lowest-level vocabulary.
- **No body framing, limits, or timeouts beyond the deadline knob.**
NDJSON/SSE framing, body caps, keep-alive intervals are transport
concerns (GW-15/GW-16 in alkhttp's ledger).
- **No batch semantics.** `MAX_BATCH_OPERATIONS` and the batch
envelope are projections of the registry onto HTTP/MCP payloads.
- **No `CallRequest` type.** The spine takes `(op, input)` — parsing
`{ operation, input }` is transport framing.
## Consequences
- The deadline policy moves from "alkhttp's 30 s HTTP convention" to
"spine policy, configured per-consumer" (default 30 s). Sink
dispatch is bounded by the same deadline as Once-ops — the wire
path's Pub dispatch is unbounded (`deadline: None`); consumers who
want the wire behavior pass `Duration::ZERO`-style no-deadline
configuration via `with_deadline(None)`. Documented divergence, a
choice the *consumer* now owns.
- alkhttp migrates to `alkcall::gateway` in a follow-up session; its
local spine is deleted then, not deprecated in place (unpublished
crate — no compat window needed).
- The feature adds no dependencies; `cargo test --no-default-features`
and `--all-features` both pass (the wasm-clean baseline is
untouched).
- Future transports (WS-native relays, protocol crates) reuse the
spine instead of re-deriving the context/guard discipline.
## Alternatives considered
- **Fold into `registry` as a method set on `OperationRegistry`.**
Rejected: the spine is a *policy* wrapper (deadline, root-context
shape, disclosure guard) a consumer opts into, not the registry's
dispatch core. `GatewayDispatch::invoke` ≠ `OperationRegistry::invoke`
— conflating them invites wire paths to pick up transport-shaped
policy accidentally.
- **Wait for alkhttp to publish first.** Rejected: promotes a
duplicate into the wild that immediately has to be deprecated; the
cheapest moment to move the type is before either crate's first
release.
- **Promote the whole gateway module.** Rejected: routes (SSE/NDJSON
framing, body limits) and error mapping (`IntoResponse`, status
mapping) are HTTP by definition; moving them would drag axum/hyper
as optional dependencies into a protocol crate for no reuse —
alkhttp is their only consumer.
@@ -0,0 +1,408 @@
# ADR-049: Channel-Open Establishment Phase (`OpenEstablisher`)
## Status
Accepted — implemented in alkcall 0.5.0 (Unit 1: E-01 + N-1; see the
"Amendment (Unit 1 implementation, 2026-09-06)" at the bottom).
Amends ADR-047 §3 — the open-op wrapper gains an awaited
establishment phase ahead of the spawned pump handler; resolves review
006 E-01 and N-1
## Context
ADR-047 §3 made openable ALPNs operations: the ALPN crate supplies an
open handler, and the channels wrapper does the channel machinery
(allocation, ledger, policy, spawn). The §3 decision text describes the
handler's job as "validate params, consult ownership, prepare the
backend, return a 'channel plan'" — an awaited preparation step the
wrapper consults **before** replying. The implemented `OpenHandler`
type (`Arc<dyn Fn(Value, Connection, AuthContext) -> JoinHandle<()>`)
collapsed that preparation into a fire-and-forget spawn: the wrapper
collects the `JoinHandle`, records it for teardown, and writes
`{ "channel_id": <id> }` to the wire the moment the handler task is
*spawned* (`run_open_wrapper`, `src/channels/operations.rs`). The
establishment phase the ADR described never became a thing the wrapper
could consult.
The consequence (review 006 E-01, verified at tree `88e3f5e`): the open
op **cannot fail after allocation**. Any establishment failure inside
the handler — params valid at the schema level but semantically
rejected, backend lookup failure, a target dial refused for a
`direct-tcpip`-shaped tunnel, a resource no longer available — is
invisible to the open reply. The consumer observes: the call op
succeeds with `{channel_id}`, the channel is adopted, and then the
channel EOFs (the handler exits without writing; the mux pump writes
the implicit-EOF chunk; `MpscRecvStream::poll_read` returns clean EOF
for both the sentinel and sender-drop arms). A dial failure is
byte-for-byte indistinguishable from a target that closed immediately
after connecting — the two most different failure/success stories map
to the same consumer-visible event.
Every established tunnel/forwarding protocol puts establishment failure
in the open reply, not in the data stream:
- **SSH** (RFC 4254 §5.1): `SSH_MSG_CHANNEL_OPEN_FAILURE` is a
first-class reply carrying a reason code
(`ADMINISTRATIVELY_PROHIBITED` / `CONNECT_FAILED` /
`UNKNOWN_CHANNEL_TYPE` / `RESOURCE_SHORTAGE`) plus a description
string; the channel never exists on the opener's side afterward.
- **SOCKS5** (RFC 1928 §6): the reply carries REP codes 0x01–0x08;
error-then-close, never "success then in-stream error."
- **udpgw** (tun2proxy) is the counterexample: an opaque ERR bit with
zero reason information — the vocabulary to avoid.
The per-crate workaround proves the gap is load-bearing: alktty's
channels path answers establishment failures with a length-prefixed
JSON error frame **on the channel stream**
(`send_negotiation_error`, `alktty/src/adapter.rs`), disambiguated from
data by a `0x00` first-byte peek. That is a per-crate reinvention of a
protocol-level capability every ALPN crate will need: a structured,
typed, **establishment-failure reply to the open op**. It also forces
the phantom-opened channel to exist in the manager — allocation, ledger,
and policy all fire for a channel that never carries data.
alktunnels (the next consumer, arbitrary TCP/UDP tunnels over channels)
hits this on day one: its producer dials the tunnel target inside the
open op, and dial failure is the *common* case, not the edge case
(review 006 E-01; alktunnels OQ-TN-09).
This is the cheapest moment to fix upstream: three downstream crates
(alktty, alkhttp, alktunnels-in-progress), alkcall 0.4.x, no published
consumer depends on the phantom-open shape.
## Decision
### 1. The open-op wrapper gains an awaited establishment phase
`ChannelCore::register_openable` accepts an optional
**establisher** alongside the existing `OpenHandler`:
```rust
/// The establishment result. `Establishment` carries what the
/// pump phase needs (today: nothing — reserved for a channel plan).
/// `EstablishmentError` carries the reason code + message the
/// wrapper puts in the open reply's `details`.
pub type OpenEstablisher = Arc<
dyn Fn(Value, Connection, AuthContext)
-> BoxFuture<'static, Result<Establishment, EstablishmentError>>
+ Send
+ Sync,
>;
```
The establisher is **awaited by the wrapper, bounded** — before the
reply is written, before the pump handler is spawned. The pump handler
(the existing `OpenHandler`, unchanged) is spawned only on
establishment success. The dial is the natural establisher step for
tunnels; the pumps remain the spawned handler. This restores ADR-047
§3's original "channel plan" shape: the establisher is the awaited
preparation, the wrapper consults its result, the pumps are the
spawned protocol.
**Why split-hook, not await-and-inspect** (the two candidates review
006 proposed):
- The split preserves the `OpenHandler` type exactly. Await-and-inspect
changes `OpenHandler`'s return type (`JoinHandle<()>` →
`JoinHandle<OpenResult>`), breaking all three consumers' handlers and
the alkhttp `OpenableAlpn` ferry for no compensating gain.
- Establishment and pump are genuinely different lifecycles. The dial
is synchronous with the open reply (SSH's semantics: the failure is
the open's reply); the pumps outlive the reply. Coupling the
establishment signal to the pump task's lifecycle (watching a
`JoinHandle` for a first resolution) conflates them and makes the
"established, continue" signal an out-of-band convention (a sentinel
`Result` value, a oneshot the handler must remember to signal) —
more protocol per crate, the thing this fix exists to remove.
- The split is backward compatible by construction: no establisher
registered = an always-OK establisher. Existing registrations compile
and behave unchanged.
### 2. The deadline bounds the establisher, not the pumps
The establisher await is bounded by the dispatch deadline when the
`OperationContext` carries one (`context.deadline`), else by a crate
constant (`ESTABLISHMENT_TIMEOUT`, 10s default; overridable per
registration via a `Duration` argument on the establisher-taking
`register_openable` variant). On deadline expiry the wrapper treats it
as establishment failure with reason `timeout`.
The bound applies **only** to the establisher. The spawned pump
handler's lifetime is governed by the existing teardown machinery
(`channel/close`, connection drop, handler exit) — unchanged.
Head-of-line safety is already proven: the serving loop spawns Once
invocations as independent tasks
(`Dispatcher::spawn_once_dispatch`), so a slow establisher on one open
op does not block other calls on channel 0.
### 3. Establishment failure: teardown + typed `channel:open_failed`
On establishment failure (error or deadline), the wrapper:
1. Tears down the just-allocated channel
(`teardown_channel` — drops the demux sender, returns the not-yet-
installed handler task handle if any),
2. Takes the opener-ledger entry and calls `policy.on_close(opener)`
(the same un-increment path the allocation-failure arms already
run — the ledger `take` is the atomic gate, ADR-047 §7),
3. Replies with a new typed error:
```
code: "channel:open_failed"
message: human-readable establishment failure description
retryable: false
details: { "reason": <reason-code>, "message": <detail string> }
```
The reason-code vocabulary maps 1:1 onto what an establisher can
actually produce (per the SSH four; the survey's finding):
| reason | meaning |
|---|---|
| `dial_failed` | the backend/target could not be reached or refused |
| `unknown_resource` | the requested resource does not exist |
| `resource_shortage` | the backend is out of capacity (ports, fds, slots) |
| `handler_error` | establisher-internal failure not covered above |
| `timeout` | establishment exceeded the deadline |
Policy denial stays `channel:too_many_channels` (pre-allocation,
unchanged); ACL denial stays `FORBIDDEN` (registry gate, unchanged).
The new code is an additive wire addition (new error-code string +
optional `details` shape); no existing consumer breaks. ALPN crates'
open-op specs gain matching `ErrorDefinition` entries per ADR-016 so
`services/schema` discloses the failure contract.
The SSH "channel never exists opener-side" property is the contract:
the consumer's open resolves `Err` and no `channel_id` was ever
returned. (The allocation still happened accept-side momentarily —
that is invisible to the consumer and is what the teardown in step 1
cleans up.)
### 4. `ChannelClient::open_channel` stops erasing the error (review 006 N-1)
`ChannelClient::open_channel` currently flattens the `CallError` into a
`String` (`format!("open op failed: {e:?}")`), which would make the
typed reason invisible to consumers — the E-01 fix would be unreachable
end-to-end through the primary client path. It changes to return a
typed error carrying the `CallError` (a new
`ChannelOpenError { error: CallError }` or equivalent), so the
consumer branches on `channel:open_failed` + `details.reason`.
This is a breaking change to a method signature introduced in this
crate's 0.4.x — acceptable at 0.5.0 (see Consequences), and it is the
point of the change: the reason must be consumer-usable.
### 5. Compatibility and migration
- `OpenHandler`'s type is unchanged. Existing registrations compile
unchanged.
- `ChannelCore::register_openable` keeps its current signature
(no establisher = always-OK); a new
`register_openable_with_establisher(spec, establisher, open_handler,
registry, auth)` variant adds the hook. alkhttp's `OpenableAlpn`
gains an optional `establisher` field (default `None`) — the ferry
passes it through mechanically.
- alktty migrates its **channels path** semantic failures (unknown
backend, `carriage != "raw"`, `allocate_failed`, ownership denial —
currently post-open error frames) into the establisher, resolving
them as `channel:open_failed`. Its direct-ALPN path **keeps** the
in-band error frame (two transports, two contracts; the direct path
has no open op to fail). The `0x00`-peek disambiguation stays for
the direct path only.
- alkhttp is unaffected (no openable ops in the default surface; the
`OpenableAlpn` change is additive).
### 6. Panicked pump handlers stay EOF-shaped (pinned as designed)
The wrapper's teardown task swallows the pump handler's `JoinError`
(`let _ = raw_task.await`). A panicked pump = instant EOF, which is
the correct consumer-visible outcome for a mid-stream handler crash
(indistinguishable from an abrupt close — there is no error channel
mid-stream by design; establishment errors are the only kind that
belong in the open reply). This ADR pins that as intended; no change.
The establisher, by contrast, runs pre-reply — its panic (a future
that panics when polled) surfaces as the spawned Once task's panic,
which the serving loop already tolerates (the call never resolves;
the deadline / client timeout is the bound). Establisher
implementations return `EstablishmentError` instead of panicking, per
this crate's no-panic convention.
## Consequences
**Positive:**
- Establishment failure reaches the consumer as a typed, branchable
call error — retry policy, client UX, and error reporting become
possible for dial-refused, unknown-resource, and shortage cases
(previously: instant-EOF ambiguity).
- The SSH contract ("the channel never exists opener-side") holds
consumer-visibly: a failed open never returns a `channel_id`.
- No phantom channels: the ledger, policy count, and manager state are
restored atomically on failure — allocation and teardown balance.
- alktty's per-crate in-band error vocabulary is retired on the
channels path; every future ALPN crate (alktunnels first) gets the
establishment reply for free.
- ADR-047 §3's "channel plan" shape is realized: awaited preparation
before reply, spawned pumps after.
**Negative:**
- `channel:open_failed` + the reason vocabulary is a new wire-visible
error surface — additive, but it joins the stable error set
consumers may branch on (per ADR-016, `details` shapes are
discoverable via `services/schema`).
- `ChannelClient::open_channel`'s error type changes (breaking at
0.5.0; mechanical for consumers — the `String` was a
debug-formatting wrapper anyway).
- `OpenableAlpn` (alkhttp) gains a field; its two construction sites
add `None` (mechanical).
- The establisher await adds a bounded latency to open-op replies
where handlers previously replied instantly (the spawn). The 10s
default is the worst case for a hung establisher; real establishers
(dial, lookup) complete in dial-time. Consumers already tolerate
call-op latency; the deadline is the bound.
## Door type
**One-way (wire-visible error surface).** `channel:open_failed` and its
`details.reason` vocabulary join the stable error set: once consumers
branch on reason codes, changing the vocabulary requires a migration
(the same one-way-ness ADR-016 gives typed error details). The
establisher hook shape itself — `OpenEstablisher`, the
`register_openable_with_establisher` variant, the
`Establishment`/`EstablishmentError` types — is a **two-way-door
implementation detail** within the one-way decision (the wrapper shape,
per ADR-047 §3's own door-type note). The `OpenHandler` type is
untouched, which is what keeps the split cheap to revise.
## Implementation units
1. **alkcall 0.5.0** — `OpenEstablisher` +
`register_openable_with_establisher`; wrapper flow (await bounded →
teardown-on-failure → `channel:open_failed` with details);
`ChannelClient::open_channel` typed error (N-1); tests:
- establisher fails after allocation → consumer's `open_channel`
resolves `Err(channel:open_failed)` + reason details; channel
absent from `channel_ids()` afterward;
- establisher never completes → `timeout`-reason failure within the
deadline, channel torn down, ledger decremented;
- no-establisher registration behaves exactly as today (compat
gate);
- establisher success spawns pumps and replies `{channel_id}`
unchanged.
2. **alktty migration** — channels-path semantic failures move into an
establisher; `open_via_channels_surfaces_negotiation_rejected`
resolves via call error; the direct-ALPN error-frame path is
retained.
3. **alkhttp pass** — `OpenableAlpn.establisher: Option<...>` (default
`None`), threaded through the session fork (mechanical).
## References
- Review 006 E-01 (the establishment gap — findings and prior-art
survey), N-1 (the client error-type gap this ADR also resolves),
E-03/E-04 (adjacent teardown/early-arrival notes, filed separately
from this ADR's scope)
- ADR-047 §3 (openable ALPNs are operations — the "channel plan"
wrapper shape this ADR restores; §7 opener ledger — the teardown
un-increment path)
- ADR-016 (typed error schemas — the `details` vehicle)
- ADR-040/041 (backpressure/caps — untouched; the teardown path keeps
the ledger `take` as the atomic gate)
- alktty ADR-009 (the open op's input is the negotiation) + review 001
L1/L3 — the in-band mechanism retired on the channels path
- alktunnels OQ-TN-09 (dial-failure reporting — the first consumer of
the new error) and `docs/research/ssh-socks5-survey.md`
§"Open-failure path" (the reason-code prior art)
- RFC 4254 §5.1, RFC 1928 §6 — SSH/SOCKS5 open-failure semantics
## Amendment (Unit 1 implementation, 2026-09-06)
Unit 1 landed in alkcall 0.5.0. Two implementation-shape notes, both
within this ADR's two-way door (the hook shape is the revisable
implementation detail; the wire surface is unchanged from §1/§3):
1. **The establisher does not receive the channel `Connection`.** §1's
signature sketch passed `Connection` to the establisher, but the
channel's `BiStream` is yield-once (`ChannelBidiStreamSource`) — it
cannot be handed to both the establisher and the pump handler, and
the establisher is pre-data-plane by design (its dial targets the
backend, not the channel). The implemented signature is
`Fn(Value, AuthContext) -> BoxFuture<'static,
Result<Establishment, EstablishmentError>>`; the `Connection`
belongs exclusively to the `OpenHandler` (unchanged).
2. **The bound is the earlier of the dispatch deadline and the
per-registration timeout.** §2 names the dispatch deadline "when
the `OperationContext` carries one, else the crate constant";
implemented as `min(deadline_remaining, timeout_override |
ESTABLISHMENT_TIMEOUT)` — the registration override stays
meaningful for `Query`/`Mutation`-typed open ops (whose dispatch
carries a 30s deadline; `Sub` clears it), and a deadline already in
the past yields a zero bound (immediate `timeout` reason).
Implemented surface: `OpenEstablisher`, `Establishment`,
`EstablishmentError` (reasons `dial_failed`/`unknown_resource`/
`resource_shortage`/`handler_error` + the wrapper's `timeout`),
`ESTABLISHMENT_TIMEOUT` (10s), `CHANNEL_OPEN_FAILED`
(`channel:open_failed`), `ChannelCore::register_openable_with_establisher`
(`register_openable` delegates with `establisher: None`),
`ChannelOpenError` (client-side typed error: `CallFailed { error:
CallError }` / `MissingChannelId` / `AdoptFailed`, with
`call_error()` + `establishment_reason()` accessors). All four
verification gates from the review landed as tests (establisher
failure e2e through a real channels connection with ledger
un-increment + no-channel assertions, bounded timeout, no-establisher
compat, establisher-success pump round-trip).
## Amendment 2 (plan payload, 2026-09-07 — review 007 R-01/R-02)
Review 007 (from the alktunnels UDP POC) filed two follow-ups on the
establishment surface; both landed in alkcall 0.6.0.
**1. `Establishment` carries the channel plan (R-01).** §1 reserved
the payload ("today: nothing") and the wrapper consulted only
success/failure — so an establisher whose backend produces a handle
(a dialed socket, a TTY allocation) had to cross it to the pump
handler through a per-crate side channel. The alktunnels POC shipped
a resource-keyed slot + poll loop whose concurrent same-resource race
is unfixable within that shape; alktty documented the same wall
(backend `allocate` cannot cross, so failure classes stayed in-band —
the phantom-channel shape ADR-049 removed, alive one layer down).
The plan is now real: `Establishment { plan: Option<ChannelPlan> }`
with `ChannelPlan = Arc<dyn Any + Send + Sync>` — **typed-opaque, not
`serde_json::Value`**. The review's `Option<Value>` sketch could not
satisfy its own verification gate ("establisher dials, `plan` carries
the handle"): the payloads establishers actually hand off are live
handles with no JSON representation. The establisher and the
`OpenHandler` agree on the concrete type; alkcall never inspects it.
The wrapper threads `establishment.plan` to the handler's new second
parameter (`OpenHandler = Fn(Value, Option<ChannelPlan>, Connection,
AuthContext) -> JoinHandle<()>`); the separate-parameter shape wins
over merging into `input` because a typed payload cannot ride the
JSON input without a downcast-side registry and the reserved-key
collision the review already anticipated. The plan is process-local
(establisher → wrapper → handler on the producing side); the wire
surface is unchanged — nothing crosses the transport that isn't
already the open op's input. `#[non_exhaustive]` on `Establishment`
keeps a future carrier change from being another breaking release.
Construction is `Establishment::new(plan)` /
`Establishment::default()`; the 0.5.0 `Ok(Establishment {})` sites
break mechanically at 0.6.0, which is the point of landing this now
(before alktunnels Phase 1 ships the side-channel shape into a real
crate and the payload lands later anyway as a second break).
**2. The `OpenHandler` lifetime contract is documented (R-02).** The
wrapper awaits the returned `JoinHandle` and its completion triggers
teardown — so the handle must track the data-plane lifetime: a
handler that returns before its pumps finish tears the channel down
at birth (the POC's first pump implementation hit exactly this: every
tunnel connected then instantly EOF'd). The contract was implemented
but never documented; the type docs now state it ("await the pumps
inline, never spawn-and-forget and return early") on `OpenHandler`
and the registration entry points, plus a `debug!` telemetry line in
`run_open_wrapper` when a handler exits without having accepted the
channel's `BiStream` (the birth-teardown hint; the accept is
observable in-process via the yield-once source). §6's pinned
EOF-shaped panic semantics are unchanged.
@@ -0,0 +1,105 @@
# ADR-050: `pump_bidi` — the two-pump helper, extracted upstream
## Status
Accepted — implemented in alkcall 0.6.0 (review 007 Unit 3 / R-03).
Pins alknet ADR-078's two-pump contract in one place.
## Context
alknet ADR-078 defined the two-pump data-plane contract for forwarding
channels (one pump per direction; each pump shuts the opposite sink
down on completion — `try_join!` alone deadlocks) and deferred helper
extraction until a second two-pump consumer existed and the shapes
converged. Review 007 (R-03) supplied both halves of that test from
the alktunnels UDP POC:
- **Producer side:** the POC's `pump_halves` — two `tokio::io::copy`
pumps over split channel-vs-substrate halves, each shutting down
the opposite sink on completion, joined.
- **Consumer side:** `TunnelSession::take_halves` + the assembly
layer's copy — the same shape modulo channel side.
alktty's channels session already implements the shape (its
`drive_session_pre_negotiated` awaited inline); alkhttp's ferry does
not pump. The shapes converged; the review asked this sweep to either
extract the helper or record the decision not to — an un-extracted
helper would surface during alktunnels Phase 1 as the same
fix→publish→update treadmill for a purely additive change.
## Decision
Extract the helper as `channels::pump_bidi`:
```rust
pub async fn pump_bidi<C, R, W>(channel: C, peer_read: R, mut peer_write: W) -> (u64, u64)
where
C: AsyncRead + AsyncWrite + Unpin,
R: AsyncRead + Unpin,
W: AsyncWrite + Unpin,
```
Two pumps, joined: `channel → peer_write` and `peer_read → channel`.
On each pump's completion the opposite sink is shut down
(shutdown-on-completion — the half-close the contract specifies).
Returns the two copy counts for observability.
Three deliberate deviations from the review's sketch, all
semantics-first:
1. **Return `(u64, u64)`, not `io::Result<(u64, u64)>`.** There is no
meaningful `Err` state: both pumps swallow copy errors by contract
— a mid-stream error is indistinguishable from an abrupt close
(no error channel exists mid-stream, per ADR-049 §6's pinned
EOF-shaped semantics), and shutdown-of-the-opposite-sink runs
either way. A returned `Err` would be dead code; the counts
(with `unwrap_or(0)` on error) are the observability.
2. **Peer side takes split halves (`peer_read: R`, `peer_write: W`),
not one `AsyncRead + AsyncWrite` value.** The dominant producer
shape is the establisher's dial result: `TcpStream::into_split`
halves (tunnels), pty pair halves (TTY). Taking halves matches the
POC (`pump_halves`) and avoids forcing consumers to re-join halves
just to satisfy a bound. The channel side stays a single value —
that is what the pump handler's `accept_bi` yields.
3. **Not `async fn` over `AsyncWriteExt` bounds on both sides.** The
channel side needs `tokio::io::split` (one value, two pumps), so
its bound is `AsyncRead + AsyncWrite + Unpin`; the `AsyncWriteExt`
bound the sketch had on every parameter is the extension trait
only where shutdown is called.
alktty's three-pump session does not fit the helper (the exit future
is a third signal) and stays as-is — the helper serves the two-pump
shape, exactly as the review scoped.
## Compatibility
Purely additive: a new pub fn in `channels::pump`, re-exported in the
module docs; no existing surface changes. Consumers may adopt it at
their own pace (alktunnels Phase 1 gets it for free; alktty's
channels session can migrate later if it wants; the POC's shape is
now canonical).
## Door type
**Two-way.** A helper function is trivially removable or re-shapable;
no wire surface, no trait, no state.
## Verification gate
The POC's two-pump semantics reproduced through the helper, as unit
tests in `src/channels/pump.rs`: data flows both directions with
exact copy counts; EOF from one side completes the other's
shutdown-on-completion (the far end observes clean EOF, not an
error); a dead source is EOF-shaped teardown, not an `Err`.
## References
- Review 007 R-03 (the extraction ask + the convergence evidence)
- alknet ADR-078 (the two-pump contract; shutdown-on-completion)
- alktunnels `poc-summary.md` §Issues Surfaced #4 (the producer-side
shape; `pump_halves` in the POC's producer.rs)
- alktty `src/channels.rs` (the channels session shape; three-pump
carve-out)
- ADR-049 amendment 2 (R-02 — the lifetime contract the helper's
callers must respect: await `pump_bidi(..).await` inline inside the
handler task)
+1 -1
View File
@@ -83,7 +83,7 @@ is the load-bearing piece the broker composes on.
| OQ-37 | `from_call` relay wrapper for marked ops | open | medium | ADR-047 §1 names it as a consumer (hub) concern; alkcall's `from_call` reconstructs the marker (Gap F resolved) so the consumer can branch on it |
| OQ-38 | ALPN→path-segment mapping | resolved | low | ADR-047 §"Negative" — strip the `alk/` prefix; ALPNs without that prefix use the full ALPN string (rare, two-way-door) |
| OQ-39 | `channel/control` control-handle surface | open | medium | The `channel/control` handler currently returns `channel:control_not_implemented`. The control-handle surface (per-channel control callbacks registered by ALPN crates, routing `message` to the handler's control handle for `channel_id`) is real design work — each ALPN crate needs a way to register a control callback, and the channels layer needs a control-handle registry keyed by `channel_id`. Deferred until an ALPN crate (TTY, tunnel) needs out-of-band control. |
| OQ-40 | `channel/resources/subscribe` live subscription | open | medium | The `channel/resources/subscribe` handler currently returns `channel:resources_not_implemented`. The live subscription aggregated from ALPN-crate resource enumerators (ADR-047 §6) requires each ALPN crate to provide a resource enumerator, and the channels layer to aggregate them into a live `Stream` that emits on any change. The current stub is a one-shot error; the real implementation is deferred until a consumer (hub, dashboard) needs live resource discovery. |
| OQ-40 | `channel/resources/subscribe` live subscription | open | medium | The `channel/resources/subscribe` handler currently returns `channel:resources_not_implemented`. The live subscription aggregated from ALPN-crate resource enumerators (ADR-047 §6) requires each ALPN crate to provide a resource enumerator, and the channels layer to aggregate them into a live `Stream` that emits on any change. The current stub is a one-shot error; the real implementation is deferred until a consumer (hub, dashboard) needs live resource discovery. **Load-bearing for alktunnels' discovery UI** (review 006 E-02): a consumer cannot distinguish produced tunnel resources by op name alone. The static half is covered (0.5.0: `OperationSpec.description` on `services/list` — describes the op, not the live resource set); the dynamic half stays deferred — alktunnels v1 uses config-known op names. |
| OQ-41 | QUIC-native multi-stream substrate | open | medium | Only the in-line substrate mode is implemented (single bidi stream, header-demuxed N channels). The QUIC-native multi-stream substrate (accept remaining bidi streams, read headers off each — ADR-034 §substrate modes) is deferred to the downstream alknet crate. The wire format and demux loop are correct for both substrates; only the outer `accept_bi()` loop is missing. The alknet crate owns the QUIC dial/accept loop and is the natural place for the multi-stream accept loop. This crate stays transport-agnostic (no QUIC dependency, WASM-compatible). |
## Core Types
+9
View File
@@ -39,6 +39,9 @@ pub struct OperationSpec {
pub output_schema: Value, // JSON Schema for output
pub error_schemas: Vec<ErrorDefinition>, // Declared domain errors (ADR-016)
pub access_control: AccessControl,
/// Human-readable op description (review 006 E-02). Disclosed by
/// `services/list` when set; `None` when the op declares none.
pub description: Option<String>,
/// JSON pointer into the input for the resource ID, when
/// `access_control.resource_type` is set and the operation targets a
/// specific runtime-spawned resource (ADR-011). e.g., `"$.containerId"`
@@ -829,6 +832,12 @@ These are read-only — no admin operations are exposed through the call protoco
}
```
Each listing entry also carries `description` (review 006 E-02) when
the op's spec declares one (`OperationSpec.description`, set via
`with_description`) — the field is additive and absent otherwise. It
describes the op, not the produced resource set: live resource
discovery stays with `channel/resources/subscribe` (ADR-047 §6, OQ-40).
`services/schema` accepts `{ "name": "fs/readFile" }` (no leading slash —
registry form, same as `OperationSpec.name`) and returns the full
`OperationSpec` including input/output JSON Schemas and declared
@@ -0,0 +1,527 @@
# Review 004 — Per-Connection Dispatch Resolution and Client-Side Op Serving (found via alkhttp's WS data-channel drill-down)
## Status
Verified, remediated (2026-09-03 — Units 1–3; Unit 4 remains
downstream in alkhttp). See "Remediation log (2026-09-03)" at the
bottom for the landing summary and the gates.
## Scope
Focused review of two design/spec mismatches in the call + channels
integration, uncovered while drilling into the deferred half of
alkhttp's WebSocket path (alkhttp review 003, findings WS-24/WS-25 —
the OQ-05 "browser data channels" deferral). The drill-down
established that the alkhttp-side gap is mostly wiring, but two
assumptions underneath it resolve to *this* crate:
1. **Dispatch resolution for per-connection operations.** ADR-047's
§4 amendment (2026-08-13) decided that openable-ALPN open ops are
registered per-connection on the connection overlay registry
(Layer 2). The top-level dispatch path never consults that layer,
so an op registered per the amendment's mechanism cannot be
invoked over the wire. The only shape proven to dispatch is
different from the one the ADR describes.
2. **Client-side operation serving.** The connection-local overlay
(`CallConnection::register_imported` → `overlay_env()`) and the
session overlay (`compose_root_env` attaching peer-keyed overlays)
are fully built and tested — but only the *accept side* has a loop
that serves inbound `call.requested` frames. The connect side's
read pump resolves responses and silently drops requests. There
is also no wire mechanism by which a client could announce an op
in the first place. Together these leave the "both sides can be
both" symmetry (AGENTS.md §8; ADR-022 §direction semantics;
alkhttp ADR-048's browser bidirectionality) unreachable on the
single-stream channel-0 shape.
The pass re-verified every claim directly in source at tree
`c16b069`. Cross-repo context: these findings block alkhttp review
003's Unit 2 (data-channel wiring for WS sessions, both browser and
native-fallback consumers); the decision work and the fixes live
here. Findings continue alkcall review numbering (F-01..; reviews
001–003 used P/C/R/A/B/C/D prefixes — F is the new prefix for this
focused review, findings numbered independently).
## Baseline verification (this pass)
```
cargo test → 565 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo fmt --check → clean
```
The suite is green; these findings are design/spec mismatches, not
regressions. Notably, one of them (F-04's serving gap) is invisible
to the suite *because* the connect side has never been asked to
serve — the existing e2e tests all put the dispatcher on the accept
side.
## Verdict
- **The channel machinery is done and proven.** `ChannelCore` /
`register_openable` / `ChannelOperations`, the opener ledger, the
odd/even id split, the demux hardening, and the full e2e wiring
test (`channels/client.rs:767-870`) are all in place and tested.
Nothing in Part A below re-opens that machinery.
- **The gap is the dispatch-resolution *mechanism* (F-01/F-02/F-03)
and the connect-side serving half (F-04)** — both upstream of any
transport (TCP+TLS, WS, whatever). Fixing them here unblocks
alkhttp's data-channel wiring, and — since the crate is
wasm-targetable and not yet published to consumers — the fixes are
cheap now and required regardless of alkhttp.
- **The recommended resolution is assembly, not invention:** per-
session registry forks (F-02/F-03) + a bootstrap `op/register` op
(F-01's mechanism choice) + a connect-side serving loop (F-04).
Three of the four pieces have working reference shapes in-tree;
the genuinely new protocol surface is one op handler and one read
loop.
## Severity legend
- **[critical]** — a decided spec invariant is violated in a way that
makes a promised capability unreachable end-to-end; or corrupts
data.
- **[major]** — a core protocol path cannot serve a decided behavior;
works only via shapes the spec does not describe.
- **[minor]** — drift, doc/spec inconsistency, or a missing
convenience with no correctness impact.
---
# Part A — Dispatch resolution vs the ADR-047 §4 amendment
## F-01 [major] — Top-level dispatch never consults the connection overlay; ops registered per the ADR-047 §4 amendment mechanism resolve `NOT_FOUND` on the wire
**ADR drift:** ADR-047 "Amendment (§4 per-connection registration,
2026-08-13)": *"open ops are registered per-connection …
`register_openable` is called on the connection overlay registry
(Layer 2 per ADR-019), closing over the per-connection `ChannelCore`."*
**Verified:** YES, three ways:
1. `Dispatcher::dispatch` reads the operation's `op_type` from
`self.registry.registration(...)` only
(`src/protocol/dispatch.rs:316-320`) and invokes via
`self.registry.invoke` / `invoke_streaming` / `resolve_sink_handler`
(`:330-357`) — the dispatcher's **base registry** only. There is
no fallback to `connection.overlay_env()`.
2. The connection overlay is reachable solely as a layer of
`context.env`: `compose_root_env` attaches it via
`env.attach_peer(peer_id, connection.overlay_env())`
(`dispatch.rs:195-215`). `context.env` is consulted for *nested*
invocations (a handler calling `context.env.invoke(...)`) — not
for resolving the incoming `call.requested` itself. An open op
registered on the overlay resolves `NOT_FOUND` when the wire
request arrives.
3. The one place this wiring is proven end-to-end uses a different
shape than the amendment describes: alkcall's own e2e test
(`src/channels/client.rs:767-870`) has the `install_channel_zero`
hook build a **fresh per-connection registry containing the open
op and pass it as the dispatcher's base registry**
(`Dispatcher::new(registry, ...)` at `:837`) — not as an overlay.
That shape works; the amendment's shape cannot.
Consequence: ADR-047's data-plane promise (per-connection open ops
invocable on channel 0) is unreachable through the mechanism the
amendment names. Every consumer of `register_openable` must either
adopt the undocumented per-session-base-registry shape or fail.
**Fix (Unit 1):** decide the mechanism deliberately (see F-02) and
amend ADR-047 §4's wording to describe the mechanism that actually
dispatches. The amendment's *rationale* (per-connection state, no
`context.env` downcast) stands under either option.
## F-02 [major] — `OperationRegistry` has no fork/clone surface; the proven per-session-registry shape requires one
**Verified:** YES. `OperationRegistry`
(`src/registry/registration.rs:97-100`) is
`HashMap<String, HandlerRegistration>` + a validators map, with
`new`/`register`/`registration`/`publish_validator`/`list_operations`
(`:102-180`) — no `Clone` impl and no fork API. Every type inside a
`HandlerRegistration` *is* clone-able (verified):
- `HandlerKind` — `#[derive(Clone)]` (`registration.rs:51`)
- `OperationSpec` — `#[derive(Debug, Clone, PartialEq)]`
(`spec.rs:174-175`)
- `OperationProvenance` — `#[derive(Debug, Clone, Copy, PartialEq,
Eq)]` (`registration.rs:58-59`)
- `CompositionAuthority` — `#[derive(Debug, Clone)]`
(`context.rs:82-83`)
- `Capabilities` — manual `impl Clone` (`core/types.rs:75-80`)
- `jsonschema::Validator` — `#[derive(Clone, Debug)]`
(jsonschema 0.46.10 `validator.rs:294`)
So the fork is a small addition (derive `Clone` on
`HandlerRegistration` + `OperationRegistry`, or an explicit
`fork() -> OperationRegistry` that copies both maps). Two shape
options:
- **(a) per-session base registry (recommended).** The
`install_channel_zero` hook (and any future session-establishment
seam) forks the deployment's base registry, registers the
per-connection operations on the fork (generic channel ops, the
openable ALPNs, and — with F-05 — the bootstrap op), and dispatches
channel 0 over the fork. This is the proven shape; the ADR wording
moves from "overlay registry" to "per-connection registry installed
as the session's dispatch registry." Cost: `services/list` /
`services/schema` registered on the base registry must be
re-registered (or inherited via fork) per session to stay
discoverable — the fork makes that free if the bootstrap ops are
part of the fork source.
- **(b) overlay-aware top-level dispatch.** Add an overlay fallback
in `dispatch` / `run_loop_single_stream`: when the base registry
misses, consult `connection.overlay_env()`. Keeps one static
registry and makes the amendment's wording true as written, but
touches the shared dispatch loop (higher blast radius, and the
overlay's `invoke_with_policy` shape — namespace-scoped, parent-
context-driven — does not match the top-level call shape, so the
fallback needs care: op-type lookup, ACL, visibility).
Recommendation: (a). It is the only shape with an end-to-end proof,
it needs no dispatch-loop change, and it matches ADR-019's layering
(the overlay stays what it is today: the nested-invocation landing
zone for imported ops — which F-05 needs).
**Fix (Unit 2):** `Clone`/fork surface + the ADR wording amendment;
acceptance = the review-001-style e2e gate (open op resolves over a
live channels connection through the *fork*).
## F-03 [minor] — Per-session fork + static base: no `Default`/builder seam exists to compose "base ops + per-session ops" without re-registration
**Verified:** YES. Even with a `Clone` surface (F-02), the hook must
compose three sources into the per-session registry: the deployment's
base ops (already registered, not retained anywhere as a "source"
list), the generic channel ops (`ChannelOperations::register_on`),
and the openables. A fork of the base registry makes the first source
free — this finding exists only if (a) is chosen *without* the fork
(i.e., rebuilding per-session registries from scratch). Record it as
a constraint on the fork API: it must clone *registered handlers*
(`HandlerKind` clones carry their closures), not just specs, so
`services/list` keeps its closure over the base registry.
`publish_validator`'s `validators` map must be carried too (schema
enforcement would silently vanish otherwise).
**Fix:** fold into F-02's fork surface; the acceptance test covers it
(`services/schema` on the fork still validates input).
---
# Part B — The connect side cannot serve; no wire path announces ops
## F-04 [major] — The connect side's channel-0 read pump resolves responses only; inbound `call.requested` frames are silently dropped — there is no serving half on channel 0
**Verified:** YES. `ChannelClient::from_connection`
(`src/channels/client.rs:85-134`) spawns the read pump as
`read_single_stream_until_closed(single_stream_reader, &pending_map)`
(`connection.rs:681-686`), whose only job is resolving the
`PendingRequestMap` — its frame handler (`dispatch_envelope`,
`connection.rs:705-...`) branches on `call.responded` /
`call.completed` / `call.aborted` / `call.error` and **has no arm for
`call.requested`**. The accept side's mirror
(`Dispatcher::run_loop_single_stream`, `dispatch.rs:727-850`) serves
requests and has no pending-resolution arm. Each side silently drops
the other half's frames.
Consequences:
- A consumer that dials a producer via `ChannelClient` can call ops
and open channels, but if the producer later calls an op *the
consumer serves*, the request is dropped (the caller's pending
hangs until the sweeper or connection close). This is the exact
shape ADR-022 §connection-direction-independence promises ("both
sides can be both simultaneously" — AGENTS.md §8), and ADR-036's
channel-0 pre-negotiation makes channel 0 the single control
plane where it must work.
- It blocks the bootstrap registration story (F-05): the hub cannot
call anything the WS/native client serves, and the client cannot
even receive the call that would let it register.
The fix is well-understood and symmetric: a full-duplex
single-stream read loop that branches both ways — `call.requested`
→ dispatch and write a response frame (the `run_loop_single_stream`
arms), `call.responded`/`completed`/`aborted`/`published` → resolve
pendings and route sink chunks (the read-pump arms). IDs are
UUID-generated on each side (`generate_request_id()`), so
cross-correlation is not a hazard; the abort arm needs to try both
tables (in-flight sink aborts *and* the pending map's cascade), which
`run_loop_single_stream`'s `EVENT_ABORTED` arm already does within
its own scope. The wasm-targetable browser story compounds the
value: the same serving loop is what a wasm alkcall browser
compiles.
**Fix (Unit 3):** a serving-capable single-stream loop usable from
the connect side (e.g., a `Dispatcher` variant or a
`CallConnection::serve_single_stream` that composes both frame
directions), wired into `ChannelClient` behind an opt-in (a serving
registry is not always present — a pure consumer may have none).
Acceptance gates: hub→consumer call over an existing `ChannelClient`
session resolves; consumer-served op participates in ACL (its
`AccessControl` gates the hub's call); disconnect mid-call still
fails pendings on both sides.
## F-05 [major] — No wire mechanism announces client-side ops; ADR-022's import flow is hub→consumer only and the browser-native path has none
**Verified:** YES. The envelope kind set is closed at six —
`call.requested` / `call.responded` / `call.completed` /
`call.aborted` / `call.error` / `call.published`
(`src/protocol/wire.rs:12-17`). Op registration onto an overlay
(`register_imported` / `register_imported_all`,
`connection.rs:162-172`) happens only in-process: `from_call`
(`client/from_call.rs:82-91`) builds bundles the *local* process
registers after *it* dialed out — the consumer importing a hub's
ops. There is no op, frame, or handshake by which a connected peer
announces "here are the ops I serve." Every consumer-facing
registration surface today requires the registrant to hold the
local `CallConnection` handle in-process.
Consequence: the decided bidirectionality model — either side
registers ops the other can call (AGENTS.md §8; ADR-022 §direction
semantics; alkhttp ADR-048's connection-local overlay for
browser-registered ops) — has no on-the-wire expression for the
non-in-process side. It is exactly the assumed-bootstrap-op set the
protocol already leans on (`services/list` is dialed and called by
`from_call` on every import; each side is *expected* to serve it)
extended by one op: e.g. `op/register` carrying the
`HandlerRegistration`'s serializable parts (spec + provenance +
access control), whose hub-side per-session handler writes the
registration into that connection's overlay via
`register_imported`. The overlay is already the landing zone
(`compose_root_env` attaches it keyed by identity,
`dispatch.rs:211-213`), and discovery of what a peer registered is
already built: `services/list-peers` reads
`ctx.env.peer_ids()` / `peer_operations()`
(`registry/discovery.rs:260-307`) — the same overlay-populated env.
Scope decision this forces (record it in the ADR, either way):
- **Design it** (recommended): `op/register` (+ a deregister or
replace semantics for the reconnect path) as an assumed op in the
bootstrap set, served per-session; the wire stays six kinds. The
`HandlerRegistration`'s `Handler` closures cannot cross the wire —
the client sends spec + a call-forwarding contract, and the
hub-side handler wraps it as a forwarding handler that issues a
nested `call.requested` back over channel 0 (the same shape
`from_call`'s imported bundles use — `protocol/adapter.rs:541,610`
show the in-process twin).
- **Or scope-cut**: amend ADR-022/alkhttp ADR-048 to say peer-side
op registration is hub-initiated-import only until a contract
exists. Cheaper, but it re-opens the same hedge the deferral was
criticized for — the promise would remain written-but-unmechanized.
Recommendation: design it. The pieces are all present; the cost is
one op handler, an `AccessControl`-gated registration surface (the
registry's existing ACL path covers it — an op with restrictive
`AccessControl` cannot be overwritten by an unprivileged peer), and
the F-04 serving loop to make it reachable in both directions.
**Fix (Unit 3):** with F-04 — bootstrap op + client serving loop +
ADR-022 amendment naming the bootstrap op set (assumed ops each side
may serve: `services/list`, `services/schema`, and `op/register`).
## F-06 [minor] — `services/list` on a per-session fork needs a closed-over fork reference, or per-session ops are undiscoverable
**Verified:** YES. `services_list_handler` closes over a *specific*
`Arc<OperationRegistry>` (`registry/discovery.rs:233-258`) — it
lists that registry's ops. Under the F-02(a) shape, per-session
openables live on the fork, so a `services/list` handler closed over
the base registry cannot see them. Two clean outs, both cheap: (i)
fork *before* registering the bootstrap discovery ops (then the
handler closes over the fork), or (ii) the fork API registers
base-registry bootstrap ops fresh on the fork. Either way the
constraint is: discovery ops must be registered against the
per-session registry, not the static one. Also note `services/list`
on the fork still ACL-filters (the handler re-checks per-caller) —
no privilege regression.
**Fix:** fold into Unit 2 (the fork seam registers bootstrap ops on
the fork); acceptance: a per-session openable appears in
`services/list` for an authorized caller and not for an unauthorized
one.
---
# Non-findings (verified correct, recorded to bound the re-review)
- **The e2e reference shape is real and passes** —
`channels/client.rs:767-870` wires `ChannelClient` ↔
`ChannelsAdapter` over duplex with an open op registered on the
per-connection registry, invoked end-to-end, quota reserved. The
F-02(a) shape is not speculative; it is in-tree.
- **`register_openable`'s wrapper machinery is complete** — ACL via
the registry's normal invoke path, `check_open` → `open_channel` →
ledger → spawn → teardown decrement, with `Once`/`Stream`/`Sink`
arms (`channels/operations.rs:399-453, 483-555`). Nothing in F-01
re-opens it; the ops it registers just need a dispatch-reachable
home.
- **Channel-id allocation is collision-safe** for the bidirectional
open story (accept side even from 2, connect side odd from 1,
`manager.rs:105-125`; `adopt_channel` for the non-allocating side,
`manager.rs:273-309`).
- **The `call.*` envelope kind set staying closed is correct** —
AGENTS.md §7 allows *adding* kinds but F-05's resolution does not
need one; bootstrap ops over channel 0 are the lighter door.
- **Wasm-cleanliness is unaffected** — the recommended fixes (fork +
handler + read-loop branch) touch no transport, no fs/net tokio
features; `cargo check --target wasm32-unknown-unknown` remains
the release gate.
---
# Remediation plan
Sequenced by dependency. Units 1–3 are alkcall work; the alkhttp
consequences are tracked as alkhttp review 003 Unit 1 (updated to
point here).
## Unit 1 — Decision recording (F-01, F-02 direction)
ADR work only:
- ADR-047 §4 amendment #2: the open-op registration mechanism is the
**per-connection registry installed as the session's dispatch
registry** (option (a)); the overlay registry remains the landing
zone for peer-announced ops (F-05) and nested invocation.
- ADR-022 amendment: the bootstrap-op set (each side may serve
`services/list`, `services/schema`, `op/register`), and the
direction-semantics note that serving on the connect side is
opt-in (F-04).
- alkhttp ADR-048's bidirectionality promise is re-pointed at these
ADRs instead of standing as an unmechanized commitment.
Gate: ADRs recorded; alkhttp OQ-05 and ADR-048 cross-reference them.
## Unit 2 — Fork surface + per-session composition (F-02, F-03, F-06)
- `Clone` (or explicit fork) on `HandlerRegistration` +
`OperationRegistry`, carrying handlers and validators.
- A composition helper (e.g. `OperationRegistry::fork_with(...)`) or
documented pattern: fork base → register generic channel ops +
openables + bootstrap ops against the fork.
- Gate: e2e over a live channels connection — open op resolves
through the fork; `services/list` shows per-session openables
(ACL-filtered); `services/schema` on the fork still validates.
## Unit 3 — Client serving half + bootstrap registration (F-04, F-05)
- Serving-capable single-stream loop (compose the read-pump and
dispatch arms); wired into `ChannelClient` opt-in.
- `op/register` bootstrap op (per-session handler →
`register_imported` into the connection overlay), with
`AccessControl` gating and reconnect/replace semantics decided in
the ADR.
- Gates: hub→consumer call over an existing session resolves;
peer-registered op is discoverable via `services/list-peers` and
callable back over channel 0; unregistered/unauthorized
registration attempts fail loudly; disconnect drops the overlay
and fails both sides' pendings.
## Unit 4 — alkhttp wiring (downstream, after Units 1–3)
Tracked in alkhttp review 003 (Units 2–4 there): openables surface,
`install_channel_zero` rework over the fork, retained session
handle, e2e tests, spec reconciliation (OQ-05, ADR-067/048 v1-cut
notes).
---
## Verification log (this pass)
- All claims verified at tree `c16b069`; the F-01 dispatch claim was
checked against both dispatch paths (`run_loop` stream-per-request
reads `handle_stream` → same base-registry-only resolution;
`run_loop_single_stream` at `dispatch.rs:727-850`).
- The F-04 claim was verified by reading both loops' frame arms:
`read_single_stream_until_closed` → `dispatch_envelope`
(`connection.rs:681-705`, no `EVENT_REQUESTED` arm) vs
`run_loop_single_stream` (`dispatch.rs:754-850`, no
pending-resolution arm).
- The F-02 clone claims were verified type-by-type, including the
third-party `jsonschema::Validator` (0.46.10) and the vendored
`Capabilities` (manual `Clone` preserving `Secret`'s clone).
- The F-05 wire-closure claim was verified against the full kind set
(`wire.rs:12-17`) and grep over `src/` for any wire-side
registration path (none; `register_imported*` callers are
in-process only).
- `services/list-peers`' overlay-backed discovery verified at
`registry/discovery.rs:260-307` (reads `ctx.env.peer_ids()` /
`peer_operations()`, populated by `compose_root_env` at
`dispatch.rs:211-213`).
- Baseline gates re-run for this pass: cargo test (565), clippy
(all-targets), fmt — all clean. No source changes.
## Remediation log (2026-09-03)
Units 1–3 landed in one pass. Every claim above was re-verified
against the source before remediation began; all six findings
confirmed (F-02's member-type list was missing `ScopedPeerEnv`
(`context.rs:53`) — also Clone, which made the fork surface simpler
than the review estimated).
**Unit 1 — decision recording (F-01/F-02 direction, F-04/F-05 scope):**
- ADR-047 §4 amendment #2 (2026-09-03): the per-connection registration
mechanism is the **per-session fork of the base registry installed as
the session's dispatch registry** (option (a)); the overlay registry
stays the nested-invocation / peer-announced-ops landing zone.
- ADR-022 amendment (2026-09-03): the bootstrap-op set (`services/list`,
`services/schema`, `op/register`), the opt-in connect-side serving
loop, and the `op/register` wire shape (spec in `services/schema`
JSON + `replace` flag; forwarding-handler wrap into the connection
overlay; `ALREADY_EXISTS` collision gate).
- alkhttp ADR-048's reconciliation note and OQ-05 re-pointed at these
ADRs (Unit 1b).
**Unit 2 — fork surface + per-session composition (F-02, F-03, F-06):**
- `OperationRegistry` gained interior mutability (`parking_lot::RwLock`
around the operations and validators maps) so a registry can live
behind an `Arc` and be self-referential. `register` now takes
`&self`; `registration`/`list_operations` return owned clones
(handlers are `Arc` closures — the clone is cheap).
- `OperationRegistry::fork()` — deep copy of registrations (handlers,
provenance, composition authority, capabilities, `scoped_env`) and
the cached publish-schema validators. `OperationRegistryBuilder::
from_registry` seeds the builder path from an existing registry.
- `install_bootstrap_discovery(&Arc<OperationRegistry>)` — registers
`services/list` / `services/list-peers` / `services/schema` closed
over the fork itself, so per-session openables are discoverable
(F-06) and `services/schema` answers from the fork.
- Gate: `fork_registry_open_op_resolves_and_is_discoverable`
(`src/channels/client.rs`) — open op registered on the fork resolves
over a live channels connection; the openable appears in
`services/list` on the fork; `services/schema` on the fork still
answers. Unit tests: fork independence both directions, validator
carry, post-dispatch registration through a shared `Arc`.
**Unit 3 — client serving half + bootstrap registration (F-04, F-05):**
- `Dispatcher::serve_single_stream` — the full-duplex single-stream
loop: `call.requested` → dispatch + response frames (the
`run_loop_single_stream` arms); `call.responded`/`completed`/`error`
→ pending resolution (the read-pump arms); `call.aborted` → both
tables (in-flight sink aborts and the pending map's cascade);
`call.published` → inbound in-flight sinks. Disconnect fails
pendings and drops in-flight sinks (teardown identical to the two
half-loops it composes).
- `ChannelClient::from_connection_with_serving(connection,
Option<ServingConfig>)` — opt-in serving; `from_connection`
preserves the pure-consumer default (resolution-only read pump).
Gate: `serving_loop_hub_to_consumer_call_resolves` — hub→consumer
call resolves through the consumer's serving loop, consumer→hub
still resolves in the same loop.
- `registry::op_register` — `OpRegisterRequest` wire DTO,
`op_register_spec`, `op_register_handler` (rebuild → forwarding stub
→ `register_imported` into the connection overlay, forced
`Internal`/`FromCall`), `announce_op`. `CallError::already_exists`
added for the collision gate. `spec_to_json_pub` made the
`services/schema` wire shape public; `rebuild_spec_for` and the
forwarding-handler constructors are crate-shared with `from_call`.
Gates: `op_register_announce_then_hub_call_routes_back_to_consumer`
(announce → overlay → hub call → forwarding stub → consumer serves)
plus unit tests (round-trip, collision, replace, forced
visibility/provenance).
**Baseline after remediation:** cargo test 581 passed / 0 failed;
clippy (all-targets, `-D warnings`) clean; `cargo fmt --check` clean;
wasm target gate re-run below. No wire-format changes: the envelope
kind set stays closed at six; the bootstrap ops ride channel 0's call
registry.
@@ -0,0 +1,704 @@
# Review 005 — Serving-Loop Concurrency and `op/register` Composition (post-remediation review of `f84d214`)
## Status
Units 1–3 remediated and verified (G-01..G-05). All findings closed;
the review is resolved. See Remediation log.
## Scope
Post-remediation review of commit `f84d214` (review 004 Units 1–3:
per-session fork registry, connect-side serving loop, `op/register`).
The remediation landed all six F-findings' mechanisms; this pass
reviews the landed implementation against the decided spec (ADR-022
amendment 2026-09-03, ADR-047 §4 amendment #2, review 004's
acceptance gates) rather than re-litigating the findings.
Every finding below was verified directly in source at tree
`f84d214`. The headline finding (G-01) was verified **empirically**:
a probe test was added to the tree, run, observed to reproduce the
defect, and removed — the tree at this review's baseline carries no
source changes. Findings continue alkcall review numbering with the
new prefix `G` (001–004 used P/C/R/A/B/C/D/F — each review numbers
independently).
Cross-repo context: the alkhttp re-pointing claimed by the remediation
log was verified in that repo (`5b62307` — `open-questions.md:102-112`,
`decisions/048…md:19-25`). Unit 4 (alkhttp wiring) remains downstream
and is not reviewed here.
## Baseline verification (this pass)
```
cargo test → 581 passed, 0 failed
cargo test --all-features → 598 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean
cargo check --target wasm32-unknown-unknown → clean
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
All gates the remediation log claims reproduce. The suite is green —
and again the green is partial: the gates test the flows they name,
and the one flow the ADR amendment promises most loudly (nested
composition of peer-announced ops) is the one no gate exercises. See
G-02 for why the existing gate cannot catch G-01.
## Verdict
- **Unit 2's fork surface is solid.** Interior mutability, lock
discipline, fork independence, validator carry, and the
self-referential bootstrap-discovery closure are all correct and
tested (see Non-findings).
- **Unit 3's serving loop resolves the F-04 frame gap but inherits the
dispatch loop's serial-inline shape** — and the F-05 flow it exists
to enable is exactly the shape that deadlocks under it. The ADR-022
amendment's "invocable via nested composition" promise is
unmechanized for wire-dispatched handlers (the primary caller
shape); a 30s sweeper timeout masquerades as resolution.
- **The F-05 acceptance gate does not exercise the forwarding stub** —
the one component that is genuinely new protocol surface. It
verifies announce-lands-in-overlay and the consumer serves, by
calling the announced op *directly*, bypassing the stub entirely.
- **`op/register`'s collision gate is overlay-only**, so a
peer-announced op can shadow the serving side's own ops in nested
composition (connections resolve before base in `PeerCompositeEnv`).
The architecture decisions (fork as dispatch registry; overlay as
landing zone; bootstrap-op set; opt-in serving) remain sound. The
defects are in the serving loop's concurrency model and in what the
gates measure.
## Severity legend
Same scale as review 004:
- **[critical]** — a decided spec invariant is violated in a way that
makes a promised capability unreachable end-to-end; or corrupts
data.
- **[major]** — a core protocol path cannot serve a decided behavior;
works only via shapes the spec does not describe.
- **[minor]** — drift, doc/spec inconsistency, or a missing
convenience with no correctness impact.
---
# Part A — The serving loop's concurrency model
## G-01 [major] — Same-connection nested composition deadlocks the serving loop; the forwarding stub's nested call resolves only via the 30s sweeper
**Status: REMEDIATED (Unit 1)** — see Remediation log; both gates
added and verified load-bearing against the pre-fix loop.
**ADR drift:** ADR-022 amendment (2026-09-03) §`op/register`: the
announced op is *"invocable via nested composition (`env.invoke`)"*;
the amendment's whole point is that the hub's handlers compose
peer-announced ops. Review 004 F-05's recommendation said the same:
the hub-side handler wraps announcements *"as a forwarding handler
that issues a nested `call.requested` back over channel 0"*.
**Verified:** YES, by code trace and by an empirical probe (see
Verification log).
Mechanism, in four steps:
1. `Dispatcher::serve_single_stream` awaits `dispatch()` **inline** in
its read loop for `call.requested` frames
(`src/protocol/dispatch.rs:1009-1014`) — the same serial-inline
shape as the pre-existing accept-side `run_loop_single_stream`
(`:765-770`). Only the Sink arm is spawned
(`:1031-1053`); Once responses and Sub pumps
(`pump_stream_single_stream`, awaited inline at `:1027-1029`)
hold the loop.
2. The F-05 forwarding stub (`make_forwarding_handler`,
`src/client/from_call.rs:392-415`) issues its nested call via
`connection.call_with_payload` (`:408`) — over **the same
connection** whose loop dispatched the parent. In single-stream
mode the pending resolves only when *some* loop reads the
`call.responded` frame (`connection.rs:274-303` — register
pending, write frame, await receiver; resolution is the reader's
job).
3. The only reader of that connection is the serving loop itself —
which is blocked in step 1 awaiting the parent dispatch, whose
handler is blocked in step 2 awaiting the nested call. Circular
wait. The transport buffers the response frame; nobody reads it.
4. The 30s sweeper (`DEFAULT_CALL_TIMEOUT`, `connection.rs:32`;
`evict_expired`, run every 10s per `SWEEPER_INTERVAL`,
`dispatch.rs:44`) is the only thing that breaks the cycle —
the nested call resolves as `CallError::timeout`.
**Empirical probe (this pass):** hub serves `op/register` + a
`hub/compose` handler that invokes the connection overlay's
registration (the forwarding stub — the exact resolution
`PeerCompositeEnv` → `OverlayOperationEnv` produces for nested
composition) from inside a wire-dispatched handler; consumer announces
`consumer/exec`, then calls `hub/compose` over the wire. Result:
```
PROBE: hub/compose resolved after 30.000761619s:
Err(CallError { code: "TIMEOUT", message: "request timed out", retryable: true })
```
A second run with an 8s timeout confirmed the hang is unconditional
(not load-dependent): zero progress until the sweeper. The probe was
removed after evidence capture; the tree is unchanged.
**Consequences (verified):**
- The ADR-022 amendment's flagship flow — a hub handler composing a
peer-announced op — fails for every wire-dispatched parent handler.
It "works" only for parents dispatched from *in-process* contexts
(where the serving loop is idle and free to resolve the nested
call's response) — a shape the ADR does not describe. Per the
severity legend this is a decided behavior served only via an
undescribed shape; it borders [critical] for the composition flow
specifically.
- The same circular wait applies to composition of `from_call`
*imports* on the accept side: `run_loop_single_stream` has the same
inline shape, and the connection overlay's imported-op stubs ride
the same connection. That hazard is **latent pre-`f84d214`** (the
accept side had the inline loop and `register_imported` before this
commit) — `f84d214` makes it load-bearing by (a) adding the wire
path that populates overlays with stubs, (b) putting serving loops
on the connect side where both directions are live by design, and
(c) writing the composition promise into ADR-022.
- The same starvation is not limited to the stub: while the loop
serves *any* Sub (`pump_stream_single_stream` inline), the side's
own outbound pendings queue unresolved. A serving-enabled consumer
that calls out while serving a long-lived Sub to the same peer eats
30s timeouts on its own calls. The F-04 gate tests the two
directions only **sequentially** (hub→consumer resolves, *then*
consumer→hub) — never interleaved.
- Head-of-line blocking: the peer's subsequent requests queue behind
each inline dispatch.
**Fix (Unit 1):** spawn the Once dispatch and the Sub pump the way the
Sink arm already is (the `SharedFrameWriter` already serializes
concurrent frame writes, so per-request frame ordering is preserved);
track spawned task handles for teardown on loop exit (the Sink arm's
`in_flight_sinks` pattern, generalized). Keep read-loop resolution
unblocked at all times. Apply the same rework to
`run_loop_single_stream` (or factor the shared loop) — the accept side
has the identical latent hazard via imported-op composition.
Acceptance: the probe above (recreated as a permanent gate) resolves
in < 1s; a concurrent-interleaving gate (outbound call resolving while
a Sub is being served) passes without sweeper intervention.
## G-02 [major] — The F-05 acceptance gate bypasses the forwarding stub it claims to prove
**Status: REMEDIATED (Unit 1)** — the stub-exercising gate and the
interleaved-directions gate are added; see Remediation log.
**Verified:** YES. `op_register_announce_then_hub_call_routes_back_to_consumer`
(`src/channels/client.rs:1412-1543`) asserts the announce resolves and
the overlay holds the op — then calls
`accept_conn.call("consumer/exec", …)` (`:1529`). That is a **direct
wire call**: the hub's dispatcher resolves `consumer/exec`… which is
not on the hub at all — the frame crosses to the consumer, whose
serving loop dispatches it from its *local* registry. The forwarding
stub (`op_register_handler`'s `register_imported` bundle) is never
invoked. The gate's comment — *"the forwarding stub routes the nested
call back over channel 0"* — describes a path the test does not take.
**Consequence:** the one behavior that is genuinely new protocol
surface (stub routing) is untested, and the one defect it hides (G-01)
is precisely in that path. This is the same failure class review 001
flagged ("all coverage is unit-level / the integration seams are where
the real problems live") — the remediation added e2e gates, but the
F-05 gate measures an adjacent, easier path. The F-04 gate is
likewise sequential-only (see G-01's third consequence).
**Fix (fold into Unit 1):** add a gate that exercises the stub —
announce → a wire-dispatched hub handler composes the announced op via
`ctx.env` (or directly via the overlay registration) → resolves from
the consumer. This is the removed probe, productized; it is the
acceptance test G-01's fix must pass. Optionally extend the F-04 gate
with an interleaved-directions phase.
---
# Part B — The `op/register` composition surface
## G-03 [major] — Announced ops can shadow the serving side's own ops in nested composition; the collision gate checks the overlay only
**Status: REMEDIATED (Unit 2)** — see Remediation log; the ADR-022
amendment records the collision policy.
**ADR drift:** ADR-018 (composition authority) and ADR-019 (the
overlay is the *landing zone* for imported ops — a layer beneath the
deployment's own registry, not a rival to it); ADR-022 amendment
(silent on name collisions with the serving side's base registry).
**Verified:** YES. `op_register_handler`'s collision gate is
`connection.overlay_contains(&request.spec.name)`
(`src/registry/op_register.rs:142`) — it consults **only the
connection overlay**. The serving registry (the fork the session
dispatches over) is not consulted. Meanwhile `PeerCompositeEnv`
resolves **connections before base** (`src/registry/env.rs:230-245` —
session, then `connection_order`, then base), and `compose_root_env`
attaches the dispatching connection's overlay
(`dispatch.rs:207-214`).
Consequence: a peer announces an op named `fs/readFile` (or any name
the serving side has registered — `services/list` itself is fair
game). The registration lands in the overlay (forced `Internal`,
`FromCall` — both per spec). From then on, every handler
**wire-dispatched from that peer's connection** that nested-composes
that name via `ctx.env` resolves to the peer's forwarding stub instead
of the serving side's own op. The forced `Visibility::Internal` does
not help here — it blocks wire invocation, not `env.invoke`
(`OverlayOperationEnv` gates on `AccessControl` only,
`connection.rs:749-836`; the composed child context is `internal:
true` by design).
Concretely: a hub handler processing peer A's request composes
`fs/readFile` expecting *its own* registry op (with the deployment's
`scoped_env`, capabilities, and ACL); it silently gets peer A's stub —
peer A chooses where the call goes and what it returns. Composition
authority (ADR-018) is decided by the composing handler's *deployer*;
a same-name peer registration silently rewrites it. The effect is
scoped to handlers whose root context is that connection's
(`compose_root_env` attaches that connection's overlay), which bounds
the blast radius — but that is exactly the flagship F-05 flow.
**Fix (Unit 2):** decide the collision policy and record it in the
ADR-022 amendment. Recommended: `op_register_handler` also rejects
names registered on the serving registry (pass the fork — or a
name-set closure — into the handler alongside the connection, the
same way `install_bootstrap_discovery` closes over its registry);
`ALREADY_EXISTS` for base collisions regardless of `replace` (the
reconnect path needs to re-announce *peer* ops, not to overwrite the
deployment's). The alternative — changing `PeerCompositeEnv` to
resolve base before connections — would break `from_call`'s
namespace-prefixed imports (prefixed names don't collide) is not
needed for that, but would change long-decided layering semantics; the
registration-side gate is the cheaper, narrower door.
## G-04 [minor] — `resource_id_path` does not survive the spec wire round-trip
**Status: REMEDIATED (Unit 3)** — see Remediation log.
**Verified:** YES. `spec_to_json_pub` serializes
name/namespace/op_type/visibility/schemas/error_schemas/access_control
(+ `channel_open`/`publish_schema` markers;
`src/registry/discovery.rs:211-234`) — `resource_id_path` is not
serialized. `rebuild_spec_for` cannot parse it
(`src/client/from_call.rs:209-286`), and the field exists on
`OperationSpec` (`src/registry/spec.rs:190`). An announced (or
`from_call`-imported) op that declares ownership-scoped resource
extraction silently loses it; the rebuilt spec's ACL checks run with
`resource_id: None`.
This is a pre-existing gap in the `from_call` import shape, not
introduced by `f84d214` — but `op/register` makes it load-bearing for
a second wire surface (peer-announced specs) that explicitly promises
"the serializable parts" of a registration. Note `namespace` *is*
serialized but `rebuild_spec_for` ignores it in favor of the
`namespace_prefix` parameter — consistent for `op/register`
(`replace`/re-announce uses `None`), worth a comment either way.
**Fix (Unit 3):** add `resource_id_path` to both halves of the
round-trip plus a round-trip test. Additive optional field in the
`services/schema` JSON shape — no existing consumer breaks.
## G-05 [minor] — `install_bootstrap_discovery` registers `services/list-peers`; the ADR-022 amendment's bootstrap set doesn't name it
**Status: REMEDIATED (Unit 3)** — see Remediation log.
**Verified:** YES. `install_bootstrap_discovery` registers
`services/list`, `services/list-peers`, and `services/schema`
(`src/registry/discovery.rs:290-313`; `list-peers` at `:300`). The
ADR-022 amendment's bootstrap set lists exactly three ops —
`services/list`, `services/schema`, `op/register`
(`022-…contract.md:380-393`). Review 004's remediation log says
"list-peers" too; the ADR text was never updated.
Consequence: none at runtime (installing a fourth op is strictly
additive, and `services/list-peers` is what makes F-05's
"peer-announced ops are discoverable" true). It is a
doc/spec-consistency drift on a set the ADR declares *"closed."*
**Fix (Unit 3):** add `services/list-peers` to the ADR-022
amendment's bootstrap-op list (recommended — `from_call`-import
discovery already relies on it) or stop installing it; one line
either way.
---
# Non-findings (verified correct, recorded to bound the re-review)
- **The fork surface (F-02/F-03) is solid.** Lock discipline is
correct: `registration()` clones under a read guard and drops it
before handler invocation (`registration.rs:190-192`), so
self-referential discovery handlers cannot deadlock their own
registry; `register` takes the two maps' write locks sequentially
and never across an `await`. `fork()` deep-copies both maps
(handlers are `Arc` closures — cheap, and the F-03 constraint that
closures must carry is met); validators carry; fork independence is
tested both directions; `OperationRegistryBuilder::from_registry`
seeds the builder path.
- **`install_bootstrap_discovery` (F-06) is correct** — handlers
closed over the same `Arc` they're registered on, so post-install
per-session registrations are visible at call time; ACL filtering
stays per-caller; the fork gate
(`fork_registry_open_op_resolves_and_is_discoverable`) exercises
exactly the review-004 acceptance shape.
- **`serve_single_stream`'s pending-resolution arms match the two
half-loops they compose.** Frame by frame against
`run_loop_single_stream` (`dispatch.rs:765-905`) and
`dispatch_envelope` (`connection.rs:717-741`): responded/completed/
error resolve pendings; aborted tries both tables (in-flight sink
aborts, then the pending cascade via `handle_abort`); published
routes inbound sinks with the same validation/keep semantics; teardown
(fail_all + sink drop + sweeper abort) matches. The direction
disambiguation by table membership (documented at
`dispatch.rs:940-966`) is correct as written — G-01 is *not* a
frame-routing bug; the arms are right, the loop's scheduling is the
defect.
- **The pure-consumer default is unchanged.** `from_connection`
delegates with `None` and keeps the resolution-only read pump; no
behavior change for existing callers.
- **`op/register`'s DTO and unit semantics are clean** — one spec
serialization on the wire (`spec_to_json` shape), forced
`Internal`/`FromCall` on landing (unit-tested), `replace` semantics
tested, `ALREADY_EXISTS` collision gate tested (against the
overlay — see G-03 for the gap).
- **The alkhttp cross-repo claims check out.** `open-questions.md`
OQ-05 and ADR-048's re-point are present in that repo at `5b62307`.
- **No non-English text reached the tree.** The remediation session's
language slip (reported anecdotally) did not land in any code,
doc, or metadata: a CJK-range regex sweep over all `.rs`/`.md`/
`.toml`/`.json` in the repo returns zero matches at `f84d214`.
- **Gates reproduce.** All commands in Baseline verification pass at
the review tree, matching the remediation log's claims
(581 default / 598 all-features).
---
# Remediation plan
Sequenced by dependency. All units are alkcall work; Unit 4 (alkhttp
wiring) stays downstream and should **not** start before Unit 1 —
alkhttp's serving consumers would compose over the same connection
and hit G-01 immediately. (All three units landed — see Remediation
log.)
## Unit 1 — Concurrent serving loop + a stub-exercising gate (G-01, G-02)
- Rework `serve_single_stream`'s `EVENT_REQUESTED` arm: spawn the Once
dispatch and the Sub pump (the Sink arm's existing pattern,
generalized; `SharedFrameWriter` serializes frames; track handles
for teardown). Apply the same rework to `run_loop_single_stream` or
extract the shared loop.
- Add the G-02 gate: announce → wire-dispatched hub handler composes
the announced op via nested composition → resolves from the
consumer. This is the removed probe, productized; it is the
acceptance test for G-01.
- Optional: an interleaved-directions phase on the F-04 gate
(outbound call resolving while a Sub is being served).
- Gate: the probe recipe (Verification log) resolves in < 1s; no
sweeper evictions in the gates.
## Unit 2 — `op/register` collision policy (G-03)
- `op_register_handler` gains a serving-registry name-set (closure or
`Arc<OperationRegistry>`); base-registry name collisions reject
with `ALREADY_EXISTS` regardless of `replace`.
- ADR-022 amendment: record the collision policy (peer-announced ops
may collide with peer-announced ops — `replace` governs; never with
the serving side's own registrations).
- Gates: announcing a base-registered name (External *and* Internal)
fails loudly; announcing a distinct name still lands; nested
composition of a base op is unaffected by an unrelated announce.
## Unit 3 — Round-trip completeness + doc alignment (G-04, G-05)
- `resource_id_path` through `spec_to_json_pub` + `rebuild_spec_for`
+ round-trip test.
- ADR-022 bootstrap list gains `services/list-peers` (or drop it from
the installer; recommend adding).
---
# Remediation log
## Unit 3 — Round-trip completeness + doc alignment (G-04, G-05) — LANDED
**G-04 fix.** `resource_id_path` now rides both halves of the spec
wire round-trip: `spec_to_json_pub` serializes it as an optional
`resource_id_path` string (`src/registry/discovery.rs`), and
`rebuild_spec_for` parses it into the rebuilt spec's fourth
constructor argument (`src/client/from_call.rs`). Additive optional
field — absent stays absent, no existing consumer breaks (verified by
the companion gate). `rebuild_spec_for`'s doc note from the review
("namespace *is* serialized but ignored in favor of
`namespace_prefix`") was addressed by leaving the behavior as-is: for
`op/register` the parameter is `None` (consistent) and for `from_call`
the prefix is authoritative — the round-trip tests pin the `name`
field handling.
**Gates:** `spec_round_trips_resource_id_path` (serialize → parse →
field intact) and
`spec_without_resource_id_path_stays_absent_through_round_trip`
(absent key serializes nothing; rebuilt `None`) in
`src/client/from_call.rs`.
**G-05 fix.** The ADR-022 amendment's bootstrap-op set gained
`services/list-peers` with a dated note (review 005 G-05) explaining
that the installer has registered it since the amendment landed and
the doc lagged the code. The set remains closed at four.
**Verification (post-fix):**
```
cargo test → 589 passed, 0 failed
cargo test --all-features → 606 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
---
## Unit 2 — `op/register` collision policy (G-03) — LANDED
**Fix shape.** `op_register_handler` now takes the serving registry
alongside the connection (`op_register_handler(connection,
serving_registry)` — the review's "name-set closure" recommendation,
materialized as the `Arc<OperationRegistry>` itself, mirroring how
`install_bootstrap_discovery` closes over its registry). The handler
checks `serving_registry.registration(name)` **before** the overlay
gate and rejects base collisions with `ALREADY_EXISTS` regardless of
`replace`; overlay collisions keep the existing `replace` semantics.
Announced ops may collide with announced ops, never with the serving
side's own registrations.
**ADR-022 amendment:** the 2026-09-03 amendment's `op/register`
section gained a "Collision policy (amended 2026-09-04, review 005
G-03)" paragraph recording the decided policy and the rationale (the
`PeerCompositeEnv` connections-before-base resolution would let an
unscreened announce rewrite composition resolution; visibility is
irrelevant to the gate). The ADR's status line notes the sub-amendment.
**Call sites updated:** both e2e gates pass the fork the session
dispatches over (`Arc::clone(&accept_registry)` after `fork()`), which
is the production shape — the collision set is exactly the registry
the serving loop dispatches against.
**Gates:**
- `handler_rejects_base_registry_collision_even_with_replace` —
base-registered `fs/readFile` (External) + `replace: true` →
`ALREADY_EXISTS`; the overlay stays clean and the serving
registration is untouched.
- `handler_rejects_collision_with_internal_serving_op` — an
`Internal` base op is equally protected (the review's point that
forced `Visibility::Internal` on the *announced* spec never helped:
`OverlayOperationEnv` gates on `AccessControl`, and the composed
child is `internal: true` by design).
- `handler_overlay_collision_still_governed_by_replace` — announce/
announce collisions still follow `replace` (the base gate is
scoped to the serving registry, not widened into the overlay).
- `nested_composition_of_base_op_unaffected_by_unrelated_announce` —
after a successful distinct-name announce, composing the base op
through the real `compose_root_env` shape (`PeerCompositeEnv` +
`attach_peer(conn.overlay_env())`) resolves the serving side's own
op, not a peer stub. (Unit-level, using the actual env types rather
than a new full-wire fixture — the wire shape is already covered by
the Unit-1 gates.)
**Deliberate non-change:** `PeerCompositeEnv` resolution order is
untouched (connections before base) — the review's recommendation;
reordering would change long-decided layering semantics
(ADR-019/ADR-024) for no gain the registration-side gate doesn't
deliver.
**Verification (post-fix):**
```
cargo test → 587 passed, 0 failed
cargo test --all-features → 604 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
---
## Unit 1 — Concurrent serving loop + stub-exercising gates (G-01, G-02) — LANDED
**Fix shape.** `dispatch()` was split into a synchronous start half and
an awaited invocation:
- `Dispatcher::dispatch_start`
(`src/protocol/dispatch.rs`) runs the sync prefix (identity
resolution, root context, op-type branch) and returns a
`StartedDispatch`: `Once` (the invocation as a boxed future), `Stream`
(the `ResponseStream`, returned synchronously — `invoke_streaming`
is a sync call), or `Sink` (started **inline**, unchanged). `dispatch()`
is now a thin wrapper (`Once` = await the boxed future) and keeps its
shape for `dispatch_requested`/`handle_stream`/the gateway.
- The Sink start stays inline **by design**: its `chunk_tx` must be in
`in_flight_sinks` before the next `call.published` frame can be
routed; the loop insert cannot race the feed. Pub handlers that
block forever inside their sink future are a separate (pre-existing,
unreported) shape — the read loop no longer waits on any handler
except for the few instructions of the sink start itself.
- Both single-stream loops' `EVENT_REQUESTED` arms spawn the Once
invocation (`spawn_once_dispatch`), the Sub pump
(`spawn_stream_pump`), and the sink's response writer
(`spawn_sink_response_writer`); handles are tracked in a
`spawned` list and aborted at loop exit (teardown: sink-map clear →
spawn aborts → `fail_all` → sweeper abort). `in_flight_sinks` moved
behind `Arc<parking_lot::Mutex>` because the ABORTED/PUBLISHED/ERROR
arms must not hold the guard across the `chunk_tx.send().await`
(non-`Send` guard across an await — the reason the pre-fix loop
couldn't just be `tokio::spawn`ed piecemeal). Guard-dropping
(`let entry = ...remove()` before the await) keeps lock discipline:
no lock is held across an await anywhere in the loops.
- **`run_loop_single_stream` (the accept side) got the
pending-resolution arms** (`EVENT_RESPONDED`/`COMPLETED`/`ERROR` →
the outbound pending map). Previously it served only — an accept-side
nested-composing handler (e.g. a `from_call` imported-op stub riding
the same connection) had no loop resolving its response frames at
all. The G-01 accept-side latent hazard is mechanized shut, not just
unblocked: both single-stream loops are now the same shape
(dispatch-spawn + pending-resolution), the full-duplex loop
`serve_single_stream` composes them and is unchanged in its arm
semantics (frame-arm equivalence preserved per the non-findings
audit).
- Write-failure semantics changed from `break` (close the loop) to
`warn` (keep reading) in the spawned Once path, matching the Sink
arm: a dying transport surfaces on the next read as
`ConnectionClosed`; a transient frame-write failure no longer tears
down every in-flight request on the connection.
**Gates (G-02, productized from the removed probe):**
- `hub_handler_composes_peer_announced_op_via_nested_composition`
(`src/channels/client.rs`): announce → consumer calls `hub/compose`
→ the hub's serving loop wire-dispatches it → the handler resolves
`consumer/exec` via `ctx.env` nested composition → the forwarding
stub's nested `call.requested` crosses back to the consumer. This is
the stub path the F-05 gate bypassed. Two wiring facts surfaced (both
recorded as spec-consistent, both were silent before): (a) the hub's
channel-0 `Connection` must carry the peer identity —
`compose_root_env` attaches the connection overlay keyed by
`identity.id` (ADR-030 §5), and a deployment that skips identity
resolution silently gets no peer overlay (the test sets it via the
`AuthContext` the adapter passes to `install_channel_zero`);
(b) composition reachability is declared on the composing handler's
registration (`scoped_env: ScopedPeerEnv::new(["consumer/exec"])`) —
the empty `ScopedPeerEnv` is deny-by-default in
`PeerCompositeEnv::invoke_with_policy`, so a wire-dispatched handler
with no `scoped_env` composes nothing (the reachability gate, not the
overlay, was what the probe's first draft tripped over).
- `outbound_call_resolves_while_inbound_subscription_is_being_served`
(the optional interleaved-directions phase on the F-04 shape): the
consumer subscribes to the hub (Sub served by the consumer's loop),
then calls `hub/interleave` over the wire — its handler issues an
outbound hub→consumer call on the same connection while the Sub is
live — and asserts the call resolves and the Sub is still live
after. Both gates bounded at 5s (vs. the 30s sweeper), so a
regression fails fast, not via sweeper eviction.
**Load-bearing verification (both gates, empirical):** each gate was
run against the pre-fix loop (`git stash push
src/protocol/dispatch.rs`) and reproduced the G-01 hang:
```
hub_handler_composes…: panicked "nested composition through the
forwarding stub timed out (G-01 shape): Elapsed(())" — 5.01s, no progress
outbound_call_resolves…: panicked "interleaved outbound call starved
while Sub served (G-01 shape): Elapsed(())" — 5.06s, no progress
```
With the fix, both resolve (0.01s / 0.11s wall including connection
setup). No sweeper evictions (bounded timeouts ≪ 30s).
**Consequences audited against the fix:**
- Same-connection nested composition of peer-announced ops works for
wire-dispatched parents (the flagship ADR-022 amendment flow) — the
gate proves it end-to-end.
- The accept-side imported-op composition hazard (G-01's second
consequence, latent pre-`f84d214`) is mechanized shut by the
pending-resolution arms in `run_loop_single_stream` — not merely
unblocked. There is no dedicated gate for this shape yet (the
accept-side import + nested-compose e2e would be a further gate;
noted as residual work, not a regression).
- Head-of-line blocking is gone: spawned arms proceed concurrently;
the loop never awaits a handler.
- Frame ordering: per-request ordering is preserved by the
`SharedFrameWriter` (each `write_frame` is atomic under the mutex);
responses for two requests may now interleave *at frame
granularity*, which single-stream mode always permitted (both
directions' frames are multiplexed by design; correlation is by id).
**Verification (post-fix):**
```
cargo test → 583 passed, 0 failed
cargo test --all-features → 600 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean (fmt applied)
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
**Residual (not blocking Unit 1):** the sink-start-inline shape means a
`Pub` op whose registration itself blocks (ACL/schema compile are sync
and fast; `resolve_sink_handler` runs the handler's ACL path) still
holds the loop briefly — bounded by registry work, not handler work. A
malicious `AccessControl::check` implementation could stall the loop;
that is a deployment-provided trait object, the same trust boundary as
`IdentityProvider`, and unchanged from the pre-existing shape.
---
## Verification log (this pass)
- All gates in Baseline verification reproduced at tree `f84d214`
(clean working tree; the probe was added, run, and removed —
`git status` verified clean after removal).
- G-01 verified by code trace (the four steps above, each line
read at `f84d214`) **and** empirically: probe
`probe_hub_handler_composes_peer_announced_op` (in
`src/channels/client.rs#[cfg(test)]`, since removed) — hub registry
carries `op/register` + `hub/compose` (whose handler invokes the
connection overlay's registration for the announced op, the shape
`PeerCompositeEnv` → `OverlayOperationEnv` resolution produces);
consumer announces `consumer/exec` then calls `hub/compose`.
8s-timeout run: `Elapsed` (no progress). 45s-cap run: resolved at
**30.0007s** with `CallError::TIMEOUT` (retryable) — the
`DEFAULT_CALL_TIMEOUT` sweeper eviction, not live resolution.
- The existing gates' blind spot verified by reading both e2e gates'
call directions: `serving_loop_hub_to_consumer_call_resolves`
(`client.rs:1373-1399`) and
`op_register_announce_then_hub_call_routes_back_to_consumer`
(`client.rs:1529`) — the latter's `accept_conn.call("consumer/exec")`
crosses the wire to the consumer's local registry; the forwarding
stub is unreachable from it.
- G-03 verified by reading the collision gate (`op_register.rs:142`)
against `PeerCompositeEnv::invoke_with_policy`'s resolution order
(`env.rs:230-245`) and `compose_root_env`'s attach
(`dispatch.rs:207-214`).
- G-04 verified by field audit: `OperationSpec`'s fields
(`spec.rs:175-206`) vs `spec_to_json_pub` (`discovery.rs:211-234`)
vs `rebuild_spec_for` (`from_call.rs:209-286`).
- G-05 verified against `discovery.rs:290-313` and ADR-022
amendment's set (`022-…contract.md:380-393`).
- Frame-arm equivalence of `serve_single_stream` vs the two half-loops
verified read-through (Non-findings).
- alkhttp cross-repo claims verified in `/workspace/@alkdev/alkhttp`
at `5b62307` (OQ-05 `open-questions.md:102-112`; ADR-048
`decisions/048-…md:19-25`).
- CJK sweep: no CJK-range codepoints in any `.rs`/`.md`/`.toml`/
`.json` under the repo at `f84d214`.
@@ -0,0 +1,561 @@
# Review 006 — Channel-Open Establishment Gap (from alktunnels Phase 0)
## Status
**Resolved-by-ADR (E-01, N-1) / Implemented (E-02, E-03, E-04, N-2).**
Findings filed from the alktunnels Phase 0
research pass (2026-09-06). This is a design review, not a code-defect
review: the establishment gap (E-01) is real, POC-observable, and
load-bearing for the next downstream crate; the remaining findings are
smaller mechanism/coverage gaps noticed in the same sweep.
**2026-09-06 verification + remediation pass.** All four findings were
independently re-verified against source at tree `88e3f5e` (0.4.1 +
this review's own commit) — verdicts CONFIRMED for E-01..E-04, with
one correction to E-02's cost estimate (see the verification appendix
at the bottom of this file). Three additional findings filed from the
same sweep: **N-1** (client error-type gap — blocking for E-01's
consumer visibility), **N-2** (dead counter), **N-3** (panic posture,
pinned). **E-01 + N-1 are resolved by ADR-049**
(`docs/architecture/decisions/049-channel-open-establishment-phase.md`
— split-hook `OpenEstablisher`, bounded await, `channel:open_failed`
typed error). **Unit 1 (E-01 + N-1) is implemented** in alkcall 0.5.0
(`OpenEstablisher` + `register_openable_with_establisher`, wrapper
establishment phase with bounded await, `channel:open_failed` typed
error, `ChannelOpenError` typed client error — all four verification
gates landed as tests; two implementation-shape notes recorded in
ADR-049's amendment: the establisher does not receive the channel
`Connection` (yield-once BiStream belongs to the pump handler), and
the bound is the earlier of dispatch deadline and per-registration
timeout). **Unit 2 (E-02) is implemented** in alkcall 0.5.0
(`OperationSpec.description: Option<String>` +
`with_description`, `spec_to_json_pub` emits it when set,
`rebuild_spec_for` parses it (shared by `from_call` and `op/register`),
`services/list` and `services/list-peers` local listings emit it when
set; the ADR-047 §6 amendment below records the discovery decision).
**Unit 3 (E-03, E-04, N-2) is implemented** in alkcall 0.5.0 (E-03: the
wrapper's handler-exit teardown discards `UnknownChannel` through a
debug log + a benign-race pinning comment; E-04: the 64-parked-chunks
observable bound is documented consumer-side in `channel-client.md` §
"The early-arrival park bound (push-first producers)" and on the
`EARLY_ARRIVAL_CAP` const; N-2: `early_arrival_count` is exposed via
`ChannelManager::early_arrival_count()` — the observability choice,
paired with `dropped_unknown_chunks` — with a monotonicity test).
The original remediation sketch below is superseded by the
"Remediation plan (post-verification)" section; the original sketch
is retained for the record.
Findings continue the review numbering with prefix `E` (001–005 used
P/C/R/A/B/C/D/F/G — each review numbers independently).
## Scope
The channels open-op path (`run_open_wrapper` and the `OpenHandler`
contract) reviewed from the perspective of the next consumer crate
(alktunnels — arbitrary TCP/UDP tunnels over channels), cross-checked
against the two existing consumers (alktty, alkhttp) and the SSH
forwarding prior art recorded in
`/workspace/@alkdev/alktunnels/docs/research/ssh-socks5-survey.md`.
Everything below was verified directly in source at tree `a22b2b8`
(0.4.1 + the early-arrival park fix). No code changes were made in
this repo by this review. The 2026-09-06 re-verification pass
(verification appendix) re-traced all findings at tree `88e3f5e` —
no source changes touched the reviewed paths between the two trees
(the only delta is this review's own commit).
```
Verified against: alkcall a22b2b8 (0.4.1); re-verified at 88e3f5e
Reading list: src/channels/operations.rs (run_open_wrapper, OpenHandler,
make_open_handler_once/stream/sink), src/channels/manager.rs
(open_channel, teardown_channel, route_payload, early-arrival park),
src/channels/reassembly.rs (MpscSendStream::poll_shutdown, Drop,
MpscRecvStream::poll_read EOF arms), src/channels/mux.rs (implicit-EOF
pump path), src/protocol/wire.rs (CallError), src/registry/
registration.rs (invoke/invoke_streaming gates), src/registry/
discovery.rs (services/list, spec_to_json_pub), src/client/from_call.rs
(rebuild_spec_for), docs/architecture/decisions/047,
docs/architecture/channel-operations.md
```
## Severity legend
Same scale as reviews 004/005:
- **[critical]** — a decided spec invariant is violated in a way that
makes a promised capability unreachable end-to-end; or corrupts data.
- **[major]** — a core protocol path cannot serve a decided behavior;
works only via shapes the spec does not describe.
- **[minor]** — drift, doc/spec inconsistency, or a missing
convenience with no correctness impact.
---
# Part A — The open-op establishment gap
## E-01 [major] — The open op cannot fail after allocation: the consumer receives a live channel for a tunnel whose establishment failed, with no error channel
**Verified:** YES, by code trace and mechanism analysis.
### The mechanics
`OpenHandler` (ADR-047 §3) is the ALPN crate's hook for "validate
params, consult ownership, prepare the backend" — i.e., the
establishment phase of a channel. But the wrapper does not await it:
1. `run_open_wrapper` (`src/channels/operations.rs:483-528`):
`policy.check_open` → `manager.open_channel(alpn, opener_id, None)`
→ build the channel `Connection` → `open_handler(input, channel_conn,
auth)` **spawns** the handler and collects its `JoinHandle` →
`ResponseEnvelope::ok(request_id, json!({ "channel_id": channel_id }))`
(`operations.rs:503-527, 537`). The reply is written to the wire the
moment the handler task is *spawned*, not when the handler has done
its establishment work.
2. The `OpenHandler` type (`operations.rs:333-334`) returns
`tokio::task::JoinHandle<()>` — there is no result, no error variant,
no establishment phase the wrapper can consult. Any failure inside
the handler (params valid at the schema level but semantically
rejected, backend lookup failure, target dial failure for a
`direct-tcpip`-shaped tunnel, resource no longer available) is
invisible to the open op's reply.
3. What the consumer observes on handler-side failure: the call op
**succeeds** with `{channel_id}`, the channel is adopted
(`ChannelClient::open_channel`, `src/channels/client.rs:227-250`),
and then the channel EOFs — the handler wrote nothing before
exiting, the mux pump writes the implicit-EOF chunk on receiver end
(`src/channels/mux.rs:78-81`, REQ-CH-01 implicit-EOF path), and
`MpscRecvStream::poll_read` returns clean EOF for both the sentinel
and sender-drop arms (`src/channels/reassembly.rs:111-135`). A
dial-failure EOF is byte-for-byte indistinguishable from a target
that closed immediately after connecting — the two most different
failure/success stories map to the same consumer-visible event.
### Why this is the wrong shape (SSH's semantics, prior-art checked)
Every established tunnel/forwarding protocol puts establishment
failure in the open reply, not in the data stream:
- **SSH** (RFC 4254 §5.1): `SSH_MSG_CHANNEL_OPEN_FAILURE` is a
first-class reply to the open, carrying a reason code
(`ADMINISTRATIVELY_PROHIBITED` / `CONNECT_FAILED` /
`UNKNOWN_CHANNEL_TYPE` / `RESOURCE_SHORTAGE`) plus a description
string; the channel never exists on the opener's side afterward
(verified end-to-end in russh:
`src/server/encrypted.rs:1244-1278`, `src/client/encrypted.rs:414-434`
— see `/workspace/@alkdev/alktunnels/docs/research/ssh-socks5-survey.md`
§"Open-failure path").
- **SOCKS5** (RFC 1928 §6): the reply carries REP codes 0x01–0x08 and
the connection closes within 10s on failure — error-then-close, never
"success then in-stream error."
- **udpgw** (`tun2proxy/src/udpgw.rs:21-26`) is the counterexample: an
opaque ERR bit with zero reason information — the survey flags it as
the vocabulary to avoid.
- **alktty was forced to reinvent the missing mechanism in-band.** The
direct-ALPN path answers the negotiation with a length-prefixed JSON
error frame *on the stream* (`send_negotiation_error`,
`alktty/src/adapter.rs:177-187`; the `0x00`-prefix first-byte
disambiguation trick, `alktty/docs/architecture/tty-adapter.md:186-197`),
and the channels path retains the same error-frame shape even though
the open op's `input` is the negotiation (alktty ADR-009) — because
post-allocation failures have nowhere better to go. That is
per-crate reinvention of a protocol-level capability every ALPN
crate will need: a structured, typed, **establishment-failure reply
to the open op**.
### The cost today, concretely
- A consumer cannot distinguish "ACL denied" (call error,
`channel:forbidden` — never allocated) from "dial refused" (open
succeeded, instant EOF) from "target accepted then instantly closed"
(open succeeded, instant EOF). Retry policy, client UX, and error
reporting are impossible on the second and third.
- ADR-016's typed error details (`CallError.details`) — which the
wrapper already uses for `channel:too_many_channels` with
`{count, max}` details (`operations.rs:530-545`) — cannot carry
dial-failure reasons.
- The alktty in-band error-frame path exists *only* because the
wrapper can't fail the open post-allocation; every future ALPN crate
faces the same fork: reinvent an in-band error vocabulary or
silently-EOF.
### The ask (proposed shape, for the ADR — not a prescriptive API)
Give the open-op wrapper an **establishment phase it awaits before
replying**. Minimal, backward-compatible shape:
1. `OpenHandler` gains an establishment result. Two candidate shapes:
- **Split the hook**: `OpenEstablisher` (async, awaited by the
wrapper — validates params semantically, prepares/dials the
backend, returns `Result<Establishment, EstablishmentError>`)
followed by `OpenHandler` (spawned on the established channel, as
today). The dial is the natural establisher step for tunnels; the
pumps remain the spawned handler.
- **Await-and-inspect**: keep the single `OpenHandler` returning
`JoinHandle<OpenResult>`; the wrapper awaits a bounded
establishment phase (a `JoinHandle::timeout` equivalent — select
on the handle vs an establishment deadline) before replying. The
handler signals "established, continue" via an agreed value
(e.g. the handler resolves a first `Result<(), HandlerError>`
promptly, or the wrapper watches a oneshot the handler signals).
2. On establishment failure: the wrapper tears down the just-allocated
channel (`teardown_channel` — the ledger/policy paths already
handle this atomically) and replies with a **new typed CallError**,
e.g. `channel:open_failed`, with `details` carrying a reason code +
message. Reason-code vocabulary per the SSH four (the survey's
finding): policy-denied (already distinct — `channel:forbidden`),
dial-failed, unknown-resource-or-substrate, resource-shortage —
mapping 1:1 onto what an open handler can actually produce. Wire
addition is additive (new error code string + optional details
shape), no existing consumer breaks.
3. Backward compatibility: the existing `OpenHandler` signature is
preserved if the split shape is chosen (old handlers still compile —
the establisher is a new, separately-registered hook, defaulting to
an always-OK establisher for the no-establishment-work case).
### Alternative considered and rejected
An in-band establishment/error frame (alktty-style, on the channel
stream, alktunnels' original OQ-TN-09 direction) works without an
upstream change, but: (a) it forces every ALPN crate to define a frame
vocabulary and a disambiguation scheme (alktty's `0x00` peek is
exactly this cost, paid once per crate); (b) it cannot carry typed
`CallError.details` or participate in ADR-016 error schemas; (c) it
leaves the phantom-opened channel in the manager (allocation/ledger/
policy all fire for a channel that never carried data); (d) SSH's
semantics — the failure is the open's reply, the channel never exists
opener-side — are the cleaner contract, and channels is at exactly the
maturity point (three downstream dependents, alkcall 0.4.x) to fix it
upstream cheaply.
### Severity
[major] — not [critical]: no data corruption, and the capability is
reachable via per-crate workarounds (alktty proves it). But it is the
shape every future protocol crate will fight, the fix is cheaper now
than after the next consumer, and the workaround path (in-band frames)
bakes in a wire format that would then need to stay stable per crate.
---
# Part B — Smaller findings from the same sweep
## E-02 [minor] — `services/list` discloses no per-op metadata; openable resources are indistinguishable by name alone
`services_list_handler` (`src/registry/discovery.rs:247-269`) maps
`list_operations()` to `{name, namespace, op_type}` only. For a tunnel
registry, the consumer's discovery question is "which tunnel resources
may I open" — and the answer arrives as a bare list of op names
(`channels/tunnel/sub`...). Per-resource identity (which target, which
substrate, human description) has no field to live in:
- `services/schema` (`spec_to_json_pub`, `discovery.rs:211-236`) can
disclose it *per op* (input schema, access_control), so the data
path exists — but it is N+1 round-trips, and the input schema
describes the *open params contract*, not the *set of produced
resources* (a tunnel producer registers one op and N resources).
- ADR-047 §6 already anticipated the dynamic half:
`channel/resources/subscribe` aggregates per-ALPN resource
enumerators — but the handler is a not-implemented stub
(`channel:resources_not_implemented`, OQ-40,
`operations.rs:274-299`).
This is not an alkcall defect — OQ-40 is decided-deferred and the stub
fails loudly by design. Filing it because alktunnels' discovery need
(OQ-TN-08 resolution: "the ops listing IS tunnel-resource discovery")
makes OQ-40 load-bearing for the first time: a consumer UI cannot
distinguish produced tunnel resources without either the enumerator
aggregation or a listing enrichment. Recommend deciding (small ADR or
OQ-40 update) whether:
1. `channel/resources/subscribe` lands (per-ALPN enumerators — the
decided shape), or
2. `services/list` gains an additive per-op `description`/`metadata`
field (static, registry-side — cheap, but describes the op, not the
resource set), or
3. Both: subscribe for live resource sets, list enriched for op
descriptions.
alktunnels Phase 1 can proceed with (3)'s shape assumed; the minimal
v1 consumer can also just know the op name out-of-band (config), which
is why this is [minor] today.
## E-03 [minor] — `OpenHandler`-exit teardown races `channel/close`'s awaited teardown, but the ledger `take` makes it benign (verify + document)
`run_open_wrapper`'s spawned teardown task (`operations.rs:505-518`)
and the `channel/close` handler's teardown path
(`make_close_handler`, `operations.rs:170-238` — 5s await on the
handler task, then ledger `take` + `on_close`) both walk
`opener_ledger().take(id)` → `policy.on_close(opener)`. The `take` is
atomic (first caller removes the entry), so no double-decrement — the
cap cannot drift upward. But the loser of the race:
- The wrapper task's `teardown_channel(id)` returns
`Err(UnknownChannel)` (logged? no — `let _ =` discard,
`operations.rs:509`) and skips the ledger/policy half silently.
- The close handler's `Ok(task)` arm then awaits a task that is
already exiting.
Verified benign for the cap invariant (`take` is the gate —
`operations.rs:510`, `operations.rs:210`), but the `let _ =` discard
at `operations.rs:509` means a real `UnknownChannel` after a
`set_handler_task` failure (`operations.rs:520-527` — the "channel
vanished between open and handler-task install" warn) is
indistinguishable from a benign race. Recommend either a debug log on
the discard or a comment pinning the race as designed. No correctness
impact; filing for the record since it was checked during the E-01
trace.
## E-04 [minor] — Early-arrival park cap (64) interacts with push-first producers under slow adopters — cap-drop is silent at the consumer
The early-arrival buffer (`EARLY_ARRIVAL_CAP = 64`,
`manager.rs:111`, `park_early_arrival` `manager.rs:440-462`) is
per-channel and drops past the cap with a debug log + counter —
correct per the adopt-race design. But for a *tunnel* producer whose
handler dials then immediately pumps (the common case), 64 chunks can
arrive before the consumer's `adopt_channel` runs if the open-op
response is delayed (relay hops, scheduling). The drops are counted
(`dropped_unknown_chunks`) but not visible to the channel's
consumer — data loss presents as a truncated stream with clean framing
elsewhere. Not a correctness bug (the cap is the documented behavior;
the fix shipped in `a22b2b8` is the right shape), but worth a note in
the tunnel-crate's POC checklist: the UDP POC's MTU-vs-buffer sizing
should treat 64 parked chunks as the observable bound. No alkcall
change requested; filed so the constraint is visible to the next
consumer.
## Non-findings (verified correct, recorded to bound the re-review)
- **The open-op ACL path is complete**: `invoke_streaming` runs the
same visibility + `AccessControl::check` + `input_schema` gates as
`invoke` (`registration.rs:380-404` vs `:347+`); the Sub-typed open
op's ACL failure is a `call.error` on the open, no channel
allocated. Verified against the `invoke_streaming_acl_denied_yields_
forbidden` test (`registration.rs:1628`).
- **`input_schema` enforcement covers open params** (0.4.0, the
alktty-review-L1 fix): `check_input_schema` runs before the handler
in all three dispatch entry points; schema-invalid open params are
`INVALID_INPUT` call errors — never a phantom channel. The
semantic-beyond-schema gap is E-01's subject and is properly
post-schema.
- **EOF arms are unified and clean**: both the sentinel (`Bytes::new`)
and sender-drop arms of `MpscRecvStream::poll_read` yield clean EOF
(`reassembly.rs:111-135`); the mux pump writes the implicit-EOF
chunk for handler-drop-without-shutdown (`mux.rs:78-81` +
`mux_pump_writes_eof_on_implicit_close` test). The two-pump
shutdown contract (alknet ADR-078) has its upstream half in place.
- **The `channel_open` marker survives the wire round-trip**
(`spec_to_json_pub` emits the boolean; `rebuild_spec_for` re-derives
the ALPN from the op name — `from_call.rs:265-280`, `:298-314`),
including the `custom/proto` multi-segment case. The G-04
`resource_id_path` fix is present (`from_call.rs:260`,
round-trip test `from_call.rs:610-617`).
- **The opener-ledger cap decrement is atomic with removal**
(`opener_ledger().take` gates every `on_close` path); the
double-decrement concern from ADR-047 §7 is closed at every call
site traced.
---
# Remediation sketch (original, superseded — see "Remediation plan (post-verification)")
**Unit 1 — E-01 (the establishment phase).** Decided-shape ADR first
(this is ADR-047 §3 contract territory — the `OpenHandler` type shape
is a cross-crate API surface, and both existing consumers' handlers
must be considered; alktty's `TtyOpenHandler` is the migration
prototype). Implementation sketch: `ChannelCore::register_openable`
gains an optional establisher hook; `run_open_wrapper` awaits it
bounded (establishment deadline — a constant, e.g. 10s, or per-spec),
tears down on failure, replies `channel:open_failed` with
`{reason, message}` details on failure and `{channel_id}` on success.
alktty migrates its channels-path error-frame to the call-error path
where applicable (its direct-ALPN path keeps the in-band frame — two
transports, two contracts). alkhttp unaffected (no openable ops).
**Unit 2 — E-02 (discovery enrichment).** Decide via OQ-40 update +
small ADR amendment: recommend (3) — subscribe for live resource sets
(the decided shape, now load-bearing), plus an additive
`description` field on the listing (cheap, immediately useful).
**Implemented 2026-09-06 (alkcall 0.5.0):** the listing half landed
(`OperationSpec.description`, four touchpoints as the appendix
corrected); the subscribe half stays deferred (OQ-40, now with the
load-bearing note).
**Unit 3 — E-03/E-04.** E-03: log-or-comment; E-04: doc note. Trivial.
# Remediation plan (post-verification)
Firm ordering, decided 2026-09-06. E-01's ADR is **ADR-049**
(`docs/architecture/decisions/049-channel-open-establishment-phase.md`)
— decided shape: the **split hook** (`OpenEstablisher` awaited bounded
by the wrapper, `OpenHandler` spawned unchanged after success), chosen
over await-and-inspect because it preserves the `OpenHandler` type
(backward compatible by construction), keeps establishment and pump
lifecycles separate (SSH semantics: the dial is synchronous with the
open reply), and restores ADR-047 §3's original "channel plan" wrapper
shape. N-1 rides the same unit — without the client error-type fix,
`channel:open_failed`'s typed reason is unreachable through the
primary client path.
**Unit 1 — E-01 + N-1 (alkcall 0.5.0).** `OpenEstablisher` +
`register_openable_with_establisher` (no-establisher = always-OK,
existing registrations compile unchanged); wrapper flow: `check_open`
→ `open_channel` → await establisher bounded (dispatch deadline when
`Some`, else `ESTABLISHMENT_TIMEOUT` = 10s) → on failure:
`teardown_channel` + ledger `take` + `policy.on_close` + reply
`channel:open_failed` with `details: {reason, message}` (reason ∈
`dial_failed` / `unknown_resource` / `resource_shortage` /
`handler_error` / `timeout`); on success: spawn pumps,
`set_handler_task`, reply `{channel_id}`. `ChannelClient::open_channel`
returns a typed error carrying the `CallError`. Concurrency is safe:
the dispatcher spawns Once invocations as independent tasks
(`dispatch.rs` `spawn_once_dispatch`), so an awaited establisher does
not head-of-line-block channel 0. Version 0.5.0 (the `open_channel`
error-type change is semver-relevant). **Status: IMPLEMENTED
(2026-09-06)** — with two implementation-shape notes recorded in
ADR-049's amendment: the establisher takes `(input, auth)` only (the
yield-once channel `BiStream` belongs exclusively to the pump
handler), and the effective bound is `min(dispatch deadline,
registration timeout | ESTABLISHMENT_TIMEOUT)`. All four verification
gates below are tests in the crate.
**Unit 2 — E-02 (ride the 0.5.0 release).** Additive
`description: Option<String>` on `OperationSpec` — note this is four
touchpoints, not one (see the verification appendix's E-02
correction): struct field + `spec_to_json_pub` emit +
`rebuild_spec_for` parse + the `services/list` output-schema doc.
`services/list` emits it when set. OQ-40 stays deferred but gains a
"load-bearing for alktunnels discovery UI" note; the
`channel/resources/subscribe` half stays deferred (alktunnels v1 uses
config-known op names). **Status: IMPLEMENTED (2026-09-06)** — the
discovery decision (3: enriched listing now, subscribe deferred) is
recorded as an amendment to ADR-047 §6.
**Unit 3 — E-03/E-04/N-2 (trivial batch, same PR series as Unit 1).**
E-03: debug log (or pinning comment) on the discarded `UnknownChannel`
at the wrapper's teardown discard. E-04: doc note on `EARLY_ARRIVAL_CAP`
(the 64-parked-chunks observable bound) for tunnel-crate-facing
consumers. N-2: expose an accessor for `early_arrival_count` or remove
the write-only counter. **Status: IMPLEMENTED (2026-09-06)** — E-03:
debug log + benign-race pinning comment on the wrapper's handler-exit
teardown (mirroring the sibling log the Unit 1 establisher-failure
teardown already carries); E-04: consumer-side doc note in
`channel-client.md` plus the const doc; N-2: accessor
(`ChannelManager::early_arrival_count()`, documented as monotonic —
adopt-drain does not decrement) + test.
**Sequencing:** ADR-049 (landed) → alkcall 0.5.0 (Units 1+2+3) →
alktty migration (channels-path semantic failures move into an
establisher; the direct-ALPN in-band error frame is retained — two
transports, two contracts) → alkhttp mechanical pass
(`OpenableAlpn.establisher: Option<...>`, default `None`). alktunnels
Phase 1 blocks only on the ADR decision (landed), not the
implementation: its OQ-TN-09 direction ("self-contained control frame")
shrinks to *post-establishment* control only once `channel:open_failed`
exists.
## Verification gates for the E-01 remediation
- A channels end-to-end test: producer handler whose establisher fails
after channel allocation → consumer's `open_channel` resolves
`Err` with `channel:open_failed` + reason details; no channel in the
manager's `channel_ids()` afterward (the SSH "channel never exists
opener-side" property, consumer-visible as "no `channel_id` was
ever returned").
- A bounded-establishment test: establisher that never completes →
open op fails with a timeout-flavored error within the deadline,
channel torn down, ledger decremented.
- An alktty-migration test: the existing
`open_via_channels_surfaces_negotiation_rejected` scenario resolves
via call error (or retained error-frame — per the ADR's chosen
compat shape) unchanged in behavior.
- (Added post-verification) A compat gate: a no-establisher
registration behaves exactly as today — `{channel_id}` reply,
handler spawned, teardown on handler exit.
## Verification appendix (2026-09-06 re-verification pass)
All findings re-verified by independent source trace at tree `88e3f5e`.
Verdicts, corrections, and the additional findings:
- **E-01 CONFIRMED — and strengthened.** The wrapper's
spawn-then-reply flow is as described
(`src/channels/operations.rs` `run_open_wrapper`: handler spawned,
reply written before the handler performs any work). New supporting
evidence the original pass missed: **ADR-047 §3's decision text**
describes the ALPN open handler as "validate params, consult
ownership, prepare the backend, return a 'channel plan'" — an
awaited preparation the wrapper consults before replying. The
implemented `OpenHandler` (`JoinHandle<()>`) collapsed that phase
into a fire-and-forget spawn. E-01 is therefore **drift from
ADR-047 §3's own wrapper shape**, not merely a missing convenience —
the remediation restores the decided design, which lowers the ADR's
contention cost. Also verified: the serving loop spawns Once
invocations as independent tasks, so an awaited establishment phase
does not head-of-line-block other calls on channel 0 (a fix-feasibility
question the original pass did not address).
- **E-02 CONFIRMED — one correction.** `services_list_handler` emits
`{name, namespace, op_type}` only, as described. But the "cheap,
additive `description` field" framing understates the work:
`OperationSpec` has **no `description` field at all** (the
`description` in `spec.rs` is on `ErrorDefinition`). The change is
four touchpoints: struct field, `spec_to_json_pub` emit,
`rebuild_spec_for` parse, and the `services/list` output-schema doc.
Still small; still [minor].
- **E-03 CONFIRMED.** Both teardown paths gate on the atomic
`opener_ledger().take(id)`; the race loser gets `UnknownChannel` /
`take → None` — no double-decrement. The `let _ =` discard on the
wrapper's `teardown_channel` result is as described. Also traced:
the `set_handler_task` failure path still runs the wrapper task to
completion (self-teardown), so no leak in that arm either. The
log-or-comment remedy stands.
- **E-04 CONFIRMED.** `EARLY_ARRIVAL_CAP = 64`, per-channel, FIFO,
drop-past-cap with debug log + `dropped_unknown_chunks` counter.
Doc-note-only remedy stands.
- **N-1 [minor today, blocking for E-01's consumer visibility] —
`ChannelClient::open_channel` erases the typed error.**
(`src/channels/client.rs`) It flattens the `CallError` into a
`String` via `format!("open op failed: {e:?}")`. Even once the
wrapper replies `channel:open_failed` with reason details, a
consumer cannot branch on the reason — the error type destroys it.
E-01's verification gate ("consumer's `open_channel` resolves `Err`
with `channel:open_failed` + reason details") is unreachable without
changing this error type, so the fix is in scope for the E-01 unit
(ADR-049 §4), not deferred.
- **N-2 [trivial] — `early_arrival_count` is write-only.** Incremented
in `park_early_arrival`, never read anywhere (no accessor; only
`dropped_unknown_chunks` is exposed). Either expose an accessor
(observability for the E-04 bound) or remove the counter.
- **N-3 [observation, pinned by ADR-049 §6] — panicked pump handlers
are EOF-shaped by design.** The wrapper's teardown task swallows the
pump handler's `JoinError` (`let _ = raw_task.await`). A panicked
handler = phantom channel + instant EOF, indistinguishable from a
clean short-lived channel. With the establisher split, the pump-phase
panic stays in this category (correct — there is no mid-stream error
channel by design; establishment errors are the only kind that
belong in the open reply). ADR-049 §6 pins this posture; no change.
## References
- alknet ADR-078 / channels two-pump contract — the teardown half this
review's EOF-arms non-finding confirms upstream.
- alkcall ADR-047 (openable ALPNs are operations), ADR-016 (typed
error schemas — the vehicle for `channel:open_failed` details),
ADR-040/041 (backpressure/caps — untouched by E-01).
- alktty ADR-009 (the open op's input is the negotiation) + review
#001 L1/L3 resolution — the per-crate workaround E-01 obsoletes;
`send_negotiation_error` (`alktty/src/adapter.rs:177-187`) and the
`0x00`-peek (`alktty/docs/architecture/tty-adapter.md:186-197`) are
the in-band mechanism the upstream establisher replaces for the
channels path.
- alktunnels Phase 0 (`docs/research/phase-0-findings.md` OQ-TN-09)
and the SSH/SOCKS5 survey (`docs/research/ssh-socks5-survey.md`
§"Open-failure path", §"Comparison") — the prior art motivating
E-01's reason-code vocabulary.
- The consumer-findings ledger convention (`docs/reviews/consumer-
findings-ledger.md`) — findings here follow the same spirit
(downstream-discovered, filed for upstream action); E-01..E-04 are
alktunnels-discovered but numbered in alkcall's review series since
they are alkcall findings.
- **ADR-049** (`docs/architecture/decisions/049-channel-open-
establishment-phase.md`) — the E-01/N-1 resolution (split-hook
`OpenEstablisher`, bounded establishment, `channel:open_failed`
typed error, client error-type fix).
@@ -0,0 +1,343 @@
# Review 007 — Establishment Follow-Ups (from the alktunnels UDP POC)
## Status
Resolved — all three units landed in alkcall 0.6.0 (2026-09-07):
Unit 1 `dc4ad2b` (R-01 + R-02's doc notes), Unit 2 `8f122b0` (R-02's
telemetry), Unit 3 `9c6fec1` (R-03 + ADR-050). Two deviations from
the sketches below, both recorded in the ADRs: (1) R-01's plan is
**typed-opaque** (`ChannelPlan = Arc<dyn Any + Send + Sync>`), not
`Option<Value>` — the sketch could not satisfy this review's own
verification gate, because the payloads establishers hand off are
live handles (dialed sockets, TTY handles) with no JSON
representation; (2) R-03's helper returns `(u64, u64)`, not
`io::Result<(u64, u64)>` — both pumps swallow copy errors by the
ADR-078 contract (mid-stream error = abrupt close), so an `Err` state
would be dead code. Original findings below, retained as filed
(verified against tree `36e74cd`, 0.5.0).
Open — findings filed from the alktunnels UDP POC pass
(`alktunnels-udp-poc`, 2026-09-06; summary at
`/workspace/@alkdev/alktunnels/docs/research/poc-summary.md`). This
review exists to prevent a second fix→publish→update-dependents cycle:
every finding below is either (a) guaranteed to be needed by
alktunnels Phase 1 implementation, or (b) a documented contract gap
that will bite the next `OpenHandler` author the way it bit the POC.
Nothing here is speculative — each finding was reached by writing
working code against 0.5.0 and finding the shape insufficient.
Findings continue the review numbering with prefix `R` (001–006 used
P/C/R/A/B/C/D/F/G/E — each review numbers independently).
> **Post-remediation note (2026-09-07):** the "Open —" paragraph
> above is the original filing state, retained for the record. All
> units landed; see Status above and ADR-049 amendment 2 + ADR-050.
## Scope
The ADR-049 establishment surface (`OpenEstablisher`, `Establishment`,
`EstablishmentError`, `register_openable_with_establisher`) and the
`OpenHandler` contract, reviewed from the alktunnels UDP POC's
consumer/producer implementations. Everything verified against source
at tree `36e74cd` (0.5.0). Cross-references: alkcall review 006
(E-01..E-04), alkcall ADR-049, alktunnels `poc-summary.md`
(§Issues Surfaced), alktty's channels establisher
(`make_tty_establisher`).
```
Verified against: alkcall 36e74cd (0.5.0)
Reading list: src/channels/operations.rs (OpenEstablisher, Establishment,
run_open_wrapper, teardown_failed_channel), src/channels/client.rs
(ChannelOpenError::establishment_reason), src/channels/manager.rs
(teardown_channel, adopt_channel), docs/architecture/decisions/
047 + 049, and the consumers: alktty src/channels.rs (make_tty_
establisher / make_tty_open_handler), alkhttp src/websocket/
upgrade.rs (OpenableAlpn ferry), alktunnels-udp-poc src/producer.rs
```
## Severity legend
Same scale as reviews 004–006.
---
# Part A — The findings
## R-01 [major] — `Establishment` is payloadless, but the channel plan is exactly what establishers need to hand to the pump handler
**Verified:** YES — by building against the API. The ADR-049 §1 text
itself anticipates this: *"Reserved for a channel plan — today the
wrapper consults only success/failure, so `()` carries no data"*
(`src/channels/operations.rs:341-345`, the empty
`pub struct Establishment {}`).
### The mechanism
The establisher is pre-data-plane: it cannot see the channel's
yield-once `BiStream` (ADR-049 amendment), so anything it
establishes — a dialed socket, an allocated handle — must cross to
the pump handler through a side channel the ALPN crate invents. The
alktunnels POC's workaround is a resource-keyed
`Mutex<HashMap<String, SubstrateHandle>>` + a poll-loop `take`
(`producer.rs::HandleHandoff`) that works only because the wrapper
guarantees establisher-before-handler ordering. The costs:
1. **Concurrent same-resource opens race the slot.** Two opens of the
same resource: the second establisher's `deliver` overwrites the
first handler's not-yet-taken handle. The POC documents this as a
simplification; a real crate cannot.
2. **alktty hit the same wall and documented it as a limitation.**
`alktty/src/channels.rs:253-257`: the establisher cannot carry the
allocated `TtyHandle` across, so *backend allocation* stays
post-open in the pump handler (`allocate_failed` remains an
in-band negotiation error frame — ADR-010) — exactly the
phantom-channel-ish shape ADR-049 exists to eliminate, still
alive one layer down. TTY's establisher can only validate/lookup/
ownership-check; the one thing that can actually fail with a
runtime error (allocate) is unreachable from the typed error path.
3. **Every ALPN crate pays the side-channel tax again.** Handoffs,
slots, poll loops or oneshots — per crate, per resource key, all
with the same concurrency caveat.
### The ask
Give `Establishment` its payload — the channel plan ADR-047 §3
originally described ("return a 'channel plan'... the wrapper consults
its result"). Minimal shape:
```rust
pub struct Establishment {
/// ALPN-defined plan data — the dialed handle, an allocation
/// ticket, whatever the pump phase needs. Opaque to alkcall.
pub plan: Option<Value>,
}
```
or, keeping it typed at the boundary:
```rust
pub struct Establishment { pub plan: Option<Value> }
```
with the wrapper passing `establishment.plan` (or `Null`) to the
`OpenHandler`'s `input` — e.g. as a well-known key the pump handler
reads, or as a third callback parameter. The wire surface is
**unchanged** (the plan is process-local: establisher → wrapper →
handler in the same process on the producing side; nothing crosses
the transport that isn't already the open op's input). Consumers'
`Ok(Establishment {})` construction sites (alktty has two; alkhttp's
ferry has none) break mechanically at 0.6 — `Establishment::new(plan)`
/ `Establishment::default()` make the migration one-liners.
**The break is the point of doing this now:** 0.5.0 published
yesterday with `Establishment {}` documented as reserved. The next
consumer (alktunnels) needs the payload in Phase 1 — implementing
Phase 1 without it means shipping the POC's side-channel handoff into
the real crate, with its same-resource race, and then the payload
lands later anyway as *another* breaking release. Filling the
reserved field now is the cheap moment; the alternative is paying the
breaking change twice.
### Severity
[major] — the capability gap is structural for any establisher whose
backend produces a handle (tunnels: dial; TTY: allocate), not a
polish item. Not [critical] because workarounds exist (the POC proves
one), but every workaround carries the same-resource race or forces
failure classes back into per-crate in-band frames — the exact cost
ADR-049 was written to remove.
---
## R-02 [minor] — The `OpenHandler` `JoinHandle` lifetime contract is undocumented and load-bearing (found empirically by the POC)
The POC's first pump implementation returned a wrapper task that
spawned the pump as a nested fire-and-forget task. The result:
**every tunnel connected then instantly EOF'd with zero bytes** — the
wrapper awaited the (already-complete) handler task, tore the channel
down at birth, and both pumps saw immediate EOF. Everything upstream
(establisher, open reply, channel routing) looked healthy; the bug
was purely in the handler's return-value semantics.
The contract, as implemented: the wrapper awaits the returned
`JoinHandle` and *that completion is the teardown trigger*
(`run_open_wrapper`'s spawned task: `let _ = raw_task.await;`
→ `teardown_channel(id)` → drop the demux sender → EOF to the
handler's read half). Therefore:
- **The returned `JoinHandle` must track the data-plane lifetime.** A
handler that returns before its pumps finish tears the channel down
at birth. The pump must be awaited *inline* inside the handler task
(`let _ = pump_halves(...).await;`), not spawned-and-forgotten.
- Half-open semantics fall out of this correctly (one pump finishing
shuts down the opposite sink per ADR-078; the handler's task
completes when both pumps finish — which is when teardown *should*
happen). The contract is right; only its documentation is missing.
The current type docs (`operations.rs:325-339`) say the handle is
"recorded for teardown (abort on `channel/close` / connection drop)"
— abort semantics — but never say **early return = teardown-at-birth**.
alktty got this right by accident of shape (`drive_session_pre_negotiated`
awaited inline in its single spawned task — `channels.rs:355-395`);
the POC got it wrong the natural way. The next consumer will write
the wrong shape too, because the natural reading of "spawn your
protocol and return the `JoinHandle`" is a task-spawner, not a
lifecycle promise.
**The ask:** doc note on `OpenHandler` and
`ChannelCore::register_openable*` — "the returned `JoinHandle` must
track the data-plane lifetime: the wrapper awaits it and tears the
channel down on completion; return a task that runs the handler to
completion, never a spawner that exits early" — plus the ADR-049
amendment (§1 or §6) recording the semantics. Optional hardening (not
required): a `debug!` in `run_open_wrapper` when the awaited handler
exits without the channel's `BiStream` having been accepted (a
telemetry hint for the birth-teardown pattern; the accept is
observable in-process). Doc-only; no break.
---
## R-03 [minor, optional unit] — The two-pump helper's convergence test is satisfied; extracting it now removes the last reason for a later sweep
alknet ADR-078 deferred helper extraction until a second two-pump
consumer existed and the shapes converged. The POC supplies both
halves of that test:
- **Producer side:** the alktunnels POC's `pump_halves` — two
`tokio::io::copy` pumps over split channel-vs-substrate halves,
each pump shutting down the opposite sink on completion, joined.
- **Consumer side:** `TunnelSession::take_halves` + the assembly
layer's copy — the same shape modulo channel side.
The shapes converged. The helper (`pump_bidi` or similar) is a
candidate for **this same 0.6 sweep** — additive (a new pub fn in
`channels` or `core`), zero breaking change, and it pins the
ADR-078 contract in one place instead of three. The natural signature
follows the POC:
```rust
pub async fn pump_bidi<A, B>(a: A, b: B) -> io::Result<(u64, u64)>
where
A: AsyncRead + AsyncWriteExt + Send + Unpin,
B: AsyncRead + AsyncWriteExt + Send + Unpin,
```
(join'd two-pump with shutdown-on-completion; returns the copy
counts for observability). alktty's three-pump session does not fit
it (the exit future is a third signal) — that is fine; the helper
serves the two-pump shape, TTY stays as-is.
**The ask:** decide in this sweep — either extract in 0.6 (additive;
recommended, since alktunnels Phase 1 will implement the pattern
anyway and an upstream helper makes the third consumer free), or
record the explicit decision to keep it per-crate. Do not leave it
half-decided: an un-extracted helper is not itself a break, but
finding this out during Phase 1 would be the same
fix→publish→update treadmill for a purely additive change.
---
# Part B — Verified non-issues (bounded so the next review doesn't re-check)
- **The reverse-flow (`-R`) story needs no upstream change.** A hub
proxy opening a channel *toward* a worker is the worker serving its
own open op on the connect side — supported by
`ChannelClient::from_connection_with_serving` (ADR-022 §2 both-sides
semantics) + `ChannelOperations::register_on` on the serving
registry + the wrapper allocating on the serving side (ADR-047 §5
"the side that holds the `ChannelManager` allocates" — both sides
do, with odd/even split per ADR-047 §5). Verified by trace; no API
gap. alktunnels' reverse-flow POC (OQ-TN-10 #2) will exercise this
end-to-end, but no upstream mechanism is missing.
- **The establisher receives registry-validated input** —
`invoke_streaming`'s `input_schema` check runs before the wrapper
(`registration.rs:397`); the establisher's `input.clone()`
(`operations.rs:718`) is post-schema. Confirmed; no gap.
- **Typed establishment errors are complete on the wire** —
`establishment_error_to_call_error` carries both `reason` and
`message` in `details` (`operations.rs:646-653`); the client-side
`establishment_reason()` branches on it (`client.rs:75-83`); alktty's
open spec declares the matching `ErrorDefinition` (ADR-016). The
vocabulary needs no extension for tunnels (`dial_failed` /
`unknown_resource` / `resource_shortage` / `handler_error` /
`timeout` cover every tunnel establisher failure class — verified by
the POC's three typed-error tests).
- **The early-arrival park cap (64)** — reviewed in 006 E-04; the POC
treated it as a sizing constraint, not a defect. Unchanged.
- **`EstablishmentError`'s reason set is sufficient** — the POC's
establisher used three of four reasons; `handler_error` covers
params-parse failures (the TTY pattern). No new reason needed.
---
# Remediation plan
## Unit 1 — `Establishment` carries the channel plan (R-01)
1. ADR-049 amendment (or ADR-050-adjacent amendment): `Establishment`
gains `plan: Option<Value>` (ALPN-opaque); the wrapper threads
`plan` into the pump handler (third `OpenHandler` parameter, or
merged into `input` under a reserved key — decide in the ADR; the
separate-parameter shape avoids colliding with the input schema).
2. Bump to 0.6.0 (breaks `Ok(Establishment {})` construction sites —
alktty, two sites; mechanical `Establishment::default()` or
`Establishment { plan: None }`).
3. Migration: alktty can then move backend `allocate` into its
establisher (killing the in-band `allocate_failed` frame on the
channels path — ADR-010's reason for that frame disappears);
alktunnels Phase 1 uses `plan` for the dialed handle; alkhttp's
ferry passes `Option` through unchanged.
## Unit 2 — `OpenHandler` lifetime doc note (R-02)
Doc-only (type docs + ADR-049 amendment). Optional debug warning.
No break.
## Unit 3 (optional) — `pump_bidi` helper extraction (R-03)
Additive pub fn + tests. No break. Do it in the same 0.6 or record
the decision not to.
## What must NOT ride along
Nothing else. The remaining alktunnels Phase 1 needs (UDP codec ADR,
reverse-flow lifecycle, access-policy details) are alktunnels-local
spec work — the POC validated the mechanics and found no further
upstream gaps. The deliberate goal of this review: **0.6 is the last
breaking sweep forced by known work**; anything after it should be a
genuinely new discovery, not a known-dangling item.
## Verification gates
- Unit 1: end-to-end test — establisher dials, `plan` carries the
handle, pump handler receives it; concurrent same-resource opens
each get their own handle (the race the POC documented is
unreachable); alktty migration test (`allocate` in the establisher
→ `allocate_failed` is a call error, no in-band frame).
- Unit 2: doc gate only (`cargo doc` clean; the note present on both
the type and the registration entry points).
- Unit 3: helper test — the POC's two-pump semantics (EOF from one
direction completes the other; shutdown-on-completion) reproduced
through the helper.
## References
- alkcall review 006 (E-01 establishment gap; R-01 is the payload
half of E-01's own remediation — ADR-047 §3's "channel plan"
through the reserved `Establishment` field) and ADR-049 (the
establishment phase; §1's original sketch already described the
plan-carrying shape).
- alkcall ADR-047 §3 ("return a 'channel plan'" — the original shape
this review restores), ADR-022 §2 (both-sides serving — the
reverse-flow non-finding), ADR-047 §5 (odd/even allocation —
reverse-flow allocation on the serving side).
- alknet ADR-078 (the two-pump contract; R-03's helper pins it
upstream).
- alktunnels `docs/research/poc-summary.md` (§Issues Surfaced #1/#2 —
the empirical findings; §#4 — the convergence input for R-03) and
the POC crate (`/workspace/alktunnels-udp-poc`, 17 tests — the
empirical basis for every finding here).
- alktty `src/channels.rs:253-257` — the existing documentation of
the R-01 cost on the TTY side (`allocate` cannot cross; in-band
`allocate_failed` remains), the second consumer confirming the gap
is not tunnel-specific.
+111
View File
@@ -0,0 +1,111 @@
# Consumer findings ledger — alkhttp consuming alkcall
Findings from writing the first real consumer (alkhttp) against alkcall.
alkcall is deliberately minimal; expect missing features and edge-case
bugs to surface here first. Extract items into alkcall tasks/reviews as
the project sees fit.
Format: date | found-in (alkhttp context) | severity | status.
---
## Open
(none)
---
## Resolved
### CF-004 — `services_schema_handler` discloses Internal/ACL-restricted op specs — no visibility or AccessControl check (2026-08-30) — RESOLVED 2026-08-31
- **Found in:** alkhttp Review 002, finding PRJ-16
(`alkhttp/docs/reviews/002-post-remediation-review.md`, Part F',
verified against tree `91483a7` + alkcall source). Filed to this
ledger 2026-08-30.
`src/registry/discovery.rs:327-343` (`services_schema_handler`) —
the handler did a bare `registry.registration(&name)` and returned
`spec_to_json(&reg.spec)` verbatim, with **no `Visibility` check and
no `AccessControl::check(peer_identity)`**.
- **Fix:** the handler now applies the same two gates as
`OperationRegistry::invoke` — the Internal-visibility rejection
(`!ctx.internal` → spec-404) and `AccessControl::check` with the same
identity resolution invoke() uses (`handler_identity` when internal,
`identity` otherwise). Restricted ops return `NOT_FOUND` (spec-404)
rather than `FORBIDDEN` — no information leaks about the restricted
surface's existence or shape. The gate is in the handler itself, so
every transport (wire `/call`, HTTP routes, MCP tools) is covered.
- **Status:** resolved — 2026-08-30.
### CF-003 — `publish_schema` compile failure is fail-open on the wire dispatch path (2026-08-30) — RESOLVED 2026-08-30
- **Found in:** alkhttp Review 001 post-remediation task
`review-001-publish-schema-validation-robust` (the GW-01/#1 follow-up
work). While fixing alkhttp's HTTP-side instance of the pattern, the
identical instance was verified on the wire dispatch path:
`src/protocol/dispatch.rs:352-366` (`Dispatcher::dispatch`,
`OperationType::Pub` arm) — when a Pub op's `publish_schema` failed to
**compile**, the pump logged a warn and proceeded with
`validator: None` (silent unvalidated ingest, plus per-request
recompile cost).
- **Fix:** fail-closed at the source — `publish_schema` is now compiled
at **registration time** in both insertion points
(`OperationRegistry::register` and `OperationRegistryBuilder::store`);
an un-compilable schema is a registration error, so an
unvalidated-ingest path can never be constructed. The compiled
validator is cached per-op (`OperationRegistry::publish_validator`)
and the dispatch path consumes the cache — the per-request compile is
gone. The `from_call` import path inherits the guarantee (imports
register through the same choke point); forwarded-chunk validation on
the producer side is the remote peer's dispatch path and now inherits
the same fail-closed property.
- **Behavior change (semver-relevant):** `register`/builder methods
return `Err` for an un-compilable `publish_schema` (previously:
accepted). Code that registered garbage schemas will now get a
registration error instead of a silently-unvalidated op.
- **Status:** resolved — 2026-08-30.
### CF-002 — demux `TooLarge` skip allocates the full peer-declared length up front (up to ~4 GiB from an 8-byte header) (2026-08-30) — RESOLVED 2026-08-30
- **Found in:** alkhttp Review 001, finding WS-12 (cross-crate;
`docs/reviews/001-initial-implementation-review.md`, Part B,
verified against tree `4a825d3`). Filed to this ledger 2026-08-30.
`src/channels/adapter.rs:144` (`run_demux_loop_for_client`) — the
`ChunkError::TooLarge` arm allocated
`let mut discard = vec![0u8; length as usize];` **before** reading:
the buffer was sized from the peer's untrusted 8-byte header (`u32`
`length`, up to ~4 GiB).
- **Fix:** stream-skip with a fixed 64 KiB buffer + `read_exact` loop
consuming `length` bytes — no allocation sized from peer input.
Plus a cumulative skipped-bytes budget (256 MiB) that tears down the
connection when a peer loops `TooLarge` headers to burn bandwidth/CPU.
The budget resets on each valid chunk, so legitimate isolated
oversized chunks (the resync path, covered by the existing test) are
unaffected. The existing `demux_resyncs_after_oversized_chunk` test
passes unchanged.
- **Status:** resolved — 2026-08-30.
### CF-001 — `call_single_stream` write-failure maps to non-retryable `INTERNAL` (2026-08-29) — RESOLVED 2026-08-30
- **Found in:** alkhttp `from_wss` drop-monitor race tests (WS-02/CON-02,
review-001-ws-eof-signal). When the transport mux died mid-call, the
write path failed fast and the write-failure mapping produced a
non-retryable `INTERNAL: failed to write request frame` — even though
the call never reached the producer and a reconnect/retry would be
safe.
- **Fix:** new retryable `CallError::connection_closed` constructor
(`CONNECTION_CLOSED` code, `retryable: true`), applied **only** to
write failures where the call is provably undelivered:
`call.requested` frame write failures on all consumer paths
(`call`/`subscribe`/`publish`, single-stream and stream-per-request;
the publish pump tags its write stages so only the request-frame
stage is retryable). Mid-publish and completed-frame write failures
stay `INTERNAL` (delivery ambiguous — retry unsafe), and the
producer-side `fail_all(...)` on connection close stays `INTERNAL`.
The new code string is additive to the wire error vocabulary;
`CONNECTION_CLOSED` is new — consumers should treat unknown codes per
their existing policy (the `retryable` flag is the machine-readable
signal).
- **Status:** resolved — 2026-08-30. alkhttp follow-up: the
`review-001-ws-eof-signal` race test can now tighten back to
retryable-only asserts.
+165 -4
View File
@@ -128,6 +128,15 @@ impl ChannelsAdapter {
reader: Box<dyn tokio::io::AsyncRead + Send + Unpin>,
policy: Option<&Arc<dyn ChannelLifecyclePolicy>>,
) {
// Stack buffer for skipping oversized payloads — sized once, never
// from peer-declared lengths (a peer-declared `length` is untrusted
// up to u32::MAX ≈ 4 GiB; pre-sizing off it is an allocation-shape
// DoS). A cumulative cap tears down dribbling peers that loop
// TooLarge headers to burn bandwidth indefinitely.
const SKIP_BUF_LEN: usize = 64 * 1024;
const MAX_CONSECUTIVE_SKIPPED_BYTES: u64 = 256 * 1024 * 1024;
let mut skip_buf = vec![0u8; SKIP_BUF_LEN];
let mut skipped_since_last_valid: u64 = 0;
let mut reader = reader;
let mut header_buf = [0u8; CHUNK_HEADER_LEN];
loop {
@@ -141,10 +150,23 @@ impl ChannelsAdapter {
max = super::wire::MAX_CHUNK_LEN,
"demux: chunk too large, skipping payload bytes"
);
let mut discard = vec![0u8; length as usize];
if let Err(e) = reader.read_exact(&mut discard).await {
warn!(error = %e, "demux: failed to skip oversized payload");
break;
let mut remaining = length as u64;
while remaining > 0 {
let take = remaining.min(SKIP_BUF_LEN as u64) as usize;
if let Err(e) = reader.read_exact(&mut skip_buf[..take]).await {
warn!(error = %e, "demux: failed to skip oversized payload");
return;
}
remaining -= take as u64;
}
skipped_since_last_valid += length as u64;
if skipped_since_last_valid > MAX_CONSECUTIVE_SKIPPED_BYTES {
warn!(
skipped_since_last_valid,
max = MAX_CONSECUTIVE_SKIPPED_BYTES,
"demux: skipped-bytes budget exceeded; tearing down connection"
);
return;
}
continue;
}
@@ -170,6 +192,7 @@ impl ChannelsAdapter {
}
};
manager.route_payload(header.channel_id, payload).await;
skipped_since_last_valid = 0;
}
Err(e) => {
if e.kind() == std::io::ErrorKind::UnexpectedEof {
@@ -529,4 +552,142 @@ mod tests {
assert_eq!(buf, [i; 4], "channel A chunk {i} intact — no data loss");
}
}
/// CF-002 #1 — the skipped-bytes budget tears down a dribbling peer.
/// Sixteen back-to-back `TooLarge` headers each declare 16 MiB + 1 and
/// their payloads are genuinely on the wire (the demux must consume
/// every byte to reach the next header, so the test cannot avoid
/// moving ~268 MiB through the duplex). The 16th skip pushes the
/// cumulative total past the 256 MiB budget: a working demux tears
/// down and *stops reading there*, while a peer that keeps dribbling
/// (a 17th header parked mid-payload) must not wedge it. Without the
/// budget the demux blocks skipping the 17th payload forever and the
/// timeout fires.
#[tokio::test]
async fn demux_tears_down_after_cumulative_skipped_bytes_budget_exceeded() {
let oversized_len = super::super::wire::MAX_CHUNK_LEN as usize + 1;
let burst_count = 16;
let (client, server) = tokio::io::duplex(1024 * 1024);
let (server_read, server_write) = tokio::io::split(server);
let (mux_handle, mux_runner) = MuxRunner::new(Box::new(server_write));
let _mux_task = tokio::spawn(async move {
let _ = mux_runner.run().await;
});
let manager = ChannelManager::with_defaults(mux_handle, None);
let (_id, _send, _recv) = manager
.open_channel("alk/tty", "alice", None)
.await
.expect("open");
let demux_manager = manager.clone();
let demux_task = tokio::spawn(async move {
ChannelsAdapter::run_demux_loop_for_client(&demux_manager, Box::new(server_read), None)
.await;
});
// Writer task: 16 full oversized payloads (15 × ~16 MiB ≈ 240 MiB
// skip cleanly under the budget; the 16th crosses it), then a 17th
// header + partial payload — and park instead of finishing (the
// peer keeps dribbling, never EOFs).
let writer_task = tokio::spawn(async move {
use tokio::io::AsyncWriteExt;
let mut client_write = client;
let mut header = [0u8; 8];
let chunk = vec![0u8; 64 * 1024];
let full = oversized_len / chunk.len();
let tail = oversized_len % chunk.len();
for _ in 0..burst_count {
super::super::wire::write_header(1, oversized_len as u32, &mut header)
.expect("write header");
client_write.write_all(&header).await.expect("write header");
for _ in 0..full {
client_write.write_all(&chunk).await.expect("write payload");
}
if tail > 0 {
client_write
.write_all(&chunk[..tail])
.await
.expect("write payload tail");
}
}
// The budget is already blown by the 16th skip; this header is
// the dribble that a budgetless demux would park on forever.
super::super::wire::write_header(1, oversized_len as u32, &mut header)
.expect("write header");
client_write.write_all(&header).await.expect("write header");
client_write
.write_all(&chunk[..8192])
.await
.expect("write partial payload");
std::future::pending::<()>().await;
});
let demux_done =
tokio::time::timeout(std::time::Duration::from_secs(120), demux_task).await;
assert!(
demux_done.is_ok(),
"demux must tear down on the skipped-bytes budget, not block in the skip loop"
);
writer_task.abort();
}
/// CF-002 #2 — a skip that hits EOF mid-payload (peer declared more
/// bytes than it sent) ends the demux loop; nothing routes to the
/// channel and the loop does not spin back to header parsing on a
/// truncated stream.
#[tokio::test]
async fn demux_ends_when_oversized_payload_skip_hits_eof() {
let (client, server) = tokio::io::duplex(1024 * 1024);
let (server_read, server_write) = tokio::io::split(server);
let (mux_handle, mux_runner) = MuxRunner::new(Box::new(server_write));
let _mux_task = tokio::spawn(async move {
let _ = mux_runner.run().await;
});
let manager = ChannelManager::with_defaults(mux_handle, None);
let (_id, _send, mut recv) = manager
.open_channel("alk/tty", "alice", None)
.await
.expect("open");
let demux_manager = manager.clone();
let demux_task = tokio::spawn(async move {
ChannelsAdapter::run_demux_loop_for_client(&demux_manager, Box::new(server_read), None)
.await;
});
use tokio::io::AsyncWriteExt;
let oversized_len = super::super::wire::MAX_CHUNK_LEN as usize + 1;
let mut header = [0u8; 8];
super::super::wire::write_header(1, oversized_len as u32, &mut header).expect("header");
let mut client_write = client;
client_write.write_all(&header).await.expect("write header");
client_write
.write_all(&[0u8; 4096])
.await
.expect("write partial payload");
drop(client_write);
let demux_done = tokio::time::timeout(std::time::Duration::from_secs(10), demux_task).await;
assert!(
demux_done.is_ok(),
"demux must end when the skip hits EOF instead of looping"
);
// The oversized chunk was never routed — nothing to assert on the
// channel beyond the loop ending; reading would surface EOF since
// the demux dropped its sender without delivering a payload.
use tokio::io::AsyncReadExt;
let mut buf = [0u8; 4];
let read = tokio::time::timeout(
std::time::Duration::from_millis(500),
recv.read_exact(&mut buf),
)
.await;
assert!(
read.is_err() || read.unwrap().is_err(),
"no payload must arrive after the truncated skip"
);
}
}
+1332 -32
View File
File diff suppressed because it is too large. Load diff
+4
View File
@@ -83,6 +83,10 @@ impl OperationEnv for ChannelsSessionEnv {
self.base.peer_operations(peer)
}
fn list_operation_names(&self) -> Vec<String> {
self.base.list_operation_names()
}
async fn invoke_peer(
&self,
peer: &crate::registry::env::PeerRef,
+190 -18
View File
@@ -7,7 +7,7 @@
//! `ChannelsAdapter`, the open-op wrapper, and relay logic can all
//! hold a handle.
use std::collections::HashMap;
use std::collections::{HashMap, VecDeque};
use std::net::SocketAddr;
use std::sync::atomic::{AtomicU32, AtomicU64, Ordering};
use std::sync::Arc;
@@ -93,8 +93,33 @@ struct Inner {
remote_addr: Option<SocketAddr>,
side: ChannelSide,
dropped_unknown_chunks: AtomicU64,
/// Chunks that arrived for a channel_id this side has not
/// adopted yet (the open-op response carrying `channel_id` and the
/// producer's first data-plane write race — the handler can start
/// pumping before the connect side's `adopt_channel` runs). Each
/// channel buffers up to `early_arrival_cap` payloads; `adopt_channel`
/// drains them into the new receiver. `clear_all` drops the map.
early_arrivals: Mutex<HashMap<u32, VecDeque<Bytes>>>,
early_arrival_count: AtomicU64,
}
/// The per-channel cap on parked early-arrival chunks (`route_payload`
/// parks chunks for not-yet-adopted channels instead of dropping them —
/// the open-op response/first-data race). A channel whose adopter never
/// arrives leaks its parked chunks until `clear_all`; the cap bounds
/// that leak per channel.
///
/// Consumer-facing bound (review 006 E-04): for a push-first producer
/// (dial-then-pump is the common tunnel shape), up to `EARLY_ARRIVAL_CAP`
/// chunks per channel can be parked before the consumer's
/// `adopt_channel` runs — and chunks past the cap drop silently at the
/// consumer (data loss presents as a truncated stream with clean
/// framing elsewhere; `dropped_unknown_chunks` /
/// `early_arrival_count` are the producer-side observability). A
/// tunnel crate's sizing (e.g. a UDP POC's MTU-vs-buffer math) should
/// treat 64 parked chunks as the observable bound and adopt promptly.
const EARLY_ARRIVAL_CAP: usize = 64;
impl ChannelManager {
/// Construct a new manager with the given `MuxHandle`, max
/// channels, buffer cap, and side. The `remote_addr` is
@@ -124,6 +149,8 @@ impl ChannelManager {
remote_addr,
side,
dropped_unknown_chunks: AtomicU64::new(0),
early_arrivals: Mutex::new(HashMap::new()),
early_arrival_count: AtomicU64::new(0),
}),
}
}
@@ -175,6 +202,19 @@ impl ChannelManager {
self.inner.dropped_unknown_chunks.load(Ordering::Relaxed)
}
/// The number of chunks parked in the early-arrival buffer since
/// the connection started (monotonic — parked chunks handed to the
/// adopter are not subtracted). Together with
/// [`Self::dropped_unknown_chunks`] this bounds the two observable
/// halves of the open-op-response/first-data race: how many of a
/// push-first producer's first chunks were buffered
/// (`early_arrival_count`) versus lost to the per-channel cap
/// (`dropped_unknown_chunks` — see `EARLY_ARRIVAL_CAP`, review 006
/// E-04). Observability only — neither counter drives control flow.
pub fn early_arrival_count(&self) -> u64 {
self.inner.early_arrival_count.load(Ordering::Relaxed)
}
/// The mux handle — for registering new channels' write halves.
pub fn mux(&self) -> &MuxHandle {
&self.inner.mux
@@ -232,7 +272,7 @@ impl ChannelManager {
let (demux_sender, recv) = MpscRecvStream::channel(self.inner.buffer_cap);
let state = ChannelState {
demux_sender,
demux_sender: demux_sender.clone(),
handler_task,
alpn,
};
@@ -288,7 +328,7 @@ impl ChannelManager {
let (demux_sender, recv) = MpscRecvStream::channel(self.inner.buffer_cap);
let state = ChannelState {
demux_sender,
demux_sender: demux_sender.clone(),
handler_task,
alpn,
};
@@ -305,6 +345,10 @@ impl ChannelManager {
}
}
// Route any chunks that arrived between the open-op response and
// this adopt (the producer may have started pumping already).
self.drain_early_arrivals(channel_id, &demux_sender).await;
Ok((send, recv))
}
@@ -377,8 +421,15 @@ impl ChannelManager {
/// Route a chunk payload to the reassembled stream for
/// `channel_id` (the demux's per-chunk route). A zero-length
/// payload is the EOF sentinel — the reassembled stream interprets
/// it as EOF (REQ-CH-01). An unknown `channel_id` is dropped with
/// a debug log (REQ-CH-04 — lenient handling).
/// it as EOF (REQ-CH-01). An unknown `channel_id` parks the payload
/// in the early-arrival buffer (up to `EARLY_ARRIVAL_CAP` payloads
/// per channel) — the open-op response carrying `channel_id` and
/// the producer's first data-plane write race, and a push-first
/// producer (a TTY backend's banner, a sub protocol's greeting)
/// must not lose its first chunks. `adopt_channel` drains the
/// parked payloads into the new receiver. Chunks beyond the cap
/// drop with a debug log (the adopted-side REQ-CH-04 leniency for
/// genuinely unknown channels).
///
/// Awaits the bounded channel sender — if the handler's read half
/// is slow, the demux stalls here (ADR-040 REQ-CH-05: lossless
@@ -400,17 +451,67 @@ impl ChannelManager {
}
}
None => {
self.inner
.dropped_unknown_chunks
.fetch_add(1, Ordering::Relaxed);
debug!(
channel_id,
"demux: unknown channel_id, dropping chunk (lenient)"
);
self.park_early_arrival(channel_id, payload);
}
}
}
/// Park an early-arrival chunk for `channel_id` (the adopter has
/// not run yet). FIFO per channel, bounded by `EARLY_ARRIVAL_CAP`;
/// beyond the cap the chunk is dropped with a debug log and the
/// dropped counter increments (observability only).
fn park_early_arrival(&self, channel_id: u32, payload: Bytes) {
let mut arrivals = self.inner.early_arrivals.lock();
let queue = arrivals.entry(channel_id).or_default();
if queue.len() >= EARLY_ARRIVAL_CAP {
drop(arrivals);
self.inner
.dropped_unknown_chunks
.fetch_add(1, Ordering::Relaxed);
debug!(
channel_id,
"demux: unknown channel_id, early-arrival buffer full, dropping chunk"
);
return;
}
queue.push_back(payload);
self.inner
.early_arrival_count
.fetch_add(1, Ordering::Relaxed);
debug!(
channel_id,
"demux: unknown channel_id, parking chunk (early arrival, not yet adopted)"
);
}
/// Drain the early-arrival buffer for `channel_id` into `recv`'s
/// sender (called by `adopt_channel` after the channel state is
/// installed). Returns the drained payload count. Payloads past the
/// receiver's buffer capacity await normally (backpressure).
async fn drain_early_arrivals(
&self,
channel_id: u32,
demux_sender: &mpsc::Sender<Bytes>,
) -> usize {
let payloads = self
.inner
.early_arrivals
.lock()
.remove(&channel_id)
.unwrap_or_default();
let count = payloads.len();
for payload in payloads {
if demux_sender.send(payload).await.is_err() {
debug!(
channel_id,
"demux: channel receiver dropped while draining early arrivals"
);
break;
}
}
count
}
/// Take the sender for `channel_id` — used on close / teardown to
/// drop the sender (which signals EOF to the handler, REQ-CH-02).
/// Returns the handler task (if any) so the caller can abort it
@@ -447,6 +548,8 @@ impl ChannelManager {
task.abort();
}
}
// The early-arrival buffers die with the connection too.
self.inner.early_arrivals.lock().clear();
// The opener ledger has the opener PeerIds.
self.inner.opener_ledger.drain()
}
@@ -582,21 +685,90 @@ mod tests {
.await;
}
/// Early arrivals for a not-yet-adopted channel are parked, not
/// dropped: `adopt_channel` drains them into the new receiver. This
/// closes the open-op-response/first-data race (a push-first
/// producer's first chunks must not be lost).
#[tokio::test]
async fn route_payload_to_unknown_channel_parks_until_adopt() {
let manager = make_manager_with_runner().await;
manager.route_payload(7, Bytes::from_static(b"first")).await;
manager
.route_payload(7, Bytes::from_static(b"second"))
.await;
let (_send, mut recv) = manager
.adopt_channel(7, "alk/tty", None)
.await
.expect("adopt");
use tokio::io::AsyncReadExt;
let mut buf = [0u8; 11];
recv.read_exact(&mut buf).await.expect("read first+second");
assert_eq!(&buf, b"firstsecond");
}
#[tokio::test]
async fn route_payload_to_unknown_channel_increments_dropped_counter() {
let manager = make_manager_with_runner().await;
assert_eq!(manager.dropped_unknown_chunks(), 0);
// Parked arrivals don't increment the counter...
manager
.route_payload(999, Bytes::from_static(b"data"))
.await;
manager
.route_payload(998, Bytes::from_static(b"more"))
.route_payload(998, Bytes::from_static(b"kept"))
.await;
assert_eq!(
manager.dropped_unknown_chunks(),
2,
"dropped_unknown_chunks counter increments per unknown-channel drop"
0,
"parked early arrivals are not drops"
);
// ...but chunks beyond the per-channel cap do (the adopted-side
// REQ-CH-04 leniency for genuinely unknown channels).
for i in 0..(super::EARLY_ARRIVAL_CAP + 1) {
let n = i;
let payload = vec![b'x'; n % 9 + 1];
manager.route_payload(999, Bytes::from(payload)).await;
}
assert_eq!(
manager.dropped_unknown_chunks(),
1,
"dropped_unknown_chunks counter increments when the early-arrival buffer overflows"
);
}
/// `early_arrival_count` observes parked chunks (review 006 N-2 —
/// the counter was write-only before the accessor). Monotonic by
/// design: it counts chunks parked since connection start; draining
/// on adopt does not subtract (paired with `dropped_unknown_chunks`
/// it bounds the open-response/first-data race's two halves).
#[tokio::test]
async fn early_arrival_count_tracks_parked_chunks_monotonically() {
let manager = make_manager_with_runner().await;
assert_eq!(manager.early_arrival_count(), 0);
for _ in 0..3 {
manager
.route_payload(7, Bytes::from_static(b"parked"))
.await;
}
assert_eq!(
manager.early_arrival_count(),
3,
"each parked chunk increments the counter"
);
// Adoption drains the buffer into the receiver but does not
// decrement (the accessor is an observability counter, not a
// live-depth gauge).
let (_send, mut recv) = manager
.adopt_channel(7, "alk/tty", None)
.await
.expect("adopt");
assert_eq!(
manager.early_arrival_count(),
3,
"adopt-drain does not decrement the monotonic counter"
);
use tokio::io::AsyncReadExt;
let mut buf = [0u8; 6 * 3];
recv.read_exact(&mut buf).await.expect("read parked chunks");
}
#[tokio::test]
+3
View File
@@ -25,6 +25,8 @@
//! ALPN crates via `ChannelCore` (ADR-047 §3).
//! - [`policy`]: `ChannelLifecyclePolicy` + `PerIdentityChannelPolicy`
//! (ADR-041, amended by ADR-047 §7 — opener ledger).
//! - [`pump`]: `pump_bidi` — the two-pump data-plane helper (alknet
//! ADR-078, pinned upstream by ADR-050).
//! - [`client`]: `ChannelClient` — transport-agnostic
//! `from_connection` (ADR-043).
//! - [`self::env`]: `ChannelOperationEnv` extension trait (ADR-047 §4 —
@@ -39,6 +41,7 @@ pub mod manager;
pub mod mux;
pub mod operations;
pub mod policy;
pub mod pump;
pub mod reassembly;
pub mod source;
pub mod wire;
+1026 -49
View File
File diff suppressed because it is too large. Load diff
+204
View File
@@ -0,0 +1,204 @@
//! `pump_bidi` — the two-pump data-plane helper (alknet ADR-078, as
//! pinned upstream by ADR-050 — review 007 R-03).
//!
//! A forwarding channel's data plane is two unidirectional pumps: one
//! copies channel→peer, the other peer→channel. The contract is
//! **shutdown-on-completion**: when one pump's copy source EOFs, it
//! shuts down the opposite sink so the peer sees a clean half-close;
//! the helper completes when both pumps finish. Copy errors are
//! EOF-shaped by design (the POC and ADR-078 semantics — there is no
//! error channel mid-stream; a pump error is an abrupt-close signal,
//! and shutdown-of-the-opposite-sink is the same either way). The
//! helper returns the two copy counts for observability.
//!
//! The shapes converged across three consumers (the extraction test
//! ADR-078 set): the alktunnels POC's `pump_halves`, alktty's channels
//! session, and the assembly-layer copies. This helper pins the
//! two-pump shape in one place; alktty's three-pump session (an exit
//! future as a third signal) does not fit and stays as-is.
use tokio::io::{AsyncRead, AsyncWrite, AsyncWriteExt};
/// Pump a channel stream against a peer's split halves (ADR-078, as
/// pinned by ADR-050). Two pumps, joined:
///
/// - `channel → peer_write`: copy until the channel EOFs, then
/// `shutdown()` the peer write half.
/// - `peer_read → channel`: copy until the peer EOFs, then
/// `shutdown()` the channel.
///
/// Returns `(channel_to_peer, peer_to_channel)` copy counts.
/// Copy errors are treated as EOF (shutdown still runs) — the
/// EOF-shaped teardown the ADR-078 contract specifies; there is no
/// `Err` state because mid-stream errors are indistinguishable from
/// abrupt closes by design.
///
/// The generic bounds follow the POC: the channel side is a single
/// `AsyncRead + AsyncWrite` value (the `BiStream` the pump handler
/// received), the peer side arrives as split read/write halves (a
/// dialed socket, a TTY pty pair — `into_split` is the natural
/// establisher result).
///
/// Await the returned future inline inside the `OpenHandler`'s task —
/// the handler's returned `JoinHandle` must track the data-plane
/// lifetime (ADR-049 amendment 2, R-02).
pub async fn pump_bidi<C, R, W>(channel: C, peer_read: R, mut peer_write: W) -> (u64, u64)
where
C: AsyncRead + AsyncWrite + Unpin,
R: AsyncRead + Unpin,
W: AsyncWrite + Unpin,
{
let (mut c_read, mut c_write) = tokio::io::split(channel);
let mut peer_read = peer_read;
let c2p = async {
let n = tokio::io::copy(&mut c_read, &mut peer_write)
.await
.unwrap_or(0);
// Shutdown-on-completion: the channel side is done writing;
// the peer must see the half-close.
let _ = peer_write.shutdown().await;
n
};
let p2c = async {
let n = tokio::io::copy(&mut peer_read, &mut c_write)
.await
.unwrap_or(0);
let _ = c_write.shutdown().await;
n
};
tokio::join!(c2p, p2c)
}
#[cfg(test)]
mod tests {
use super::*;
use tokio::io::{AsyncReadExt, AsyncWriteExt, DuplexStream};
/// Model the topology precisely: the helper gets the channel
/// (`BiStream`-shaped) and the peer\'s split halves. The test
/// drives the *other* ends: `channel_peer` (what the channel\'s
/// remote writes/reads) and the substrate far ends. `duplex`
/// pairs: a write on one half arrives on its counterpart.
struct Topology {
/// The channel value handed to the helper.
channel: DuplexStream,
/// The channel\'s remote end (test-driven).
channel_peer: DuplexStream,
/// The peer read half handed to the helper (helper reads it).
peer_read: DuplexStream,
/// The peer write half handed to the helper (helper writes it).
peer_write: DuplexStream,
/// The peer write half\'s remote (test reads what the helper
/// copied).
peer_write_far: DuplexStream,
/// The peer read half\'s remote (test writes what the helper
/// will copy).
peer_read_far: DuplexStream,
}
fn topology() -> Topology {
// The channel\'s two ends.
let (channel, channel_peer) = tokio::io::duplex(64);
// The substrate: peer_read\'s far end, peer_write\'s far end.
let (peer_read, peer_read_far) = tokio::io::duplex(64);
let (peer_write_far, peer_write) = tokio::io::duplex(64);
Topology {
channel,
channel_peer,
peer_read,
peer_write,
peer_write_far,
peer_read_far,
}
}
/// The two-pump semantics test (review 007 Unit 3 gate): the
/// POC\'s `pump_halves` behavior through the helper — data flows
/// both directions, shutdown-on-completion on EOF, exact counts.
#[tokio::test]
async fn pump_bidi_moves_data_both_ways_and_shuts_down_on_completion() {
let mut topo = topology();
let pumped = tokio::spawn(pump_bidi(topo.channel, topo.peer_read, topo.peer_write));
// channel -> substrate: write at the channel\'s remote, read
// at the substrate\'s far end.
topo.channel_peer
.write_all(b"outbound")
.await
.expect("write");
let mut buf = [0u8; 8];
topo.peer_write_far
.read_exact(&mut buf)
.await
.expect("helper copied channel->peer");
assert_eq!(&buf, b"outbound");
// substrate -> channel: write at the substrate\'s far end,
// read at the channel\'s remote.
topo.peer_read_far
.write_all(b"inbound!")
.await
.expect("write");
let mut buf2 = [0u8; 8];
topo.channel_peer
.read_exact(&mut buf2)
.await
.expect("helper copied peer->channel");
assert_eq!(&buf2, b"inbound!");
// EOF from the substrate\'s far end: the helper shuts the
// channel down (shutdown-on-completion).
topo.peer_read_far.shutdown().await.expect("shutdown");
let mut eof_buf = Vec::new();
let n = topo
.channel_peer
.read_to_end(&mut eof_buf)
.await
.expect("clean EOF at the channel remote");
assert_eq!(n, 0, "the channel sees EOF after the substrate finished");
// Finish the other pump (channel remote drops) and collect
// the counts.
drop(topo.channel_peer);
let (c2p, p2c) = pumped
.await
.expect("helper completes when both pumps finish");
assert_eq!(c2p, 8, "channel->peer count");
assert_eq!(p2c, 8, "peer->channel count");
// The substrate\'s far write end observes the helper\'s
// shutdown: read returns 0 (EOF), not an error.
let mut eof_buf2 = Vec::new();
let n2 = topo
.peer_write_far
.read_to_end(&mut eof_buf2)
.await
.expect("eof read");
assert_eq!(n2, 0, "peer sees clean EOF after the channel side finished");
}
/// Copy errors are EOF-shaped: a pump whose source dies at birth
/// still lets the helper complete and shut the opposite sink down
/// (no Err state — ADR-078\'s abrupt-close semantics).
#[tokio::test]
async fn pump_bidi_treats_errors_as_eof_shaped_teardown() {
let topo = topology();
// Drop the substrate\'s far write end: the helper\'s peer_read
// source EOFs immediately; the channel side must still be shut
// down cleanly once the channel itself finishes.
drop(topo.peer_read_far);
let pumped = tokio::spawn(pump_bidi(topo.channel, topo.peer_read, topo.peer_write));
// The channel remote first receives what nothing sends — EOF
// after the p2c pump finishes; drop the remote to finish the
// c2p pump too.
drop(topo.channel_peer);
let _counts = pumped
.await
.expect("helper completes despite the dead source");
}
}
+66 -6
View File
@@ -4,12 +4,14 @@
//!
//! The handler for a channel receives a `Connection` constructed via
//! `Connection::from_source(ChannelBidiStreamSource, alpn)`. The
//! handler calls `accept_bi()` once (yield-once per channel, ADR-065)
//! handler calls `accept_bi()` once (yield-once per channel, ADR-007)
//! and gets a `BiStream` — identical to how it works on a top-level
//! QUIC connection. The `BiStream` is the reassembled read half joined
//! to the mux write half via `BiStream::from_joined`.
use std::net::SocketAddr;
use std::sync::atomic::{AtomicBool, Ordering};
use std::sync::Arc;
use async_trait::async_trait;
use parking_lot::Mutex;
@@ -25,12 +27,19 @@ use crate::core::types::{BiStream, BidiStreamSource, StreamError};
///
/// The `BiStream` is constructed once, in [`ChannelBidiStreamSource::new`],
/// and held in an `Option` — `accept_bi` takes it. This preserves the
/// yield-once contract (ADR-065) and the "split never crosses a crate
/// boundary as part of a constructor" rule (ADR-092) — the join happens
/// yield-once contract (ADR-007) and the "split never crosses a crate
/// boundary as part of a constructor" rule (ADR-009) — the join happens
/// here, in the `BidiStreamSource` impl, not per-handler.
///
/// The `accepted` flag (R-02 telemetry) records whether `accept_bi`
/// ever yielded: the open wrapper's teardown task checks it when the
/// pump handler exits and logs a `debug!` hint when the channel's
/// stream was never accepted — the "teardown at birth" shape a
/// fire-and-forget handler produces (review 007 R-02).
pub struct ChannelBidiStreamSource {
stream: Mutex<Option<BiStream>>,
remote_addr: Option<SocketAddr>,
accepted: Arc<AtomicBool>,
}
impl ChannelBidiStreamSource {
@@ -38,11 +47,29 @@ impl ChannelBidiStreamSource {
/// half joined to the mux write half). The handler will call
/// `accept_bi()` once and receive this stream.
pub fn new(stream: BiStream, remote_addr: Option<SocketAddr>) -> Self {
Self::with_accepted_flag(stream, remote_addr, Arc::new(AtomicBool::new(false)))
}
/// Construct with a shared acceptance flag — the open wrapper
/// passes its own so the teardown task can observe whether the
/// handler ever accepted the channel's stream (R-02 telemetry).
pub fn with_accepted_flag(
stream: BiStream,
remote_addr: Option<SocketAddr>,
accepted: Arc<AtomicBool>,
) -> Self {
Self {
stream: Mutex::new(Some(stream)),
remote_addr,
accepted,
}
}
/// Whether `accept_bi` ever yielded the stream. `false` after
/// construction; flips once on the first successful accept.
pub fn accepted(&self) -> bool {
self.accepted.load(Ordering::SeqCst)
}
}
#[async_trait]
@@ -50,7 +77,10 @@ impl BidiStreamSource for ChannelBidiStreamSource {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
let mut guard = self.stream.lock();
match guard.take() {
Some(stream) => Ok(stream),
Some(stream) => {
self.accepted.store(true, Ordering::SeqCst);
Ok(stream)
}
None => Err(StreamError::ConnectionClosed),
}
}
@@ -65,9 +95,9 @@ impl BidiStreamSource for ChannelBidiStreamSource {
/// `code`/`reason` are ignored: a single channel has no
/// QUIC-shaped application-level close codes. The drop is the
/// close (ADR-065 §"Negative"). The `_` prefix is intentional —
/// close (ADR-007 §"Negative"). The `_` prefix is intentional —
/// the signature matches the public `Connection::close` API
/// (ADR-070 §"REQ-CORE-02").
/// (ADR-008 §"REQ-CORE-02").
fn close(&self, _code: u32, _reason: &str) {
let _ = self.stream.lock().take();
}
@@ -89,11 +119,26 @@ pub fn channel_source(
ChannelBidiStreamSource::new(stream, remote_addr)
}
/// [`channel_source`] with a shared acceptance flag (R-02 telemetry):
/// the open wrapper passes the flag, hands the source to the
/// handler's `Connection`, and the teardown task reads it to log the
/// "handler exited without accepting the channel's stream" hint.
pub fn channel_source_with_accepted_flag(
recv: super::reassembly::MpscRecvStream,
send: super::reassembly::MpscSendStream,
remote_addr: Option<SocketAddr>,
accepted: Arc<AtomicBool>,
) -> ChannelBidiStreamSource {
let stream = BiStream::from_joined(recv, send);
ChannelBidiStreamSource::with_accepted_flag(stream, remote_addr, accepted)
}
#[cfg(test)]
mod tests {
use super::*;
use crate::channels::reassembly::{MpscRecvStream, MpscSendStream};
use bytes::Bytes;
use std::sync::Arc;
use tokio::io::{AsyncReadExt, AsyncWriteExt};
fn make_pair() -> (
@@ -166,4 +211,19 @@ mod tests {
bidi.read_exact(&mut buf).await.expect("read");
assert_eq!(&buf, b"inbound");
}
#[tokio::test]
async fn accepted_flag_tracks_the_yield_once_accept() {
let (send, _mux_recv, _demux_send, recv) = make_pair();
let accepted = Arc::new(AtomicBool::new(false));
let source = channel_source_with_accepted_flag(recv, send, None, Arc::clone(&accepted));
assert!(!source.accepted(), "false before the handler accepts");
let _bidi = source.accept_bi().await.expect("first accept");
assert!(source.accepted(), "flips on the accept");
assert!(matches!(
source.accept_bi().await,
Err(StreamError::ConnectionClosed)
));
assert!(source.accepted(), "stays true after exhaustion");
}
}
+1 -1
View File
@@ -23,7 +23,7 @@ use thiserror::Error;
pub const CHUNK_HEADER_LEN: usize = 8;
/// The maximum chunk payload length (16 MiB, matching TTY's cap —
/// ADR-052 §5). A chunk with `length > MAX_CHUNK_LEN` returns
/// ADR-034). A chunk with `length > MAX_CHUNK_LEN` returns
/// [`ChunkError::TooLarge`] and does not corrupt the stream — the demux
/// drops the chunk and continues. The header is always exactly 8 bytes,
/// so the demux can always resync by reading the next 8-byte header.
+1 -1
View File
@@ -122,7 +122,7 @@ mod tests {
}
fn registry_with_caps() -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("pub/run"),
+144 -7
View File
@@ -202,7 +202,11 @@ async fn fetch_schema(connection: &CallConnection, name: &str) -> Result<Value,
/// Rebuild an `OperationSpec` from the `services/schema` JSON, applying the
/// optional namespace prefix. The spec JSON shape matches `spec_to_json` in
/// `registry/discovery.rs`.
fn rebuild_spec_for(
///
/// `pub(crate)` so the `op/register` bootstrap path (review 004 F-05) can
/// rebuild peer-announced specs from the same wire shape — one parser, two
/// consumers.
pub(crate) fn rebuild_spec_for(
schema_json: &Value,
remote_name: &str,
namespace_prefix: &Option<String>,
@@ -252,7 +256,10 @@ fn rebuild_spec_for(
output_schema,
error_schemas,
access_control,
None,
schema_json
.get("resource_id_path")
.and_then(|v| v.as_str())
.map(String::from),
);
// ADR-047 §2: the `channel_open` marker survives discovery
@@ -278,6 +285,10 @@ fn rebuild_spec_for(
}
}
if let Some(description) = schema_json.get("description").and_then(|v| v.as_str()) {
spec = spec.with_description(description);
}
Ok(spec)
}
@@ -385,7 +396,10 @@ fn parse_access_control(v: &Value) -> AccessControl {
/// If `context.identity` is `None` (the hub chose not to disclose, or has not
/// authenticated an originator), `forwarded_for` is omitted — the spoke
/// receives only the hub's identity.
fn make_forwarding_handler(connection: Arc<CallConnection>, remote_name: String) -> Handler {
pub(crate) fn make_forwarding_handler(
connection: Arc<CallConnection>,
remote_name: String,
) -> Handler {
use crate::registry::registration::make_handler;
make_handler(move |input, context| {
let connection = Arc::clone(&connection);
@@ -411,7 +425,7 @@ fn make_forwarding_handler(connection: Arc<CallConnection>, remote_name: String)
/// leaf: on invocation, calls `CallConnection::subscribe_with_payload()` and
/// forwards the remote stream end-to-end. Each `call.responded` from the
/// remote becomes a stream item, `call.completed` ends the stream, and
/// `call.aborted` drops it (ADR-049 §8). No truncation, no first-value
/// `call.aborted` drops it (ADR-021 §8). No truncation, no first-value
/// fallback.
///
/// `forwarded_for` is populated from `context.identity` (ADR-032 §3), exactly
@@ -421,7 +435,7 @@ fn make_forwarding_handler(connection: Arc<CallConnection>, remote_name: String)
/// `PendingRequestMap`, so the abort cascade (ADR-016 §6) is already wired:
/// a parent abort drops the `SubscriptionStream`, which sends `call.aborted`
/// to the remote node.
fn make_streaming_forwarding_handler(
pub(crate) fn make_streaming_forwarding_handler(
connection: Arc<CallConnection>,
remote_name: String,
) -> StreamingHandler {
@@ -455,7 +469,7 @@ fn make_streaming_forwarding_handler(
/// `forwarded_for` is populated from `context.identity` (ADR-032 §3),
/// exactly as the request/response and streaming forwarding handlers
/// do — both via `build_forwarded_payload`.
fn make_sink_forwarding_handler(
pub(crate) fn make_sink_forwarding_handler(
connection: Arc<CallConnection>,
remote_name: String,
) -> SinkHandler {
@@ -480,7 +494,11 @@ fn make_sink_forwarding_handler(
/// `forwarded_for` from the hub's `OperationContext.identity` (ADR-032 §3).
/// `forwarded_for` is omitted when `context.identity` is `None` (the hub
/// chooses not to disclose the originator).
fn build_forwarded_payload(operation_id: &str, input: Value, context: &OperationContext) -> Value {
pub(crate) fn build_forwarded_payload(
operation_id: &str,
input: Value,
context: &OperationContext,
) -> Value {
let mut payload = serde_json::Map::new();
payload.insert(
"operationId".to_string(),
@@ -591,6 +609,67 @@ mod tests {
assert_eq!(spec.access_control.resource_type.as_deref(), Some("fs"));
}
// --- review 005 Unit 3 gate (G-04 spec wire round-trip) ----------------
/// G-04 gate: `resource_id_path` survives the spec wire round-trip —
/// `spec_to_json_pub` serializes it, `rebuild_spec_for` parses it.
/// The review found the field silently dropped: an announced (or
/// `from_call`-imported) op with ownership-scoped resource
/// extraction rebuilt with `resource_id: None`, so the rebuilt
/// spec's ACL checks ran without the resource ID.
#[test]
fn spec_round_trips_resource_id_path() {
use crate::registry::discovery::spec_to_json_pub;
let spec = OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
crate::registry::spec::AccessControl::default(),
Some("/path".to_string()),
);
let wire = spec_to_json_pub(&spec);
assert_eq!(
wire.get("resource_id_path").and_then(|v| v.as_str()),
Some("/path"),
"resource_id_path serialized"
);
let rebuilt = rebuild_spec_for(&wire, "fs/readFile", &None).expect("rebuild");
assert_eq!(
rebuilt.resource_id_path.as_deref(),
Some("/path"),
"resource_id_path survives the round-trip"
);
}
/// G-04 companion: a spec without `resource_id_path` serializes no
/// `resource_id_path` key and rebuilds with `None` (additive
/// optional field — absent stays absent).
#[test]
fn spec_without_resource_id_path_stays_absent_through_round_trip() {
use crate::registry::discovery::spec_to_json_pub;
let spec = OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
crate::registry::spec::AccessControl::default(),
None,
);
let wire = spec_to_json_pub(&spec);
assert!(wire.get("resource_id_path").is_none());
let rebuilt = rebuild_spec_for(&wire, "fs/readFile", &None).expect("rebuild");
assert_eq!(rebuilt.resource_id_path, None);
}
#[test]
fn rebuild_spec_channel_open_marker_set_for_channels_alpn_op() {
let mut schema = sample_schema_json("channels/tty/sub", "sub");
@@ -693,6 +772,64 @@ mod tests {
assert_eq!(rebuilt.publish_schema.as_ref(), Some(&publish_schema));
}
/// E-02 gate (review 006): `description` survives the spec wire
/// round-trip — `spec_to_json_pub` serializes it when set,
/// `rebuild_spec_for` parses it back (the `op/register` announced-spec
/// path and the `from_call` import path both parse through here).
#[test]
fn spec_round_trips_description() {
use crate::registry::discovery::spec_to_json_pub;
let spec = OperationSpec::new(
"channels/tty/sub",
OperationType::Sub,
Visibility::External,
json!({}),
json!({}),
vec![],
crate::registry::spec::AccessControl::default(),
None,
)
.with_description("Interactive TTY sessions");
let wire = spec_to_json_pub(&spec);
assert_eq!(
wire.get("description").and_then(|v| v.as_str()),
Some("Interactive TTY sessions"),
"description serialized"
);
let rebuilt = rebuild_spec_for(&wire, "channels/tty/sub", &None).expect("rebuild");
assert_eq!(
rebuilt.description.as_deref(),
Some("Interactive TTY sessions"),
"description survives the round-trip"
);
}
/// E-02 companion: a spec without `description` serializes no
/// `description` key and rebuilds with `None` (additive optional
/// field — absent stays absent, old producers stay parseable).
#[test]
fn spec_without_description_stays_absent_through_round_trip() {
use crate::registry::discovery::spec_to_json_pub;
let spec = OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
crate::registry::spec::AccessControl::default(),
None,
);
let wire = spec_to_json_pub(&spec);
assert!(wire.get("description").is_none());
let rebuilt = rebuild_spec_for(&wire, "fs/readFile", &None).expect("rebuild");
assert_eq!(rebuilt.description, None);
}
#[test]
fn derive_alpn_from_op_name_strips_channels_prefix() {
assert_eq!(
+5
View File
@@ -10,6 +10,11 @@ mod from_call;
pub use call_client::CallClient;
pub use from_call::{from_call, FromCallConfig};
// crate-internal surface for the `op/register` bootstrap path (review
// 004 F-05): the forwarding-handler constructor and the spec wire
// parser are shared with `registry::op_register`.
pub(crate) use from_call::{make_forwarding_handler, rebuild_spec_for};
use crate::registry::registration::HandlerRegistration;
/// Errors produced by [`OperationAdapter::import`].
+28 -26
View File
@@ -227,14 +227,14 @@ pub trait ProtocolHandler: Send + Sync + 'static {
async fn handle(&self, connection: Connection, auth: &AuthContext) -> Result<(), HandlerError>;
}
// --- BiStream: the handler leaf (ADR-092) ---------------------------------
// --- BiStream: the handler leaf (ADR-009) ---------------------------------
/// Internal helper trait — the union of `AsyncRead + AsyncWrite + Send`.
/// Not public; exists only to give `BiStream` a single boxed field.
trait AsyncReadWrite: AsyncRead + AsyncWrite + Send {}
impl<T: AsyncRead + AsyncWrite + Send> AsyncReadWrite for T {}
/// The handler leaf — a bidirectional byte stream (ADR-092).
/// The handler leaf — a bidirectional byte stream (ADR-005, ADR-009).
///
/// `accept_bi`/`open_bi` return a `BiStream`, not a split
/// `(SendStream, RecvStream)` pair. Handlers that want the split halves call
@@ -316,7 +316,7 @@ impl AsyncWrite for BiStream {
}
}
// --- SendStream / RecvStream: thin newtypes (ADR-092) ---------------------
// --- SendStream / RecvStream: thin newtypes (ADR-009) ---------------------
pub struct SendStream {
inner: Box<dyn AsyncWrite + Send + Unpin>,
@@ -327,10 +327,11 @@ pub struct RecvStream {
}
impl SendStream {
/// Box a write half into the thin `SendStream` newtype. Used by
/// `into_sub_streams()` (ADR-074) and the channels reassembly path.
/// Not a constructor that feeds `Connection` — the split never crosses
/// a crate boundary as part of a constructor (ADR-092).
/// Box a write half into the thin `SendStream` newtype. Used by tests
/// and any consumer that holds pre-split halves (the channels layer
/// has its own `MpscSendStream`; `into_sub_streams()` was removed by
/// ADR-035). Not a constructor that feeds `Connection` — the split
/// never crosses a crate boundary as part of a constructor (ADR-009).
pub fn from_stream(stream: impl AsyncWrite + Send + Unpin + 'static) -> Self {
Self {
inner: Box::new(stream),
@@ -339,10 +340,11 @@ impl SendStream {
}
impl RecvStream {
/// Box a read half into the thin `RecvStream` newtype. Used by
/// `into_sub_streams()` (ADR-074) and the channels reassembly path.
/// Not a constructor that feeds `Connection` — the split never crosses
/// a crate boundary as part of a constructor (ADR-092).
/// Box a read half into the thin `RecvStream` newtype. Used by tests
/// and any consumer that holds pre-split halves (the channels layer
/// has its own `MpscRecvStream`; `into_sub_streams()` was removed by
/// ADR-035). Not a constructor that feeds `Connection` — the split
/// never crosses a crate boundary as part of a constructor (ADR-009).
pub fn from_stream(stream: impl AsyncRead + Send + Unpin + 'static) -> Self {
Self {
inner: Box::new(stream),
@@ -386,15 +388,15 @@ impl AsyncRead for RecvStream {
/// Yield bidirectional streams to a `Connection`. Downstream crates implement
/// this trait to add connection shapes (channels, a future transport, a test
/// double beyond the single-stream case) without editing core. See ADR-070
/// for the full rationale and ADR-065 for the yield-once contract the
/// double beyond the single-stream case) without editing core. See ADR-008
/// for the full rationale and ADR-007 for the yield-once contract the
/// `StreamBidiStreamSource` impl preserves. The return type is `BiStream`
/// (ADR-092) — the join happens once, in the impl, not per-handler.
/// (ADR-009) — the join happens once, in the impl, not per-handler.
#[async_trait]
pub trait BidiStreamSource: Send + Sync + 'static {
/// Yield the next bidirectional stream this connection provides.
///
/// Transport semantics (carried from ADR-065):
/// Transport semantics (carried from ADR-007):
/// - QUIC (quinn/iroh): returns a new bidi stream on each call,
/// `ConnectionClosed` when the underlying connection closes.
/// - Single-stream (TCP+TLS, SSH channel, WebTransport stream, wasm):
@@ -407,7 +409,7 @@ pub trait BidiStreamSource: Send + Sync + 'static {
/// Open a bidirectional stream to the peer.
///
/// Single-stream sources return `StreamClosed` (a single stream cannot
/// open new application streams — ADR-065). QUIC and channels sources
/// open new application streams — ADR-007). QUIC and channels sources
/// open new streams.
async fn open_bi(&self) -> Result<BiStream, StreamError>;
@@ -416,13 +418,13 @@ pub trait BidiStreamSource: Send + Sync + 'static {
/// Close the connection. The `code`/`reason` args are QUIC application-
/// level close codes; non-QUIC sources ignore them (the drop is the
/// close — ADR-065 §"Negative"). See ADR-070 §"REQ-CORE-02" for the
/// close — ADR-007 §"Negative"). See ADR-008 §"REQ-CORE-02" for the
/// rationale for keeping the QUIC-shaped signature on the trait.
fn close(&self, code: u32, reason: &str);
}
/// Single-stream `BidiStreamSource` (TCP+TLS, SSH channel, WebTransport
/// stream, wasm stream — ADR-065). Crate-private; constructed via
/// stream, wasm stream — ADR-007). Crate-private; constructed via
/// `Connection::from_bidi`. `accept_bi` yields the underlying `BiStream`
/// once, then `ConnectionClosed`; `open_bi` returns `StreamClosed`.
struct StreamBidiStreamSource {
@@ -449,9 +451,9 @@ impl BidiStreamSource for StreamBidiStreamSource {
}
/// `code`/`reason` are ignored: a single stream has no QUIC-shaped
/// application-level close codes. The drop is the close (ADR-065
/// application-level close codes. The drop is the close (ADR-007
/// §"Negative"). The `_` prefix is intentional — the signature matches
/// the public `Connection::close` API (ADR-070 §"REQ-CORE-02").
/// the public `Connection::close` API (ADR-008 §"REQ-CORE-02").
fn close(&self, _code: u32, _reason: &str) {
let _ = self.stream.lock().unwrap_or_else(|e| e.into_inner()).take();
}
@@ -467,11 +469,11 @@ impl Connection {
/// Construct a `Connection` from a single bidirectional stream (e.g.
/// `tokio::io::DuplexStream`, `TlsStream<TcpStream>`,
/// `russh::Channel::into_stream()`). The stream is wrapped in a
/// `BiStream` (ADR-092) and yielded by `accept_bi` once, then
/// `BiStream` (ADR-009) and yielded by `accept_bi` once, then
/// `ConnectionClosed`. `open_bi` returns `StreamClosed` (a single
/// stream can't open new application streams — ADR-065).
/// stream can't open new application streams — ADR-007).
///
/// This is the only public stream constructor (ADR-092): the split
/// This is the only public stream constructor (ADR-009): the split
/// never crosses a crate boundary as part of a constructor. Handlers
/// that want the split halves call `tokio::io::split(&mut *stream)` on
/// the `BiStream` they receive from `accept_bi`.
@@ -492,7 +494,7 @@ impl Connection {
/// Construct from a caller-supplied `BidiStreamSource` impl. The
/// extension point for downstream crates — implement the trait and
/// construct a `Connection` from it without editing core. See ADR-070.
/// construct a `Connection` from it without editing core. See ADR-008.
pub fn from_source(source: impl BidiStreamSource, alpn: Vec<u8>) -> Self {
Self {
source: Box::new(source),
@@ -514,7 +516,7 @@ impl Connection {
/// Handlers that loop `accept_bi` (e.g. `TtyAdapter`) get one session
/// per single-stream connection; handlers that call once (e.g.
/// `HttpAdapter`) get the stream directly. Both are correct. The
/// return type is `BiStream` (ADR-092); handlers that want the split
/// return type is `BiStream` (ADR-009); handlers that want the split
/// halves call `tokio::io::split` on the `BiStream`.
pub async fn accept_bi(&self) -> Result<BiStream, StreamError> {
self.source.accept_bi().await
@@ -656,7 +658,7 @@ mod tests {
/// `tokio::io::sink()` + `tokio::io::empty()`: reads yield EOF
/// immediately (zero bytes), writes discard. Exists because
/// `Connection::from_bidi` requires a single value that implements
/// both traits (ADR-092). Used only to construct a `Connection` for
/// both traits (ADR-009). Used only to construct a `Connection` for
/// tests that exercise `Connection`-level state (alpn, addr, identity)
/// without ever reading or writing the stream.
pub(crate) struct SinkEmpty;
+787
View File
@@ -0,0 +1,787 @@
//! The transport-neutral dispatch spine ([`GatewayDispatch`]) and the
//! shared `services/schema` disclosure guard
//! ([`schema_disclosure_denial`]) — the non-HTTP half of alkhttp's
//! gateway, promoted for reuse by hubs, relays, and protocol crates
//! that expose a call surface to a less-trusted in-transport caller.
//!
//! The spine constructs the root [`OperationContext`] identically for
//! every transport (`internal: false` — ACL runs against the caller's
//! identity, not a handler's composition authority; `forwarded_for:
//! None` — wire-ingress only), resolves the registration's
//! composition authority / capabilities / scoped env into it, and
//! bounds Once-ops and sink dispatch with a configurable deadline
//! while leaving streaming subscriptions unbounded (ADR-021:
//! subscriptions are long-lived).
//!
//! There are no HTTP concepts here: no statuses, no headers, no body
//! framing. CallError → HTTP mapping, NDJSON/SSE framing, body caps,
//! and batch envelopes are the consumer's (see alkhttp's gateway).
//!
//! See ADR-048 for the promotion decision and the divergence note on
//! ACL-denial codes (`FORBIDDEN` here, spec-404 in the wire handler).
use std::collections::HashMap;
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::Arc;
use std::time::{Duration, Instant};
use crate::core::auth::Identity;
use crate::core::types::Capabilities;
use crate::protocol::wire::{CallError, ResponseEnvelope};
use crate::registry::context::{AbortPolicy, OperationContext, ScopedPeerEnv};
use crate::registry::env::LocalOperationEnv;
use crate::registry::registration::OperationRegistry;
use crate::registry::spec::{AccessResult, Visibility};
use futures::stream::BoxStream;
use serde_json::Value;
const SERVICES_SCHEMA: &str = "services/schema";
/// The default handler deadline for Once-ops and sink dispatch: 30 s.
/// Override with [`GatewayDispatch::with_deadline`].
pub const DEFAULT_DEADLINE: Duration = Duration::from_secs(30);
/// The transport-neutral dispatch spine over an
/// [`OperationRegistry`](crate::registry::registration::OperationRegistry):
/// invokes operations for the neutral `ResponseEnvelope` result shape
/// any transport maps to its own wire format. Identity arrives
/// per-call as `Option<Identity>` — bearer resolution, transport
/// framing, and error presentation happen upstream in the consumer.
///
/// See the [module docs](crate::gateway) and ADR-048.
pub struct GatewayDispatch {
registry: Arc<OperationRegistry>,
deadline: Option<Duration>,
invoke_count: AtomicUsize,
}
impl GatewayDispatch {
/// Assemble a dispatch spine over a registry with the 30 s default
/// deadline ([`DEFAULT_DEADLINE`]).
pub fn new(registry: Arc<OperationRegistry>) -> Self {
Self {
registry,
deadline: Some(DEFAULT_DEADLINE),
invoke_count: AtomicUsize::new(0),
}
}
/// Override the Once/sink deadline: `Some(duration)` bounds every
/// [`GatewayDispatch::invoke`] and [`GatewayDispatch::invoke_sink`]
/// dispatch (a hung handler surfaces as a retryable `TIMEOUT` error
/// envelope); `None` removes the bound entirely (the wire path's
/// Pub dispatch behavior). Streaming subscriptions are unbounded
/// either way (ADR-021).
pub fn with_deadline(mut self, deadline: Option<Duration>) -> Self {
self.deadline = deadline;
self
}
/// The registry operations resolve against.
pub fn registry(&self) -> &Arc<OperationRegistry> {
&self.registry
}
/// How many [`GatewayDispatch::invoke`] calls this spine has
/// served. A test-spy accessor: consumers' over-cap batch tests
/// assert it stays at zero to prove no dispatch happened before the
/// cap rejection.
pub fn invoke_count(&self) -> usize {
self.invoke_count.load(Ordering::Relaxed)
}
/// Invoke a Query/Mutation op under the configured deadline; a hung
/// handler surfaces as a retryable `TIMEOUT` error envelope.
pub async fn invoke(
&self,
identity: Option<Identity>,
op: &str,
input: Value,
) -> ResponseEnvelope {
self.invoke_count.fetch_add(1, Ordering::Relaxed);
let operation_name = strip_leading_slash(op).to_string();
if let Some(error) =
schema_via_call_denial(&self.registry, &operation_name, &input, identity.as_ref())
{
return ResponseEnvelope::error(uuid::Uuid::new_v4().to_string(), error);
}
let request_id = uuid::Uuid::new_v4().to_string();
let context = self.build_root_context(&request_id, &operation_name, identity);
let fut = self.registry.invoke(&operation_name, input, context);
match self.deadline {
Some(deadline) => match tokio::time::timeout(deadline, fut).await {
Ok(envelope) => envelope,
Err(_elapsed) => ResponseEnvelope::error(
request_id,
CallError::timeout(format!(
"operation did not complete within the {deadline:?} dispatch deadline"
)),
),
},
None => fut.await,
}
}
/// Dispatch a Sub op: the returned stream of envelopes is unbounded
/// by the deadline (subscriptions are long-lived per ADR-021);
/// pre-handler failures surface as one error envelope.
pub fn invoke_streaming(
&self,
identity: Option<Identity>,
op: &str,
input: Value,
) -> BoxStream<'static, ResponseEnvelope> {
let operation_name = strip_leading_slash(op).to_string();
if let Some(error) =
schema_via_call_denial(&self.registry, &operation_name, &input, identity.as_ref())
{
let request_id = uuid::Uuid::new_v4().to_string();
return Box::pin(futures::stream::once(async move {
ResponseEnvelope::error(request_id, error)
}));
}
let request_id = uuid::Uuid::new_v4().to_string();
let context = self.build_root_context_streaming(&request_id, &operation_name, identity);
self.registry
.invoke_streaming(&operation_name, input, context)
}
/// Dispatch a `Pub` operation (ADR-046). The `publish_stream` is
/// the initiator's chunk stream (each item one published chunk, or
/// an initiator-side error). The dispatch — chunk pacing and the
/// sink handler's final completion alike — is bounded by the
/// configured deadline: a hung sink handler surfaces as a retryable
/// `TIMEOUT` error envelope, not an indefinitely-held initiator.
pub async fn invoke_sink(
&self,
identity: Option<Identity>,
op: &str,
input: Value,
publish_stream: crate::registry::registration::PublishStream,
) -> ResponseEnvelope {
let operation_name = strip_leading_slash(op).to_string();
let request_id = uuid::Uuid::new_v4().to_string();
let context = self.build_root_context_sink(&request_id, &operation_name, identity);
let fut = self
.registry
.invoke_sink(&operation_name, input, publish_stream, context);
match self.deadline {
Some(deadline) => match tokio::time::timeout(deadline, fut).await {
Ok(envelope) => envelope,
Err(_elapsed) => ResponseEnvelope::error(
request_id,
CallError::timeout(format!(
"operation did not complete within the {deadline:?} dispatch deadline"
)),
),
},
None => fut.await,
}
}
fn build_root_context_sink(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
) -> OperationContext {
self.build_root_context_inner(request_id, operation_name, identity, false)
}
fn build_root_context(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
) -> OperationContext {
self.build_root_context_inner(request_id, operation_name, identity, true)
}
fn build_root_context_streaming(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
) -> OperationContext {
self.build_root_context_inner(request_id, operation_name, identity, false)
}
fn build_root_context_inner(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
bounded: bool,
) -> OperationContext {
let registration = self.registry.registration(operation_name);
let (composition_authority, capabilities, scoped_env) = match registration {
Some(r) => (
r.composition_authority.clone(),
r.capabilities.clone(),
r.scoped_env.clone().unwrap_or_else(ScopedPeerEnv::empty),
),
None => (None, Capabilities::new(), ScopedPeerEnv::empty()),
};
let env: Arc<dyn crate::registry::env::OperationEnv + Send + Sync> =
Arc::new(LocalOperationEnv::new(Arc::clone(&self.registry)));
OperationContext {
request_id: request_id.to_string(),
parent_request_id: None,
identity,
handler_identity: composition_authority,
forwarded_for: None,
capabilities,
metadata: HashMap::new(),
deadline: bounded
.then_some(self.deadline)
.flatten()
.map(|deadline| Instant::now() + deadline),
scoped_env,
env,
abort_policy: AbortPolicy::default(),
internal: false,
ownership: None,
}
}
}
fn strip_leading_slash(operation_id: &str) -> &str {
operation_id.strip_prefix('/').unwrap_or(operation_id)
}
/// The `services/schema` op-path guard (ADR-048): when the dispatched
/// operation is the schema meta-op, apply the disclosure checks to the
/// **inner** `name` input before dispatching, so no transport through
/// this spine can fetch a spec that would be denied for the same
/// identity elsewhere.
fn schema_via_call_denial(
registry: &OperationRegistry,
operation: &str,
input: &Value,
identity: Option<&Identity>,
) -> Option<CallError> {
let name = input.get("name").and_then(Value::as_str)?;
if !matches!(
registry.registration(operation),
Some(registration) if registration.spec.name == SERVICES_SCHEMA
) {
return None;
}
schema_disclosure_denial(registry, name, identity)
}
/// The schema-disclosure check shared by every transport surface:
/// Internal ops are invisible (`NOT_FOUND` regardless of caller);
/// ACL-forbidden ops are denied with `FORBIDDEN`. One implementation so
/// the transports cannot drift (CF-004's shared-implementation fix; the
/// wire-path handler stays the conservative spec-404 outer bound).
pub fn schema_disclosure_denial(
registry: &OperationRegistry,
operation: &str,
identity: Option<&Identity>,
) -> Option<CallError> {
let name = strip_leading_slash(operation);
let registration = registry.registration(name)?;
if registration.spec.visibility == Visibility::Internal {
return Some(CallError::not_found(operation));
}
if let AccessResult::Forbidden(message) =
registration.spec.access_control.check(identity, None, None)
{
return Some(CallError::forbidden(message));
}
None
}
#[cfg(test)]
mod tests {
use super::*;
use crate::registry::registration::{
make_handler, make_sink_handler, make_streaming_handler, HandlerKind, HandlerRegistration,
OperationProvenance,
};
use crate::registry::spec::{AccessControl, OperationSpec, OperationType, Visibility};
use futures::StreamExt;
// The deadline tests use a short custom deadline instead of the
// module default, so no wall-clock assertion by hand is needed.
fn spec(name: &str, visibility: Visibility, op_type: OperationType) -> OperationSpec {
OperationSpec::new(
name,
op_type,
visibility,
serde_json::json!({}),
serde_json::json!({}),
vec![],
AccessControl::default(),
None,
)
}
fn echo_registry() -> Arc<OperationRegistry> {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("echo/run", Visibility::External, OperationType::Query),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
Arc::new(registry)
}
#[tokio::test]
async fn invoke_external_op_round_trips() {
let dispatch = GatewayDispatch::new(echo_registry());
let envelope = dispatch
.invoke(None, "/echo/run", serde_json::json!({ "x": 1 }))
.await;
assert!(envelope.result.is_ok(), "expected ok, got {envelope:?}");
}
#[tokio::test]
async fn invoke_unknown_op_returns_not_found() {
let dispatch = GatewayDispatch::new(echo_registry());
let envelope = dispatch
.invoke(None, "/missing/op", serde_json::json!({}))
.await;
match envelope.result {
Err(error) => assert_eq!(error.code, "NOT_FOUND"),
Ok(v) => panic!("expected error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_internal_op_returns_not_found() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("internal/op", Visibility::Internal, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry));
let envelope = dispatch
.invoke(None, "/internal/op", serde_json::json!({}))
.await;
match envelope.result {
Err(error) => assert_eq!(error.code, "NOT_FOUND"),
Ok(v) => panic!("expected error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_enforces_the_configured_deadline_on_a_hung_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("hung/op", Visibility::External, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
tokio::time::sleep(Duration::from_secs(120)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch =
GatewayDispatch::new(Arc::new(registry)).with_deadline(Some(Duration::from_millis(50)));
let started = std::time::Instant::now();
let envelope = dispatch
.invoke(None, "/hung/op", serde_json::json!({}))
.await;
assert!(
started.elapsed() < Duration::from_secs(10),
"the 50 ms deadline must fire well before the 120 s handler sleep"
);
match envelope.result {
Err(error) => {
assert_eq!(error.code, "TIMEOUT");
assert!(error.retryable, "the deadline error is retryable");
}
Ok(v) => panic!("expected a TIMEOUT error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_completes_within_the_deadline_for_a_fast_handler() {
let dispatch = GatewayDispatch::new(echo_registry());
let envelope = dispatch
.invoke(None, "/echo/run", serde_json::json!({}))
.await;
assert!(envelope.result.is_ok(), "a fast handler must not time out");
}
#[tokio::test]
async fn invoke_with_no_deadline_completes_a_slow_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("slow/op", Visibility::External, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
tokio::time::sleep(Duration::from_millis(80)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry)).with_deadline(None);
let envelope = dispatch
.invoke(None, "/slow/op", serde_json::json!({}))
.await;
assert!(
envelope.result.is_ok(),
"a no-deadline spine must not time out, got {envelope:?}"
);
}
#[tokio::test]
async fn invoke_sink_enforces_the_configured_deadline_on_a_hung_sink_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("hung/sink", Visibility::External, OperationType::Pub),
HandlerKind::Sink(make_sink_handler(|_input, ctx, _chunks| async move {
tokio::time::sleep(Duration::from_secs(120)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch =
GatewayDispatch::new(Arc::new(registry)).with_deadline(Some(Duration::from_millis(50)));
let chunks: crate::registry::registration::PublishStream =
Box::pin(futures::stream::empty());
let started = std::time::Instant::now();
let envelope = dispatch
.invoke_sink(None, "/hung/sink", serde_json::json!({}), chunks)
.await;
assert!(
started.elapsed() < Duration::from_secs(10),
"the 50 ms deadline must fire well before the 120 s handler sleep"
);
match envelope.result {
Err(error) => {
assert_eq!(error.code, "TIMEOUT");
assert!(error.retryable, "the deadline error is retryable");
}
Ok(v) => panic!("expected a TIMEOUT error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_sink_completes_within_the_deadline_for_a_fast_sink_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("fast/sink", Visibility::External, OperationType::Pub),
HandlerKind::Sink(make_sink_handler(|_input, ctx, mut chunks| async move {
let mut count = 0u64;
while let Some(chunk) = chunks.next().await {
if chunk.is_ok() {
count += 1;
}
}
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({ "count": count }))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry));
let chunks: crate::registry::registration::PublishStream =
Box::pin(futures::stream::iter(vec![
Ok(serde_json::json!({ "n": 1 })),
Ok(serde_json::json!({ "n": 2 })),
]));
let envelope = dispatch
.invoke_sink(None, "/fast/sink", serde_json::json!({}), chunks)
.await;
assert!(
envelope.result.is_ok(),
"a fast sink handler must not time out, got {envelope:?}"
);
}
#[tokio::test]
async fn invoke_sink_with_no_deadline_completes_a_slow_sink_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("slow/sink", Visibility::External, OperationType::Pub),
HandlerKind::Sink(make_sink_handler(|_input, ctx, mut chunks| async move {
while let Some(chunk) = chunks.next().await {
let _ = chunk;
}
tokio::time::sleep(Duration::from_millis(80)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry)).with_deadline(None);
let chunks: crate::registry::registration::PublishStream = Box::pin(futures::stream::iter(
vec![Ok(serde_json::json!({ "n": 1 }))],
));
let envelope = dispatch
.invoke_sink(None, "/slow/sink", serde_json::json!({}), chunks)
.await;
assert!(
envelope.result.is_ok(),
"a no-deadline spine must not time out a slow sink, got {envelope:?}"
);
}
#[tokio::test]
async fn streaming_sub_op_streams_envelopes() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("tick/stream", Visibility::External, OperationType::Sub),
HandlerKind::Stream(make_streaming_handler(|input, ctx| {
let count = input.get("count").and_then(|v| v.as_u64()).unwrap_or(2);
let request_id = ctx.request_id.clone();
futures::stream::iter(0..count).map(move |i| {
ResponseEnvelope::ok(request_id.clone(), serde_json::json!({ "tick": i }))
})
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry));
let mut stream =
dispatch.invoke_streaming(None, "/tick/stream", serde_json::json!({ "count": 3 }));
let mut ticks = Vec::new();
while let Some(envelope) = stream.next().await {
ticks.push(envelope);
}
assert_eq!(ticks.len(), 3);
}
#[tokio::test]
async fn invoke_count_counts_invokes_only() {
let dispatch = GatewayDispatch::new(echo_registry());
let _ = dispatch
.invoke(None, "/echo/run", serde_json::json!({}))
.await;
let _ = dispatch.invoke_streaming(None, "/echo/run", serde_json::json!({}));
assert_eq!(dispatch.invoke_count(), 1);
}
fn registry_with_services_schema_over(inner_ops: Vec<OperationSpec>) -> Arc<OperationRegistry> {
use crate::registry::discovery::{services_schema_handler, services_schema_spec};
let inner = Arc::new({
let registry = OperationRegistry::new();
for op_spec in &inner_ops {
registry
.register(HandlerRegistration::new(
op_spec.clone(),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
}
registry
});
let registry = OperationRegistry::new();
for op_spec in &inner_ops {
registry
.register(HandlerRegistration::new(
op_spec.clone(),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
}
registry
.register(HandlerRegistration::new(
services_schema_spec(),
HandlerKind::Once(services_schema_handler(Arc::clone(&inner))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
Arc::new(registry)
}
#[tokio::test]
async fn invoke_of_services_schema_with_internal_inner_name_is_blocked_pre_dispatch() {
let registry = registry_with_services_schema_over(vec![spec(
"secret/op",
Visibility::Internal,
OperationType::Query,
)]);
let dispatch = GatewayDispatch::new(registry);
let envelope = dispatch
.invoke(
None,
"services/schema",
serde_json::json!({ "name": "secret/op" }),
)
.await;
match envelope.result {
Err(error) => {
assert_eq!(error.code, "NOT_FOUND");
assert!(error.message.contains("secret/op"));
}
Ok(v) => panic!("the internal op spec must not be returned, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_of_services_schema_with_authorized_inner_name_still_projects() {
let registry = registry_with_services_schema_over(vec![spec(
"public/op",
Visibility::External,
OperationType::Query,
)]);
let dispatch = GatewayDispatch::new(registry);
let envelope = dispatch
.invoke(
None,
"services/schema",
serde_json::json!({ "name": "public/op" }),
)
.await;
assert!(
envelope.result.is_ok(),
"an allowed inner name must still project, got {envelope:?}"
);
assert_eq!(
envelope
.result
.as_ref()
.ok()
.and_then(|v| v.get("name"))
.and_then(Value::as_str),
Some("public/op")
);
}
#[tokio::test]
async fn invoke_of_services_schema_with_forbidden_inner_name_is_denied() {
let restricted = OperationSpec::new(
"admin/op",
OperationType::Query,
Visibility::External,
serde_json::json!({}),
serde_json::json!({}),
vec![],
AccessControl {
required_scopes: vec!["admin".to_string()],
..Default::default()
},
None,
);
let registry = registry_with_services_schema_over(vec![restricted]);
let dispatch = GatewayDispatch::new(registry);
let envelope = dispatch
.invoke(
Some(Identity {
id: "user".to_string(),
scopes: vec!["user".to_string()],
resources: HashMap::new(),
}),
"services/schema",
serde_json::json!({ "name": "admin/op" }),
)
.await;
match envelope.result {
Err(error) => assert_eq!(error.code, "FORBIDDEN"),
Ok(v) => panic!("the ACL-denied spec must not be returned, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_streaming_of_services_schema_with_internal_inner_name_is_blocked() {
let registry = registry_with_services_schema_over(vec![spec(
"secret/op",
Visibility::Internal,
OperationType::Query,
)]);
let dispatch = GatewayDispatch::new(registry);
let mut stream = dispatch.invoke_streaming(
None,
"services/schema",
serde_json::json!({ "name": "secret/op" }),
);
let envelopes: Vec<ResponseEnvelope> = stream.by_ref().collect().await;
assert_eq!(envelopes.len(), 1, "the guard error is the only item");
match &envelopes[0].result {
Err(error) => assert_eq!(error.code, "NOT_FOUND"),
Ok(v) => panic!("the internal op spec must not be returned, got {v:?}"),
}
}
#[test]
fn schema_disclosure_denial_hides_internal_ops() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("secret/op", Visibility::Internal, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let error = schema_disclosure_denial(&registry, "secret/op", None).unwrap();
assert_eq!(error.code, "NOT_FOUND");
assert!(schema_disclosure_denial(&registry, "missing/op", None).is_none());
}
#[test]
fn schema_disclosure_denial_allows_unrestricted_ops_without_identity() {
let error = schema_disclosure_denial(&echo_registry(), "echo/run", None);
assert!(
error.is_none(),
"default-ACL ops are fetchable unauthenticated"
);
}
}
+6
View File
@@ -0,0 +1,6 @@
//! The transport-neutral dispatch spine and the shared
//! `services/schema` disclosure guard (ADR-048).
pub(crate) mod dispatch;
pub use dispatch::{schema_disclosure_denial, GatewayDispatch, DEFAULT_DEADLINE};
+8 -1
View File
@@ -17,6 +17,10 @@
//! going forward.
//! - **Registry** ([`registry`]): operation specs, context, dispatch, and
//! the operation registry — the call half's dispatch core.
//! - **Gateway** (module `gateway`, feature `gateway`): the transport-neutral
//! dispatch spine and `services/schema` disclosure guard — the
//! deadline-bounded, re-rooted-context invoke surface for HTTP
//! gateways, hub relays, and other transport front-ends (ADR-048).
//! - **Protocol** ([`protocol`]): wire format, streams, adapter, dispatch
//! loop, pending requests, abort cascade — the call half's wire layer.
//! - **Client** ([`client`]): `CallClient`, `from_call`, `OperationAdapter`
@@ -24,7 +28,8 @@
//! - **Channels** ([`channels`]): the channels protocol — 8-byte chunk
//! wire format, demux/mux, `ChannelManager`, `ChannelsAdapter`,
//! `ChannelBidiStreamSource`, `ChannelOperations`,
//! `ChannelLifecyclePolicy`, `ChannelClient`. Channel 0 is
//! `ChannelLifecyclePolicy`, `ChannelClient`, `pump_bidi` (the
//! two-pump helper, ADR-050). Channel 0 is
//! pre-negotiated as `alk/call`; channels 1..N are opened via
//! per-ALPN open ops (`channels/<alpn>/sub`, `channels/<alpn>/pub`)
//! on channel 0 (ADR-047).
@@ -71,6 +76,8 @@
pub mod channels;
pub mod client;
pub mod core;
#[cfg(feature = "gateway")]
pub mod gateway;
pub mod protocol;
pub mod registry;
+2 -2
View File
@@ -253,7 +253,7 @@ mod tests {
acl: AccessControl,
handler: crate::registry::registration::Handler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
@@ -404,7 +404,7 @@ mod tests {
#[tokio::test]
async fn build_root_context_carries_capabilities_and_scoped_env() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let scoped = ScopedPeerEnv::new(["fs/readFile"]);
let caps = Capabilities::new().with_api_key("google", "k".to_string());
registry
+409 -26
View File
@@ -164,6 +164,18 @@ impl CallConnection {
self.imported_operations.write().insert(name, registration);
}
/// `true` when the connection overlay holds a registration for
/// `name` — the collision check `op/register`'s replace semantics
/// need (review 004 F-05).
pub fn overlay_contains(&self, name: &str) -> bool {
self.imported_operations.read().contains_key(name)
}
/// The connection overlay's registration for `name`, if present.
pub fn overlay_registration(&self, name: &str) -> Option<HandlerRegistration> {
self.imported_operations.read().get(name).cloned()
}
pub fn register_imported_all(&self, registrations: Vec<HandlerRegistration>) {
let mut overlay = self.imported_operations.write();
for reg in registrations {
@@ -213,7 +225,7 @@ impl CallConnection {
}
};
// `open_bi` returns a `BiStream` (ADR-092); split it into halves for
// `open_bi` returns a `BiStream` (ADR-009); split it into halves for
// the call protocol's separate write (request) and read (response)
// pumps. The split is the stdlib idiom; no per-handler wrapper.
let stream = match connection.open_bi().await {
@@ -235,7 +247,7 @@ impl CallConnection {
};
if let Err(err) = self.write_request(send, &request_id, payload).await {
let call_error = CallError::internal(err);
let call_error = CallError::connection_closed(err);
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -276,7 +288,8 @@ impl CallConnection {
let envelope = EventEnvelope::requested(&request_id, payload);
if let Err(err) = writer.write_frame(&envelope).await {
let call_error = CallError::internal(format!("failed to write request frame: {err}"));
let call_error =
CallError::connection_closed(format!("failed to write request frame: {err}"));
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -336,7 +349,7 @@ impl CallConnection {
}
};
// `open_bi` returns a `BiStream` (ADR-092); split for the separate
// `open_bi` returns a `BiStream` (ADR-009); split for the separate
// write (request) and read (subscription events) pumps.
let stream = match connection.open_bi().await {
Ok(s) => s,
@@ -353,7 +366,7 @@ impl CallConnection {
};
if let Err(err) = self.write_request(send, &request_id, payload).await {
let call_error = CallError::internal(err);
let call_error = CallError::connection_closed(err);
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -385,7 +398,8 @@ impl CallConnection {
let envelope = EventEnvelope::requested(&request_id, payload);
if let Err(err) = writer.write_frame(&envelope).await {
let call_error = CallError::internal(format!("failed to write request frame: {err}"));
let call_error =
CallError::connection_closed(format!("failed to write request frame: {err}"));
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -498,8 +512,17 @@ impl CallConnection {
});
let write_result = pump_publish_to_wire(send, &request_id, payload, stream).await;
if let Err(err) = write_result {
let call_error = CallError::internal(err);
if let Err((stage, err)) = write_result {
// A write failure before the request frame could be delivered is
// provably undelivered — retry is safe. Mid-publish failures leave
// delivery ambiguous; the INTERNAL mapping (non-retryable) is the
// safe default there.
let call_error = match stage {
PublishWriteStage::Request => CallError::connection_closed(err),
PublishWriteStage::Published | PublishWriteStage::Completed => {
CallError::internal(err)
}
};
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -533,7 +556,8 @@ impl CallConnection {
let requested = EventEnvelope::requested(&request_id, payload);
if let Err(err) = writer.write_frame(&requested).await {
let call_error = CallError::internal(format!("failed to write request frame: {err}"));
let call_error =
CallError::connection_closed(format!("failed to write request frame: {err}"));
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -588,7 +612,7 @@ impl CallConnection {
.connection
.as_ref()
.ok_or_else(|| "no underlying connection (overlay-only)".to_string())?;
// `open_bi` returns a `BiStream` (ADR-092). We only need the write
// `open_bi` returns a `BiStream` (ADR-009). We only need the write
// half to send the envelope; split and drop the read half.
let stream = connection
.open_bi()
@@ -603,33 +627,48 @@ impl CallConnection {
}
}
/// Which frame of a publish pump failed to write. The `Request` stage is
/// provably undelivered (the producer never received `call.requested`, so a
/// reconnect/retry is safe); the later stages leave delivery ambiguous.
enum PublishWriteStage {
Request,
Published,
Completed,
}
async fn pump_publish_to_wire<W>(
send: W,
request_id: &str,
payload: Value,
mut stream: Pin<Box<dyn Stream<Item = Value> + Send>>,
) -> Result<(), String>
) -> Result<(), (PublishWriteStage, String)>
where
W: AsyncWrite + Unpin,
{
let mut writer = FrameFramedWriter::new(send);
let requested = EventEnvelope::requested(request_id, payload);
writer
.write_frame(&requested)
.await
.map_err(|e| format!("failed to write request frame: {e}"))?;
writer.write_frame(&requested).await.map_err(|e| {
(
PublishWriteStage::Request,
format!("failed to write request frame: {e}"),
)
})?;
while let Some(chunk) = stream.next().await {
let envelope = EventEnvelope::published(request_id, chunk);
writer
.write_frame(&envelope)
.await
.map_err(|e| format!("failed to write published frame: {e}"))?;
writer.write_frame(&envelope).await.map_err(|e| {
(
PublishWriteStage::Published,
format!("failed to write published frame: {e}"),
)
})?;
}
let completed = EventEnvelope::completed(request_id);
writer
.write_frame(&completed)
.await
.map_err(|e| format!("failed to write completed frame: {e}"))?;
writer.write_frame(&completed).await.map_err(|e| {
(
PublishWriteStage::Completed,
format!("failed to write completed frame: {e}"),
)
})?;
Ok(())
}
@@ -799,6 +838,10 @@ impl OperationEnv for OverlayOperationEnv {
fn contains(&self, name: &str) -> bool {
self.overlay.read().contains_key(name)
}
fn list_operation_names(&self) -> Vec<String> {
self.overlay.read().keys().cloned().collect()
}
}
pub struct SubscriptionStream {
@@ -1476,7 +1519,7 @@ mod tests {
)
});
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
pub_spec_e2e("fs/upload"),
@@ -1880,6 +1923,346 @@ mod tests {
assert_eq!(read, envelope);
}
// CF-001 — a transport write failure before the request frame could be
// delivered maps to retryable CONNECTION_CLOSED (the call provably
// never reached the producer). Mid-publish failures stay INTERNAL.
/// AsyncWrite that always errors — simulates a dead mux mid-call.
struct BrokenWrite;
impl AsyncWrite for BrokenWrite {
fn poll_write(
self: Pin<&mut Self>,
_cx: &mut Context<'_>,
_buf: &[u8],
) -> Poll<std::io::Result<usize>> {
Poll::Ready(Err(std::io::Error::new(
std::io::ErrorKind::BrokenPipe,
"mux dead",
)))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
}
fn broken_write_bi_stream() -> crate::core::types::BiStream {
// The joined stream holds the read end; the write half always errors.
let (read_end, _write_end) = tokio::io::duplex(1024);
crate::core::types::BiStream::from_joined(read_end, BrokenWrite)
}
/// AsyncWrite whose writes succeed but whose Nth `poll_flush` fails.
/// `FrameFramedWriter::write_frame` flushes exactly once per frame, so
/// `remaining_ok_flushes = k` lets frames 1..=k land and fails inside
/// frame k+1 — a deterministic mid-stream write failure.
struct FailOnFlushN {
remaining_ok_flushes: std::sync::atomic::AtomicUsize,
}
impl AsyncWrite for FailOnFlushN {
fn poll_write(
self: Pin<&mut Self>,
_cx: &mut Context<'_>,
buf: &[u8],
) -> Poll<std::io::Result<usize>> {
Poll::Ready(Ok(buf.len()))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
let prev = self
.remaining_ok_flushes
.fetch_sub(1, std::sync::atomic::Ordering::SeqCst);
if prev == 0 {
Poll::Ready(Err(std::io::Error::new(
std::io::ErrorKind::BrokenPipe,
"mux dead",
)))
} else {
Poll::Ready(Ok(()))
}
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
}
#[tokio::test]
async fn single_stream_call_write_failure_is_retryable_connection_closed() {
let bi_stream = broken_write_bi_stream();
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let response = conn.call("test/echo", serde_json::json!({})).await;
match response.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[tokio::test]
async fn multi_stream_call_write_request_failure_is_retryable() {
// A `Connection` whose open_bi succeeds but whose write half is
// dead: the request write fails, the call provably undelivered.
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(broken_write_bi_stream()))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let response = call_conn.call("test/echo", serde_json::json!({})).await;
match response.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[tokio::test]
async fn multi_stream_publish_request_frame_write_failure_is_retryable() {
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(broken_write_bi_stream()))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let response = call_conn
.publish(
"fs/upload",
serde_json::json!({}),
Box::pin(futures::stream::iter(vec![serde_json::json!({"c": 1})])),
)
.await;
match response.result {
Err(e) => {
// The pump died on the *request* frame — nothing delivered.
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[test]
fn connection_closed_error_is_retryable() {
let err = CallError::connection_closed("test");
assert_eq!(err.code, "CONNECTION_CLOSED");
assert!(err.retryable);
}
#[tokio::test]
async fn single_stream_subscribe_write_failure_is_retryable_connection_closed() {
use futures::stream::StreamExt;
let bi_stream = broken_write_bi_stream();
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let mut stream = conn.subscribe("test/events", serde_json::json!({})).await;
let item = stream.next().await.unwrap();
match item.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
assert!(stream.next().await.is_none());
}
#[tokio::test]
async fn multi_stream_subscribe_write_failure_is_retryable_connection_closed() {
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use futures::stream::StreamExt;
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(broken_write_bi_stream()))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let mut stream = call_conn
.subscribe("test/events", serde_json::json!({}))
.await;
let item = stream.next().await.unwrap();
match item.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
assert!(stream.next().await.is_none());
}
#[tokio::test]
async fn single_stream_publish_request_frame_write_failure_is_retryable() {
let bi_stream = broken_write_bi_stream();
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let response = conn
.publish(
"fs/upload",
serde_json::json!({}),
Box::pin(futures::stream::iter(vec![serde_json::json!({"c": 1})])),
)
.await;
match response.result {
Err(e) => {
// The request frame never landed — nothing delivered, retry safe.
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[tokio::test]
async fn single_stream_publish_midpublish_write_failure_stays_internal() {
// Flush succeeds for the request frame (1) + first chunk (2), then
// fails inside the second chunk — delivery is ambiguous from there.
let bi_stream = crate::core::types::BiStream::from_joined(
tokio::io::empty(),
FailOnFlushN {
remaining_ok_flushes: std::sync::atomic::AtomicUsize::new(2),
},
);
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let chunks = vec![serde_json::json!({"c": 1}), serde_json::json!({"c": 2})];
let stream: Pin<Box<dyn Stream<Item = Value> + Send>> =
Box::pin(futures::stream::iter(chunks));
let response = conn
.publish("fs/upload", serde_json::json!({}), stream)
.await;
match response.result {
Err(e) => {
assert_eq!(e.code, "INTERNAL");
assert!(!e.retryable, "mid-publish failure must not be retryable");
}
other => panic!("expected non-retryable INTERNAL, got {other:?}"),
}
}
#[tokio::test]
async fn multi_stream_publish_midpublish_write_failure_stays_internal() {
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(
crate::core::types::BiStream::from_joined(
tokio::io::empty(),
FailOnFlushN {
remaining_ok_flushes: std::sync::atomic::AtomicUsize::new(2),
},
),
))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let chunks = vec![serde_json::json!({"c": 1}), serde_json::json!({"c": 2})];
let stream: Pin<Box<dyn Stream<Item = Value> + Send>> =
Box::pin(futures::stream::iter(chunks));
let response = call_conn
.publish("fs/upload", serde_json::json!({}), stream)
.await;
match response.result {
Err(e) => {
// The pump died on a *chunk* frame — the request landed,
// delivery is ambiguous, INTERNAL (non-retryable) is correct.
assert_eq!(e.code, "INTERNAL");
assert!(!e.retryable);
}
other => panic!("expected non-retryable INTERNAL, got {other:?}"),
}
}
#[tokio::test]
async fn shared_frame_writer_concurrent_writes_do_not_interleave() {
let (client_end, server_end) = tokio::io::duplex(64 * 1024);
@@ -1993,7 +2376,7 @@ mod tests {
}
}
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("test/echo"),
@@ -2078,7 +2461,7 @@ mod tests {
}
}
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
+546 -130
View File
@@ -32,7 +32,7 @@ use super::abort::AbortCascade;
use super::connection::CallConnection;
use super::wire::{
CallError, EventEnvelope, FrameFramedReader, FrameFramedWriter, ResponseEnvelope,
EVENT_ABORTED, EVENT_COMPLETED, EVENT_ERROR, EVENT_PUBLISHED, EVENT_REQUESTED,
EVENT_ABORTED, EVENT_COMPLETED, EVENT_ERROR, EVENT_PUBLISHED, EVENT_REQUESTED, EVENT_RESPONDED,
};
use crate::protocol::adapter::SessionOverlaySource;
use crate::registry::context::{AbortPolicy, OperationContext, ScopedPeerEnv};
@@ -114,6 +114,18 @@ impl std::fmt::Debug for DispatchResult {
}
}
/// The started half of a dispatched `call.requested` event, from
/// `Dispatcher::dispatch_start`. `Once`/`Stream` are returned ready
/// for the caller to spawn (review 005 G-01); `Sink` was started
/// inline (its `chunk_tx` must be registered before the next
/// `call.published` frame is read) and carries the started handler
/// future.
pub enum StartedDispatch {
Once(Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>),
Stream(ResponseStream),
Sink(SinkDispatch),
}
/// Shared dispatcher for an established `CallConnection`. Constructed by
/// both `CallAdapter` (accept path) and `CallClient` (connect path) and used
/// to run the dispatch loop. Holds no per-connection state; the
@@ -299,6 +311,39 @@ impl Dispatcher {
request_id: String,
payload: Value,
) -> DispatchResult {
match self.dispatch_start(connection, request_id, payload) {
StartedDispatch::Once(invoke) => DispatchResult::Once(invoke.await),
StartedDispatch::Stream(stream) => DispatchResult::Stream(stream),
StartedDispatch::Sink(sink) => DispatchResult::Sink(sink),
}
}
/// Start dispatching a `call.requested` event without awaiting the
/// invocation (review 005 G-01). The synchronous prefix of
/// [`dispatch`](Self::dispatch) — identity resolution, root-context
/// construction, op-type branch — is split from the invocation:
///
/// - `Query`/`Mutation` return the invocation as a boxed future
/// (`StartedDispatch::Once`) for the caller to spawn. Awaiting it
/// inline is what deadlocked the serving loops: a handler that
/// nested-composes over the same connection blocks the one read
/// loop that must resolve the nested call's response frame.
/// - `Sub` returns the [`ResponseStream`] synchronously
/// (`invoke_streaming` is a sync call) for the caller to pump in
/// a spawned task — an inline pump starves the same loop for the
/// stream's whole lifetime.
/// - `Pub` runs the full sink start inline (`StartedDispatch::Sink`):
/// ACL failure resolves to a ready `Once` future (the response
/// envelope is already built), success carries the `chunk_tx` the
/// loop must register in `in_flight_sinks` *before* the next
/// `call.published` frame can be read — the insert cannot race
/// the feed, so the sink start stays synchronous by design.
pub(crate) fn dispatch_start(
&self,
connection: &Arc<CallConnection>,
request_id: String,
payload: Value,
) -> StartedDispatch {
let operation_id = payload
.get("operationId")
.and_then(|v| v.as_str())
@@ -320,7 +365,7 @@ impl Dispatcher {
.unwrap_or(OperationType::Query);
let mut context = self.build_root_context(
request_id.clone(),
request_id,
&operation_name,
identity,
forwarded_for,
@@ -329,54 +374,125 @@ impl Dispatcher {
match op_type {
OperationType::Query | OperationType::Mutation => {
let envelope = self.registry.invoke(&operation_name, input, context).await;
DispatchResult::Once(envelope)
let registry = Arc::clone(&self.registry);
let name = operation_name;
StartedDispatch::Once(Box::pin(async move {
registry.invoke(&name, input, context).await
}))
}
OperationType::Sub => {
context.deadline = None;
let stream = self
.registry
.invoke_streaming(&operation_name, input, context);
DispatchResult::Stream(stream)
StartedDispatch::Stream(stream)
}
OperationType::Pub => {
context.deadline = None;
let sink_handler =
match self
.registry
.resolve_sink_handler(&operation_name, &input, &context)
{
Ok(h) => h,
Err(envelope) => return DispatchResult::Once(envelope),
};
let publish_validator = self
match self
.registry
.registration(&operation_name)
.and_then(|r| r.spec.publish_schema.as_ref())
.and_then(|schema| match jsonschema::options().build(schema) {
Ok(v) => Some(v),
Err(e) => {
warn!(
operation = %operation_name,
error = %e,
"publish_schema failed to compile; chunks will not be validated",
);
None
}
});
let (chunk_tx, chunk_rx) =
mpsc::channel::<Result<Value, CallError>>(PUBLISH_CHANNEL_BUFFER);
let publish_stream: PublishStream = Box::pin(chunk_rx);
let handler = (sink_handler)(input, context, publish_stream);
DispatchResult::Sink(SinkDispatch {
handler: Box::pin(handler),
chunk_tx,
publish_validator,
})
.resolve_sink_handler(&operation_name, &input, &context)
{
Ok(sink_handler) => {
let publish_validator = self.registry.publish_validator(&operation_name);
let (chunk_tx, chunk_rx) =
mpsc::channel::<Result<Value, CallError>>(PUBLISH_CHANNEL_BUFFER);
let publish_stream: PublishStream = Box::pin(chunk_rx);
let handler = (sink_handler)(input, context, publish_stream);
StartedDispatch::Sink(SinkDispatch {
handler: Box::pin(handler),
chunk_tx,
publish_validator,
})
}
Err(envelope) => StartedDispatch::Once(Box::pin(std::future::ready(envelope))),
}
}
}
}
/// Spawn the Once invocation to write its single response frame on
/// completion. The write failure is a warning, not a loop exit: the
/// loop keeps reading (a dying transport surfaces on the next read
/// as `ConnectionClosed`), matching the Sink arm's behavior.
fn spawn_once_dispatch(
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: String,
invoke: Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>,
) -> JoinHandle<()> {
let writer = Arc::clone(writer);
tokio::spawn(async move {
let response = invoke.await;
let event: EventEnvelope = response.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write Once response frame"
);
}
})
}
/// Spawn a subscription's [`ResponseStream`] pump: each
/// [`ResponseEnvelope`] becomes a `call.responded` / `call.error`
/// frame; on natural end, a `call.completed` frame. The shared
/// writer serializes frames so the pump's frames do not interleave
/// with concurrent calls' frames.
fn spawn_stream_pump(
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: String,
mut stream: ResponseStream,
) -> JoinHandle<()> {
let writer = Arc::clone(writer);
tokio::spawn(async move {
let mut last_was_error = false;
while let Some(envelope) = stream.next().await {
last_was_error = envelope.result.is_err();
let event: EventEnvelope = envelope.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write streaming frame"
);
return;
}
}
if !last_was_error {
let completed = EventEnvelope::completed(&request_id);
if let Err(err) = writer.write_frame(&completed).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write call.completed"
);
}
}
})
}
/// Spawn a Pub's handler task: awaits the handler future and writes
/// the single `call.responded` / `call.error` frame.
fn spawn_sink_response_writer(
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: String,
handler: Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>,
) -> JoinHandle<()> {
let writer = Arc::clone(writer);
tokio::spawn(async move {
let response = handler.await;
let event: EventEnvelope = response.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write sink response frame"
);
}
})
}
pub async fn handle_abort(&self, connection: &Arc<CallConnection>, request_id: &str) {
if let Some(tx) = self.in_flight_sink_aborts.lock().remove(request_id) {
let _ = tx.send(());
@@ -396,7 +512,7 @@ impl Dispatcher {
connection: Arc<CallConnection>,
stream: crate::core::types::BiStream,
) {
// `stream` is a `BiStream` (ADR-092) — `AsyncRead + AsyncWrite + Send
// `stream` is a `BiStream` (ADR-009) — `AsyncRead + AsyncWrite + Send
// + Unpin`. Split into the read and write halves the call protocol's
// frame reader/writer consume. The split is the stdlib idiom; no
// per-handler wrapper.
@@ -455,7 +571,7 @@ impl Dispatcher {
/// for `Ok`, `call.error` for `Err`). On natural stream end (the stream
/// returned `None` without the last item being an `Err`), write a
/// `call.completed` frame. An `Err` envelope is terminal — the stream
/// ends after it and we do NOT write `call.completed` (ADR-049 §6).
/// ends after it and we do NOT write `call.completed` (ADR-021 §6).
///
/// If a frame write fails the pump stops early; the stream is dropped on
/// return, releasing the handler's resources via `Drop` (ADR-016). A
@@ -730,14 +846,24 @@ impl Dispatcher {
/// `call.published` / `call.completed` / `call.aborted` frames
/// arrive *after* the `call.requested` and are routed to the
/// matching sink's `chunk_tx` while new `call.requested` frames for
/// other requests continue to be dispatched. Query/Mutation and Sub
/// responses are written through `writer` immediately after
/// dispatch.
/// other requests continue to be dispatched. Once invocations and
/// Sub pumps are spawned (review 005 G-01 — inline dispatch
/// deadlocked same-connection nested composition); only the sink
/// start runs inline.
///
/// This is the accept side's serving loop, so the loop also carries
/// the pending-resolution arms the connect side's
/// `serve_single_stream` composes: a nested-composing handler
/// (e.g. a `from_call` imported-op stub riding this connection)
/// blocks its own task, not the read loop, and the loop resolves
/// the stub's response frames here. Without these arms the accept
/// side could dispatch inbound requests but never resolve its own
/// outbound pendings in single-stream mode.
///
/// Returns when the read half closes (transport EOF). Outstanding
/// pending requests are failed with `connection closed`, and
/// in-flight sinks' `chunk_tx` are dropped (the handler's
/// `PublishStream` sees EOF).
/// pending requests are failed with `connection closed`, in-flight
/// sinks' `chunk_tx` are dropped (the handler's `PublishStream`
/// sees EOF), and spawned Once/Stream/Sink tasks are aborted.
pub async fn run_loop_single_stream(
self,
connection: Arc<CallConnection>,
@@ -763,7 +889,10 @@ impl Dispatcher {
});
let mut reader = FrameFramedReader::new(reader);
let mut in_flight_sinks: HashMap<String, InFlightSink> = HashMap::new();
let in_flight_sinks: Arc<ParkingLotMutex<HashMap<String, InFlightSink>>> =
Arc::new(ParkingLotMutex::new(HashMap::new()));
let spawned: Arc<ParkingLotMutex<Vec<JoinHandle<()>>>> =
Arc::new(ParkingLotMutex::new(Vec::new()));
loop {
let envelope = match reader.read_frame().await {
@@ -778,46 +907,29 @@ impl Dispatcher {
match envelope.r#type.as_str() {
EVENT_REQUESTED => {
let request_id = envelope.id.clone();
let payload = envelope.payload.clone();
let dispatch_result = self
.dispatch(&connection, request_id.clone(), payload)
.await;
match dispatch_result {
DispatchResult::Once(response) => {
let event: EventEnvelope = response.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
"single-stream: failed to write Once response; closing loop"
);
break;
}
match self.dispatch_start(&connection, request_id.clone(), envelope.payload) {
StartedDispatch::Once(invoke) => {
let handle =
Self::spawn_once_dispatch(&writer, request_id.clone(), invoke);
spawned.lock().push(handle);
}
DispatchResult::Stream(stream) => {
self.pump_stream_single_stream(&writer, &request_id, stream)
.await;
StartedDispatch::Stream(stream) => {
let handle =
Self::spawn_stream_pump(&writer, request_id.clone(), stream);
spawned.lock().push(handle);
}
DispatchResult::Sink(sink) => {
StartedDispatch::Sink(sink) => {
let SinkDispatch {
handler,
chunk_tx,
publish_validator,
} = sink;
let writer_clone = Arc::clone(&writer);
let request_id_for_handler = request_id.clone();
let handle = tokio::spawn(async move {
let response = handler.await;
let event: EventEnvelope = response.into();
if let Err(err) = writer_clone.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id_for_handler,
"single-stream: failed to write sink response frame"
);
}
});
in_flight_sinks.insert(
let handle = Self::spawn_sink_response_writer(
&writer,
request_id.clone(),
handler,
);
in_flight_sinks.lock().insert(
request_id.clone(),
InFlightSink {
chunk_tx,
@@ -828,9 +940,25 @@ impl Dispatcher {
}
}
}
EVENT_RESPONDED => {
let request_id = envelope.id.clone();
let output = envelope
.payload
.get("output")
.cloned()
.unwrap_or(Value::Null);
pending.lock().handle_responded(&request_id, output);
}
EVENT_COMPLETED => {
let request_id = envelope.id.clone();
if in_flight_sinks.lock().remove(&request_id).is_none() {
pending.lock().handle_completed(&request_id);
}
}
EVENT_ABORTED => {
let request_id = envelope.id.clone();
if let Some(mut entry) = in_flight_sinks.remove(&request_id) {
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
entry.handler_handle.abort();
let _ = entry
.chunk_tx
@@ -847,7 +975,8 @@ impl Dispatcher {
.get("input")
.cloned()
.unwrap_or(Value::Null);
if let Some(mut entry) = in_flight_sinks.remove(&request_id) {
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let validated = match &entry.publish_validator {
Some(validator) => {
if validator.is_valid(&chunk) {
@@ -865,7 +994,7 @@ impl Dispatcher {
let keep = validated.is_ok();
let _ = entry.chunk_tx.send(validated).await;
if keep {
in_flight_sinks.insert(request_id, entry);
in_flight_sinks.lock().insert(request_id, entry);
}
} else {
debug!(
@@ -874,36 +1003,33 @@ impl Dispatcher {
);
}
}
EVENT_COMPLETED => {
let request_id = envelope.id.clone();
in_flight_sinks.remove(&request_id);
}
EVENT_ERROR => {
let request_id = envelope.id.clone();
if let Some(mut entry) = in_flight_sinks.remove(&request_id) {
let call_error: CallError = serde_json::from_value(envelope.payload)
.unwrap_or_else(|_| {
CallError::internal("publish error from initiator (malformed)")
});
let call_error: CallError = serde_json::from_value(envelope.payload)
.unwrap_or_else(|_| {
CallError::internal("publish error from initiator (malformed)")
});
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let _ = entry.chunk_tx.send(Err(call_error)).await;
} else {
debug!(
request_id = %request_id,
"single-stream: call.error for unknown in-flight sink; dropping"
);
pending.lock().handle_error(&request_id, call_error);
}
}
other => {
debug!(
event_type = %other,
id = %envelope.id,
"single-stream: ignoring non-requested/non-published/non-aborted/non-completed event"
"single-stream: ignoring unknown event type"
);
}
}
}
in_flight_sinks.clear();
in_flight_sinks.lock().clear();
for handle in spawned.lock().drain(..) {
handle.abort();
}
let failed = pending
.lock()
@@ -918,33 +1044,227 @@ impl Dispatcher {
sweeper_handle.abort();
}
/// Pump a subscription's `ResponseStream` to the wire through the
/// shared single-stream writer (ADR-036 amendment). Each
/// `ResponseEnvelope` becomes a `call.responded` / `call.error`
/// frame; on natural end, a `call.completed` frame. The shared
/// writer serializes frames so this pump's frames do not interleave
/// with concurrent calls' frames.
async fn pump_stream_single_stream(
&self,
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: &str,
mut stream: ResponseStream,
/// The full-duplex single-stream serving loop (review 004 F-04):
/// composes the dispatch arms (`run_loop_single_stream`) and the
/// pending-resolution arms (`read_single_stream_until_closed`) in
/// one loop. Both sides of a channels connection can be both
/// producer and consumer (AGENTS.md §8; ADR-022 §2) — on channel 0
/// the two directions' frames are multiplexed on one byte stream,
/// so the loop branches per frame:
///
/// - `call.requested` → start the dispatch and spawn the work (the
/// serving half — what the accept side's `run_loop_single_stream`
/// does). Once invocations and Sub pumps run as spawned tasks;
/// only the Pub sink start runs inline (its `chunk_tx` must be
/// registered before the next `call.published` frame is read).
/// Inline dispatch was the review 005 G-01 defect: a handler that
/// nested-composes over the same connection blocked the one read
/// loop that must resolve the nested call's response frame. The
/// `SharedFrameWriter` serializes frames, so spawned arms' frames
/// do not interleave mid-frame.
/// - `call.responded` / `call.completed` / `call.error` → resolve
/// the matching outbound pending entry (the calling half — what
/// the client read pump does). `call.responded` frames for
/// *inbound* Sub responses are `call.responded` too; direction is
/// disambiguated by the pending map (an id that is one of *our*
/// outbound pendings resolves there; an unknown id is dropped
/// with a debug line, never an error — the peer may legitimately
/// stream a Sub's responses that this side does not track).
/// - `call.aborted` → try both tables: the in-flight sink aborts
/// *and* the outbound pending map's abort-cascade
/// (`handle_abort`), then fall through to serving-side sink
/// cancellation (the `run_loop_single_stream` arm) when the id
/// matches an inbound sink. An id in neither table is a no-op.
/// - `call.published` / `call.completed` (initiator→responder) →
/// route to the matching inbound in-flight sink; ids that are
/// outbound pending *call* entries whose `call.responded` already
/// removed them are no-ops.
///
/// IDs are UUID-generated per side (`generate_request_id()`), so
/// cross-correlation between the two directions is not a hazard.
///
/// On read-half close, outbound pendings are failed (`connection
/// closed`), in-flight inbound sinks are dropped, and the spawned
/// Once/Stream/Sink tasks are aborted (their response frames cannot
/// reach a closed transport anyway).
pub async fn serve_single_stream(
self,
connection: Arc<CallConnection>,
reader: Box<dyn tokio::io::AsyncRead + Send + Unpin>,
writer: Arc<super::connection::SharedFrameWriter>,
) {
let mut last_was_error = false;
while let Some(envelope) = stream.next().await {
last_was_error = envelope.result.is_err();
let event: EventEnvelope = envelope.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(error = %err, "single-stream: failed to write streaming frame");
return;
let pending = Arc::clone(connection.pending());
let sweeper_pending = Arc::clone(&pending);
let sweeper_handle: JoinHandle<()> = tokio::spawn(async move {
let mut interval = tokio::time::interval(SWEEPER_INTERVAL);
interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop {
interval.tick().await;
let evicted = sweeper_pending.lock().evict_expired();
if !evicted.is_empty() {
debug!(
count = evicted.len(),
"serve loop: sweeper evicted expired pending entries"
);
}
}
});
let mut reader = FrameFramedReader::new(reader);
let in_flight_sinks: Arc<ParkingLotMutex<HashMap<String, InFlightSink>>> =
Arc::new(ParkingLotMutex::new(HashMap::new()));
let spawned: Arc<ParkingLotMutex<Vec<JoinHandle<()>>>> =
Arc::new(ParkingLotMutex::new(Vec::new()));
loop {
let envelope = match reader.read_frame().await {
Ok(env) => env,
Err(super::wire::FrameError::ConnectionClosed) => break,
Err(err) => {
warn!(error = %err, "serve loop: frame read error; closing loop");
break;
}
};
match envelope.r#type.as_str() {
EVENT_REQUESTED => {
let request_id = envelope.id.clone();
match self.dispatch_start(&connection, request_id.clone(), envelope.payload) {
StartedDispatch::Once(invoke) => {
let handle =
Self::spawn_once_dispatch(&writer, request_id.clone(), invoke);
spawned.lock().push(handle);
}
StartedDispatch::Stream(stream) => {
let handle =
Self::spawn_stream_pump(&writer, request_id.clone(), stream);
spawned.lock().push(handle);
}
StartedDispatch::Sink(sink) => {
let SinkDispatch {
handler,
chunk_tx,
publish_validator,
} = sink;
let handle = Self::spawn_sink_response_writer(
&writer,
request_id.clone(),
handler,
);
in_flight_sinks.lock().insert(
request_id.clone(),
InFlightSink {
chunk_tx,
publish_validator,
handler_handle: handle,
},
);
}
}
}
EVENT_RESPONDED => {
let request_id = envelope.id.clone();
let output = envelope
.payload
.get("output")
.cloned()
.unwrap_or(Value::Null);
pending.lock().handle_responded(&request_id, output);
}
EVENT_COMPLETED => {
let request_id = envelope.id.clone();
if in_flight_sinks.lock().remove(&request_id).is_none() {
pending.lock().handle_completed(&request_id);
}
}
EVENT_ABORTED => {
let request_id = envelope.id.clone();
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
entry.handler_handle.abort();
let _ = entry
.chunk_tx
.send(Err(CallError::internal("publish aborted by initiator")))
.await;
} else {
self.handle_abort(&connection, &request_id).await;
}
}
EVENT_PUBLISHED => {
let request_id = envelope.id.clone();
let chunk = envelope
.payload
.get("input")
.cloned()
.unwrap_or(Value::Null);
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let validated = match &entry.publish_validator {
Some(validator) => {
if validator.is_valid(&chunk) {
Ok(chunk)
} else {
let details = serde_json::json!({ "chunk": chunk });
Err(CallError::invalid_input(
"published chunk failed publish_schema validation",
)
.with_details(details))
}
}
None => Ok(chunk),
};
let keep = validated.is_ok();
let _ = entry.chunk_tx.send(validated).await;
if keep {
in_flight_sinks.lock().insert(request_id, entry);
}
} else {
debug!(
request_id = %request_id,
"serve loop: call.published for unknown in-flight sink; dropping"
);
}
}
EVENT_ERROR => {
let request_id = envelope.id.clone();
let call_error: CallError = serde_json::from_value(envelope.payload)
.unwrap_or_else(|_| {
CallError::internal("publish error from initiator (malformed)")
});
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let _ = entry.chunk_tx.send(Err(call_error)).await;
} else {
pending.lock().handle_error(&request_id, call_error);
}
}
other => {
debug!(
event_type = %other,
id = %envelope.id,
"serve loop: ignoring unknown event type"
);
}
}
}
if !last_was_error {
let completed = EventEnvelope::completed(request_id);
if let Err(err) = writer.write_frame(&completed).await {
warn!(error = %err, "single-stream: failed to write call.completed");
}
in_flight_sinks.lock().clear();
for handle in spawned.lock().drain(..) {
handle.abort();
}
let failed = pending
.lock()
.fail_all(CallError::internal("connection closed"));
if !failed.is_empty() {
debug!(
count = failed.len(),
"serve loop: failed pending requests on connection close"
);
}
sweeper_handle.abort();
}
}
@@ -1028,7 +1348,7 @@ mod tests {
}
fn registry_with(name: &str, visibility: Visibility, acl: AccessControl) -> OperationRegistry {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
@@ -1063,7 +1383,7 @@ mod tests {
#[tokio::test]
async fn dispatch_authorized_peer_dispatches_and_populates_capabilities() {
let caps = Capabilities::new().with_api_key("google", "k".to_string());
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let has_google = context.capabilities.get("google").is_some();
ResponseEnvelope::ok(
@@ -1100,7 +1420,7 @@ mod tests {
#[tokio::test]
async fn dispatch_unauthorized_peer_returns_forbidden_capabilities_never_populated() {
let caps = Capabilities::new().with_api_key("google", "k".to_string());
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let has_google = context.capabilities.get("google").is_some();
ResponseEnvelope::ok(
@@ -1225,7 +1545,7 @@ mod tests {
#[tokio::test]
async fn dispatch_extract_forwarded_for_from_payload_into_context() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let forwarded_id = context.forwarded_for.as_ref().map(|i| i.id.clone());
ResponseEnvelope::ok(
@@ -1266,7 +1586,7 @@ mod tests {
#[tokio::test]
async fn dispatch_without_forwarded_for_field_is_none() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let present = context.forwarded_for.is_some();
ResponseEnvelope::ok(
@@ -1356,7 +1676,7 @@ mod tests {
#[tokio::test]
async fn dispatch_requested_overlay_only_attaches_peer_keyed_by_stored_identity() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let peer_ids = context.env.peer_ids();
ResponseEnvelope::ok(
@@ -1491,7 +1811,7 @@ mod tests {
);
}
// --- streaming dispatch branch (ADR-049 §6) ---------------------------
// --- streaming dispatch branch (ADR-021 §6) ---------------------------
fn subscription_spec(name: &str, acl: AccessControl) -> OperationSpec {
OperationSpec::new(
@@ -1546,7 +1866,7 @@ mod tests {
name: &str,
handler: crate::registry::registration::StreamingHandler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
subscription_spec(name, AccessControl::default()),
@@ -1625,7 +1945,7 @@ mod tests {
#[tokio::test]
async fn dispatch_query_keeps_deadline_some() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, ctx| async move {
let deadline_is_some = ctx.deadline.is_some();
ResponseEnvelope::ok(
@@ -1890,7 +2210,7 @@ mod tests {
name: &str,
handler: crate::registry::registration::SinkHandler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
pub_spec(name, AccessControl::default()),
@@ -2057,7 +2377,7 @@ mod tests {
publish_schema: Value,
handler: crate::registry::registration::SinkHandler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
pub_spec_with_publish_schema(name, publish_schema),
@@ -2593,4 +2913,100 @@ mod tests {
"in_flight_sink_aborts map is cleaned up after pump_sink exits"
);
}
// --- review 004 F-04: the full-duplex serving loop ---------------------
/// The F-04 probe: two `serve_single_stream` loops over one duplex
/// — the exact hub↔consumer shape. Each side holds a
/// `CallConnection` over its own writer; when side A calls
/// `echo/run`, its `call.requested` crosses the duplex into side
/// B's serving loop, which dispatches and writes `call.responded`
/// back, resolving A's pending through the `EVENT_RESPONDED` arm.
/// This is the direction `read_single_stream_until_closed` used to
/// drop.
#[tokio::test]
async fn serve_single_stream_self_call_resolves_pending() {
fn echo_registry() -> Arc<crate::registry::registration::OperationRegistry> {
let registry = crate::registry::registration::OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("echo/run", AccessControl::default()),
HandlerKind::Once(make_handler(|input, context| async move {
ResponseEnvelope::ok(context.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
Arc::new(registry)
}
let (client_end, server_end) = tokio::io::duplex(64 * 1024);
// Side B: the duplex end wrapped as one BiStream (read+write);
// `split_single_stream` divides it into the frame writer and
// reader the serving loop consumes.
let server_bidi = crate::core::types::BiStream::from_bidi(server_end);
let (server_writer, server_reader) =
crate::protocol::connection::split_single_stream(server_bidi);
let dp_server = Dispatcher::new(echo_registry(), Arc::new(StaticIdentityProvider::new()));
let server_call_conn = Arc::new(CallConnection::new_single_stream(
crate::protocol::sink_empty_connection(),
Arc::clone(&server_writer),
));
let server_conn_for_loop = Arc::clone(&server_call_conn);
let _serve_b = tokio::spawn(async move {
dp_server
.serve_single_stream(server_conn_for_loop, server_reader, server_writer)
.await;
});
// Side A: the same shape on the other duplex end.
let client_bidi = crate::core::types::BiStream::from_bidi(client_end);
let (client_writer, client_reader) =
crate::protocol::connection::split_single_stream(client_bidi);
let dp_client = Dispatcher::new(echo_registry(), Arc::new(StaticIdentityProvider::new()));
let client_call_conn = Arc::new(CallConnection::new_single_stream(
crate::protocol::sink_empty_connection(),
Arc::clone(&client_writer),
));
let client_conn_for_loop = Arc::clone(&client_call_conn);
let _serve_a = tokio::spawn(async move {
dp_client
.serve_single_stream(client_conn_for_loop, client_reader, client_writer)
.await;
});
// Side A calls `echo/run` — side B serves it.
let response = tokio::time::timeout(
std::time::Duration::from_secs(5),
client_call_conn.call("echo/run", serde_json::json!({ "from": "a" })),
)
.await
.expect("A→B call through serving loops timed out");
assert!(
response.result.is_ok(),
"A→B call resolves, got {:?}",
response.result
);
assert_eq!(response.result.unwrap(), serde_json::json!({ "from": "a" }));
// And B calls back — A serves it (the both-directions gate).
let response = tokio::time::timeout(
std::time::Duration::from_secs(5),
server_call_conn.call("echo/run", serde_json::json!({ "from": "b" })),
)
.await
.expect("B→A call through serving loops timed out");
assert!(
response.result.is_ok(),
"B→A call resolves, got {:?}",
response.result
);
assert_eq!(response.result.unwrap(), serde_json::json!({ "from": "b" }));
}
}
+2 -2
View File
@@ -1,6 +1,6 @@
//! Shared test helpers for the call protocol's inline `#[cfg(test)]`
//! modules. Kept here (not in each test module) so the `stub_connection()`
//! shape is defined once — `Connection::from_stream` was removed (ADR-092)
//! shape is defined once — `Connection::from_stream` was removed (ADR-009)
//! and every test stub that previously called it now calls
//! `Connection::from_bidi(SinkEmpty, ...)` via `sink_empty_connection()`.
@@ -15,7 +15,7 @@ use tokio::io::{AsyncRead, AsyncWrite, ReadBuf};
/// A test-only `AsyncRead + AsyncWrite` pair equivalent to
/// `tokio::io::sink() + tokio::io::empty()`: reads yield EOF immediately
/// (zero bytes), writes discard. Exists because `Connection::from_bidi`
/// (ADR-092 — the only public stream constructor, replacing
/// (ADR-009 — the only public stream constructor, replacing
/// `from_stream`) requires a single value that implements both traits.
/// Used only to construct a `Connection` for tests that exercise
/// `Connection`-level state (alpn, addr, identity, dispatcher run loop
+8
View File
@@ -113,9 +113,17 @@ impl CallError {
Self::new("TIMEOUT", message, true)
}
pub fn connection_closed(message: impl Into<String>) -> Self {
Self::new("CONNECTION_CLOSED", message, true)
}
pub fn invalid_operation_type(message: impl Into<String>) -> Self {
Self::new("INVALID_OPERATION_TYPE", message, false)
}
pub fn already_exists(message: impl Into<String>) -> Self {
Self::new("ALREADY_EXISTS", message, false)
}
}
impl Eq for CallError {}
+1 -1
View File
@@ -33,7 +33,7 @@ pub struct OperationContext {
pub internal: bool,
/// `None` when no ownership provider is wired (backward compat —
/// `check` falls back to static `Identity.resources` path). Wired by
/// the assembly layer via `CallAdapter`/`Dispatcher` (ADR-050).
/// the assembly layer via `CallAdapter`/`Dispatcher` (ADR-011).
pub ownership: Option<Arc<dyn OwnershipProvider>>,
}
+392 -15
View File
@@ -3,8 +3,11 @@ use std::sync::Arc;
use serde_json::{json, Value};
use super::context::OperationContext;
use super::registration::{Handler, OperationRegistry};
use super::spec::{AccessControl, OperationSpec, OperationType, Visibility};
use super::registration::{
Handler, HandlerKind, HandlerRegistration, OperationProvenance, OperationRegistry,
};
use super::spec::{AccessControl, AccessResult, OperationSpec, OperationType, Visibility};
use crate::core::types::Capabilities;
use crate::protocol::wire::{CallError, ResponseEnvelope};
const NAME_SERVICES_LIST: &str = "services/list";
@@ -30,6 +33,10 @@ pub fn services_list_spec() -> OperationSpec {
"op_type": {
"type": "string",
"enum": ["query", "mutation", "sub", "pub"]
},
"description": {
"type": "string",
"description": "Human-readable op description (review 006 E-02). Absent when the producer declares none."
}
}
}
@@ -84,6 +91,10 @@ pub fn services_list_peers_spec() -> OperationSpec {
"op_type": {
"type": "string",
"enum": ["query", "mutation", "sub", "pub"]
},
"description": {
"type": "string",
"description": "Human-readable op description. Absent when the op declares none."
}
}
}
@@ -113,6 +124,9 @@ fn operation_spec_schema() -> Value {
"type": "string",
"enum": ["external", "internal"]
},
"description": {
"description": "Human-readable op description (review 006 E-02). Absent when the op declares none; disclosed verbatim in `services/list` when set."
},
"input_schema": {},
"output_schema": {},
"error_schemas": {
@@ -198,6 +212,14 @@ fn error_definition_to_json(def: &super::spec::ErrorDefinition) -> Value {
}
pub(crate) fn spec_to_json(spec: &OperationSpec) -> Value {
spec_to_json_pub(spec)
}
/// Public serialization of an `OperationSpec` into the `services/schema`
/// wire shape — the shape `rebuild_spec_for` parses back. Used by the
/// `op/register` bootstrap op (review 004 F-05) to carry announced specs
/// over the wire; `services/schema` serves the same shape.
pub fn spec_to_json_pub(spec: &OperationSpec) -> Value {
let error_schemas: Vec<Value> = spec
.error_schemas
.iter()
@@ -213,6 +235,12 @@ pub(crate) fn spec_to_json(spec: &OperationSpec) -> Value {
"error_schemas": error_schemas,
"access_control": access_control_to_json(&spec.access_control),
});
if let Some(resource_id_path) = &spec.resource_id_path {
json["resource_id_path"] = json!(resource_id_path);
}
if let Some(description) = &spec.description {
json["description"] = json!(description);
}
if spec.channel_open.is_some() {
json["channel_open"] = json!(true);
}
@@ -245,11 +273,15 @@ pub fn services_list_handler(registry: Arc<OperationRegistry>) -> Handler {
.is_allowed()
})
.map(|s| {
json!({
let mut listing = json!({
"name": s.name,
"namespace": s.namespace,
"op_type": op_type_str(s.op_type),
})
});
if let Some(description) = &s.description {
listing["description"] = json!(description);
}
listing
})
.collect();
ResponseEnvelope::ok(ctx.request_id, json!({ "operations": ops }))
@@ -257,6 +289,53 @@ pub fn services_list_handler(registry: Arc<OperationRegistry>) -> Handler {
})
}
/// Register the bootstrap discovery ops (`services/list`,
/// `services/list-peers`, `services/schema`) against `registry` with
/// handlers closed over **that same `Arc`** (review 004 F-06): the
/// handlers see every op the registry serves at call time, including
/// per-session registrations made after this install. This is what
/// makes the per-session fork the discovery source for its own
/// openables — fork the base registry, register the generic channel
/// ops and openables, then install discovery on the fork and dispatch
/// the session over it.
///
/// `OperationRegistry` is internally mutable, so a forked registry
/// shared as an `Arc` can receive bootstrap ops after the dispatcher
/// was built. Call this once per session on the session's registry.
///
/// ACL filtering stays per-caller (each handler re-checks the calling
/// identity against every listed op's `AccessControl`) — no privilege
/// regression. Errors are per-op registration failures (e.g. a schema
/// compile failure — not reachable with the built-in specs); they
/// surface instead of being swallowed.
pub fn install_bootstrap_discovery(registry: &Arc<OperationRegistry>) -> Result<(), String> {
registry.register(HandlerRegistration::new(
services_list_spec(),
HandlerKind::Once(services_list_handler(Arc::clone(registry))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))?;
registry.register(HandlerRegistration::new(
services_list_peers_spec(),
HandlerKind::Once(services_list_peers_handler(Arc::clone(registry))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))?;
registry.register(HandlerRegistration::new(
services_schema_spec(),
HandlerKind::Once(services_schema_handler(Arc::clone(registry))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))?;
Ok(())
}
pub fn services_list_peers_handler(registry: Arc<OperationRegistry>) -> Handler {
Arc::new(move |input: Value, ctx: OperationContext| {
let registry = Arc::clone(&registry);
@@ -272,11 +351,15 @@ pub fn services_list_peers_handler(registry: Arc<OperationRegistry>) -> Handler
.is_allowed()
})
.map(|s| {
json!({
let mut listing = json!({
"name": s.name,
"namespace": s.namespace,
"op_type": op_type_str(s.op_type),
})
});
if let Some(description) = &s.description {
listing["description"] = json!(description);
}
listing
})
.collect();
let mut peers: Vec<Value> = Vec::new();
@@ -337,13 +420,35 @@ pub fn services_schema_handler(registry: Arc<OperationRegistry>) -> Handler {
);
}
};
match registry.registration(&name) {
Some(reg) => {
let spec_json = spec_to_json(&reg.spec);
ResponseEnvelope::ok(ctx.request_id, spec_json)
}
None => ResponseEnvelope::not_found(ctx.request_id, &name),
let registration = match registry.registration(&name) {
Some(reg) => reg,
None => return ResponseEnvelope::not_found(ctx.request_id, &name),
};
// Same gates as `OperationRegistry::invoke` — the schema of a
// restricted op is disclosed only to callers who could invoke
// it. Identity resolution mirrors invoke(): `handler_identity`
// (composition authority) when internal, `identity` otherwise.
if registration.spec.visibility == Visibility::Internal && !ctx.internal {
return ResponseEnvelope::not_found(ctx.request_id, &name);
}
let identity = if ctx.internal {
ctx.handler_identity
.as_ref()
.and_then(|ca| ca.as_identity())
} else {
ctx.identity.clone()
};
if let AccessResult::Forbidden(_) = registration.spec.access_control.check(
identity.as_ref(),
None,
ctx.ownership.as_deref(),
) {
return ResponseEnvelope::not_found(ctx.request_id, &name);
}
let spec_json = spec_to_json(&registration.spec);
ResponseEnvelope::ok(ctx.request_id, spec_json)
})
})
}
@@ -480,7 +585,7 @@ mod tests {
}
fn registry_with_access_controlled_ops() -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec_with_acl("public/echo", AccessControl::default()),
@@ -532,7 +637,7 @@ mod tests {
}
fn registry_with_ops() -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("fs/readFile"),
@@ -720,13 +825,88 @@ mod tests {
}
}
// CF-004 — the schema handler must apply the same visibility + ACL
// gates as invoke(). Unrestricted ops stay fetchable; Internal,
// ACL-restricted (unauthorized), and *authorized* ACL-restricted
// shapes below. Restricted disclosures return NOT_FOUND (spec-404),
// matching "restricted ops don't exist" everywhere else.
#[tokio::test]
async fn services_schema_hides_internal_op_from_external_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context("req-cf4-1");
let response = handler(serde_json::json!({ "name": "internal/hidden" }), ctx).await;
match response.result {
Err(e) => assert_eq!(e.code, "NOT_FOUND"),
other => panic!("expected NOT_FOUND for internal op, got {other:?}"),
}
}
#[tokio::test]
async fn services_schema_shows_internal_op_to_internal_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let mut ctx = root_context("req-cf4-2");
ctx.internal = true;
ctx.handler_identity = Some(CompositionAuthority::new("agent-chat", []));
let response = handler(serde_json::json!({ "name": "internal/hidden" }), ctx).await;
let spec = response.result.expect("internal caller sees internal op");
assert_eq!(spec.get("name"), Some(&json!("internal/hidden")));
}
#[tokio::test]
async fn services_schema_hides_acl_restricted_op_from_unauthorized_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context_with_identity(
"req-cf4-3",
Some(identity_with_scopes("regular-peer", &["user"])),
);
let response = handler(serde_json::json!({ "name": "admin/secret" }), ctx).await;
match response.result {
Err(e) => assert_eq!(e.code, "NOT_FOUND"),
other => panic!("expected NOT_FOUND for unauthorized ACL, got {other:?}"),
}
}
#[tokio::test]
async fn services_schema_shows_acl_restricted_op_to_authorized_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context_with_identity(
"req-cf4-4",
Some(identity_with_scopes("admin-peer", &["admin"])),
);
let response = handler(serde_json::json!({ "name": "admin/secret" }), ctx).await;
let spec = response.result.expect("authorized caller sees the spec");
assert_eq!(spec.get("name"), Some(&json!("admin/secret")));
assert_eq!(
spec.get("access_control")
.and_then(|a| a.get("required_scopes")),
Some(&json!(["admin"]))
);
}
#[tokio::test]
async fn services_schema_unrestricted_op_fetchable_without_identity() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context("req-cf4-5");
let response = handler(serde_json::json!({ "name": "public/echo" }), ctx).await;
let spec = response
.result
.expect("default-ACL op fetchable unauthenticated");
assert_eq!(spec.get("name"), Some(&json!("public/echo")));
}
#[tokio::test]
async fn services_list_handler_registered_and_invocable_via_registry() {
let registry = registry_with_ops();
let list_handler = services_list_handler(Arc::clone(&registry));
let schema_handler = services_schema_handler(Arc::clone(&registry));
let mut discovery_registry = OperationRegistry::new();
let discovery_registry = OperationRegistry::new();
discovery_registry
.register(HandlerRegistration::new(
services_list_spec(),
@@ -899,6 +1079,106 @@ mod tests {
);
}
// --- review 006 E-02: `description` on OperationSpec -------------------
#[test]
fn spec_to_json_emits_description_when_set() {
let spec = external_spec("fs/readFile").with_description("Read a file");
let json_val = spec_to_json(&spec);
assert_eq!(json_val.get("description"), Some(&json!("Read a file")));
}
#[test]
fn spec_to_json_omits_description_when_absent() {
let spec = external_spec("fs/readFile");
let json_val = spec_to_json(&spec);
assert!(
json_val.get("description").is_none(),
"description must be absent when not set (additive optional field)"
);
}
#[tokio::test]
async fn services_list_emits_description_when_set() {
let registry = Arc::new(OperationRegistry::new());
registry
.register(HandlerRegistration::new(
external_spec("channels/tty/sub").with_description("Interactive TTY sessions"),
HandlerKind::Once(echo_handler()),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
registry
.register(HandlerRegistration::new(
external_spec("fs/readFile"),
HandlerKind::Once(echo_handler()),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let handler = services_list_handler(Arc::clone(&registry));
let response = handler(json!({}), root_context("req-e02-1")).await;
let output = response.result.expect("ok response");
let ops = output
.get("operations")
.and_then(|v| v.as_array())
.expect("operations array")
.iter()
.map(|o| {
(
o.get("name").and_then(|n| n.as_str()).unwrap_or(""),
o.get("description").and_then(|d| d.as_str()),
)
})
.collect::<Vec<_>>();
assert!(
ops.contains(&("channels/tty/sub", Some("Interactive TTY sessions"))),
"described op carries its description: {ops:?}"
);
assert!(
ops.contains(&("fs/readFile", None)),
"undescribed op omits the description key: {ops:?}"
);
}
#[tokio::test]
async fn services_schema_discloses_description() {
let registry = registry_with_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let described = external_spec("fs/readFile").with_description("Read a file");
registry
.register(HandlerRegistration::new(
described,
HandlerKind::Once(echo_handler()),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let response = handler(json!({ "name": "fs/readFile" }), root_context("req-e02-2")).await;
let spec = response.result.expect("ok response");
assert_eq!(spec.get("description"), Some(&json!("Read a file")));
}
#[test]
fn operation_spec_schema_documents_description_property() {
let schema = operation_spec_schema();
let props = schema
.get("properties")
.and_then(|v| v.as_object())
.expect("properties object");
assert!(
props.contains_key("description"),
"operation_spec_schema must advertise description"
);
}
#[tokio::test]
async fn services_list_filters_by_access_control_authorized_peer() {
let registry = registry_with_access_controlled_ops();
@@ -1105,4 +1385,101 @@ mod tests {
"unauthorized peer must not see admin op in list-peers"
);
}
// --- review 004 F-06: per-fork bootstrap discovery ---------------------
fn context_for(
request_id: &str,
identity: Option<crate::core::auth::Identity>,
) -> OperationContext {
OperationContext {
request_id: request_id.to_string(),
parent_request_id: None,
identity,
handler_identity: None,
forwarded_for: None,
capabilities: Capabilities::new(),
metadata: HashMap::new(),
scoped_env: ScopedPeerEnv::empty(),
env: Arc::new(crate::registry::env::LocalOperationEnv::new(Arc::new(
OperationRegistry::new(),
))),
abort_policy: crate::registry::context::AbortPolicy::default(),
deadline: Some(std::time::Instant::now() + Duration::from_secs(30)),
internal: false,
ownership: None,
}
}
fn identity_scopes(id: &str, scopes: &[&str]) -> crate::core::auth::Identity {
crate::core::auth::Identity {
id: id.to_string(),
scopes: scopes.iter().map(|s| s.to_string()).collect(),
resources: HashMap::new(),
}
}
/// The F-06 gate: fork the base, register a per-session openable on
/// the fork, install bootstrap discovery on the fork — the openable
/// is discoverable via `services/list` for an authorized caller and
/// hidden from an unauthorized one.
#[tokio::test]
async fn bootstrap_discovery_on_fork_sees_per_session_ops() {
let base = OperationRegistry::new();
base.register(HandlerRegistration::new(
external_spec("base/op"),
HandlerKind::Once(make_handler(|input, context| async move {
ResponseEnvelope::ok(context.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let fork = Arc::new(base.fork());
fork.register(HandlerRegistration::new(
external_spec("channels/tty/sub"),
HandlerKind::Once(make_handler(|input, context| async move {
ResponseEnvelope::ok(context.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
install_bootstrap_discovery(&fork).expect("bootstrap discovery install");
let handler = fork
.registration("services/list")
.map(|r| match r.handler {
HandlerKind::Once(h) => h,
_ => panic!("services/list must be Once"),
})
.expect("services/list registered");
let authorized = context_for("req-f06-1", Some(identity_scopes("worker-a", &["tty"])));
let response = handler(json!({}), authorized).await;
let names: Vec<String> = response
.result
.expect("ok")
.get("operations")
.and_then(|v| v.as_array())
.expect("operations array")
.iter()
.filter_map(|o| o.get("name").and_then(|n| n.as_str().map(String::from)))
.collect();
assert!(
names.contains(&"channels/tty/sub".to_string()),
"per-session openable discoverable on the fork: {names:?}"
);
assert!(names.contains(&"base/op".to_string()));
let restricted = context_for("req-f06-2", None);
let response = handler(json!({}), restricted).await;
assert!(response.result.is_ok(), "list itself is callable");
}
}
+41 -1
View File
@@ -64,6 +64,16 @@ pub trait OperationEnv: Send + Sync {
Vec::new()
}
/// The operation names this env layer serves. The default returns
/// empty — single-layer envs don't need it. `OverlayOperationEnv`
/// overrides it with its overlay's registered names, and
/// `PeerCompositeEnv::peer_operations` delegates to it per peer
/// (ADR-030 — without the override `services/list-peers` shows
/// every peer with an empty operation list).
fn list_operation_names(&self) -> Vec<String> {
Vec::new()
}
/// Peer-routing composition (ADR-029 §2). Routes to a specific peer
/// (`PeerRef::Specific`) or to the first peer that serves the op
/// (`PeerRef::Any`). The default impl ignores the peer selector and
@@ -144,6 +154,14 @@ impl OperationEnv for LocalOperationEnv {
self.registry.invoke(&name, input, context).await
}
fn list_operation_names(&self) -> Vec<String> {
self.registry
.list_operations()
.into_iter()
.map(|s| s.name)
.collect()
}
}
/// Per-call composite env (ADR-024 + ADR-029 §1). Built by the `Dispatcher`
@@ -298,6 +316,28 @@ impl OperationEnv for PeerCompositeEnv {
fn peer_ids(&self) -> Vec<PeerId> {
self.connection_order.clone()
}
fn peer_operations(&self, peer: &PeerId) -> Vec<String> {
self.connections
.get(peer)
.map(|overlay| overlay.list_operation_names())
.unwrap_or_default()
}
fn list_operation_names(&self) -> Vec<String> {
let mut names: Vec<String> = self
.session
.as_ref()
.map(|s| s.list_operation_names())
.unwrap_or_default();
names.extend(
self.connections
.values()
.flat_map(|c| c.list_operation_names()),
);
names.extend(self.base.list_operation_names());
names
}
}
#[cfg(test)]
@@ -409,7 +449,7 @@ mod tests {
composition_authority: Option<CompositionAuthority>,
scoped_env: Option<ScopedPeerEnv>,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
+1
View File
@@ -8,5 +8,6 @@
pub mod context;
pub mod discovery;
pub mod env;
pub mod op_register;
pub mod registration;
pub mod spec;
+731
View File
@@ -0,0 +1,731 @@
//! `op/register` — the wire mechanism by which a connected peer
//! announces the operations it serves (review 004 F-05, ADR-022
//! amendment). The envelope kind set stays closed at six; bootstrap ops
//! over channel 0 are the door.
//!
//! Shape: the peer sends `call.requested` for `op/register` with a
//! payload of serializable registration parts — the `OperationSpec` in
//! the `services/schema` wire shape (`spec_to_json`) plus a `replace`
//! flag. The serving-side handler rebuilds the spec, wraps a
//! call-forwarding handler that issues a nested `call.requested` back
//! over channel 0 to the announcing peer (the same shape `from_call`'s
//! imported bundles use), and writes the bundle into that connection's
//! overlay via `CallConnection::register_imported`.
//!
//! The overlay is the landing zone: `compose_root_env` attaches it
//! keyed by peer identity, so nested invocations from any composed
//! handler reach the peer-announced op, and `services/list-peers`
//! discovers it (`ctx.env.peer_operations()`).
//!
//! `AccessControl` gates the surface: the `op/register` op itself
//! carries an `AccessControl` (an unprivileged peer cannot reach the
//! handler at all — the registry's normal invoke path enforces it), and
//! an announced op that collides with an existing registration is
//! rejected unless `replace` is set (the reconnect path).
//!
//! The `Handler` closures cannot cross the wire — the announcing peer
//! keeps its handler locally; the registered bundle is a forwarding
//! stub. This is the same contract `from_call` produces for the
//! hub→consumer import direction, extended to the peer→hub direction.
use std::sync::Arc;
use serde_json::{json, Value};
use crate::core::types::Capabilities;
use crate::protocol::connection::CallConnection;
use crate::protocol::wire::{CallError, ResponseEnvelope};
use crate::registry::registration::{
make_handler, Handler, HandlerKind, HandlerRegistration, OperationProvenance, OperationRegistry,
};
use crate::registry::spec::{AccessControl, OperationSpec, OperationType, Visibility};
pub const OP_REGISTER_NAME: &str = "op/register";
/// The wire DTO a peer sends as the `op/register` input. The spec
/// travels in the `services/schema` wire shape (`spec_to_json` output /
/// `rebuild_spec_for` input) so there is one spec serialization on the
/// wire.
#[derive(Debug, Clone)]
pub struct OpRegisterRequest {
pub spec: OperationSpec,
/// Replace an existing registration of the same name (the
/// reconnect path). `false` (default) rejects a collision with
/// `ALREADY_EXISTS`.
pub replace: bool,
}
impl OpRegisterRequest {
pub fn to_json(&self) -> Value {
json!({
"spec": crate::registry::discovery::spec_to_json_pub(&self.spec),
"replace": self.replace,
})
}
pub fn from_json(value: &Value) -> Result<Self, CallError> {
let spec_json = value
.get("spec")
.ok_or_else(|| CallError::invalid_input("op/register payload missing `spec`"))?;
let name = spec_json
.get("name")
.and_then(|v| v.as_str())
.ok_or_else(|| CallError::invalid_input("op/register spec missing `name`"))?
.to_string();
let spec = crate::client::rebuild_spec_for(spec_json, &name, &None).map_err(|e| {
CallError::invalid_input(format!("op/register spec rebuild failed: {e:?}"))
})?;
let replace = value
.get("replace")
.and_then(|v| v.as_bool())
.unwrap_or(false);
Ok(Self { spec, replace })
}
}
/// The `op/register` `OperationSpec`. The `access_control` here is the
/// registration surface's gate — a deployment that accepts
/// registrations only from scoped peers sets `required_scopes`; the
/// default (`AccessControl::default()`) lets any peer register. The op
/// is `Mutation`-typed (it mutates the connection overlay).
pub fn op_register_spec(access_control: AccessControl) -> OperationSpec {
OperationSpec::new(
OP_REGISTER_NAME,
OperationType::Mutation,
Visibility::External,
json!({
"type": "object",
"properties": {
"spec": { "type": "object" },
"replace": { "type": "boolean" }
},
"required": ["spec"]
}),
json!({
"type": "object",
"properties": {
"name": { "type": "string" },
"registered": { "type": "boolean" }
},
"required": ["name", "registered"]
}),
vec![],
access_control,
None,
)
}
/// Build the `op/register` handler for a connection: announces land in
/// `connection`'s overlay via `register_imported`, wrapped as
/// call-forwarding stubs that issue a nested `call.requested` back over
/// channel 0 to the announcing peer (the `from_call`-import shape).
///
/// `Visibility::Internal` is forced on the registered spec: an
/// announced op is composition material for the serving side's own
/// handlers (ADR-017), never directly callable from this side's wire —
/// the op is callable from the announcing peer's side by the peer
/// serving it there. `services/list-peers` still discovers it (the
/// overlay is peer-keyed, provenance `FromCall`).
///
/// Collision policy (review 005 G-03): announced ops may collide with
/// other *announced* ops on the same connection (`replace` governs,
/// the reconnect path) but **never** with the serving side's own
/// registrations — a name on `serving_registry` rejects with
/// `ALREADY_EXISTS` regardless of `replace`. Without this gate the
/// connection overlay shadows the base registry in `PeerCompositeEnv`
/// (connections resolve before base), so a peer could silently
/// rewrite what a wire-dispatched handler's `ctx.env.invoke` resolves
/// for any name the deployment registered — composition authority
/// (ADR-018) belongs to the composing handler's deployer, not the
/// connected peer.
///
/// Replace semantics: a registration for the same name already on the
/// overlay is rejected with `ALREADY_EXISTS` unless `replace: true`
/// (the reconnect path re-announces).
pub fn op_register_handler(
connection: Arc<CallConnection>,
serving_registry: Arc<OperationRegistry>,
) -> Handler {
make_handler(move |input, context| {
let connection = Arc::clone(&connection);
let serving_registry = Arc::clone(&serving_registry);
async move {
let request = match OpRegisterRequest::from_json(&input) {
Ok(r) => r,
Err(e) => return ResponseEnvelope::error(context.request_id, e),
};
if serving_registry.registration(&request.spec.name).is_some() {
return ResponseEnvelope::error(
context.request_id,
CallError::already_exists(format!(
"op/register: `{}` is registered by this side's own serving \
registry; peer-announced ops may not shadow it",
request.spec.name
)),
);
}
if connection.overlay_contains(&request.spec.name) && !request.replace {
return ResponseEnvelope::error(
context.request_id,
CallError::already_exists(format!(
"op/register: `{}` is already registered on this connection; \
set `replace: true` to replace it",
request.spec.name
)),
);
}
let mut spec = request.spec;
spec.visibility = Visibility::Internal;
let remote_name = spec.name.clone();
let handler = forwarding_stub_for_announced_op(
Arc::clone(&connection),
remote_name,
spec.op_type,
);
connection.register_imported(HandlerRegistration::new(
spec.clone(),
handler,
OperationProvenance::FromCall,
None,
None,
Capabilities::new(),
));
ResponseEnvelope::ok(
context.request_id,
json!({ "name": spec.name, "registered": true }),
)
}
})
}
/// The forwarding stub for a peer-announced op. Query/Mutation ops get
/// the `from_call` forwarding shape (nested `call.requested` back over
/// channel 0, `forwarded_for` populated per ADR-032 §3). Announced
/// Sub/Pub ops are registered as stubs that return
/// `INVALID_OPERATION_TYPE` on invocation — nested composition is
/// request/response-only (`OverlayOperationEnv`'s contract); the
/// streaming/sink forwarding shapes ride on the `from_call` import
/// path, which the announcing side can use in the other direction.
fn forwarding_stub_for_announced_op(
connection: Arc<CallConnection>,
remote_name: String,
op_type: OperationType,
) -> HandlerKind {
match op_type {
OperationType::Query | OperationType::Mutation => HandlerKind::Once(
crate::client::make_forwarding_handler(connection, remote_name),
),
OperationType::Sub | OperationType::Pub => {
HandlerKind::Once(make_handler(|_input, context| async move {
ResponseEnvelope::error(
context.request_id,
CallError::invalid_operation_type(
"peer-announced Sub/Pub ops are not invocable over nested \
composition (request/response only)",
),
)
}))
}
}
}
/// Announce an op to the connected peer over `connection`'s channel 0:
/// sends `call.requested` for `op/register` and awaits the response.
/// The announcing side keeps its real handler locally and serves it via
/// the serving loop (F-04) when the peer invokes the announced op back.
pub async fn announce_op(
connection: &CallConnection,
spec: OperationSpec,
replace: bool,
) -> ResponseEnvelope {
let request = OpRegisterRequest { spec, replace };
connection
.call_with_payload(serde_json::json!({
"operationId": OP_REGISTER_NAME,
"input": request.to_json(),
}))
.await
}
#[cfg(test)]
mod tests {
use super::*;
use crate::protocol::connection::CallConnection;
use crate::registry::context::OperationContext;
use crate::registry::discovery::install_bootstrap_discovery;
use crate::registry::registration::OperationRegistry;
use crate::registry::spec::Visibility;
use std::collections::HashMap;
fn announced_spec(name: &str) -> OperationSpec {
OperationSpec::new(
name,
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
)
}
fn stub_connection() -> crate::core::types::Connection {
crate::protocol::sink_empty_connection()
}
fn test_context(request_id: &str) -> OperationContext {
OperationContext {
request_id: request_id.to_string(),
parent_request_id: None,
identity: None,
handler_identity: None,
forwarded_for: None,
capabilities: Capabilities::new(),
metadata: HashMap::new(),
scoped_env: crate::registry::context::ScopedPeerEnv::empty(),
env: Arc::new(crate::registry::env::LocalOperationEnv::new(Arc::new(
OperationRegistry::new(),
))),
abort_policy: crate::registry::context::AbortPolicy::default(),
deadline: None,
internal: false,
ownership: None,
}
}
#[test]
fn request_round_trips_through_json() {
let request = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: true,
};
let json = request.to_json();
let parsed = OpRegisterRequest::from_json(&json).expect("parse");
assert_eq!(parsed.spec.name, "worker/exec");
assert_eq!(parsed.spec.op_type, OperationType::Query);
assert!(parsed.replace);
}
#[test]
fn request_missing_spec_is_invalid_input() {
let err = OpRegisterRequest::from_json(&json!({})).unwrap_err();
assert_eq!(err.code, "INVALID_INPUT");
}
#[test]
fn request_missing_name_is_invalid_input() {
let err = OpRegisterRequest::from_json(&json!({ "spec": {} })).unwrap_err();
assert_eq!(err.code, "INVALID_INPUT");
}
#[tokio::test]
async fn handler_registers_announced_op_in_overlay() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let response = handler(input, test_context("req-or-1")).await;
assert!(
response.result.is_ok(),
"register succeeded, got {:?}",
response.result
);
let registered = conn.overlay_env().contains("worker/exec");
assert!(registered, "announced op landed in the connection overlay");
assert!(conn.overlay_contains("worker/exec"));
}
#[tokio::test]
async fn handler_rejects_collision_without_replace() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let first = handler(input.clone(), test_context("req-or-2a")).await;
assert!(first.result.is_ok());
let second = handler(input, test_context("req-or-2b")).await;
let err = second.result.expect_err("collision rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
}
#[tokio::test]
async fn handler_replaces_with_replace_flag() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let original = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let first = handler(original, test_context("req-or-3a")).await;
assert!(first.result.is_ok());
let replacement = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: true,
}
.to_json();
let second = handler(replacement, test_context("req-or-3b")).await;
assert!(
second.result.is_ok(),
"replace flag permits re-registration, got {:?}",
second.result
);
}
#[tokio::test]
async fn registered_spec_forced_internal_with_fromcall_provenance() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let response = handler(input, test_context("req-or-4")).await;
assert!(response.result.is_ok());
let registration = conn
.overlay_registration("worker/exec")
.expect("registered");
assert_eq!(registration.spec.visibility, Visibility::Internal);
assert_eq!(
registration.provenance,
crate::registry::registration::OperationProvenance::FromCall
);
}
// --- review 005 Unit 2 acceptance gates (G-03 collision policy) -------
/// G-03 gate: an announce colliding with a **serving-registry**
/// name rejects with `ALREADY_EXISTS` even with `replace: true` —
/// peer-announced ops never shadow the serving side's own
/// registrations (composition authority stays with the deployer).
#[tokio::test]
async fn handler_rejects_base_registry_collision_even_with_replace() {
let serving = Arc::new(OperationRegistry::new());
serving
.register(HandlerRegistration::new(
crate::registry::spec::OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
crate::registry::registration::OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let input = OpRegisterRequest {
spec: announced_spec("fs/readFile"),
replace: true,
}
.to_json();
let response = handler(input, test_context("req-or-5")).await;
let err = response.result.expect_err("base collision rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
assert!(
!conn.overlay_contains("fs/readFile"),
"the rejected announce never lands in the overlay"
);
// The serving side's registration is untouched.
assert!(serving.registration("fs/readFile").is_some());
}
/// G-03 gate: an announce colliding with an **Internal** serving
/// op is also rejected — the visibility of the shadowed op is
/// irrelevant to composition shadowing (`OverlayOperationEnv`
/// gates on `AccessControl`, not visibility; the composed child is
/// `internal: true` by design).
#[tokio::test]
async fn handler_rejects_collision_with_internal_serving_op() {
let serving = Arc::new(OperationRegistry::new());
serving
.register(HandlerRegistration::new(
crate::registry::spec::OperationSpec::new(
"internal/vault",
OperationType::Query,
Visibility::Internal,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
crate::registry::registration::OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let input = OpRegisterRequest {
spec: announced_spec("internal/vault"),
replace: false,
}
.to_json();
let response = handler(input, test_context("req-or-6")).await;
let err = response
.result
.expect_err("internal base collision rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
}
/// G-03 gate: an announced op colliding with another *announced*
/// op still follows `replace` semantics — the base-registry gate
/// must not widen into the overlay.
#[tokio::test]
async fn handler_overlay_collision_still_governed_by_replace() {
let serving = Arc::new(OperationRegistry::new());
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let first = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
assert!(handler(first, test_context("req-or-7a"))
.await
.result
.is_ok());
let collision = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let err = handler(collision, test_context("req-or-7b"))
.await
.result
.expect_err("overlay collision without replace rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
let replacement = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: true,
}
.to_json();
assert!(
handler(replacement, test_context("req-or-7c"))
.await
.result
.is_ok(),
"overlay replace still permitted; base gate is scoped to the serving registry"
);
}
/// G-03 gate: after a successful announce of a distinct name,
/// nested composition of a **base-registered** op still resolves
/// the serving side's own op — the exact `compose_root_env` shape
/// (`PeerCompositeEnv` with the connection overlay attached). The
/// G-03 defect was that the collision gate was overlay-only; this
/// pins the invariant that survived the fix: the base layer is
/// reachable and correct whenever the name is not announced.
#[tokio::test]
async fn nested_composition_of_base_op_unaffected_by_unrelated_announce() {
let serving = Arc::new(OperationRegistry::new());
serving
.register(HandlerRegistration::new(
crate::registry::spec::OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
),
HandlerKind::Once(make_handler(|_input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, json!({ "from": "base" }))
})),
crate::registry::registration::OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
// Announce a *distinct* name; it lands in the connection overlay.
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
assert!(handler(input, test_context("req-or-8a"))
.await
.result
.is_ok());
assert!(conn.overlay_contains("worker/exec"));
// Compose the base op through the same env shape
// `compose_root_env` produces for a wire-dispatched handler.
let base = Arc::new(crate::registry::env::LocalOperationEnv::new(Arc::clone(
&serving,
)));
let mut composite = crate::registry::env::PeerCompositeEnv::new(base);
composite.attach_peer("consumer-peer".to_string(), conn.overlay_env());
let env: Arc<dyn crate::registry::env::OperationEnv + Send + Sync> = Arc::new(composite);
let mut ctx = test_context("req-or-8b");
ctx.env = env;
ctx.scoped_env = crate::registry::context::ScopedPeerEnv::new(["fs/readFile"]);
let response = ctx.env.invoke("fs", "readFile", json!({}), &ctx).await;
let out = response.result.expect("base op composes");
assert_eq!(
out,
json!({ "from": "base" }),
"the serving side's own op resolves through composition, not a peer stub"
);
}
/// UP-03 gate (alkhttp review 006): after a peer announces an op,
/// `services/list-peers` over the real `compose_root_env` shape
/// (`PeerCompositeEnv` + the connection overlay attached under the
/// peer's id) must list the announced op under that peer's
/// entry — not an empty operations array. The pre-fix failure:
/// `PeerCompositeEnv` overrode `peer_ids` only, so
/// `peer_operations` fell to the trait default (`Vec::new()`) and
/// every peer listed with `operations: []`. ADR-030's
/// `list_operation_names` override is what this exercises.
#[tokio::test]
async fn announced_op_is_discoverable_via_services_list_peers() {
use crate::registry::env::PeerCompositeEnv;
use crate::registry::{
context::ScopedPeerEnv, discovery::services_list_peers_handler, env::LocalOperationEnv,
};
let serving = Arc::new(OperationRegistry::new());
install_bootstrap_discovery(&serving).expect("bootstrap discovery install");
let peer_identity = crate::core::auth::Identity {
id: "consumer-peer".to_string(),
scopes: vec![],
resources: HashMap::new(),
};
let conn = Arc::new(CallConnection::new_overlay_only(peer_identity));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
assert!(
handler(input, test_context("req-or-9a"))
.await
.result
.is_ok(),
"announce lands in the connection overlay"
);
assert!(conn.overlay_contains("worker/exec"));
// The exact env shape `compose_root_env` produces for calls
// arriving on this connection: LocalOperationEnv base +
// the connection's overlay attached under the peer's id.
let base = Arc::new(LocalOperationEnv::new(Arc::clone(&serving)));
let mut composite = PeerCompositeEnv::new(base);
composite.attach_peer("consumer-peer".to_string(), conn.overlay_env());
let env: Arc<dyn crate::registry::env::OperationEnv + Send + Sync> = Arc::new(composite);
// Direct probe: the composite resolves the announced name from
// the attached overlay, and peer_operations lists it.
assert!(env.contains("worker/exec"));
let ops = env.peer_operations(&"consumer-peer".to_string());
assert_eq!(
ops,
vec!["worker/exec".to_string()],
"PeerCompositeEnv::peer_operations must surface the peer overlay's announced ops"
);
assert!(
env.peer_operations(&"unknown-peer".to_string()).is_empty(),
"an unattached peer has no operations"
);
// Wire-level probe: services/list-peers over the same env
// attributes the announced op to the peer.
let peers_registry = Arc::new(OperationRegistry::new());
install_bootstrap_discovery(&peers_registry).expect("bootstrap install");
let list_handler = services_list_peers_handler(Arc::clone(&peers_registry));
let mut ctx = test_context("req-or-9b");
ctx.env = env;
ctx.scoped_env = ScopedPeerEnv::empty();
let response = list_handler(json!({}), ctx).await;
let out = response.result.expect("list-peers ok");
let peers_arr = out
.get("peers")
.and_then(|v| v.as_array())
.expect("peers array");
let consumer = peers_arr
.iter()
.find(|p| p.get("peer_id").and_then(|v| v.as_str()) == Some("consumer-peer"))
.expect("consumer-peer present in list-peers output");
let names: Vec<&str> = consumer
.get("operations")
.and_then(|v| v.as_array())
.expect("consumer operations array")
.iter()
.filter_map(|o| o.get("name").and_then(|n| n.as_str()))
.collect();
assert!(
names.contains(&"worker/exec"),
"the announced op must be discoverable via services/list-peers (UP-03)"
);
// The local entry still lists the bootstrap ops (the serving
// registry's own surface is unaffected).
let local = peers_arr
.iter()
.find(|p| p.get("peer_id").and_then(|v| v.as_str()) == Some("local"))
.expect("local peer present");
let local_names: Vec<&str> = local
.get("operations")
.and_then(|v| v.as_array())
.expect("local operations array")
.iter()
.filter_map(|o| o.get("name").and_then(|n| n.as_str()))
.collect();
assert!(local_names.contains(&"services/list-peers"));
}
}
File diff suppressed because it is too large. Load diff
+38
View File
@@ -181,6 +181,13 @@ pub struct OperationSpec {
pub output_schema: Value,
pub error_schemas: Vec<ErrorDefinition>,
pub access_control: AccessControl,
/// Human-readable op description, disclosed by `services/list` when
/// set and by `services/schema` (`spec_to_json_pub`) — review 006
/// E-02. `None` (absent on the wire) for ops that declare none; the
/// field is additive and registry-side (it describes the op, not the
/// produced resource set — ADR-047 §6 keeps the live half of
/// discovery in `channel/resources/subscribe`).
pub description: Option<String>,
/// JSON pointer into the input for the resource ID, when
/// `access_control.resource_type` is set and the operation targets a
/// specific runtime-spawned resource (ADR-011). e.g. `"$.containerId"`
@@ -234,12 +241,22 @@ impl OperationSpec {
output_schema,
error_schemas,
access_control,
description: None,
resource_id_path,
publish_schema: None,
channel_open: None,
}
}
/// Set the op's `description` (review 006 E-02). Disclosed by
/// `services/list` when set; carried in the `services/schema` wire
/// shape (`spec_to_json_pub` / `rebuild_spec_for`). Builder-style;
/// returns `self` for chaining at registration sites.
pub fn with_description(mut self, description: impl Into<String>) -> Self {
self.description = Some(description.into());
self
}
/// Set the `publish_schema` (Pub ops only, ADR-046). Validates each
/// `call.published` chunk's `input`. Builder-style; returns `self`
/// for chaining at registration sites.
@@ -359,6 +376,27 @@ mod tests {
assert_eq!(spec.channel_open, None);
}
#[test]
fn description_defaults_to_none_and_builder_sets_it() {
let spec = OperationSpec::new(
"channels/tty/sub",
OperationType::Sub,
Visibility::External,
serde_json::json!({}),
serde_json::json!({}),
vec![],
AccessControl::default(),
None,
);
assert_eq!(spec.description, None);
let described = spec.with_description("Interactive TTY sessions");
assert_eq!(
described.description.as_deref(),
Some("Interactive TTY sessions")
);
}
#[test]
fn with_channel_open_sets_marker() {
let spec = OperationSpec::new(