20 Commits
Author SHA1 Message Date
glm-5.3-flash a22b2b84c9 fix: park early-arrival chunks for un-adopted channels (open/first-data race)
The connect side adopts a channel (installs local routing state) only
after the open-op response arrives, but the accept side's OpenHandler
can start pumping data the moment the channel opens — the two race and
the demux's lenient unknown-channel drop (REQ-CH-04) silently lost the
producer's first chunks (a TTY backend's banner, a sub protocol's
greeting).

route_payload now parks up to 64 payloads per unknown channel_id in a
bounded early-arrival buffer; adopt_channel drains them into the new
receiver in order. Beyond the cap the chunk drops with the existing
debug log + dropped_unknown_chunks counter (which now also counts
overflow). clear_all drops parked buffers with the connection.

Surfaced by alktty's consumer end-to-end test (review #001 L3): the
session never resolved because the producer's first chunks (stdout
sentinel + exit chunk for an immediately-resolving backend) arrived
before the adopt and were dropped. REQ-CH-04 wording updated by this
behavior; ADR-039 §demux loop describes the lenient drop for genuinely
unknown channels, which remains the case past the cap.

Verification: cargo test 597 (2 rewritten for the new semantics +
route_payload_to_unknown_channel_parks_until_adopt gate);
--all-features 614; clippy -D warnings clean; fmt clean; doc 0
warnings; publish --dry-run ok; standalone probe (handler-writes-first
e2e over one connection) shows 0 dropped chunks with the fix vs 1
without.
2026-09-05 07:04:30 +00:00
glm-5.3-flash 574f58442a feat: enforce input_schema at call time (ADR-016 INVALID_INPUT leg)
- OperationSpec.input_schema was advertise-only: services/schema
  disclosed it but no dispatch entry point consulted it (the only
  enforced schema was publish_schema per-chunk on Pub ops, P-03).
- compile input_schema once at registration, same fail-closed rule as
  publish_schema/CF-003: an un-compilable schema is a registration
  error, never a silently-skipped contract. Validator cache mirrored
  on fork and in OperationRegistryBuilder like the publish validators.
- check after the ACL gate in all three dispatch entry points:
  invoke, invoke_streaming, invoke_sink (via resolve_sink_handler,
  preserving the P-08 single-source-of-truth property). Violations
  return INVALID_INPUT with the input echoed in details.
- raw-JSON-Schema semantics (permissive on unknown keys); adapters
  wanting closed-by-default keep their own hardening (alkhttp's
  CompiledInputSchema composes unchanged).
- motivated by alktty review #001 L1: the channels open-op wrapper
  hands the registry-checked input to the OpenHandler as the
  authoritative params, which requires the registry to validate it.

Verification: cargo test 596 lib (6 new: invoke/streaming/sink
enforcement, fail-closed registration, permissive-{} compile,
fork-carries-validator); --all-features 613; clippy -D warnings
clean; fmt clean; doc 0 warnings; publish --dry-run ok. alkhttp
438+16 tests pass against 0.3.1 (registry) — re-verify against the
published 0.4.0 after upload.
2026-09-05 06:33:00 +00:00
glm-5.3-flash ba94c70eb2 chore: bump to 0.3.1, changelog for the UP-03 list-peers fix
Verification: 590 default / 607 all-features tests, clippy
(all-targets, all-features, wasm32) clean, fmt clean, publish
dry-run OK.
2026-09-04 16:01:56 +00:00
glm-5.3-flash fd212307e2 fix(up-03): PeerCompositeEnv::peer_operations override — list-peers sees peer-announced ops
Surfaced by alkhttp review 006 (UP-03): services/list-peers showed
every peer with an empty operations array. PeerCompositeEnv overrode
peer_ids only, so peer_operations fell to the trait default
(Vec::new()) and the ADR-022 amendment's "announced op is discoverable
via services/list-peers" promise never resolved on the wire. ADR-030
prescribed the fix but it had never been ported into alkcall. The
existing list-peers unit tests passed because they mock
peer_operations with hand-rolled envs.

Implements ADR-030 as specified:
- OperationEnv gains list_operation_names (default Vec::new(),
  back-compat for all existing implementors)
- OverlayOperationEnv overrides it with its overlay's registered names
- PeerCompositeEnv::peer_operations delegates to the peer overlay's
  list_operation_names; PeerCompositeEnv::list_operation_names
  aggregates session + connections + base (mirrors its contains())
- LocalOperationEnv enumerates its registry; ChannelsSessionEnv
  delegates to base

Gate: announced_op_is_discoverable_via_services_list_peers in
src/registry/op_register.rs — announces an op through op/register,
then asserts both the direct peer_operations probe and the
services/list-peers wire shape attribute the announced op to the peer,
over the exact compose_root_env shape (PeerCompositeEnv + attached
connection overlay). Verified load-bearing: reverting the
peer_operations override fails the gate.

ADR-030 status Proposed -> Accepted with the UP-03 provenance note.

Verification: 590 default / 607 all-features tests, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean,
semver-checks 196 pass against v0.3.0 (defaulted trait method is
non-breaking).
2026-09-04 16:00:23 +00:00
glm-5.3-flash 1e20bb77d0 chore: prepublish review — bump to 0.3.0, changelog, doc alignment
0.2.0 is already on crates.io (2026-08-31, c16b069); the review
004/005 remediation work is unreleased on top of it and lands as
0.3.0.

- Bump version 0.2.0 -> 0.3.0. Two OperationRegistry methods changed
  borrowed returns to owned (registration, list_operations) —
  source-breaking for annotated call sites, minor bump per 0.x
  semver rules. cargo semver-checks passes (196 checks) against the
  published baseline; the return-type changes were caught by manual
  diff review.
- CHANGELOG 0.3.0: connect-side serving (from_connection_with_serving
  + ServingConfig), OperationRegistry::fork + builder from_registry,
  registry::op_register (bootstrap op, collision policy,
  ALREADY_EXISTS), install_bootstrap_discovery, spec_to_json_pub +
  resource_id_path round-trip, overlay accessors, concurrent serving
  loops, &self registration.
- README: serving-as-consumer section, consumer role table update,
  drop the stale `mut` on the registry example.
- AGENTS.md: ADR range 001..047 -> 001..048.
- Fix rustdoc private-intra-doc-link warning on StartedDispatch.

Verification: 589 default / 606 all-features tests, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean,
publish dry-run OK.
2026-09-04 14:14:22 +00:00
glm-5.3-flash d5b2661b38 fix(review 005 Unit 3): resource_id_path wire round-trip + bootstrap-list doc alignment (G-04, G-05)
- resource_id_path rides both halves of the spec wire round-trip:
  spec_to_json_pub serializes it (optional string key), rebuild_spec_for
  parses it. Additive optional field - absent stays absent. Previously
  an announced (or from_call-imported) op declaring ownership-scoped
  resource extraction silently rebuilt with resource_id: None, so ACL
  checks ran without the resource ID.
- Gates: spec_round_trips_resource_id_path (serialize -> parse ->
  field intact) + spec_without_resource_id_path_stays_absent (additive
  field breaks no consumer).
- ADR-022 amendment: bootstrap-op set gains services/list-peers with a
  dated G-05 note (the installer has registered it since the amendment
  landed; the doc lagged the code). Set remains closed at four.

Verification: cargo test 589 / --all-features 606, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean.

Refs docs/reviews/005-...md (G-04, G-05; all findings closed).
2026-09-04 09:41:43 +00:00
glm-5.3-flash 23c9b28c6b fix(review 005 Unit 2): op/register serving-registry collision gate (G-03)
- op_register_handler takes the serving registry alongside the
  connection and rejects announced names that collide with the serving
  side's own registrations (ALREADY_EXISTS regardless of replace).
  Peer-announced ops may collide with peer-announced ops (replace
  governs, the reconnect path) but never shadow the deployment's own
  ops: the connection overlay resolves before base in PeerCompositeEnv,
  so an unscreened same-name announce would silently rewrite what a
  wire-dispatched handler's ctx.env.invoke resolves. Composition
  authority (ADR-018) stays with the deployer.
- ADR-022 amendment (2026-09-04): collision policy recorded in the
  2026-09-03 amendment's op/register section (rationale + visibility
  irrelevance); status line notes the sub-amendment.
- Gates: base-External collision rejected even with replace (overlay
  stays clean, serving registration untouched); Internal base op
  equally protected; overlay/overlay collisions still follow replace;
  nested composition of a base op resolves the serving side's own op
  after an unrelated announce (real compose_root_env env shape).
- PeerCompositeEnv resolution order deliberately unchanged.

Verification: cargo test 587 / --all-features 604, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean.

Refs docs/reviews/005-...md (G-03; Unit 3 open).
2026-09-04 09:40:11 +00:00
glm-5.3-flash 1cbb7c6536 fix(review 005 Unit 1): concurrent serving loops + stub-exercising gates (G-01, G-02)
- Split dispatch() into dispatch_start() (sync prefix) + spawned
  invocation: both single-stream loops (serve_single_stream and the
  accept-side run_loop_single_stream) spawn Once invocations, Sub
  pumps, and sink response writers; only the Pub sink start stays
  inline (chunk_tx must register before the next call.published).
  Inline dispatch deadlocked same-connection nested composition: the
  read loop awaited the parent handler, which awaited a nested call
  whose response only the same read loop could resolve (resolved only
  via the 30s sweeper). Spawned handles tracked + aborted at loop exit;
  in_flight_sinks behind an Arc<parking_lot::Mutex> with guards dropped
  before awaits.
- run_loop_single_stream gains the pending-resolution arms
  (RESPONDED/COMPLETED/ERROR): the accept side previously served only
  and had no loop resolving its own outbound pendings in single-stream
  mode — the latent accept-side imported-op composition hazard is
  mechanized shut.
- Write-failure in the spawned Once path warns instead of closing the
  loop (matches the Sink arm; dying transport still surfaces via
  ConnectionClosed on the next read).
- G-02 gate: hub_handler_composes_peer_announced_op_via_nested_composition
  — announce -> consumer calls hub/compose -> hub's serving loop
  wire-dispatches it -> handler composes via ctx.env -> forwarding
  stub's nested call crosses back to the consumer. The F-05 gate
  bypassed this path entirely.
- Interleaved-directions gate: outbound_call_resolves_while_inbound_
  subscription_is_being_served — consumer serves a live Sub while a
  wire-dispatched hub handler issues an outbound call on the same
  connection.
- Both gates verified load-bearing: run against the pre-fix loop each
  reproduces the G-01 hang (no progress, bounded-timeout failure);
  post-fix both resolve in <0.2s, no sweeper evictions.

Verification: cargo test 583 / --all-features 600, clippy
(all-targets, all-features, wasm32) clean, fmt clean, doc clean.

Refs docs/reviews/005-...md (G-01, G-02; Units 2-3 open).
2026-09-04 09:36:08 +00:00
glm-5.3-flash 435ae9da2f docs(review 005): serving-loop concurrency + op/register composition findings
Post-remediation review of f84d214 (review 004 Units 1-3). Five
findings, verified in source and (for G-01) empirically via a probe
test that was added, run, and removed:

- G-01 [major]: serve_single_stream awaits dispatch inline; a
  wire-dispatched handler composing a peer-announced op (or a
  from_call import) over the same connection deadlocks — the nested
  call resolves only via the 30s sweeper (probe: TIMEOUT at 30.0007s).
- G-02 [major]: the F-05 e2e gate calls the announced op directly,
  bypassing the forwarding stub — the one path G-01 breaks.
- G-03 [major]: op/register's collision gate is overlay-only;
  PeerCompositeEnv resolves connections before base, so an announced
  op can shadow the serving side's own ops in nested composition.
- G-04 [minor]: resource_id_path does not survive the spec wire
  round-trip (pre-existing shape, load-bearing for op/register).
- G-05 [minor]: install_bootstrap_discovery registers
  services/list-peers; ADR-022's bootstrap set doesn't name it.

Non-findings bound the re-review: fork surface lock discipline,
bootstrap discovery closure, frame-arm equivalence of the composed
loop, unchanged pure-consumer default, alkhttp cross-repo claims,
CJK sweep (none), all gates reproduce (581/598, clippy, fmt, wasm,
doc).

Remediation plan: Unit 1 (concurrent serving loop + stub-exercising
gate) gates Unit 4 downstream; Unit 2 (collision policy); Unit 3
(round-trip completeness + doc alignment).

Verification: cargo doc --no-deps clean; tree unchanged apart from
this review doc.
2026-09-04 07:25:41 +00:00
glm-5.3-flash f84d214173 feat: per-session fork registry, connect-side serving loop, op/register (review 004 Units 1-3)
Remediates all six findings of review 004 (per-connection dispatch
resolution and client-side op serving). All claims re-verified in
source before remediation; F-02's member list gains ScopedPeerEnv
(also Clone — fork surface simpler than estimated).

- OperationRegistry: interior mutability (parking_lot RwLock on both
  maps); register takes &self; registration/list_operations return
  owned clones; fork() deep-copies registrations + cached publish-schema
  validators (F-02/F-03); OperationRegistryBuilder::from_registry.
- install_bootstrap_discovery: services/list, services/list-peers,
  services/schema registered closed over the fork itself, so
  per-session openables are discoverable and services/schema answers
  from the fork (F-06).
- Dispatcher::serve_single_stream: full-duplex single-stream loop —
  call.requested dispatches inbound; responded/completed/error resolve
  outbound pendings; aborted tries both tables (in-flight sink aborts
  + pending cascade); published routes inbound sinks (F-04).
- ChannelClient::from_connection_with_serving(connection,
  Option<ServingConfig>): opt-in serving; from_connection keeps the
  pure-consumer default.
- registry::op_register: OpRegisterRequest wire DTO (spec in
  services/schema JSON + replace flag), op_register_spec,
  op_register_handler (rebuild -> forwarding stub -> register_imported,
  forced Internal/FromCall), announce_op; CallError::already_exists;
  spec_to_json_pub; from_call's rebuild_spec_for + forwarding-handler
  constructors crate-shared (F-05).
- ADR-047 §4 amendment #2: per-session fork is the dispatch-registry
  mechanism; overlay stays nested-invocation/peer-announced landing
  zone (F-01/F-02).
- ADR-022 amendment 2026-09-03: bootstrap-op set (services/list,
  services/schema, op/register), opt-in connect-side serving, op/register
  wire shape (F-04/F-05).
- alkhttp ADR-048 reconciliation note + OQ-05 re-pointed at the alkcall
  ADRs (Unit 1b).
- Review 004 status -> remediated; remediation log with gates.

Verification:
- cargo test: 581 passed, 0 failed (565 baseline + 16 new)
- cargo test --all-features: 598 passed, 0 failed
- cargo clippy --all-targets -- -D warnings: clean
- cargo clippy --all-features --all-targets -- -D warnings: clean
- cargo clippy --target wasm32-unknown-unknown -- -D warnings: clean
- cargo fmt --check: clean
- cargo doc --no-deps: clean

Gates: fork_registry_open_op_resolves_and_is_discoverable (open op via
fork + services/list shows openable + services/schema validates),
serving_loop_hub_to_consumer_call_resolves (hub->consumer call through
consumer's serving loop, consumer->hub still resolves),
op_register_announce_then_hub_call_routes_back_to_consumer (announce ->
overlay -> hub call -> forwarding stub -> consumer serves).
2026-09-03 17:32:45 +00:00
glm-5.3-flash c0dbf82518 docs(review 004): per-connection dispatch resolution + client-side op serving
Focused design-mismatch review found via alkhttp's WS data-channel
drill-down (alkhttp review 003 WS-24/WS-25). The channel machinery is
done and proven; the gaps are the dispatch-resolution mechanism and
the connect-side serving half — both upstream of any transport.

Findings:
- F-01 [major]: top-level dispatch consults only the dispatcher's
  base registry — ops registered per the ADR-047 §4 amendment's
  overlay mechanism resolve NOT_FOUND on the wire; the only proven
  shape (per-session registry as the dispatcher's base) differs from
  the ADR's wording
- F-02 [major]: no Clone/fork surface on OperationRegistry or
  HandlerRegistration — every inner payload type IS Clone-able
  (verified type-by-type, incl. jsonschema::Validator and
  Capabilities), so the fork is a small addition
- F-03 [minor]: fork must carry handlers + validators, not just specs
- F-04 [major]: connect-side channel-0 read pump resolves responses
  only; inbound call.requested frames are silently dropped — no
  serving half on the single-stream shape (ADR-022/AGENTS §8
  bidirectionality unreachable from the connect side)
- F-05 [major]: no wire mechanism announces client-side ops; the
  six call.* kinds are closed. Resolution candidate: bootstrap op
  (op/register) served per-session, handler writes into the
  connection-local overlay; discovery rides services/list-peers
- F-06 [minor]: per-session fork must carry bootstrap discovery ops
  for per-session openables to be discoverable

Includes a non-findings section (e2e reference shape,
register_openable completeness, channel-id split, envelope-kind
closure, wasm-cleanliness) and a 4-unit plan: ADR decisions (Unit 1)
-> fork surface (Unit 2) -> client serving + bootstrap op (Unit 3)
-> alkhttp wiring downstream (Unit 4, tracked in alkhttp review 003).

Verification: cargo test (565), clippy --all-targets -D warnings,
fmt, doc --no-deps — all clean at c16b069. No source changes.
2026-09-03 14:42:57 +00:00
glm-5.3-flash c16b0697e3 chore: drop dead Cargo.lock from package exclude list
cargo always includes a git-tracked lockfile in the package regardless
of the exclude entry; keeping it listed implied a behavior that does not
exist.
2026-08-31 10:01:48 +00:00
glm-5.3-flash c3d4fa1b30 chore: bump to 0.2.0
First release carrying the consumer-findings remediation (CF-001..004),
the feature-gated gateway dispatch spine (ADR-048), and the
registration-time publish_schema validation behavior change (CF-003).
Gate is version-only for existing consumers: the public 0.1.1 API
surface is unchanged (probe-verified).
2026-08-31 09:56:15 +00:00
glm-5.3-flash ae372c0c7e test+docs: prepublish hardening from review (failure paths, doc hygiene)
Coverage:
- CF-002: demux skipped-bytes budget teardown test (268 MiB skip in-memory;
  a budgetless demux wedges in the 17th skip, the real one tears down) and
  the skip-hits-EOF arm (truncated oversized payload ends the loop).
- CF-001: retryable CONNECTION_CLOSED pinned for subscribe write failures
  in both stream modes and single-stream publish request-frame failures;
  non-retryable INTERNAL pinned for mid-publish failures in both modes
  (deterministic FailOnFlushN write half).
- Gateway: invoke_sink with with_deadline(None) completes a slow sink.

Docs:
- Fix broken intra-doc link on lib.rs's feature-gated gateway mention
  (rustdoc warned on default-feature builds).
- Re-point 40 src/ references from the old alknet mono-repo ADR numbering
  (049/050/052/065/070/074/092) to this crate's numbering
  (021/011/034/007/008/009/005); drop into_sub_streams references
  removed by ADR-035.

Verification: 565 default / 582 all-features (8 new), clippy -D warnings
on default/gateway/all-features/wasm32, fmt clean, rustdoc warning-free
on default and all-features, publish dry-run clean. Consumer-facing API
continuity 0.1.1 -> 0.2.0 verified by compiling an API-surface probe
against both versions.
2026-08-31 09:56:07 +00:00
glm-5.3-flash d5fd548b8d feat: promote dispatch spine to gateway module (ADR-048, feature-gated)
Promote alkhttp's transport-neutral dispatch spine into alkcall as
alkcall::gateway behind the opt-in gateway cargo feature (default off;
adds no dependencies):

- GatewayDispatch: deadline-bounded invoke spine over OperationRegistry
  (invoke / invoke_streaming / invoke_sink) with the root-context
  discipline (internal: false, forwarded_for: None) hubs and spokes
  relaying calls (ADR-042 translate path) need identically to alkhttp's
  HTTP gateway. The 30 s deadline becomes a constructor knob
  (with_deadline).
- schema_disclosure_denial: the shared is-internal + ACL check for
  services/schema inner-name disclosure; ACL denial returns FORBIDDEN
  (identity-aware refinement), Internal visibility returns spec-404.
  One implementation so transports cannot drift (CF-004).
- MAX_BATCH_OPERATIONS / CallRequest / HTTP error mapping stay in
  alkhttp (projection + transport concerns); alkhttp migrates to this
  module in a follow-up session and drops its local copy.

Docs: ADR-048 (decision + divergence rationale), ADR index entry,
CHANGELOG.

Verification: 574 tests pass with --features gateway (16 new), 558 pass
default, clippy -D warnings clean both feature sets, --all-features
clean, fmt clean, wasm32 target clean, rustdoc warning-free.
2026-08-31 08:45:57 +00:00
glm-5.3-flash 8cb2a6eb6d fix: remediate consumer findings CF-001..004 (alkhttp ledger)
- CF-004: services_schema_handler now applies the same visibility +
  AccessControl gates as invoke() (identity resolution mirrors invoke:
  handler_identity under internal). Restricted ops return spec-404 NOT_FOUND
  — matches "restricted ops don't exist" and leaks nothing about the
  restricted surface. Closes the unauthenticated /call-path disclosure.
- CF-003: publish_schema compiled at registration time (both
  OperationRegistry::register and OperationRegistryBuilder::store);
  un-compilable schemas are a registration error — an unvalidated ingest
  path can no longer be constructed. Compiled validator cached per-op
  (publish_validator) and consumed by dispatch; per-request compile gone.
  BEHAVIOR CHANGE: register/builder reject un-compilable publish_schema.
- CF-002: demux TooLarge skip streams through a fixed 64 KiB buffer
  instead of allocating the peer-declared length (u32, up to ~4 GiB);
  cumulative 256 MiB skipped-bytes budget tears down dribbling peers.
  Existing resync test passes unchanged.
- CF-001: new retryable CallError::connection_closed (CONNECTION_CLOSED)
  applied only where the call is provably undelivered — request-frame
  write failures on all consumer paths (call/subscribe/publish, both
  stream modes; publish pump tags write stages). Mid-publish failures and
  producer-side fail_all stay non-retryable INTERNAL (delivery ambiguous).
  New code string is additive; retryable flag is the machine-readable
  signal.

Verification: cargo test (558 pass, 15 new), clippy --all-targets -D
warnings, fmt --check, wasm32-unknown-unknown check.
2026-08-31 08:22:55 +00:00
glm-5.3-flash a2d72f9737 docs(ledger): file CF-004 — services_schema_handler discloses Internal/ACL-restricted op specs
Found via alkhttp Review 002 (PRJ-16): the services/schema handler
does a bare registry.registration(name) with no Visibility and no
AccessControl check, so POST /call (and the MCP call tool) can fetch
any Internal op's complete spec unauthenticated. The GET /schema route
and MCP schema tool enforce the pre-checks; the /call path is the
hole.
2026-08-30 10:50:28 +00:00
glm-5.3-flash 1f08f1e385 docs(ledger): file CF-003 — wire-path publish_schema compile failure is fail-open
Found while fixing alkhttp's HTTP-side instance
(review-001-publish-schema-validation-robust): the identical
warn-and-skip pattern exists at src/protocol/dispatch.rs:352-366 — a
compile failure of a Pub op's publish_schema proceeds with
validator: None, so arbitrary unvalidated JSON reaches the sink
handler over the wire. Suggested direction: fail-closed + per-
registration validator cache (worked shape in alkhttp 1572a9d,
src/gateway/schema_cache.rs) + registration-time schema compilation.
Single compile site verified (both pump arms consume the one
InFlightSink.publish_validator).
2026-08-30 08:33:29 +00:00
glm-5.3-flash 84fe94c4d7 docs(ledger): file CF-002 — demux TooLarge skip allocates peer-declared length
Cross-crate finding from alkhttp Review 001 (WS-12), filed here per the
ledger's purpose (consumer-surfaced alkcall findings).

src/channels/adapter.rs:144: the ChunkError::TooLarge arm allocates
vec![0u8; length] from the peer's untrusted 8-byte header before
reading; length is u32, so ~4 GiB can be pinned per connection and held
indefinitely by a dribbling peer. Reachable via the alkhttp WS path by
any authenticated browser. Skip/resync logic is correct; the memory
shape is wrong — stream-skip with a bounded buffer instead.

The normal payload arm (:159) is safe (parse_header bounds it at
MAX_CHUNK_LEN); only the TooLarge arm is unbounded.
2026-08-30 06:16:30 +00:00
glm-5.3-flash 4c99b877e3 docs(reviews): consumer-findings ledger for alkhttp-as-consumer findings (CF-001 write-failure retryability) 2026-08-29 09:43:17 +00:00
38 changed files with 6983 additions and 339 deletions

No files matched your search

+3 -2
View File
@@ -215,13 +215,14 @@ cargo semver-checks check-release # before a release
## Architecture Context
- `docs/architecture/` — the authoritative spec. Read it before
non-trivial changes. ADRs are numbered 001..047; OQs (open questions)
non-trivial changes. ADRs are numbered 001..048; OQs (open questions)
track resolved/deferred decisions.
- This crate unifies `alknet-call` and `alknet-channels` from the
alknet mono-repo (`/workspace/@alkdev/alknet`). The source
architecture docs were ported from
`/workspace/@alkdev/alknet/docs/architecture/` and renumbered as
alkcall ADRs (001..047). The ALPN strings (`alk/call`,
alkcall ADRs (001..048; ADR-048 was authored in this crate). The ALPN
strings (`alk/call`,
`alk/channels`) are wire-stable going forward (renamed from
`alknet/` to `alk/` in v0.1.1, before the first published consumer).
- Key ADRs that inform this crate's design:
+272
View File
@@ -4,6 +4,275 @@ All notable changes to this crate are documented here. The format is
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this crate adheres to [Semantic Versioning](https://semver.org/).
## [0.4.1] - 2026-09-05
Bug-fix release: chunks arriving for a not-yet-adopted channel are
parked instead of dropped (the open-op response / first-data race).
No API changes.
### Fixed
- **Early-arrival chunks for un-adopted channels are parked, not
dropped.** The connect side adopts a channel (installs local routing
state) only after the open-op response arrives, but the accept side's
`OpenHandler` can start pumping data the moment the channel opens —
the two race. Previously the connect side's demux dropped those
chunks (REQ-CH-04's lenient unknown-channel drop), silently losing
the first chunks of any push-first producer (a TTY backend's banner
or greeting, a sub protocol's initial frame). `route_payload` now
parks up to 64 payloads per unknown `channel_id` in a bounded
early-arrival buffer and `adopt_channel` drains them into the new
receiver in order; chunks beyond the cap drop with the pre-existing
debug log and `dropped_unknown_chunks` counter (the counter now means
"early-arrival overflow or genuinely unknown channel", not just the
latter). `clear_all` drops the parked buffers with the connection.
Surfaced by alktty's consumer end-to-end test (review #001 L3): the
session never resolved because the producer's first chunks (the
stdout sentinel + exit chunk for an immediately-resolving backend)
arrived before the adopt and were dropped.
## [0.4.0] - 2026-09-05
`input_schema` is now enforced at call time (the advertise-vs-enforce
leg ADR-016 promised: `INVALID_INPUT` is the registry's
schema-mismatch error code). Minor bump — behavioral break for any
caller that was passing schema-violating inputs to ops declaring a
non-trivial `input_schema` and relying on the validation being
documentation-only.
### Changed
- **`OperationSpec.input_schema` is enforced by the registry on every
dispatch.** Until now the schema was advertise-only: `services/schema`
disclosed it, the wire and gateway paths disclosed it, but no dispatch
entry point consulted it (the only schema alkcall actually enforced
was `publish_schema` per-chunk on Pub ops, P-03). The schema now
compiles once at registration (same fail-closed rule as
`publish_schema`/CF-003: an un-compilable schema is a registration
error, never a silently-skipped contract) and is checked after the
ACL gate in all three dispatch entry points — `invoke`,
`invoke_streaming`, and `invoke_sink` (via `resolve_sink_handler`,
preserving the P-08 single-source-of-truth property). Violations
return `CallError::invalid_input("input failed input_schema
validation")` with the input echoed in `details`. Raw-JSON-Schema
semantics (permissive on unknown keys) — adapters that want
closed-by-default enforcement keep doing their own hardening, as
alkhttp's `CompiledInputSchema` already does; the two checks compose.
Built-in ops are unaffected: alkcall's own specs declare the empty
object schema (accepts anything) and `services/schema` already
returned `INVALID_INPUT` for a missing `name` before its handler ran.
Motivated by alktty's code review #001 (L1): the channels open-op
wrapper hands the registry-checked `input` to the `OpenHandler` as
the authoritative params, which requires the registry to actually
validate it.
- `OperationRegistry` gained an `input_validator` cache (compiled at
registration, mirrored on `fork` and in `OperationRegistryBuilder`
like the publish validators); `input_validator()` is `pub(crate)` —
the public surface is unchanged.
## [0.3.1] - 2026-09-04
Bug-fix release: `services/list-peers` can now list peer-announced
ops. No API breaks — `OperationEnv` gains one defaulted trait method
(non-breaking for all implementors; `cargo semver-checks` 196 checks
pass against the 0.3.0 baseline).
### Fixed
- **`services/list-peers` shows each peer's operations**
(UP-03, surfaced by alkhttp's review 006). `PeerCompositeEnv`
overrode `peer_ids` only, so `peer_operations` fell to the trait
default (`Vec::new()`) and every peer listed with an empty
operations array — the ADR-022 amendment's "announced op is
discoverable via `services/list-peers`" promise never resolved on
the wire. ADR-030 prescribed the fix but it had never been ported
into alkcall; the existing `list-peers` unit tests passed because
they mock `peer_operations` with hand-rolled envs. Implemented per
ADR-030: `OperationEnv::list_operation_names` (default `Vec::new()`)
+ overrides on `OverlayOperationEnv` (overlay map keys),
`PeerCompositeEnv` (`peer_operations` delegates to the peer's
overlay; its own aggregate mirrors `contains()`: session +
connections + base), `LocalOperationEnv` (registry names), and
`ChannelsSessionEnv` (delegates to base). Gate:
`announced_op_is_discoverable_via_services_list_peers` exercises the
exact `compose_root_env` shape (announce via `op/register`, then
assert both the `peer_operations` probe and the `services/list-peers`
wire shape attribute the op to the peer) — verified load-bearing
(reverting the override fails the gate). ADR-030 status is now
Accepted with the UP-03 provenance note.
## [0.3.0] - 2026-09-04
The connect side can serve (two-way ops over one channels connection),
per-session fork registries, and the `op/register` bootstrap op — the
remediation of reviews 004 and 005
(`docs/reviews/004-*.md`, `docs/reviews/005-*.md`). Two
`OperationRegistry` methods changed their return types from borrowed to
owned (source-breaking for annotated call sites; minor bump per 0.x
semver rules) — verified with `cargo semver-checks` (196 checks pass
against the published 0.2.0 baseline; the return-type changes were
caught by manual diff review, as the tool has no lint for that
pattern).
### Added
- **`OperationRegistry::fork`** — deep-copies a registry (handlers,
provenance, composition authority, capabilities, cached
publish-schema validators) into an independently-mutable copy. The
per-session shape: fork the deployment's base registry, register the
session's ops on the fork, dispatch the session over the fork
(ADR-047 §4 amendment, 2026-09-03). Internally mutable, so a fork
shared as an `Arc` can receive registrations after the dispatcher
was built — `install_bootstrap_discovery` relies on this.
- **`OperationRegistryBuilder::from_registry`** — seeds a builder from
an existing registry's registrations (the fork surface expressed
through the builder; registration order is not preserved — HashMap
iteration).
- **Connect-side serving (two-way ops over one connection).** The
channels connect side's channel-0 read pump previously resolved
outbound responses only and silently dropped inbound
`call.requested` frames. New `ChannelClient::from_connection_with_serving`
takes `Option<ServingConfig>` (`registry` + `identity_provider`);
with `Some(..)` the read pump becomes the full-duplex serving loop
(`Dispatcher::serve_single_stream`), so a connected peer can call the
consumer's ops. `from_connection` keeps the resolution-only pump
(pure-consumer default). Serving is opt-in: the protocol is
symmetric, the API is explicit (ADR-022 amendment, 2026-09-03).
- **`Dispatcher::serve_single_stream` and `Dispatcher::dispatch_start`
+ `StartedDispatch`** — the shared serving-loop machinery. Both
single-stream loops (accept side and connect side) now run the same
concurrency model: Once invocations, Sub pumps, and sink response
writers are spawned, so same-connection nested composition resolves
concurrently instead of deadlocking until the 30 s sweeper (review
005 G-01). Handles are tracked and aborted at loop exit; no lock is
held across an await.
- **`registry::op_register` module — the `op/register` bootstrap op**
(review 004 F-05). A connected peer announces the ops it serves over
the wire: `OP_REGISTER_NAME` (`op/register`), `OpRegisterRequest`
(wire round-trip via `to_json`/`from_json`), `op_register_spec` /
`op_register_handler`. Announcements land in the connection overlay
and are served through forwarding handlers, so hub→consumer import
(`from_call`) and consumer→hub announce are symmetric. **Collision
policy** (review 005 G-03): a peer-announced op may replace other
peer-announced ops (when `replace` is set) but never the serving
side's own registrations — a name present on the serving registry
rejects with the new `CallError::already_exists` (`ALREADY_EXISTS`,
non-retryable), regardless of `replace`. The collision gate checks
the serving registry, the session fork, and the connection overlay
itself.
- **`install_bootstrap_discovery`** (review 004 F-06) — registers
`services/list`, `services/list-peers`, and `services/schema`
against a registry with handlers closed over that same `Arc`, so
per-session forks are the discovery source for their own openables
(handlers see every op the registry serves at call time). ACL
filtering stays per-caller. The bootstrap-op set on channel 0 is
closed and now includes `services/list-peers` (review 005 G-05 —
doc alignment; the code had installed it since the amendment landed)
plus `op/register` (ADR-022 amendment).
- **`spec_to_json_pub`** — public serialization of an `OperationSpec`
into the `services/schema` wire shape (the shape `op/register`
announces with and `services/schema` serves). **`resource_id_path`
now survives the spec wire round-trip** (review 005 G-04): both the
serializer and the `op/register` parser carry the field.
- **`CallConnection::overlay_contains` /
`CallConnection::overlay_registration`** — read access to the
connection overlay, for `op/register` replace semantics and
composition reachability checks.
### Changed
- **`OperationRegistry::registration` returns
`Option<HandlerRegistration>`** (was `Option<&HandlerRegistration>`)
and **`list_operations` returns `Vec<OperationSpec>`** (was
`Vec<&OperationSpec>`). Source-breaking for call sites that annotate
the borrowed types — clone-at-the-boundary instead. Minor bump per
0.x semver rules.
- **Registration is `&self` throughout** — `OperationRegistry::register`,
`ChannelOperations::register_on`, and `ChannelCore::register_openable`
take `&OperationRegistry` (was `&mut`). Source-compatible for
callers (reborrow); this is what makes post-construction
registration on a forked, `Arc`-shared registry possible.
- **`from_call` composition reachability** — the forwarding-handler
path declares reachable namespaces on the composing handler's
`scoped_env` (empty is deny-by-default) and the connection overlay
is attached only for identity-carrying connections (ADR-030 §5).
## [0.2.0] - 2026-08-31
Consumer-findings remediation (CF-001..004 from
`docs/reviews/consumer-findings-ledger.md` — alkhttp as the first real
consumer), the promoted `gateway` dispatch spine (ADR-048), and a
docs-hygiene pass. One behavior change noted below; otherwise additive.
### Added
- **`gateway` feature: the transport-neutral dispatch spine**
(ADR-048). New `alkcall::gateway` module behind the opt-in `gateway`
cargo feature (default off; adds no dependencies). `GatewayDispatch`
is the deadline-bounded, re-rooted-context invoke spine over
`OperationRegistry` (`invoke` / `invoke_streaming` / `invoke_sink`)
promoted from alkhttp's gateway after it proved transport-agnostic —
hubs and spokes relaying calls (ADR-042 translate path) need the
identical root-context discipline (`internal: false`,
`forwarded_for: None`) without any HTTP. `schema_disclosure_denial`
is the shared is-internal + ACL check for `services/schema`
inner-op-name disclosure (one implementation so transports cannot
drift; ACL denial returns `FORBIDDEN`, Internal visibility returns
spec-404 — see ADR-048 for the split from the wire handler's
conservative spec-404). The handler deadline is a constructor knob
(`with_deadline`); the default remains 30 s. alkhttp migrates to
this module in a follow-up and drops its local copy.
### Changed
- **Un-compilable `publish_schema` values are rejected at registration
time** (CF-003). `OperationRegistry::register` and every
`OperationRegistryBuilder` method now return `Err` when a Pub op's
`publish_schema` fails to compile. Previously the bad schema was
accepted and the dispatch path failed **open** (warn log, chunks
unvalidated). The compiled validator is now cached per-op
(`OperationRegistry::publish_validator`) and consumed by the dispatch
path — the per-request compile is gone. Code that registered
un-compilable schemas will now get a registration error instead of a
silently-unvalidated op.
### Fixed
- **`services/schema` no longer discloses Internal or ACL-restricted op
specs** (CF-004). The schema handler now applies the same gates as
`invoke()` — Internal-visibility rejection and `AccessControl::check`
with matching identity resolution. Restricted ops return spec-404
(`NOT_FOUND`), consistent with "restricted ops don't exist" elsewhere;
no information leaks about the restricted surface. The gate is inside
the handler, so every transport is covered.
- **Demux `TooLarge` skip no longer allocates from the peer's header**
(CF-002). The skip now streams through a fixed 64 KiB buffer, and a
cumulative 256 MiB skipped-bytes budget tears down connections that
loop oversized headers. Previously the buffer was sized from the
untrusted `u32` length (up to ~4 GiB pinned per dribbling peer).
- **Write failures before request delivery are retryable** (CF-001).
New `CallError::connection_closed` (`CONNECTION_CLOSED`, `retryable:
true`) is returned when the `call.requested` frame write fails on any
consumer path (`call`/`subscribe`/`publish`, both stream modes) — the
call provably never reached the producer, so reconnect/retry is safe.
Mid-publish write failures and producer-side connection-teardown
failures remain non-retryable `INTERNAL` (delivery ambiguous). The
new code string is additive; the `retryable` flag is the
machine-readable signal for consumers.
### Fixed (continued)
- **Doc hygiene.** Resolved the broken intra-doc link on `lib.rs`'s
feature-gated `gateway` mention (rustdoc warned on default-feature
builds), and re-pointed 40 `src/` doc references from the old alknet
mono-repo ADR numbering (ADR-049/050/052/065/070/074/092) to this
crate's numbering (ADR-021/011/034/007/008/009/005) so published docs
reference ADRs this crate actually carries. Failure-path coverage was
added alongside: the CF-002 skipped-bytes teardown and skip-EOF arms,
CF-001's subscribe write-failure sites (both stream modes), the
publish request-frame failure in single-stream mode, and
non-retryability pinning for mid-publish failures.
## [0.1.1] - 2026-08-17
A minor release that renames the ALPN prefix from `alknet/` to `alk/`
@@ -61,5 +330,8 @@ Vendored core types (`Connection`, `ProtocolHandler`, `BiStream`,
(ADR-046), the channels protocol with openable-ALPNs-as-operations
(ADR-047), and the `ChannelClient` transport-agnostic client.
[0.3.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.3.1
[0.3.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.3.0
[0.2.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.2.0
[0.1.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.1.1
[0.1.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.1.0
Generated
+1 -1
View File
@@ -27,7 +27,7 @@ dependencies = [
[[package]]
name = "alkcall"
version = "0.1.1"
version = "0.4.1"
dependencies = [
"async-trait",
"bytes",
+3 -2
View File
@@ -1,6 +1,6 @@
[package]
name = "alkcall"
version = "0.1.1"
version = "0.4.1"
edition = "2021"
rust-version = "1.85"
license = "MIT OR Apache-2.0"
@@ -9,13 +9,14 @@ readme = "README.md"
repository = "https://git.alk.dev/alkdev/alkcall"
keywords = ["rpc", "json-rpc", "multiplexing", "wire-format", "alpn"]
categories = ["network-programming", "asynchronous", "encoding"]
exclude = [".opencode/", "AGENTS.md", "docs/reviews/", "docs/sdd_process.md", "Cargo.lock"]
exclude = [".opencode/", "AGENTS.md", "docs/reviews/", "docs/sdd_process.md"]
[lib]
name = "alkcall"
[features]
default = []
gateway = []
[dependencies]
jsonschema = { version = "0.46", default-features = false }
+30 -2
View File
@@ -21,7 +21,7 @@ use alkcall::registry::{
spec::{OperationSpec, OperationType, Visibility, AccessControl},
};
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry.register(HandlerRegistration::new(
OperationSpec::new(
"echo/run",
@@ -98,6 +98,34 @@ for reg in registrations {
let response = conn.call("remote/status", serde_json::json!({})).await;
```
### Serving your own ops as a connected consumer
The call protocol is symmetric — both sides of a connection can serve
ops. A `ChannelClient` built with `from_connection` is a pure consumer
(inbound `call.requested` frames are dropped); pass a `ServingConfig`
to also serve your registry to the peer, and use `op/register` to
announce which ops you serve:
```rust
use std::sync::Arc;
use alkcall::channels::client::{ChannelClient, ServingConfig};
use alkcall::registry::discovery::install_bootstrap_discovery;
let registry = Arc::new(OperationRegistry::new());
// ... register your ops on the registry, then:
install_bootstrap_discovery(&registry)?;
let client = ChannelClient::from_connection_with_serving(
connection,
Some(ServingConfig {
registry: Arc::clone(&registry),
identity_provider: provider,
}),
).await?;
// peer-callable ops resolve against `registry` on channel 0;
// `client.call_open_op` still works — both directions share the pump
```
## Architecture
alkcall is a pure protocol crate — no networking, no transport
@@ -107,7 +135,7 @@ Downstream crates compose on top of it in a layered dependency chain.
| Role | Call protocol | Channels protocol |
|------|---------------|-------------------|
| **Producer** | Registers ops on an `OperationRegistry`, runs a `Dispatcher` | Runs a `ChannelsAdapter`, registers openable ALPNs via `ChannelCore::register_openable` |
| **Consumer** | Uses `CallConnection` to call ops, uses `from_call` to discover/import remote ops | Uses `ChannelClient` to open channels via `call_open_op` + `open_channel` |
| **Consumer** | Uses `CallConnection` to call ops, uses `from_call` to discover/import remote ops; may also serve its own ops (`from_connection_with_serving`) | Uses `ChannelClient` to open channels via `call_open_op` + `open_channel` |
| **Hub** | Both: runs a `Dispatcher` for ops it produces, holds `CallConnection`s to spokes for ops it consumes | Both: runs a `ChannelsAdapter` for inbound connections, holds `ChannelClient`s to spokes |
| **Spoke / Worker** | Both: produces ops (its own services), consumes hub ops | Both: produces channels (TTY, tunnel), may consume hub channels |
+1
View File
@@ -100,6 +100,7 @@ are wire-stable and unchanged — see ADR-004.
| [045](decisions/045-alknetclient-native-dial-seam.md) | AlknetClient Dial Seam | spawn_dispatch / from_connection take-over; dial in consumer |
| [046](decisions/046-publish-operation-type-and-handler-kind-sink.md) | Publish Operation Type and HandlerKind::Sink | `OperationType::Pub` (producer→consumer streaming); `SinkHandler` + `HandlerKind::Sink`; `call.published` wire event; `invoke_sink()` dispatch; `Subscription` renamed to `Sub` |
| [047](decisions/047-openable-alpns-are-operations.md) | Openable ALPNs Are Operations | `channel/open` dissolves into per-ALPN ops `channels/<alpn>/sub`/`pub`; `channel_open` marker on `OperationSpec`; `ChannelCore` wrapper; extension-trait `ChannelOperationEnv`; connection-owner allocates `channel_id`; opener ledger (Gap 2 fix); ALPNs are call apps |
| [048](decisions/048-dispatch-spine-gateway-module.md) | Dispatch Spine (feature-gated `gateway` module) | `alkcall::gateway` behind the `gateway` feature; `GatewayDispatch` invoke spine (deadline knob, re-rooted context) + `schema_disclosure_denial` (FORBIDDEN for ACL deny, spec-404 for Internal); promoted from alkhttp for hub/spoke reuse |
## Relevant Open Questions
@@ -2,7 +2,7 @@
## Status
Accepted (amended 2026-06-26, 2026-07-13, and 2026-07-16 — see "Amendments" below; the 2026-07-16 amendment per ADR-045 §5 removes `CallClient::connect`)
Accepted (amended 2026-06-26, 2026-07-13, and 2026-07-16 — see "Amendments" below; the 2026-07-16 amendment per ADR-045 §5 removes `CallClient::connect`; amendment 2026-09-03 — the bootstrap-op set and the connect-side serving loop, see "Amendment (2026-09-03)" below; amendment 2026-09-04 — the `op/register` collision policy, in that amendment's "Collision policy" paragraph)
## Context
@@ -359,6 +359,123 @@ same as `from_openapi` receives HTTP credentials.
prior art
- POC at `/workspace/@alkdev/dispatch` — head/worker dispatch over SSH+axum
## Amendment (2026-09-03): bootstrap-op set and the connect-side serving loop
Review 004 (F-04/F-05) verified two gaps between this ADR's
bidirectionality promise (§2 "connection direction is independent of
call direction") and the single-stream channel-0 implementation:
1. **The connect side's read pump resolved responses only** — inbound
`call.requested` frames were silently dropped. Serving existed only
on the accept side (`run_loop_single_stream`).
2. **No wire mechanism announced client-side ops** — `from_call`
imports hub ops into the consumer, but a connected peer had no way
to announce "here are the ops I serve" over the wire.
The amendment (implemented in alkcall; the e2e gates are
`serving_loop_hub_to_consumer_call_resolves` and
`op_register_announce_then_hub_call_routes_back_to_consumer` in
`src/channels/client.rs`):
### The bootstrap-op set (one-way)
Each side of a channels connection **may serve** the bootstrap ops on
channel 0. The set is closed:
- `services/list` — discovery (the op `from_call` dials on every
import; each side is expected to serve it)
- `services/schema` — per-op schema disclosure
- `services/list-peers` — peer-keyed discovery (the op that makes
peer-announced ops discoverable; `from_call`-import discovery
relies on it). Added to this list 2026-09-04 (review 005 G-05) —
`install_bootstrap_discovery` has installed it since the amendment
landed; the set was declared closed and the doc lagged the code.
- `op/register` — peer op announcement (below)
These are the ops a peer may assume are reachable (subject to each
op's `AccessControl`); a peer that does not serve one answers
`NOT_FOUND` and the caller treats it accordingly (`from_call` already
surfaces discovery failure as `AdapterError::DiscoveryFailed`).
### Connect-side serving is opt-in (two-way door at the API level)
`ChannelClient::from_connection_with_serving(connection,
Option<ServingConfig>)` — with `None` (the pure-consumer default) the
read pump resolves outbound pendings only (previous behavior). With
`Some(ServingConfig { registry, identity_provider })` the read pump
becomes the full-duplex serving loop (`Dispatcher::serve_single_stream`):
inbound `call.requested` frames dispatch against the configured
registry and resolve back to the peer; outbound pendings still resolve
in the same loop. Serving is opt-in because a pure consumer has no
registry to serve; the *protocol* is symmetric, the *API* is explicit.
Direction disambiguation in the loop is by table membership, not
framing: an id that is one of *our* outbound pendings resolves there;
an id that is an inbound request dispatches; `call.aborted` tries both
tables (in-flight sink aborts and the pending map's cascade). IDs are
UUID-generated per side, so cross-correlation is not a hazard.
### `op/register` (one-way in wire shape)
The peer→hub direction of `from_call`'s bundle flow: a peer sends
`call.requested` for `op/register` carrying the `OperationSpec` in the
`services/schema` wire shape (`spec_to_json`) plus a `replace` flag.
The serving-side handler (`registry::op_register::op_register_handler`):
1. Rebuilds the spec (the same parser `from_call` uses — one spec
serialization on the wire).
2. Wraps a **call-forwarding handler** that issues a nested
`call.requested` back over channel 0 to the announcing peer (the
same shape `from_call`'s imported bundles use — the in-process twin
at `protocol/adapter.rs`).
3. Writes the bundle into **that connection's overlay** via
`register_imported` with `FromCall` provenance and
`Visibility::Internal` (composition material, ADR-017 — never
directly callable from the serving side's own wire).
The announced op is discoverable via `services/list-peers` (the
overlay is peer-keyed, `compose_root_env` attaches it) and invocable
via nested composition (`env.invoke`). Announced Sub/Pub ops register
as stubs that answer `INVALID_OPERATION_TYPE` — nested composition is
request/response-only (`OverlayOperationEnv`'s contract); the
streaming/sink forwarding shapes ride on the `from_call` import path.
Access control: `op/register` itself carries an `AccessControl` (an
unprivileged peer cannot reach the handler — the registry's normal
invoke path enforces it). Replace semantics: a collision with an
existing overlay registration is rejected with `ALREADY_EXISTS`
unless `replace: true` (the reconnect path re-announces). The
overlay dies with the connection (Layer 2), so reconnect re-announce
is naturally scoped.
Collision policy (amended 2026-09-04, review 005 G-03): a
peer-announced op may collide with other *peer-announced* ops on the
same connection (`replace` governs) but **never** with the serving
side's own registrations — a name present on the serving registry
rejects with `ALREADY_EXISTS` regardless of `replace`. The connection
overlay shadows the base registry in `PeerCompositeEnv` (connections
resolve before base, ADR-024 §1), so an unscreened same-name announce
would silently rewrite what a wire-dispatched handler's
`ctx.env.invoke` resolves for any name the deployment registered:
composition authority (ADR-018) belongs to the composing handler's
deployer, not the connected peer. Visibility is irrelevant to this
gate (`Internal` ops are as shadowable as `External` —
`OverlayOperationEnv` gates on `AccessControl`, not visibility; the
composed child is `internal: true` by design).
The envelope kind set stays closed at six — bootstrap ops over channel
0 are the door (AGENTS.md §7 allows adding kinds; none is needed).
### Cross-references
- Review 004 (`docs/reviews/004-per-connection-dispatch-and-client-serving-review.md`)
F-04/F-05 — the verification and the design decision
- ADR-047 §4 amendment #2 (2026-09-03) — the per-session fork the
bootstrap ops compose on
- alkhttp OQ-05 / ADR-048 — the browser data-channel wiring this
unblocks (the alkhttp-side gap was wiring; these two mechanisms are
what it wires to)
## Amendments (2026-06-26)
This ADR left four decisions as two-way doors (§1 Consequences flagged DC-1's
@@ -2,7 +2,15 @@
## Status
Proposed
Accepted (implemented 2026-09-04 — surfaced as UP-03 in alkhttp's
review 006 `docs/…/006-alkcall-0.3.0-consequence-review.md`: the
`op/register` amendment's "announced op is discoverable via
`services/list-peers`" promise did not resolve on the wire because
this override had never been ported into alkcall; the gate is
`announced_op_is_discoverable_via_services_list_peers` in
`src/registry/op_register.rs`. The `services/list-peers` unit tests
did not catch it because they mock `peer_operations` with hand-rolled
envs.)
## Context
@@ -5,7 +5,99 @@
Accepted (amends ADR-037; refines ADR-044, ADR-046; §4 amended
2026-08-13 — open ops are registered per-connection, not resolved via
`context.env` downcast — see "Amendment (§4 per-connection
registration, 2026-08-13)" below)
registration, 2026-08-13)"; amendment #2 (2026-09-03) — the
per-connection registration mechanism is the **per-session fork of the
base registry installed as the session's dispatch registry**, not the
connection overlay — see "Amendment (§4 mechanism, 2026-09-03)" below)
## Amendment (§4 mechanism, 2026-09-03)
The 2026-08-13 amendment named the registration target as "the
connection overlay registry (Layer 2 per ADR-019)". Review 004 (F-01)
verified that this shape cannot dispatch: the top-level dispatch path
(`Dispatcher::dispatch` / `run_loop_single_stream`) resolves and
invokes against the dispatcher's **base registry only**; the
connection overlay is reachable solely as a layer of `context.env`
(for nested invocations — a handler calling `env.invoke(...)), never
for resolving the incoming `call.requested` itself. An open op
registered on the overlay resolves `NOT_FOUND` on the wire.
The one shape proven end-to-end (alkcall's own e2e gate) is different:
the `install_channel_zero` hook builds a **fresh per-connection
registry containing the open op and passes it as the dispatcher's base
registry**. This amendment makes that the operative mechanism.
**The decision: per-connection registration happens on a fork of the
deployment's base registry, installed as the session's dispatch
registry.** The `install_channel_zero` hook (and any future
session-establishment seam):
1. **Forks** the deployment's base registry
(`OperationRegistry::fork` — a deep copy carrying handlers,
provenance, composition authority, capabilities, and the cached
publish-schema validators; review 004 F-02/F-03).
2. **Registers the per-session ops on the fork** — the generic channel
ops (`ChannelOperations::register_on`), the openables
(`ChannelCore::register_openable`), and the bootstrap discovery ops
(`install_bootstrap_discovery`, closed over the fork itself so
`services/list` sees the fork's per-session ops — review 004 F-06).
3. **Dispatches channel 0 over the fork** (`Dispatcher::new(fork,
...)`).
The fork is possible because `OperationRegistry` is internally
mutable (`parking_lot::RwLock` around both maps) — a fork shared as an
`Arc<OperationRegistry>` can receive bootstrap ops after the
dispatcher was built, and the self-referential discovery closure sees
every post-install registration.
The **connection overlay (Layer 2) remains what ADR-019/ADR-024
describe**: the landing zone for peer-announced ops (`op/register`,
review 004 F-05 — ADR-022 amendment) and the nested-invocation target
for imported ops. It is not the dispatch-resolution path for the
session's own ops.
Rationale for the fork shape over an overlay-aware dispatch fallback
(F-02 option (b)): the fork is the only shape with an end-to-end
proof, it needs no change to the shared dispatch loop, and it keeps
the overlay's `invoke_with_policy` shape (namespace-scoped,
parent-context-driven — built for nested composition) out of the
top-level call path, where it does not match the frame-handling
contract.
This preserves every invariant the 2026-08-13 amendment protected:
- **Layering (ADR-044):** unchanged — the open-op wrapper is in
`channels-call`; the call crate's `OperationRegistry` gains only
`fork` (and interior mutability), no channels types.
- **Per-connection resolution:** the open op gets the *right*
`ChannelManager` because the fork is built per-connection and its
openable closes over that connection's `ChannelCore`.
- **"Marked ops invoked outside a channels session" (ADR-047 §2):**
unchanged in effect — a `channels/<alpn>/sub` op registered only on
a session fork is not reachable on a bare `alk/call` connection (the
fork isn't that session's dispatch registry) — the dispatch path
returns `NOT_FOUND`.
### Door type
**Two-way (implementation detail), as before.** The registration
*target* mechanism (fork as base registry) sits within the same
wrapper-shape detail the 2026-08-13 amendment already marked two-way.
The one-way decisions (per-ALPN op names, the `channel_open` marker,
removal of `channel/open`/`direction`) are unchanged.
### References
- Review 004 F-01/F-02/F-03/F-06
(`docs/reviews/004-per-connection-dispatch-and-client-serving-review.md`)
— the verification and the mechanism decision
- ADR-022 amendment (2026-09-03) — the bootstrap-op set (`services/list`,
`services/schema`, `op/register`) and the connect-side serving loop
- ADR-019: operation registry layering (the overlay stays the nested
invocation / peer-announced-ops landing zone)
- The e2e gate: `fork_registry_open_op_resolves_and_is_discoverable`
(`src/channels/client.rs`) — open op resolves through the fork,
per-session openable in `services/list`, `services/schema` validates
## Amendment (§4 per-connection registration, 2026-08-13)
@@ -0,0 +1,151 @@
# ADR-048: Dispatch Spine (feature-gated `gateway` module)
## Status
Accepted
## Context
The first real downstream consumer of this crate — alkhttp — built its
HTTP gateway (`POST /call`, `/subscribe`, `/publish`, the MCP `call`
tool, the to_openapi projections) on a small internal component it
calls the **dispatch spine**: a thin struct over
`Arc<OperationRegistry>` that owns the *non-HTTP* half of gateway
dispatch. The HTTP layer resolves bearer tokens to an `Identity`,
frames NDJSON/SSE, maps `CallError` to HTTP statuses, and wraps axum
handlers; the spine does everything that happens after that:
- constructs the root `OperationContext` identically for every
transport (`internal: false` — ACL runs against the caller's
identity, not a handler's composition authority; `forwarded_for:
None` — wire-ingress only; identity supplied per-call),
- resolves the registration's `composition_authority` /
`capabilities` / `scoped_env` into that context,
- bounds Once-ops and sink dispatch with a deadline while leaving
streaming subscriptions unbounded (ADR-021: subscriptions are
long-lived),
- and applies the `services/schema` disclosure guard when the
dispatched operation's *input* names another operation.
This is transport-neutral work. It contains no HTTP concepts: no
statuses, no headers, no body framing. And it is not HTTP-shaped by
accident — the same shape is exactly what a **hub** needs when it
terminates channel 0 on both legs and relays calls to spokes
(ADR-042's "translate, not forward" rule): the hub must re-root the
context at itself (its own identity, `internal: false`, no
`forwarded_for`), re-resolve its capabilities for the outbound leg, and
bound the relay so a hung spoke does not wedge the browser-facing
connection. alkhttp needed it; the hub relay needs the same thing;
any protocol crate that exposes a call surface to a less-trusted
in-transport caller (a WS-native relay, a CLI bridge, a test harness)
will need it again.
Duplicating it per consumer is the failure mode this crate already
paid for once: the spine's `services/schema` guard existed because the
registry handler and the HTTP route were written against different
disclosure rules (CF-004, filed from alkhttp's consumer review). One
shared implementation is the fix; a second copy in a second crate
would re-open the drift.
alkhttp is not yet published. This is the cheapest moment to move the
component into alkcall (behind a feature, so the base crate stays lean
and the surface is opt-in) and have alkhttp consume it rather than
carry its own copy.
## Decision
Promote the dispatch spine into alkcall as a new `gateway` module,
**feature-gated**:
- `alkcall/gateway` behind the `gateway` cargo feature (default off —
the base crate stays lean; the module adds no dependencies, the gate
exists to keep the audit surface explicit and opt-in).
- `GatewayDispatch` — the spine struct: `invoke()`,
`invoke_streaming()`, `invoke_sink()`, registry access, and root
context construction. The 30 s deadline alkhttp hardcodes becomes a
constructor knob (`with_deadline`) so consumers keep their own
policy; the default remains 30 s for drop-in equivalence.
- `schema_disclosure_denial()` — the shared check that a spec the
caller could not invoke is not disclosed: `NOT_FOUND` for
Internal-visibility ops, **`FORBIDDEN` for ACL-denied ops**.
- `CallRequest` stays in alkhttp (payload framing is transport
business); `MAX_BATCH_OPERATIONS` stays in alkhttp (batch is a
projection concern).
### ACL denial returns FORBIDDEN, not spec-404
The wire-path `services/schema` handler (CF-004, discovery.rs) returns
spec-404 for both Internal visibility and ACL denial — the right
answer for an unauthenticated wire caller, where even acknowledging
the op's existence is a leak vector. The spine's guard is invoked by a
transport that has *already* resolved the caller's identity and often
already admitted the op elsewhere (the alkhttp GET `/schema` route
answers `403` for ACL-denied ops, deliberately). The spine therefore
keeps alkhttp's split:
- Internal visibility → `NOT_FOUND` (never acknowledged, at any
authority).
- ACL denial → `FORBIDDEN` (informative for a caller whose identity
the transport resolved; alkcall's own registry `invoke()` produces
the same code for the same caller state, so the guard's answer is
never *more* restrictive or *less* restrictive than the invoke that
would follow it).
The two layers are consistent by construction: the wire handler's
spec-404 is the conservative outer bound, the spine's FORBIDDEN is the
identity-aware refinement. When an unauthenticated caller hits both,
`FORBIDDEN` and `NOT_FOUND` differ only in which exists — and
`AccessControl::check` with no identity returns
`"authentication required"`, which HTTP-side mappers translate to 401.
alkhttp consumes this module in a later session and drops its local
copy; until then the two implementations coexist (byte-identical in
behavior, one in each crate).
### What the spine deliberately does NOT include
- **No HTTP mapping.** `CallError` → status/body is the consumer's
(alkhttp `gateway::error`). The neutral wire error is the module's
lowest-level vocabulary.
- **No body framing, limits, or timeouts beyond the deadline knob.**
NDJSON/SSE framing, body caps, keep-alive intervals are transport
concerns (GW-15/GW-16 in alkhttp's ledger).
- **No batch semantics.** `MAX_BATCH_OPERATIONS` and the batch
envelope are projections of the registry onto HTTP/MCP payloads.
- **No `CallRequest` type.** The spine takes `(op, input)` — parsing
`{ operation, input }` is transport framing.
## Consequences
- The deadline policy moves from "alkhttp's 30 s HTTP convention" to
"spine policy, configured per-consumer" (default 30 s). Sink
dispatch is bounded by the same deadline as Once-ops — the wire
path's Pub dispatch is unbounded (`deadline: None`); consumers who
want the wire behavior pass `Duration::ZERO`-style no-deadline
configuration via `with_deadline(None)`. Documented divergence, a
choice the *consumer* now owns.
- alkhttp migrates to `alkcall::gateway` in a follow-up session; its
local spine is deleted then, not deprecated in place (unpublished
crate — no compat window needed).
- The feature adds no dependencies; `cargo test --no-default-features`
and `--all-features` both pass (the wasm-clean baseline is
untouched).
- Future transports (WS-native relays, protocol crates) reuse the
spine instead of re-deriving the context/guard discipline.
## Alternatives considered
- **Fold into `registry` as a method set on `OperationRegistry`.**
Rejected: the spine is a *policy* wrapper (deadline, root-context
shape, disclosure guard) a consumer opts into, not the registry's
dispatch core. `GatewayDispatch::invoke` ≠ `OperationRegistry::invoke`
— conflating them invites wire paths to pick up transport-shaped
policy accidentally.
- **Wait for alkhttp to publish first.** Rejected: promotes a
duplicate into the wild that immediately has to be deprecated; the
cheapest moment to move the type is before either crate's first
release.
- **Promote the whole gateway module.** Rejected: routes (SSE/NDJSON
framing, body limits) and error mapping (`IntoResponse`, status
mapping) are HTTP by definition; moving them would drag axum/hyper
as optional dependencies into a protocol crate for no reuse —
alkhttp is their only consumer.
@@ -0,0 +1,527 @@
# Review 004 — Per-Connection Dispatch Resolution and Client-Side Op Serving (found via alkhttp's WS data-channel drill-down)
## Status
Verified, remediated (2026-09-03 — Units 1–3; Unit 4 remains
downstream in alkhttp). See "Remediation log (2026-09-03)" at the
bottom for the landing summary and the gates.
## Scope
Focused review of two design/spec mismatches in the call + channels
integration, uncovered while drilling into the deferred half of
alkhttp's WebSocket path (alkhttp review 003, findings WS-24/WS-25 —
the OQ-05 "browser data channels" deferral). The drill-down
established that the alkhttp-side gap is mostly wiring, but two
assumptions underneath it resolve to *this* crate:
1. **Dispatch resolution for per-connection operations.** ADR-047's
§4 amendment (2026-08-13) decided that openable-ALPN open ops are
registered per-connection on the connection overlay registry
(Layer 2). The top-level dispatch path never consults that layer,
so an op registered per the amendment's mechanism cannot be
invoked over the wire. The only shape proven to dispatch is
different from the one the ADR describes.
2. **Client-side operation serving.** The connection-local overlay
(`CallConnection::register_imported` → `overlay_env()`) and the
session overlay (`compose_root_env` attaching peer-keyed overlays)
are fully built and tested — but only the *accept side* has a loop
that serves inbound `call.requested` frames. The connect side's
read pump resolves responses and silently drops requests. There
is also no wire mechanism by which a client could announce an op
in the first place. Together these leave the "both sides can be
both" symmetry (AGENTS.md §8; ADR-022 §direction semantics;
alkhttp ADR-048's browser bidirectionality) unreachable on the
single-stream channel-0 shape.
The pass re-verified every claim directly in source at tree
`c16b069`. Cross-repo context: these findings block alkhttp review
003's Unit 2 (data-channel wiring for WS sessions, both browser and
native-fallback consumers); the decision work and the fixes live
here. Findings continue alkcall review numbering (F-01..; reviews
001–003 used P/C/R/A/B/C/D prefixes — F is the new prefix for this
focused review, findings numbered independently).
## Baseline verification (this pass)
```
cargo test → 565 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo fmt --check → clean
```
The suite is green; these findings are design/spec mismatches, not
regressions. Notably, one of them (F-04's serving gap) is invisible
to the suite *because* the connect side has never been asked to
serve — the existing e2e tests all put the dispatcher on the accept
side.
## Verdict
- **The channel machinery is done and proven.** `ChannelCore` /
`register_openable` / `ChannelOperations`, the opener ledger, the
odd/even id split, the demux hardening, and the full e2e wiring
test (`channels/client.rs:767-870`) are all in place and tested.
Nothing in Part A below re-opens that machinery.
- **The gap is the dispatch-resolution *mechanism* (F-01/F-02/F-03)
and the connect-side serving half (F-04)** — both upstream of any
transport (TCP+TLS, WS, whatever). Fixing them here unblocks
alkhttp's data-channel wiring, and — since the crate is
wasm-targetable and not yet published to consumers — the fixes are
cheap now and required regardless of alkhttp.
- **The recommended resolution is assembly, not invention:** per-
session registry forks (F-02/F-03) + a bootstrap `op/register` op
(F-01's mechanism choice) + a connect-side serving loop (F-04).
Three of the four pieces have working reference shapes in-tree;
the genuinely new protocol surface is one op handler and one read
loop.
## Severity legend
- **[critical]** — a decided spec invariant is violated in a way that
makes a promised capability unreachable end-to-end; or corrupts
data.
- **[major]** — a core protocol path cannot serve a decided behavior;
works only via shapes the spec does not describe.
- **[minor]** — drift, doc/spec inconsistency, or a missing
convenience with no correctness impact.
---
# Part A — Dispatch resolution vs the ADR-047 §4 amendment
## F-01 [major] — Top-level dispatch never consults the connection overlay; ops registered per the ADR-047 §4 amendment mechanism resolve `NOT_FOUND` on the wire
**ADR drift:** ADR-047 "Amendment (§4 per-connection registration,
2026-08-13)": *"open ops are registered per-connection …
`register_openable` is called on the connection overlay registry
(Layer 2 per ADR-019), closing over the per-connection `ChannelCore`."*
**Verified:** YES, three ways:
1. `Dispatcher::dispatch` reads the operation's `op_type` from
`self.registry.registration(...)` only
(`src/protocol/dispatch.rs:316-320`) and invokes via
`self.registry.invoke` / `invoke_streaming` / `resolve_sink_handler`
(`:330-357`) — the dispatcher's **base registry** only. There is
no fallback to `connection.overlay_env()`.
2. The connection overlay is reachable solely as a layer of
`context.env`: `compose_root_env` attaches it via
`env.attach_peer(peer_id, connection.overlay_env())`
(`dispatch.rs:195-215`). `context.env` is consulted for *nested*
invocations (a handler calling `context.env.invoke(...)`) — not
for resolving the incoming `call.requested` itself. An open op
registered on the overlay resolves `NOT_FOUND` when the wire
request arrives.
3. The one place this wiring is proven end-to-end uses a different
shape than the amendment describes: alkcall's own e2e test
(`src/channels/client.rs:767-870`) has the `install_channel_zero`
hook build a **fresh per-connection registry containing the open
op and pass it as the dispatcher's base registry**
(`Dispatcher::new(registry, ...)` at `:837`) — not as an overlay.
That shape works; the amendment's shape cannot.
Consequence: ADR-047's data-plane promise (per-connection open ops
invocable on channel 0) is unreachable through the mechanism the
amendment names. Every consumer of `register_openable` must either
adopt the undocumented per-session-base-registry shape or fail.
**Fix (Unit 1):** decide the mechanism deliberately (see F-02) and
amend ADR-047 §4's wording to describe the mechanism that actually
dispatches. The amendment's *rationale* (per-connection state, no
`context.env` downcast) stands under either option.
## F-02 [major] — `OperationRegistry` has no fork/clone surface; the proven per-session-registry shape requires one
**Verified:** YES. `OperationRegistry`
(`src/registry/registration.rs:97-100`) is
`HashMap<String, HandlerRegistration>` + a validators map, with
`new`/`register`/`registration`/`publish_validator`/`list_operations`
(`:102-180`) — no `Clone` impl and no fork API. Every type inside a
`HandlerRegistration` *is* clone-able (verified):
- `HandlerKind` — `#[derive(Clone)]` (`registration.rs:51`)
- `OperationSpec` — `#[derive(Debug, Clone, PartialEq)]`
(`spec.rs:174-175`)
- `OperationProvenance` — `#[derive(Debug, Clone, Copy, PartialEq,
Eq)]` (`registration.rs:58-59`)
- `CompositionAuthority` — `#[derive(Debug, Clone)]`
(`context.rs:82-83`)
- `Capabilities` — manual `impl Clone` (`core/types.rs:75-80`)
- `jsonschema::Validator` — `#[derive(Clone, Debug)]`
(jsonschema 0.46.10 `validator.rs:294`)
So the fork is a small addition (derive `Clone` on
`HandlerRegistration` + `OperationRegistry`, or an explicit
`fork() -> OperationRegistry` that copies both maps). Two shape
options:
- **(a) per-session base registry (recommended).** The
`install_channel_zero` hook (and any future session-establishment
seam) forks the deployment's base registry, registers the
per-connection operations on the fork (generic channel ops, the
openable ALPNs, and — with F-05 — the bootstrap op), and dispatches
channel 0 over the fork. This is the proven shape; the ADR wording
moves from "overlay registry" to "per-connection registry installed
as the session's dispatch registry." Cost: `services/list` /
`services/schema` registered on the base registry must be
re-registered (or inherited via fork) per session to stay
discoverable — the fork makes that free if the bootstrap ops are
part of the fork source.
- **(b) overlay-aware top-level dispatch.** Add an overlay fallback
in `dispatch` / `run_loop_single_stream`: when the base registry
misses, consult `connection.overlay_env()`. Keeps one static
registry and makes the amendment's wording true as written, but
touches the shared dispatch loop (higher blast radius, and the
overlay's `invoke_with_policy` shape — namespace-scoped, parent-
context-driven — does not match the top-level call shape, so the
fallback needs care: op-type lookup, ACL, visibility).
Recommendation: (a). It is the only shape with an end-to-end proof,
it needs no dispatch-loop change, and it matches ADR-019's layering
(the overlay stays what it is today: the nested-invocation landing
zone for imported ops — which F-05 needs).
**Fix (Unit 2):** `Clone`/fork surface + the ADR wording amendment;
acceptance = the review-001-style e2e gate (open op resolves over a
live channels connection through the *fork*).
## F-03 [minor] — Per-session fork + static base: no `Default`/builder seam exists to compose "base ops + per-session ops" without re-registration
**Verified:** YES. Even with a `Clone` surface (F-02), the hook must
compose three sources into the per-session registry: the deployment's
base ops (already registered, not retained anywhere as a "source"
list), the generic channel ops (`ChannelOperations::register_on`),
and the openables. A fork of the base registry makes the first source
free — this finding exists only if (a) is chosen *without* the fork
(i.e., rebuilding per-session registries from scratch). Record it as
a constraint on the fork API: it must clone *registered handlers*
(`HandlerKind` clones carry their closures), not just specs, so
`services/list` keeps its closure over the base registry.
`publish_validator`'s `validators` map must be carried too (schema
enforcement would silently vanish otherwise).
**Fix:** fold into F-02's fork surface; the acceptance test covers it
(`services/schema` on the fork still validates input).
---
# Part B — The connect side cannot serve; no wire path announces ops
## F-04 [major] — The connect side's channel-0 read pump resolves responses only; inbound `call.requested` frames are silently dropped — there is no serving half on channel 0
**Verified:** YES. `ChannelClient::from_connection`
(`src/channels/client.rs:85-134`) spawns the read pump as
`read_single_stream_until_closed(single_stream_reader, &pending_map)`
(`connection.rs:681-686`), whose only job is resolving the
`PendingRequestMap` — its frame handler (`dispatch_envelope`,
`connection.rs:705-...`) branches on `call.responded` /
`call.completed` / `call.aborted` / `call.error` and **has no arm for
`call.requested`**. The accept side's mirror
(`Dispatcher::run_loop_single_stream`, `dispatch.rs:727-850`) serves
requests and has no pending-resolution arm. Each side silently drops
the other half's frames.
Consequences:
- A consumer that dials a producer via `ChannelClient` can call ops
and open channels, but if the producer later calls an op *the
consumer serves*, the request is dropped (the caller's pending
hangs until the sweeper or connection close). This is the exact
shape ADR-022 §connection-direction-independence promises ("both
sides can be both simultaneously" — AGENTS.md §8), and ADR-036's
channel-0 pre-negotiation makes channel 0 the single control
plane where it must work.
- It blocks the bootstrap registration story (F-05): the hub cannot
call anything the WS/native client serves, and the client cannot
even receive the call that would let it register.
The fix is well-understood and symmetric: a full-duplex
single-stream read loop that branches both ways — `call.requested`
→ dispatch and write a response frame (the `run_loop_single_stream`
arms), `call.responded`/`completed`/`aborted`/`published` → resolve
pendings and route sink chunks (the read-pump arms). IDs are
UUID-generated on each side (`generate_request_id()`), so
cross-correlation is not a hazard; the abort arm needs to try both
tables (in-flight sink aborts *and* the pending map's cascade), which
`run_loop_single_stream`'s `EVENT_ABORTED` arm already does within
its own scope. The wasm-targetable browser story compounds the
value: the same serving loop is what a wasm alkcall browser
compiles.
**Fix (Unit 3):** a serving-capable single-stream loop usable from
the connect side (e.g., a `Dispatcher` variant or a
`CallConnection::serve_single_stream` that composes both frame
directions), wired into `ChannelClient` behind an opt-in (a serving
registry is not always present — a pure consumer may have none).
Acceptance gates: hub→consumer call over an existing `ChannelClient`
session resolves; consumer-served op participates in ACL (its
`AccessControl` gates the hub's call); disconnect mid-call still
fails pendings on both sides.
## F-05 [major] — No wire mechanism announces client-side ops; ADR-022's import flow is hub→consumer only and the browser-native path has none
**Verified:** YES. The envelope kind set is closed at six —
`call.requested` / `call.responded` / `call.completed` /
`call.aborted` / `call.error` / `call.published`
(`src/protocol/wire.rs:12-17`). Op registration onto an overlay
(`register_imported` / `register_imported_all`,
`connection.rs:162-172`) happens only in-process: `from_call`
(`client/from_call.rs:82-91`) builds bundles the *local* process
registers after *it* dialed out — the consumer importing a hub's
ops. There is no op, frame, or handshake by which a connected peer
announces "here are the ops I serve." Every consumer-facing
registration surface today requires the registrant to hold the
local `CallConnection` handle in-process.
Consequence: the decided bidirectionality model — either side
registers ops the other can call (AGENTS.md §8; ADR-022 §direction
semantics; alkhttp ADR-048's connection-local overlay for
browser-registered ops) — has no on-the-wire expression for the
non-in-process side. It is exactly the assumed-bootstrap-op set the
protocol already leans on (`services/list` is dialed and called by
`from_call` on every import; each side is *expected* to serve it)
extended by one op: e.g. `op/register` carrying the
`HandlerRegistration`'s serializable parts (spec + provenance +
access control), whose hub-side per-session handler writes the
registration into that connection's overlay via
`register_imported`. The overlay is already the landing zone
(`compose_root_env` attaches it keyed by identity,
`dispatch.rs:211-213`), and discovery of what a peer registered is
already built: `services/list-peers` reads
`ctx.env.peer_ids()` / `peer_operations()`
(`registry/discovery.rs:260-307`) — the same overlay-populated env.
Scope decision this forces (record it in the ADR, either way):
- **Design it** (recommended): `op/register` (+ a deregister or
replace semantics for the reconnect path) as an assumed op in the
bootstrap set, served per-session; the wire stays six kinds. The
`HandlerRegistration`'s `Handler` closures cannot cross the wire —
the client sends spec + a call-forwarding contract, and the
hub-side handler wraps it as a forwarding handler that issues a
nested `call.requested` back over channel 0 (the same shape
`from_call`'s imported bundles use — `protocol/adapter.rs:541,610`
show the in-process twin).
- **Or scope-cut**: amend ADR-022/alkhttp ADR-048 to say peer-side
op registration is hub-initiated-import only until a contract
exists. Cheaper, but it re-opens the same hedge the deferral was
criticized for — the promise would remain written-but-unmechanized.
Recommendation: design it. The pieces are all present; the cost is
one op handler, an `AccessControl`-gated registration surface (the
registry's existing ACL path covers it — an op with restrictive
`AccessControl` cannot be overwritten by an unprivileged peer), and
the F-04 serving loop to make it reachable in both directions.
**Fix (Unit 3):** with F-04 — bootstrap op + client serving loop +
ADR-022 amendment naming the bootstrap op set (assumed ops each side
may serve: `services/list`, `services/schema`, and `op/register`).
## F-06 [minor] — `services/list` on a per-session fork needs a closed-over fork reference, or per-session ops are undiscoverable
**Verified:** YES. `services_list_handler` closes over a *specific*
`Arc<OperationRegistry>` (`registry/discovery.rs:233-258`) — it
lists that registry's ops. Under the F-02(a) shape, per-session
openables live on the fork, so a `services/list` handler closed over
the base registry cannot see them. Two clean outs, both cheap: (i)
fork *before* registering the bootstrap discovery ops (then the
handler closes over the fork), or (ii) the fork API registers
base-registry bootstrap ops fresh on the fork. Either way the
constraint is: discovery ops must be registered against the
per-session registry, not the static one. Also note `services/list`
on the fork still ACL-filters (the handler re-checks per-caller) —
no privilege regression.
**Fix:** fold into Unit 2 (the fork seam registers bootstrap ops on
the fork); acceptance: a per-session openable appears in
`services/list` for an authorized caller and not for an unauthorized
one.
---
# Non-findings (verified correct, recorded to bound the re-review)
- **The e2e reference shape is real and passes** —
`channels/client.rs:767-870` wires `ChannelClient` ↔
`ChannelsAdapter` over duplex with an open op registered on the
per-connection registry, invoked end-to-end, quota reserved. The
F-02(a) shape is not speculative; it is in-tree.
- **`register_openable`'s wrapper machinery is complete** — ACL via
the registry's normal invoke path, `check_open` → `open_channel` →
ledger → spawn → teardown decrement, with `Once`/`Stream`/`Sink`
arms (`channels/operations.rs:399-453, 483-555`). Nothing in F-01
re-opens it; the ops it registers just need a dispatch-reachable
home.
- **Channel-id allocation is collision-safe** for the bidirectional
open story (accept side even from 2, connect side odd from 1,
`manager.rs:105-125`; `adopt_channel` for the non-allocating side,
`manager.rs:273-309`).
- **The `call.*` envelope kind set staying closed is correct** —
AGENTS.md §7 allows *adding* kinds but F-05's resolution does not
need one; bootstrap ops over channel 0 are the lighter door.
- **Wasm-cleanliness is unaffected** — the recommended fixes (fork +
handler + read-loop branch) touch no transport, no fs/net tokio
features; `cargo check --target wasm32-unknown-unknown` remains
the release gate.
---
# Remediation plan
Sequenced by dependency. Units 1–3 are alkcall work; the alkhttp
consequences are tracked as alkhttp review 003 Unit 1 (updated to
point here).
## Unit 1 — Decision recording (F-01, F-02 direction)
ADR work only:
- ADR-047 §4 amendment #2: the open-op registration mechanism is the
**per-connection registry installed as the session's dispatch
registry** (option (a)); the overlay registry remains the landing
zone for peer-announced ops (F-05) and nested invocation.
- ADR-022 amendment: the bootstrap-op set (each side may serve
`services/list`, `services/schema`, `op/register`), and the
direction-semantics note that serving on the connect side is
opt-in (F-04).
- alkhttp ADR-048's bidirectionality promise is re-pointed at these
ADRs instead of standing as an unmechanized commitment.
Gate: ADRs recorded; alkhttp OQ-05 and ADR-048 cross-reference them.
## Unit 2 — Fork surface + per-session composition (F-02, F-03, F-06)
- `Clone` (or explicit fork) on `HandlerRegistration` +
`OperationRegistry`, carrying handlers and validators.
- A composition helper (e.g. `OperationRegistry::fork_with(...)`) or
documented pattern: fork base → register generic channel ops +
openables + bootstrap ops against the fork.
- Gate: e2e over a live channels connection — open op resolves
through the fork; `services/list` shows per-session openables
(ACL-filtered); `services/schema` on the fork still validates.
## Unit 3 — Client serving half + bootstrap registration (F-04, F-05)
- Serving-capable single-stream loop (compose the read-pump and
dispatch arms); wired into `ChannelClient` opt-in.
- `op/register` bootstrap op (per-session handler →
`register_imported` into the connection overlay), with
`AccessControl` gating and reconnect/replace semantics decided in
the ADR.
- Gates: hub→consumer call over an existing session resolves;
peer-registered op is discoverable via `services/list-peers` and
callable back over channel 0; unregistered/unauthorized
registration attempts fail loudly; disconnect drops the overlay
and fails both sides' pendings.
## Unit 4 — alkhttp wiring (downstream, after Units 1–3)
Tracked in alkhttp review 003 (Units 2–4 there): openables surface,
`install_channel_zero` rework over the fork, retained session
handle, e2e tests, spec reconciliation (OQ-05, ADR-067/048 v1-cut
notes).
---
## Verification log (this pass)
- All claims verified at tree `c16b069`; the F-01 dispatch claim was
checked against both dispatch paths (`run_loop` stream-per-request
reads `handle_stream` → same base-registry-only resolution;
`run_loop_single_stream` at `dispatch.rs:727-850`).
- The F-04 claim was verified by reading both loops' frame arms:
`read_single_stream_until_closed` → `dispatch_envelope`
(`connection.rs:681-705`, no `EVENT_REQUESTED` arm) vs
`run_loop_single_stream` (`dispatch.rs:754-850`, no
pending-resolution arm).
- The F-02 clone claims were verified type-by-type, including the
third-party `jsonschema::Validator` (0.46.10) and the vendored
`Capabilities` (manual `Clone` preserving `Secret`'s clone).
- The F-05 wire-closure claim was verified against the full kind set
(`wire.rs:12-17`) and grep over `src/` for any wire-side
registration path (none; `register_imported*` callers are
in-process only).
- `services/list-peers`' overlay-backed discovery verified at
`registry/discovery.rs:260-307` (reads `ctx.env.peer_ids()` /
`peer_operations()`, populated by `compose_root_env` at
`dispatch.rs:211-213`).
- Baseline gates re-run for this pass: cargo test (565), clippy
(all-targets), fmt — all clean. No source changes.
## Remediation log (2026-09-03)
Units 1–3 landed in one pass. Every claim above was re-verified
against the source before remediation began; all six findings
confirmed (F-02's member-type list was missing `ScopedPeerEnv`
(`context.rs:53`) — also Clone, which made the fork surface simpler
than the review estimated).
**Unit 1 — decision recording (F-01/F-02 direction, F-04/F-05 scope):**
- ADR-047 §4 amendment #2 (2026-09-03): the per-connection registration
mechanism is the **per-session fork of the base registry installed as
the session's dispatch registry** (option (a)); the overlay registry
stays the nested-invocation / peer-announced-ops landing zone.
- ADR-022 amendment (2026-09-03): the bootstrap-op set (`services/list`,
`services/schema`, `op/register`), the opt-in connect-side serving
loop, and the `op/register` wire shape (spec in `services/schema`
JSON + `replace` flag; forwarding-handler wrap into the connection
overlay; `ALREADY_EXISTS` collision gate).
- alkhttp ADR-048's reconciliation note and OQ-05 re-pointed at these
ADRs (Unit 1b).
**Unit 2 — fork surface + per-session composition (F-02, F-03, F-06):**
- `OperationRegistry` gained interior mutability (`parking_lot::RwLock`
around the operations and validators maps) so a registry can live
behind an `Arc` and be self-referential. `register` now takes
`&self`; `registration`/`list_operations` return owned clones
(handlers are `Arc` closures — the clone is cheap).
- `OperationRegistry::fork()` — deep copy of registrations (handlers,
provenance, composition authority, capabilities, `scoped_env`) and
the cached publish-schema validators. `OperationRegistryBuilder::
from_registry` seeds the builder path from an existing registry.
- `install_bootstrap_discovery(&Arc<OperationRegistry>)` — registers
`services/list` / `services/list-peers` / `services/schema` closed
over the fork itself, so per-session openables are discoverable
(F-06) and `services/schema` answers from the fork.
- Gate: `fork_registry_open_op_resolves_and_is_discoverable`
(`src/channels/client.rs`) — open op registered on the fork resolves
over a live channels connection; the openable appears in
`services/list` on the fork; `services/schema` on the fork still
answers. Unit tests: fork independence both directions, validator
carry, post-dispatch registration through a shared `Arc`.
**Unit 3 — client serving half + bootstrap registration (F-04, F-05):**
- `Dispatcher::serve_single_stream` — the full-duplex single-stream
loop: `call.requested` → dispatch + response frames (the
`run_loop_single_stream` arms); `call.responded`/`completed`/`error`
→ pending resolution (the read-pump arms); `call.aborted` → both
tables (in-flight sink aborts and the pending map's cascade);
`call.published` → inbound in-flight sinks. Disconnect fails
pendings and drops in-flight sinks (teardown identical to the two
half-loops it composes).
- `ChannelClient::from_connection_with_serving(connection,
Option<ServingConfig>)` — opt-in serving; `from_connection`
preserves the pure-consumer default (resolution-only read pump).
Gate: `serving_loop_hub_to_consumer_call_resolves` — hub→consumer
call resolves through the consumer's serving loop, consumer→hub
still resolves in the same loop.
- `registry::op_register` — `OpRegisterRequest` wire DTO,
`op_register_spec`, `op_register_handler` (rebuild → forwarding stub
→ `register_imported` into the connection overlay, forced
`Internal`/`FromCall`), `announce_op`. `CallError::already_exists`
added for the collision gate. `spec_to_json_pub` made the
`services/schema` wire shape public; `rebuild_spec_for` and the
forwarding-handler constructors are crate-shared with `from_call`.
Gates: `op_register_announce_then_hub_call_routes_back_to_consumer`
(announce → overlay → hub call → forwarding stub → consumer serves)
plus unit tests (round-trip, collision, replace, forced
visibility/provenance).
**Baseline after remediation:** cargo test 581 passed / 0 failed;
clippy (all-targets, `-D warnings`) clean; `cargo fmt --check` clean;
wasm target gate re-run below. No wire-format changes: the envelope
kind set stays closed at six; the bootstrap ops ride channel 0's call
registry.
@@ -0,0 +1,704 @@
# Review 005 — Serving-Loop Concurrency and `op/register` Composition (post-remediation review of `f84d214`)
## Status
Units 1–3 remediated and verified (G-01..G-05). All findings closed;
the review is resolved. See Remediation log.
## Scope
Post-remediation review of commit `f84d214` (review 004 Units 1–3:
per-session fork registry, connect-side serving loop, `op/register`).
The remediation landed all six F-findings' mechanisms; this pass
reviews the landed implementation against the decided spec (ADR-022
amendment 2026-09-03, ADR-047 §4 amendment #2, review 004's
acceptance gates) rather than re-litigating the findings.
Every finding below was verified directly in source at tree
`f84d214`. The headline finding (G-01) was verified **empirically**:
a probe test was added to the tree, run, observed to reproduce the
defect, and removed — the tree at this review's baseline carries no
source changes. Findings continue alkcall review numbering with the
new prefix `G` (001–004 used P/C/R/A/B/C/D/F — each review numbers
independently).
Cross-repo context: the alkhttp re-pointing claimed by the remediation
log was verified in that repo (`5b62307` — `open-questions.md:102-112`,
`decisions/048…md:19-25`). Unit 4 (alkhttp wiring) remains downstream
and is not reviewed here.
## Baseline verification (this pass)
```
cargo test → 581 passed, 0 failed
cargo test --all-features → 598 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean
cargo check --target wasm32-unknown-unknown → clean
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
All gates the remediation log claims reproduce. The suite is green —
and again the green is partial: the gates test the flows they name,
and the one flow the ADR amendment promises most loudly (nested
composition of peer-announced ops) is the one no gate exercises. See
G-02 for why the existing gate cannot catch G-01.
## Verdict
- **Unit 2's fork surface is solid.** Interior mutability, lock
discipline, fork independence, validator carry, and the
self-referential bootstrap-discovery closure are all correct and
tested (see Non-findings).
- **Unit 3's serving loop resolves the F-04 frame gap but inherits the
dispatch loop's serial-inline shape** — and the F-05 flow it exists
to enable is exactly the shape that deadlocks under it. The ADR-022
amendment's "invocable via nested composition" promise is
unmechanized for wire-dispatched handlers (the primary caller
shape); a 30s sweeper timeout masquerades as resolution.
- **The F-05 acceptance gate does not exercise the forwarding stub** —
the one component that is genuinely new protocol surface. It
verifies announce-lands-in-overlay and the consumer serves, by
calling the announced op *directly*, bypassing the stub entirely.
- **`op/register`'s collision gate is overlay-only**, so a
peer-announced op can shadow the serving side's own ops in nested
composition (connections resolve before base in `PeerCompositeEnv`).
The architecture decisions (fork as dispatch registry; overlay as
landing zone; bootstrap-op set; opt-in serving) remain sound. The
defects are in the serving loop's concurrency model and in what the
gates measure.
## Severity legend
Same scale as review 004:
- **[critical]** — a decided spec invariant is violated in a way that
makes a promised capability unreachable end-to-end; or corrupts
data.
- **[major]** — a core protocol path cannot serve a decided behavior;
works only via shapes the spec does not describe.
- **[minor]** — drift, doc/spec inconsistency, or a missing
convenience with no correctness impact.
---
# Part A — The serving loop's concurrency model
## G-01 [major] — Same-connection nested composition deadlocks the serving loop; the forwarding stub's nested call resolves only via the 30s sweeper
**Status: REMEDIATED (Unit 1)** — see Remediation log; both gates
added and verified load-bearing against the pre-fix loop.
**ADR drift:** ADR-022 amendment (2026-09-03) §`op/register`: the
announced op is *"invocable via nested composition (`env.invoke`)"*;
the amendment's whole point is that the hub's handlers compose
peer-announced ops. Review 004 F-05's recommendation said the same:
the hub-side handler wraps announcements *"as a forwarding handler
that issues a nested `call.requested` back over channel 0"*.
**Verified:** YES, by code trace and by an empirical probe (see
Verification log).
Mechanism, in four steps:
1. `Dispatcher::serve_single_stream` awaits `dispatch()` **inline** in
its read loop for `call.requested` frames
(`src/protocol/dispatch.rs:1009-1014`) — the same serial-inline
shape as the pre-existing accept-side `run_loop_single_stream`
(`:765-770`). Only the Sink arm is spawned
(`:1031-1053`); Once responses and Sub pumps
(`pump_stream_single_stream`, awaited inline at `:1027-1029`)
hold the loop.
2. The F-05 forwarding stub (`make_forwarding_handler`,
`src/client/from_call.rs:392-415`) issues its nested call via
`connection.call_with_payload` (`:408`) — over **the same
connection** whose loop dispatched the parent. In single-stream
mode the pending resolves only when *some* loop reads the
`call.responded` frame (`connection.rs:274-303` — register
pending, write frame, await receiver; resolution is the reader's
job).
3. The only reader of that connection is the serving loop itself —
which is blocked in step 1 awaiting the parent dispatch, whose
handler is blocked in step 2 awaiting the nested call. Circular
wait. The transport buffers the response frame; nobody reads it.
4. The 30s sweeper (`DEFAULT_CALL_TIMEOUT`, `connection.rs:32`;
`evict_expired`, run every 10s per `SWEEPER_INTERVAL`,
`dispatch.rs:44`) is the only thing that breaks the cycle —
the nested call resolves as `CallError::timeout`.
**Empirical probe (this pass):** hub serves `op/register` + a
`hub/compose` handler that invokes the connection overlay's
registration (the forwarding stub — the exact resolution
`PeerCompositeEnv` → `OverlayOperationEnv` produces for nested
composition) from inside a wire-dispatched handler; consumer announces
`consumer/exec`, then calls `hub/compose` over the wire. Result:
```
PROBE: hub/compose resolved after 30.000761619s:
Err(CallError { code: "TIMEOUT", message: "request timed out", retryable: true })
```
A second run with an 8s timeout confirmed the hang is unconditional
(not load-dependent): zero progress until the sweeper. The probe was
removed after evidence capture; the tree is unchanged.
**Consequences (verified):**
- The ADR-022 amendment's flagship flow — a hub handler composing a
peer-announced op — fails for every wire-dispatched parent handler.
It "works" only for parents dispatched from *in-process* contexts
(where the serving loop is idle and free to resolve the nested
call's response) — a shape the ADR does not describe. Per the
severity legend this is a decided behavior served only via an
undescribed shape; it borders [critical] for the composition flow
specifically.
- The same circular wait applies to composition of `from_call`
*imports* on the accept side: `run_loop_single_stream` has the same
inline shape, and the connection overlay's imported-op stubs ride
the same connection. That hazard is **latent pre-`f84d214`** (the
accept side had the inline loop and `register_imported` before this
commit) — `f84d214` makes it load-bearing by (a) adding the wire
path that populates overlays with stubs, (b) putting serving loops
on the connect side where both directions are live by design, and
(c) writing the composition promise into ADR-022.
- The same starvation is not limited to the stub: while the loop
serves *any* Sub (`pump_stream_single_stream` inline), the side's
own outbound pendings queue unresolved. A serving-enabled consumer
that calls out while serving a long-lived Sub to the same peer eats
30s timeouts on its own calls. The F-04 gate tests the two
directions only **sequentially** (hub→consumer resolves, *then*
consumer→hub) — never interleaved.
- Head-of-line blocking: the peer's subsequent requests queue behind
each inline dispatch.
**Fix (Unit 1):** spawn the Once dispatch and the Sub pump the way the
Sink arm already is (the `SharedFrameWriter` already serializes
concurrent frame writes, so per-request frame ordering is preserved);
track spawned task handles for teardown on loop exit (the Sink arm's
`in_flight_sinks` pattern, generalized). Keep read-loop resolution
unblocked at all times. Apply the same rework to
`run_loop_single_stream` (or factor the shared loop) — the accept side
has the identical latent hazard via imported-op composition.
Acceptance: the probe above (recreated as a permanent gate) resolves
in < 1s; a concurrent-interleaving gate (outbound call resolving while
a Sub is being served) passes without sweeper intervention.
## G-02 [major] — The F-05 acceptance gate bypasses the forwarding stub it claims to prove
**Status: REMEDIATED (Unit 1)** — the stub-exercising gate and the
interleaved-directions gate are added; see Remediation log.
**Verified:** YES. `op_register_announce_then_hub_call_routes_back_to_consumer`
(`src/channels/client.rs:1412-1543`) asserts the announce resolves and
the overlay holds the op — then calls
`accept_conn.call("consumer/exec", …)` (`:1529`). That is a **direct
wire call**: the hub's dispatcher resolves `consumer/exec`… which is
not on the hub at all — the frame crosses to the consumer, whose
serving loop dispatches it from its *local* registry. The forwarding
stub (`op_register_handler`'s `register_imported` bundle) is never
invoked. The gate's comment — *"the forwarding stub routes the nested
call back over channel 0"* — describes a path the test does not take.
**Consequence:** the one behavior that is genuinely new protocol
surface (stub routing) is untested, and the one defect it hides (G-01)
is precisely in that path. This is the same failure class review 001
flagged ("all coverage is unit-level / the integration seams are where
the real problems live") — the remediation added e2e gates, but the
F-05 gate measures an adjacent, easier path. The F-04 gate is
likewise sequential-only (see G-01's third consequence).
**Fix (fold into Unit 1):** add a gate that exercises the stub —
announce → a wire-dispatched hub handler composes the announced op via
`ctx.env` (or directly via the overlay registration) → resolves from
the consumer. This is the removed probe, productized; it is the
acceptance test G-01's fix must pass. Optionally extend the F-04 gate
with an interleaved-directions phase.
---
# Part B — The `op/register` composition surface
## G-03 [major] — Announced ops can shadow the serving side's own ops in nested composition; the collision gate checks the overlay only
**Status: REMEDIATED (Unit 2)** — see Remediation log; the ADR-022
amendment records the collision policy.
**ADR drift:** ADR-018 (composition authority) and ADR-019 (the
overlay is the *landing zone* for imported ops — a layer beneath the
deployment's own registry, not a rival to it); ADR-022 amendment
(silent on name collisions with the serving side's base registry).
**Verified:** YES. `op_register_handler`'s collision gate is
`connection.overlay_contains(&request.spec.name)`
(`src/registry/op_register.rs:142`) — it consults **only the
connection overlay**. The serving registry (the fork the session
dispatches over) is not consulted. Meanwhile `PeerCompositeEnv`
resolves **connections before base** (`src/registry/env.rs:230-245` —
session, then `connection_order`, then base), and `compose_root_env`
attaches the dispatching connection's overlay
(`dispatch.rs:207-214`).
Consequence: a peer announces an op named `fs/readFile` (or any name
the serving side has registered — `services/list` itself is fair
game). The registration lands in the overlay (forced `Internal`,
`FromCall` — both per spec). From then on, every handler
**wire-dispatched from that peer's connection** that nested-composes
that name via `ctx.env` resolves to the peer's forwarding stub instead
of the serving side's own op. The forced `Visibility::Internal` does
not help here — it blocks wire invocation, not `env.invoke`
(`OverlayOperationEnv` gates on `AccessControl` only,
`connection.rs:749-836`; the composed child context is `internal:
true` by design).
Concretely: a hub handler processing peer A's request composes
`fs/readFile` expecting *its own* registry op (with the deployment's
`scoped_env`, capabilities, and ACL); it silently gets peer A's stub —
peer A chooses where the call goes and what it returns. Composition
authority (ADR-018) is decided by the composing handler's *deployer*;
a same-name peer registration silently rewrites it. The effect is
scoped to handlers whose root context is that connection's
(`compose_root_env` attaches that connection's overlay), which bounds
the blast radius — but that is exactly the flagship F-05 flow.
**Fix (Unit 2):** decide the collision policy and record it in the
ADR-022 amendment. Recommended: `op_register_handler` also rejects
names registered on the serving registry (pass the fork — or a
name-set closure — into the handler alongside the connection, the
same way `install_bootstrap_discovery` closes over its registry);
`ALREADY_EXISTS` for base collisions regardless of `replace` (the
reconnect path needs to re-announce *peer* ops, not to overwrite the
deployment's). The alternative — changing `PeerCompositeEnv` to
resolve base before connections — would break `from_call`'s
namespace-prefixed imports (prefixed names don't collide) is not
needed for that, but would change long-decided layering semantics; the
registration-side gate is the cheaper, narrower door.
## G-04 [minor] — `resource_id_path` does not survive the spec wire round-trip
**Status: REMEDIATED (Unit 3)** — see Remediation log.
**Verified:** YES. `spec_to_json_pub` serializes
name/namespace/op_type/visibility/schemas/error_schemas/access_control
(+ `channel_open`/`publish_schema` markers;
`src/registry/discovery.rs:211-234`) — `resource_id_path` is not
serialized. `rebuild_spec_for` cannot parse it
(`src/client/from_call.rs:209-286`), and the field exists on
`OperationSpec` (`src/registry/spec.rs:190`). An announced (or
`from_call`-imported) op that declares ownership-scoped resource
extraction silently loses it; the rebuilt spec's ACL checks run with
`resource_id: None`.
This is a pre-existing gap in the `from_call` import shape, not
introduced by `f84d214` — but `op/register` makes it load-bearing for
a second wire surface (peer-announced specs) that explicitly promises
"the serializable parts" of a registration. Note `namespace` *is*
serialized but `rebuild_spec_for` ignores it in favor of the
`namespace_prefix` parameter — consistent for `op/register`
(`replace`/re-announce uses `None`), worth a comment either way.
**Fix (Unit 3):** add `resource_id_path` to both halves of the
round-trip plus a round-trip test. Additive optional field in the
`services/schema` JSON shape — no existing consumer breaks.
## G-05 [minor] — `install_bootstrap_discovery` registers `services/list-peers`; the ADR-022 amendment's bootstrap set doesn't name it
**Status: REMEDIATED (Unit 3)** — see Remediation log.
**Verified:** YES. `install_bootstrap_discovery` registers
`services/list`, `services/list-peers`, and `services/schema`
(`src/registry/discovery.rs:290-313`; `list-peers` at `:300`). The
ADR-022 amendment's bootstrap set lists exactly three ops —
`services/list`, `services/schema`, `op/register`
(`022-…contract.md:380-393`). Review 004's remediation log says
"list-peers" too; the ADR text was never updated.
Consequence: none at runtime (installing a fourth op is strictly
additive, and `services/list-peers` is what makes F-05's
"peer-announced ops are discoverable" true). It is a
doc/spec-consistency drift on a set the ADR declares *"closed."*
**Fix (Unit 3):** add `services/list-peers` to the ADR-022
amendment's bootstrap-op list (recommended — `from_call`-import
discovery already relies on it) or stop installing it; one line
either way.
---
# Non-findings (verified correct, recorded to bound the re-review)
- **The fork surface (F-02/F-03) is solid.** Lock discipline is
correct: `registration()` clones under a read guard and drops it
before handler invocation (`registration.rs:190-192`), so
self-referential discovery handlers cannot deadlock their own
registry; `register` takes the two maps' write locks sequentially
and never across an `await`. `fork()` deep-copies both maps
(handlers are `Arc` closures — cheap, and the F-03 constraint that
closures must carry is met); validators carry; fork independence is
tested both directions; `OperationRegistryBuilder::from_registry`
seeds the builder path.
- **`install_bootstrap_discovery` (F-06) is correct** — handlers
closed over the same `Arc` they're registered on, so post-install
per-session registrations are visible at call time; ACL filtering
stays per-caller; the fork gate
(`fork_registry_open_op_resolves_and_is_discoverable`) exercises
exactly the review-004 acceptance shape.
- **`serve_single_stream`'s pending-resolution arms match the two
half-loops they compose.** Frame by frame against
`run_loop_single_stream` (`dispatch.rs:765-905`) and
`dispatch_envelope` (`connection.rs:717-741`): responded/completed/
error resolve pendings; aborted tries both tables (in-flight sink
aborts, then the pending cascade via `handle_abort`); published
routes inbound sinks with the same validation/keep semantics; teardown
(fail_all + sink drop + sweeper abort) matches. The direction
disambiguation by table membership (documented at
`dispatch.rs:940-966`) is correct as written — G-01 is *not* a
frame-routing bug; the arms are right, the loop's scheduling is the
defect.
- **The pure-consumer default is unchanged.** `from_connection`
delegates with `None` and keeps the resolution-only read pump; no
behavior change for existing callers.
- **`op/register`'s DTO and unit semantics are clean** — one spec
serialization on the wire (`spec_to_json` shape), forced
`Internal`/`FromCall` on landing (unit-tested), `replace` semantics
tested, `ALREADY_EXISTS` collision gate tested (against the
overlay — see G-03 for the gap).
- **The alkhttp cross-repo claims check out.** `open-questions.md`
OQ-05 and ADR-048's re-point are present in that repo at `5b62307`.
- **No non-English text reached the tree.** The remediation session's
language slip (reported anecdotally) did not land in any code,
doc, or metadata: a CJK-range regex sweep over all `.rs`/`.md`/
`.toml`/`.json` in the repo returns zero matches at `f84d214`.
- **Gates reproduce.** All commands in Baseline verification pass at
the review tree, matching the remediation log's claims
(581 default / 598 all-features).
---
# Remediation plan
Sequenced by dependency. All units are alkcall work; Unit 4 (alkhttp
wiring) stays downstream and should **not** start before Unit 1 —
alkhttp's serving consumers would compose over the same connection
and hit G-01 immediately. (All three units landed — see Remediation
log.)
## Unit 1 — Concurrent serving loop + a stub-exercising gate (G-01, G-02)
- Rework `serve_single_stream`'s `EVENT_REQUESTED` arm: spawn the Once
dispatch and the Sub pump (the Sink arm's existing pattern,
generalized; `SharedFrameWriter` serializes frames; track handles
for teardown). Apply the same rework to `run_loop_single_stream` or
extract the shared loop.
- Add the G-02 gate: announce → wire-dispatched hub handler composes
the announced op via nested composition → resolves from the
consumer. This is the removed probe, productized; it is the
acceptance test for G-01.
- Optional: an interleaved-directions phase on the F-04 gate
(outbound call resolving while a Sub is being served).
- Gate: the probe recipe (Verification log) resolves in < 1s; no
sweeper evictions in the gates.
## Unit 2 — `op/register` collision policy (G-03)
- `op_register_handler` gains a serving-registry name-set (closure or
`Arc<OperationRegistry>`); base-registry name collisions reject
with `ALREADY_EXISTS` regardless of `replace`.
- ADR-022 amendment: record the collision policy (peer-announced ops
may collide with peer-announced ops — `replace` governs; never with
the serving side's own registrations).
- Gates: announcing a base-registered name (External *and* Internal)
fails loudly; announcing a distinct name still lands; nested
composition of a base op is unaffected by an unrelated announce.
## Unit 3 — Round-trip completeness + doc alignment (G-04, G-05)
- `resource_id_path` through `spec_to_json_pub` + `rebuild_spec_for`
+ round-trip test.
- ADR-022 bootstrap list gains `services/list-peers` (or drop it from
the installer; recommend adding).
---
# Remediation log
## Unit 3 — Round-trip completeness + doc alignment (G-04, G-05) — LANDED
**G-04 fix.** `resource_id_path` now rides both halves of the spec
wire round-trip: `spec_to_json_pub` serializes it as an optional
`resource_id_path` string (`src/registry/discovery.rs`), and
`rebuild_spec_for` parses it into the rebuilt spec's fourth
constructor argument (`src/client/from_call.rs`). Additive optional
field — absent stays absent, no existing consumer breaks (verified by
the companion gate). `rebuild_spec_for`'s doc note from the review
("namespace *is* serialized but ignored in favor of
`namespace_prefix`") was addressed by leaving the behavior as-is: for
`op/register` the parameter is `None` (consistent) and for `from_call`
the prefix is authoritative — the round-trip tests pin the `name`
field handling.
**Gates:** `spec_round_trips_resource_id_path` (serialize → parse →
field intact) and
`spec_without_resource_id_path_stays_absent_through_round_trip`
(absent key serializes nothing; rebuilt `None`) in
`src/client/from_call.rs`.
**G-05 fix.** The ADR-022 amendment's bootstrap-op set gained
`services/list-peers` with a dated note (review 005 G-05) explaining
that the installer has registered it since the amendment landed and
the doc lagged the code. The set remains closed at four.
**Verification (post-fix):**
```
cargo test → 589 passed, 0 failed
cargo test --all-features → 606 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
---
## Unit 2 — `op/register` collision policy (G-03) — LANDED
**Fix shape.** `op_register_handler` now takes the serving registry
alongside the connection (`op_register_handler(connection,
serving_registry)` — the review's "name-set closure" recommendation,
materialized as the `Arc<OperationRegistry>` itself, mirroring how
`install_bootstrap_discovery` closes over its registry). The handler
checks `serving_registry.registration(name)` **before** the overlay
gate and rejects base collisions with `ALREADY_EXISTS` regardless of
`replace`; overlay collisions keep the existing `replace` semantics.
Announced ops may collide with announced ops, never with the serving
side's own registrations.
**ADR-022 amendment:** the 2026-09-03 amendment's `op/register`
section gained a "Collision policy (amended 2026-09-04, review 005
G-03)" paragraph recording the decided policy and the rationale (the
`PeerCompositeEnv` connections-before-base resolution would let an
unscreened announce rewrite composition resolution; visibility is
irrelevant to the gate). The ADR's status line notes the sub-amendment.
**Call sites updated:** both e2e gates pass the fork the session
dispatches over (`Arc::clone(&accept_registry)` after `fork()`), which
is the production shape — the collision set is exactly the registry
the serving loop dispatches against.
**Gates:**
- `handler_rejects_base_registry_collision_even_with_replace` —
base-registered `fs/readFile` (External) + `replace: true` →
`ALREADY_EXISTS`; the overlay stays clean and the serving
registration is untouched.
- `handler_rejects_collision_with_internal_serving_op` — an
`Internal` base op is equally protected (the review's point that
forced `Visibility::Internal` on the *announced* spec never helped:
`OverlayOperationEnv` gates on `AccessControl`, and the composed
child is `internal: true` by design).
- `handler_overlay_collision_still_governed_by_replace` — announce/
announce collisions still follow `replace` (the base gate is
scoped to the serving registry, not widened into the overlay).
- `nested_composition_of_base_op_unaffected_by_unrelated_announce` —
after a successful distinct-name announce, composing the base op
through the real `compose_root_env` shape (`PeerCompositeEnv` +
`attach_peer(conn.overlay_env())`) resolves the serving side's own
op, not a peer stub. (Unit-level, using the actual env types rather
than a new full-wire fixture — the wire shape is already covered by
the Unit-1 gates.)
**Deliberate non-change:** `PeerCompositeEnv` resolution order is
untouched (connections before base) — the review's recommendation;
reordering would change long-decided layering semantics
(ADR-019/ADR-024) for no gain the registration-side gate doesn't
deliver.
**Verification (post-fix):**
```
cargo test → 587 passed, 0 failed
cargo test --all-features → 604 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
---
## Unit 1 — Concurrent serving loop + stub-exercising gates (G-01, G-02) — LANDED
**Fix shape.** `dispatch()` was split into a synchronous start half and
an awaited invocation:
- `Dispatcher::dispatch_start`
(`src/protocol/dispatch.rs`) runs the sync prefix (identity
resolution, root context, op-type branch) and returns a
`StartedDispatch`: `Once` (the invocation as a boxed future), `Stream`
(the `ResponseStream`, returned synchronously — `invoke_streaming`
is a sync call), or `Sink` (started **inline**, unchanged). `dispatch()`
is now a thin wrapper (`Once` = await the boxed future) and keeps its
shape for `dispatch_requested`/`handle_stream`/the gateway.
- The Sink start stays inline **by design**: its `chunk_tx` must be in
`in_flight_sinks` before the next `call.published` frame can be
routed; the loop insert cannot race the feed. Pub handlers that
block forever inside their sink future are a separate (pre-existing,
unreported) shape — the read loop no longer waits on any handler
except for the few instructions of the sink start itself.
- Both single-stream loops' `EVENT_REQUESTED` arms spawn the Once
invocation (`spawn_once_dispatch`), the Sub pump
(`spawn_stream_pump`), and the sink's response writer
(`spawn_sink_response_writer`); handles are tracked in a
`spawned` list and aborted at loop exit (teardown: sink-map clear →
spawn aborts → `fail_all` → sweeper abort). `in_flight_sinks` moved
behind `Arc<parking_lot::Mutex>` because the ABORTED/PUBLISHED/ERROR
arms must not hold the guard across the `chunk_tx.send().await`
(non-`Send` guard across an await — the reason the pre-fix loop
couldn't just be `tokio::spawn`ed piecemeal). Guard-dropping
(`let entry = ...remove()` before the await) keeps lock discipline:
no lock is held across an await anywhere in the loops.
- **`run_loop_single_stream` (the accept side) got the
pending-resolution arms** (`EVENT_RESPONDED`/`COMPLETED`/`ERROR` →
the outbound pending map). Previously it served only — an accept-side
nested-composing handler (e.g. a `from_call` imported-op stub riding
the same connection) had no loop resolving its response frames at
all. The G-01 accept-side latent hazard is mechanized shut, not just
unblocked: both single-stream loops are now the same shape
(dispatch-spawn + pending-resolution), the full-duplex loop
`serve_single_stream` composes them and is unchanged in its arm
semantics (frame-arm equivalence preserved per the non-findings
audit).
- Write-failure semantics changed from `break` (close the loop) to
`warn` (keep reading) in the spawned Once path, matching the Sink
arm: a dying transport surfaces on the next read as
`ConnectionClosed`; a transient frame-write failure no longer tears
down every in-flight request on the connection.
**Gates (G-02, productized from the removed probe):**
- `hub_handler_composes_peer_announced_op_via_nested_composition`
(`src/channels/client.rs`): announce → consumer calls `hub/compose`
→ the hub's serving loop wire-dispatches it → the handler resolves
`consumer/exec` via `ctx.env` nested composition → the forwarding
stub's nested `call.requested` crosses back to the consumer. This is
the stub path the F-05 gate bypassed. Two wiring facts surfaced (both
recorded as spec-consistent, both were silent before): (a) the hub's
channel-0 `Connection` must carry the peer identity —
`compose_root_env` attaches the connection overlay keyed by
`identity.id` (ADR-030 §5), and a deployment that skips identity
resolution silently gets no peer overlay (the test sets it via the
`AuthContext` the adapter passes to `install_channel_zero`);
(b) composition reachability is declared on the composing handler's
registration (`scoped_env: ScopedPeerEnv::new(["consumer/exec"])`) —
the empty `ScopedPeerEnv` is deny-by-default in
`PeerCompositeEnv::invoke_with_policy`, so a wire-dispatched handler
with no `scoped_env` composes nothing (the reachability gate, not the
overlay, was what the probe's first draft tripped over).
- `outbound_call_resolves_while_inbound_subscription_is_being_served`
(the optional interleaved-directions phase on the F-04 shape): the
consumer subscribes to the hub (Sub served by the consumer's loop),
then calls `hub/interleave` over the wire — its handler issues an
outbound hub→consumer call on the same connection while the Sub is
live — and asserts the call resolves and the Sub is still live
after. Both gates bounded at 5s (vs. the 30s sweeper), so a
regression fails fast, not via sweeper eviction.
**Load-bearing verification (both gates, empirical):** each gate was
run against the pre-fix loop (`git stash push
src/protocol/dispatch.rs`) and reproduced the G-01 hang:
```
hub_handler_composes…: panicked "nested composition through the
forwarding stub timed out (G-01 shape): Elapsed(())" — 5.01s, no progress
outbound_call_resolves…: panicked "interleaved outbound call starved
while Sub served (G-01 shape): Elapsed(())" — 5.06s, no progress
```
With the fix, both resolve (0.01s / 0.11s wall including connection
setup). No sweeper evictions (bounded timeouts ≪ 30s).
**Consequences audited against the fix:**
- Same-connection nested composition of peer-announced ops works for
wire-dispatched parents (the flagship ADR-022 amendment flow) — the
gate proves it end-to-end.
- The accept-side imported-op composition hazard (G-01's second
consequence, latent pre-`f84d214`) is mechanized shut by the
pending-resolution arms in `run_loop_single_stream` — not merely
unblocked. There is no dedicated gate for this shape yet (the
accept-side import + nested-compose e2e would be a further gate;
noted as residual work, not a regression).
- Head-of-line blocking is gone: spawned arms proceed concurrently;
the loop never awaits a handler.
- Frame ordering: per-request ordering is preserved by the
`SharedFrameWriter` (each `write_frame` is atomic under the mutex);
responses for two requests may now interleave *at frame
granularity*, which single-stream mode always permitted (both
directions' frames are multiplexed by design; correlation is by id).
**Verification (post-fix):**
```
cargo test → 583 passed, 0 failed
cargo test --all-features → 600 passed, 0 failed
cargo clippy --all-targets -- -D warnings → clean
cargo clippy --all-features --all-targets -- -D warnings → clean
cargo fmt --check → clean (fmt applied)
cargo clippy --target wasm32-unknown-unknown -- -D warnings → clean
cargo doc --no-deps → clean
```
**Residual (not blocking Unit 1):** the sink-start-inline shape means a
`Pub` op whose registration itself blocks (ACL/schema compile are sync
and fast; `resolve_sink_handler` runs the handler's ACL path) still
holds the loop briefly — bounded by registry work, not handler work. A
malicious `AccessControl::check` implementation could stall the loop;
that is a deployment-provided trait object, the same trust boundary as
`IdentityProvider`, and unchanged from the pre-existing shape.
---
## Verification log (this pass)
- All gates in Baseline verification reproduced at tree `f84d214`
(clean working tree; the probe was added, run, and removed —
`git status` verified clean after removal).
- G-01 verified by code trace (the four steps above, each line
read at `f84d214`) **and** empirically: probe
`probe_hub_handler_composes_peer_announced_op` (in
`src/channels/client.rs#[cfg(test)]`, since removed) — hub registry
carries `op/register` + `hub/compose` (whose handler invokes the
connection overlay's registration for the announced op, the shape
`PeerCompositeEnv` → `OverlayOperationEnv` resolution produces);
consumer announces `consumer/exec` then calls `hub/compose`.
8s-timeout run: `Elapsed` (no progress). 45s-cap run: resolved at
**30.0007s** with `CallError::TIMEOUT` (retryable) — the
`DEFAULT_CALL_TIMEOUT` sweeper eviction, not live resolution.
- The existing gates' blind spot verified by reading both e2e gates'
call directions: `serving_loop_hub_to_consumer_call_resolves`
(`client.rs:1373-1399`) and
`op_register_announce_then_hub_call_routes_back_to_consumer`
(`client.rs:1529`) — the latter's `accept_conn.call("consumer/exec")`
crosses the wire to the consumer's local registry; the forwarding
stub is unreachable from it.
- G-03 verified by reading the collision gate (`op_register.rs:142`)
against `PeerCompositeEnv::invoke_with_policy`'s resolution order
(`env.rs:230-245`) and `compose_root_env`'s attach
(`dispatch.rs:207-214`).
- G-04 verified by field audit: `OperationSpec`'s fields
(`spec.rs:175-206`) vs `spec_to_json_pub` (`discovery.rs:211-234`)
vs `rebuild_spec_for` (`from_call.rs:209-286`).
- G-05 verified against `discovery.rs:290-313` and ADR-022
amendment's set (`022-…contract.md:380-393`).
- Frame-arm equivalence of `serve_single_stream` vs the two half-loops
verified read-through (Non-findings).
- alkhttp cross-repo claims verified in `/workspace/@alkdev/alkhttp`
at `5b62307` (OQ-05 `open-questions.md:102-112`; ADR-048
`decisions/048-…md:19-25`).
- CJK sweep: no CJK-range codepoints in any `.rs`/`.md`/`.toml`/
`.json` under the repo at `f84d214`.
+111
View File
@@ -0,0 +1,111 @@
# Consumer findings ledger — alkhttp consuming alkcall
Findings from writing the first real consumer (alkhttp) against alkcall.
alkcall is deliberately minimal; expect missing features and edge-case
bugs to surface here first. Extract items into alkcall tasks/reviews as
the project sees fit.
Format: date | found-in (alkhttp context) | severity | status.
---
## Open
(none)
---
## Resolved
### CF-004 — `services_schema_handler` discloses Internal/ACL-restricted op specs — no visibility or AccessControl check (2026-08-30) — RESOLVED 2026-08-31
- **Found in:** alkhttp Review 002, finding PRJ-16
(`alkhttp/docs/reviews/002-post-remediation-review.md`, Part F',
verified against tree `91483a7` + alkcall source). Filed to this
ledger 2026-08-30.
`src/registry/discovery.rs:327-343` (`services_schema_handler`) —
the handler did a bare `registry.registration(&name)` and returned
`spec_to_json(&reg.spec)` verbatim, with **no `Visibility` check and
no `AccessControl::check(peer_identity)`**.
- **Fix:** the handler now applies the same two gates as
`OperationRegistry::invoke` — the Internal-visibility rejection
(`!ctx.internal` → spec-404) and `AccessControl::check` with the same
identity resolution invoke() uses (`handler_identity` when internal,
`identity` otherwise). Restricted ops return `NOT_FOUND` (spec-404)
rather than `FORBIDDEN` — no information leaks about the restricted
surface's existence or shape. The gate is in the handler itself, so
every transport (wire `/call`, HTTP routes, MCP tools) is covered.
- **Status:** resolved — 2026-08-30.
### CF-003 — `publish_schema` compile failure is fail-open on the wire dispatch path (2026-08-30) — RESOLVED 2026-08-30
- **Found in:** alkhttp Review 001 post-remediation task
`review-001-publish-schema-validation-robust` (the GW-01/#1 follow-up
work). While fixing alkhttp's HTTP-side instance of the pattern, the
identical instance was verified on the wire dispatch path:
`src/protocol/dispatch.rs:352-366` (`Dispatcher::dispatch`,
`OperationType::Pub` arm) — when a Pub op's `publish_schema` failed to
**compile**, the pump logged a warn and proceeded with
`validator: None` (silent unvalidated ingest, plus per-request
recompile cost).
- **Fix:** fail-closed at the source — `publish_schema` is now compiled
at **registration time** in both insertion points
(`OperationRegistry::register` and `OperationRegistryBuilder::store`);
an un-compilable schema is a registration error, so an
unvalidated-ingest path can never be constructed. The compiled
validator is cached per-op (`OperationRegistry::publish_validator`)
and the dispatch path consumes the cache — the per-request compile is
gone. The `from_call` import path inherits the guarantee (imports
register through the same choke point); forwarded-chunk validation on
the producer side is the remote peer's dispatch path and now inherits
the same fail-closed property.
- **Behavior change (semver-relevant):** `register`/builder methods
return `Err` for an un-compilable `publish_schema` (previously:
accepted). Code that registered garbage schemas will now get a
registration error instead of a silently-unvalidated op.
- **Status:** resolved — 2026-08-30.
### CF-002 — demux `TooLarge` skip allocates the full peer-declared length up front (up to ~4 GiB from an 8-byte header) (2026-08-30) — RESOLVED 2026-08-30
- **Found in:** alkhttp Review 001, finding WS-12 (cross-crate;
`docs/reviews/001-initial-implementation-review.md`, Part B,
verified against tree `4a825d3`). Filed to this ledger 2026-08-30.
`src/channels/adapter.rs:144` (`run_demux_loop_for_client`) — the
`ChunkError::TooLarge` arm allocated
`let mut discard = vec![0u8; length as usize];` **before** reading:
the buffer was sized from the peer's untrusted 8-byte header (`u32`
`length`, up to ~4 GiB).
- **Fix:** stream-skip with a fixed 64 KiB buffer + `read_exact` loop
consuming `length` bytes — no allocation sized from peer input.
Plus a cumulative skipped-bytes budget (256 MiB) that tears down the
connection when a peer loops `TooLarge` headers to burn bandwidth/CPU.
The budget resets on each valid chunk, so legitimate isolated
oversized chunks (the resync path, covered by the existing test) are
unaffected. The existing `demux_resyncs_after_oversized_chunk` test
passes unchanged.
- **Status:** resolved — 2026-08-30.
### CF-001 — `call_single_stream` write-failure maps to non-retryable `INTERNAL` (2026-08-29) — RESOLVED 2026-08-30
- **Found in:** alkhttp `from_wss` drop-monitor race tests (WS-02/CON-02,
review-001-ws-eof-signal). When the transport mux died mid-call, the
write path failed fast and the write-failure mapping produced a
non-retryable `INTERNAL: failed to write request frame` — even though
the call never reached the producer and a reconnect/retry would be
safe.
- **Fix:** new retryable `CallError::connection_closed` constructor
(`CONNECTION_CLOSED` code, `retryable: true`), applied **only** to
write failures where the call is provably undelivered:
`call.requested` frame write failures on all consumer paths
(`call`/`subscribe`/`publish`, single-stream and stream-per-request;
the publish pump tags its write stages so only the request-frame
stage is retryable). Mid-publish and completed-frame write failures
stay `INTERNAL` (delivery ambiguous — retry unsafe), and the
producer-side `fail_all(...)` on connection close stays `INTERNAL`.
The new code string is additive to the wire error vocabulary;
`CONNECTION_CLOSED` is new — consumers should treat unknown codes per
their existing policy (the `retryable` flag is the machine-readable
signal).
- **Status:** resolved — 2026-08-30. alkhttp follow-up: the
`review-001-ws-eof-signal` race test can now tighten back to
retryable-only asserts.
+165 -4
View File
@@ -128,6 +128,15 @@ impl ChannelsAdapter {
reader: Box<dyn tokio::io::AsyncRead + Send + Unpin>,
policy: Option<&Arc<dyn ChannelLifecyclePolicy>>,
) {
// Stack buffer for skipping oversized payloads — sized once, never
// from peer-declared lengths (a peer-declared `length` is untrusted
// up to u32::MAX ≈ 4 GiB; pre-sizing off it is an allocation-shape
// DoS). A cumulative cap tears down dribbling peers that loop
// TooLarge headers to burn bandwidth indefinitely.
const SKIP_BUF_LEN: usize = 64 * 1024;
const MAX_CONSECUTIVE_SKIPPED_BYTES: u64 = 256 * 1024 * 1024;
let mut skip_buf = vec![0u8; SKIP_BUF_LEN];
let mut skipped_since_last_valid: u64 = 0;
let mut reader = reader;
let mut header_buf = [0u8; CHUNK_HEADER_LEN];
loop {
@@ -141,10 +150,23 @@ impl ChannelsAdapter {
max = super::wire::MAX_CHUNK_LEN,
"demux: chunk too large, skipping payload bytes"
);
let mut discard = vec![0u8; length as usize];
if let Err(e) = reader.read_exact(&mut discard).await {
warn!(error = %e, "demux: failed to skip oversized payload");
break;
let mut remaining = length as u64;
while remaining > 0 {
let take = remaining.min(SKIP_BUF_LEN as u64) as usize;
if let Err(e) = reader.read_exact(&mut skip_buf[..take]).await {
warn!(error = %e, "demux: failed to skip oversized payload");
return;
}
remaining -= take as u64;
}
skipped_since_last_valid += length as u64;
if skipped_since_last_valid > MAX_CONSECUTIVE_SKIPPED_BYTES {
warn!(
skipped_since_last_valid,
max = MAX_CONSECUTIVE_SKIPPED_BYTES,
"demux: skipped-bytes budget exceeded; tearing down connection"
);
return;
}
continue;
}
@@ -170,6 +192,7 @@ impl ChannelsAdapter {
}
};
manager.route_payload(header.channel_id, payload).await;
skipped_since_last_valid = 0;
}
Err(e) => {
if e.kind() == std::io::ErrorKind::UnexpectedEof {
@@ -529,4 +552,142 @@ mod tests {
assert_eq!(buf, [i; 4], "channel A chunk {i} intact — no data loss");
}
}
/// CF-002 #1 — the skipped-bytes budget tears down a dribbling peer.
/// Sixteen back-to-back `TooLarge` headers each declare 16 MiB + 1 and
/// their payloads are genuinely on the wire (the demux must consume
/// every byte to reach the next header, so the test cannot avoid
/// moving ~268 MiB through the duplex). The 16th skip pushes the
/// cumulative total past the 256 MiB budget: a working demux tears
/// down and *stops reading there*, while a peer that keeps dribbling
/// (a 17th header parked mid-payload) must not wedge it. Without the
/// budget the demux blocks skipping the 17th payload forever and the
/// timeout fires.
#[tokio::test]
async fn demux_tears_down_after_cumulative_skipped_bytes_budget_exceeded() {
let oversized_len = super::super::wire::MAX_CHUNK_LEN as usize + 1;
let burst_count = 16;
let (client, server) = tokio::io::duplex(1024 * 1024);
let (server_read, server_write) = tokio::io::split(server);
let (mux_handle, mux_runner) = MuxRunner::new(Box::new(server_write));
let _mux_task = tokio::spawn(async move {
let _ = mux_runner.run().await;
});
let manager = ChannelManager::with_defaults(mux_handle, None);
let (_id, _send, _recv) = manager
.open_channel("alk/tty", "alice", None)
.await
.expect("open");
let demux_manager = manager.clone();
let demux_task = tokio::spawn(async move {
ChannelsAdapter::run_demux_loop_for_client(&demux_manager, Box::new(server_read), None)
.await;
});
// Writer task: 16 full oversized payloads (15 × ~16 MiB ≈ 240 MiB
// skip cleanly under the budget; the 16th crosses it), then a 17th
// header + partial payload — and park instead of finishing (the
// peer keeps dribbling, never EOFs).
let writer_task = tokio::spawn(async move {
use tokio::io::AsyncWriteExt;
let mut client_write = client;
let mut header = [0u8; 8];
let chunk = vec![0u8; 64 * 1024];
let full = oversized_len / chunk.len();
let tail = oversized_len % chunk.len();
for _ in 0..burst_count {
super::super::wire::write_header(1, oversized_len as u32, &mut header)
.expect("write header");
client_write.write_all(&header).await.expect("write header");
for _ in 0..full {
client_write.write_all(&chunk).await.expect("write payload");
}
if tail > 0 {
client_write
.write_all(&chunk[..tail])
.await
.expect("write payload tail");
}
}
// The budget is already blown by the 16th skip; this header is
// the dribble that a budgetless demux would park on forever.
super::super::wire::write_header(1, oversized_len as u32, &mut header)
.expect("write header");
client_write.write_all(&header).await.expect("write header");
client_write
.write_all(&chunk[..8192])
.await
.expect("write partial payload");
std::future::pending::<()>().await;
});
let demux_done =
tokio::time::timeout(std::time::Duration::from_secs(120), demux_task).await;
assert!(
demux_done.is_ok(),
"demux must tear down on the skipped-bytes budget, not block in the skip loop"
);
writer_task.abort();
}
/// CF-002 #2 — a skip that hits EOF mid-payload (peer declared more
/// bytes than it sent) ends the demux loop; nothing routes to the
/// channel and the loop does not spin back to header parsing on a
/// truncated stream.
#[tokio::test]
async fn demux_ends_when_oversized_payload_skip_hits_eof() {
let (client, server) = tokio::io::duplex(1024 * 1024);
let (server_read, server_write) = tokio::io::split(server);
let (mux_handle, mux_runner) = MuxRunner::new(Box::new(server_write));
let _mux_task = tokio::spawn(async move {
let _ = mux_runner.run().await;
});
let manager = ChannelManager::with_defaults(mux_handle, None);
let (_id, _send, mut recv) = manager
.open_channel("alk/tty", "alice", None)
.await
.expect("open");
let demux_manager = manager.clone();
let demux_task = tokio::spawn(async move {
ChannelsAdapter::run_demux_loop_for_client(&demux_manager, Box::new(server_read), None)
.await;
});
use tokio::io::AsyncWriteExt;
let oversized_len = super::super::wire::MAX_CHUNK_LEN as usize + 1;
let mut header = [0u8; 8];
super::super::wire::write_header(1, oversized_len as u32, &mut header).expect("header");
let mut client_write = client;
client_write.write_all(&header).await.expect("write header");
client_write
.write_all(&[0u8; 4096])
.await
.expect("write partial payload");
drop(client_write);
let demux_done = tokio::time::timeout(std::time::Duration::from_secs(10), demux_task).await;
assert!(
demux_done.is_ok(),
"demux must end when the skip hits EOF instead of looping"
);
// The oversized chunk was never routed — nothing to assert on the
// channel beyond the loop ending; reading would surface EOF since
// the demux dropped its sender without delivering a payload.
use tokio::io::AsyncReadExt;
let mut buf = [0u8; 4];
let read = tokio::time::timeout(
std::time::Duration::from_millis(500),
recv.read_exact(&mut buf),
)
.await;
assert!(
read.is_err() || read.unwrap().is_err(),
"no payload must arrive after the truncated skip"
);
}
}
+1007 -24
View File
File diff suppressed because it is too large. Load diff
+4
View File
@@ -83,6 +83,10 @@ impl OperationEnv for ChannelsSessionEnv {
self.base.peer_operations(peer)
}
fn list_operation_names(&self) -> Vec<String> {
self.base.list_operation_names()
}
async fn invoke_peer(
&self,
peer: &crate::registry::env::PeerRef,
+131 -18
View File
@@ -7,7 +7,7 @@
//! `ChannelsAdapter`, the open-op wrapper, and relay logic can all
//! hold a handle.
use std::collections::HashMap;
use std::collections::{HashMap, VecDeque};
use std::net::SocketAddr;
use std::sync::atomic::{AtomicU32, AtomicU64, Ordering};
use std::sync::Arc;
@@ -93,8 +93,23 @@ struct Inner {
remote_addr: Option<SocketAddr>,
side: ChannelSide,
dropped_unknown_chunks: AtomicU64,
/// Chunks that arrived for a channel_id this side has not
/// adopted yet (the open-op response carrying `channel_id` and the
/// producer's first data-plane write race — the handler can start
/// pumping before the connect side's `adopt_channel` runs). Each
/// channel buffers up to `early_arrival_cap` payloads; `adopt_channel`
/// drains them into the new receiver. `clear_all` drops the map.
early_arrivals: Mutex<HashMap<u32, VecDeque<Bytes>>>,
early_arrival_count: AtomicU64,
}
/// The per-channel cap on parked early-arrival chunks (`route_payload`
/// parks chunks for not-yet-adopted channels instead of dropping them —
/// the open-op response/first-data race). A channel whose adopter never
/// arrives leaks its parked chunks until `clear_all`; the cap bounds
/// that leak per channel.
const EARLY_ARRIVAL_CAP: usize = 64;
impl ChannelManager {
/// Construct a new manager with the given `MuxHandle`, max
/// channels, buffer cap, and side. The `remote_addr` is
@@ -124,6 +139,8 @@ impl ChannelManager {
remote_addr,
side,
dropped_unknown_chunks: AtomicU64::new(0),
early_arrivals: Mutex::new(HashMap::new()),
early_arrival_count: AtomicU64::new(0),
}),
}
}
@@ -232,7 +249,7 @@ impl ChannelManager {
let (demux_sender, recv) = MpscRecvStream::channel(self.inner.buffer_cap);
let state = ChannelState {
demux_sender,
demux_sender: demux_sender.clone(),
handler_task,
alpn,
};
@@ -288,7 +305,7 @@ impl ChannelManager {
let (demux_sender, recv) = MpscRecvStream::channel(self.inner.buffer_cap);
let state = ChannelState {
demux_sender,
demux_sender: demux_sender.clone(),
handler_task,
alpn,
};
@@ -305,6 +322,10 @@ impl ChannelManager {
}
}
// Route any chunks that arrived between the open-op response and
// this adopt (the producer may have started pumping already).
self.drain_early_arrivals(channel_id, &demux_sender).await;
Ok((send, recv))
}
@@ -377,8 +398,15 @@ impl ChannelManager {
/// Route a chunk payload to the reassembled stream for
/// `channel_id` (the demux's per-chunk route). A zero-length
/// payload is the EOF sentinel — the reassembled stream interprets
/// it as EOF (REQ-CH-01). An unknown `channel_id` is dropped with
/// a debug log (REQ-CH-04 — lenient handling).
/// it as EOF (REQ-CH-01). An unknown `channel_id` parks the payload
/// in the early-arrival buffer (up to `EARLY_ARRIVAL_CAP` payloads
/// per channel) — the open-op response carrying `channel_id` and
/// the producer's first data-plane write race, and a push-first
/// producer (a TTY backend's banner, a sub protocol's greeting)
/// must not lose its first chunks. `adopt_channel` drains the
/// parked payloads into the new receiver. Chunks beyond the cap
/// drop with a debug log (the adopted-side REQ-CH-04 leniency for
/// genuinely unknown channels).
///
/// Awaits the bounded channel sender — if the handler's read half
/// is slow, the demux stalls here (ADR-040 REQ-CH-05: lossless
@@ -400,17 +428,67 @@ impl ChannelManager {
}
}
None => {
self.inner
.dropped_unknown_chunks
.fetch_add(1, Ordering::Relaxed);
debug!(
channel_id,
"demux: unknown channel_id, dropping chunk (lenient)"
);
self.park_early_arrival(channel_id, payload);
}
}
}
/// Park an early-arrival chunk for `channel_id` (the adopter has
/// not run yet). FIFO per channel, bounded by `EARLY_ARRIVAL_CAP`;
/// beyond the cap the chunk is dropped with a debug log and the
/// dropped counter increments (observability only).
fn park_early_arrival(&self, channel_id: u32, payload: Bytes) {
let mut arrivals = self.inner.early_arrivals.lock();
let queue = arrivals.entry(channel_id).or_default();
if queue.len() >= EARLY_ARRIVAL_CAP {
drop(arrivals);
self.inner
.dropped_unknown_chunks
.fetch_add(1, Ordering::Relaxed);
debug!(
channel_id,
"demux: unknown channel_id, early-arrival buffer full, dropping chunk"
);
return;
}
queue.push_back(payload);
self.inner
.early_arrival_count
.fetch_add(1, Ordering::Relaxed);
debug!(
channel_id,
"demux: unknown channel_id, parking chunk (early arrival, not yet adopted)"
);
}
/// Drain the early-arrival buffer for `channel_id` into `recv`'s
/// sender (called by `adopt_channel` after the channel state is
/// installed). Returns the drained payload count. Payloads past the
/// receiver's buffer capacity await normally (backpressure).
async fn drain_early_arrivals(
&self,
channel_id: u32,
demux_sender: &mpsc::Sender<Bytes>,
) -> usize {
let payloads = self
.inner
.early_arrivals
.lock()
.remove(&channel_id)
.unwrap_or_default();
let count = payloads.len();
for payload in payloads {
if demux_sender.send(payload).await.is_err() {
debug!(
channel_id,
"demux: channel receiver dropped while draining early arrivals"
);
break;
}
}
count
}
/// Take the sender for `channel_id` — used on close / teardown to
/// drop the sender (which signals EOF to the handler, REQ-CH-02).
/// Returns the handler task (if any) so the caller can abort it
@@ -447,6 +525,8 @@ impl ChannelManager {
task.abort();
}
}
// The early-arrival buffers die with the connection too.
self.inner.early_arrivals.lock().clear();
// The opener ledger has the opener PeerIds.
self.inner.opener_ledger.drain()
}
@@ -582,20 +662,53 @@ mod tests {
.await;
}
/// Early arrivals for a not-yet-adopted channel are parked, not
/// dropped: `adopt_channel` drains them into the new receiver. This
/// closes the open-op-response/first-data race (a push-first
/// producer's first chunks must not be lost).
#[tokio::test]
async fn route_payload_to_unknown_channel_parks_until_adopt() {
let manager = make_manager_with_runner().await;
manager.route_payload(7, Bytes::from_static(b"first")).await;
manager
.route_payload(7, Bytes::from_static(b"second"))
.await;
let (_send, mut recv) = manager
.adopt_channel(7, "alk/tty", None)
.await
.expect("adopt");
use tokio::io::AsyncReadExt;
let mut buf = [0u8; 11];
recv.read_exact(&mut buf).await.expect("read first+second");
assert_eq!(&buf, b"firstsecond");
}
#[tokio::test]
async fn route_payload_to_unknown_channel_increments_dropped_counter() {
let manager = make_manager_with_runner().await;
assert_eq!(manager.dropped_unknown_chunks(), 0);
// Parked arrivals don't increment the counter...
manager
.route_payload(999, Bytes::from_static(b"data"))
.await;
manager
.route_payload(998, Bytes::from_static(b"more"))
.route_payload(998, Bytes::from_static(b"kept"))
.await;
assert_eq!(
manager.dropped_unknown_chunks(),
2,
"dropped_unknown_chunks counter increments per unknown-channel drop"
0,
"parked early arrivals are not drops"
);
// ...but chunks beyond the per-channel cap do (the adopted-side
// REQ-CH-04 leniency for genuinely unknown channels).
for i in 0..(super::EARLY_ARRIVAL_CAP + 1) {
let n = i;
let payload = vec![b'x'; n % 9 + 1];
manager.route_payload(999, Bytes::from(payload)).await;
}
assert_eq!(
manager.dropped_unknown_chunks(),
1,
"dropped_unknown_chunks counter increments when the early-arrival buffer overflows"
);
}
+14 -14
View File
@@ -62,7 +62,7 @@ impl ChannelOperations {
/// Register the three generic ops on the call `OperationRegistry`.
/// The per-ALPN open ops are registered separately by the ALPN
/// crates via [`ChannelCore::register_openable`] (ADR-047 §3).
pub fn register_on(&self, registry: &mut OperationRegistry) -> Result<(), String> {
pub fn register_on(&self, registry: &OperationRegistry) -> Result<(), String> {
let manager = self.manager.clone();
let policy = Arc::clone(&self.policy);
registry.register(HandlerRegistration::new(
@@ -400,7 +400,7 @@ impl ChannelCore {
&self,
spec: OperationSpec,
open_handler: OpenHandler,
registry: &mut OperationRegistry,
registry: &OperationRegistry,
auth: AuthContext,
) -> Result<(), String> {
let op_type = spec.op_type;
@@ -765,8 +765,8 @@ mod tests {
async fn register_on_registers_three_ops() {
let manager = make_manager().await;
let ops = ChannelOperations::with_default_policy(manager);
let mut registry = OperationRegistry::new();
ops.register_on(&mut registry).expect("register");
let registry = OperationRegistry::new();
ops.register_on(&registry).expect("register");
assert!(registry.registration(OP_CHANNEL_CLOSE).is_some());
assert!(registry.registration(OP_CHANNEL_CONTROL).is_some());
assert!(registry
@@ -899,8 +899,8 @@ mod tests {
let manager = make_manager().await;
let policy = super::super::policy::default_policy();
let ops = ChannelOperations::new(manager, policy);
let mut registry = OperationRegistry::new();
ops.register_on(&mut registry).expect("register");
let registry = OperationRegistry::new();
ops.register_on(&registry).expect("register");
assert!(registry.registration(OP_CHANNEL_CLOSE).is_some());
}
@@ -1129,11 +1129,11 @@ mod tests {
spawned_clone.store(true, Ordering::SeqCst);
tokio::spawn(async {})
});
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
core.register_openable(
spec,
open_handler,
&mut registry,
&registry,
AuthContext::anonymous(b"alk/call"),
)
.expect("register");
@@ -1172,11 +1172,11 @@ mod tests {
spawned_clone.store(true, Ordering::SeqCst);
tokio::spawn(async {})
});
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
core.register_openable(
spec,
open_handler,
&mut registry,
&registry,
AuthContext::anonymous(b"alk/call"),
)
.expect("register");
@@ -1209,11 +1209,11 @@ mod tests {
)
.with_channel_open(ChannelOpenSpec::new("alk/tty"));
let open_handler: OpenHandler = Arc::new(|_input, _conn, _auth| tokio::spawn(async {}));
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
core.register_openable(
spec,
open_handler,
&mut registry,
&registry,
AuthContext::anonymous(b"alk/call"),
)
.expect("register");
@@ -1239,11 +1239,11 @@ mod tests {
None,
);
let open_handler: OpenHandler = Arc::new(|_input, _conn, _auth| tokio::spawn(async {}));
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let result = core.register_openable(
spec,
open_handler,
&mut registry,
&registry,
AuthContext::anonymous(b"alk/call"),
);
assert!(result.is_err());
+5 -5
View File
@@ -4,7 +4,7 @@
//!
//! The handler for a channel receives a `Connection` constructed via
//! `Connection::from_source(ChannelBidiStreamSource, alpn)`. The
//! handler calls `accept_bi()` once (yield-once per channel, ADR-065)
//! handler calls `accept_bi()` once (yield-once per channel, ADR-007)
//! and gets a `BiStream` — identical to how it works on a top-level
//! QUIC connection. The `BiStream` is the reassembled read half joined
//! to the mux write half via `BiStream::from_joined`.
@@ -25,8 +25,8 @@ use crate::core::types::{BiStream, BidiStreamSource, StreamError};
///
/// The `BiStream` is constructed once, in [`ChannelBidiStreamSource::new`],
/// and held in an `Option` — `accept_bi` takes it. This preserves the
/// yield-once contract (ADR-065) and the "split never crosses a crate
/// boundary as part of a constructor" rule (ADR-092) — the join happens
/// yield-once contract (ADR-007) and the "split never crosses a crate
/// boundary as part of a constructor" rule (ADR-009) — the join happens
/// here, in the `BidiStreamSource` impl, not per-handler.
pub struct ChannelBidiStreamSource {
stream: Mutex<Option<BiStream>>,
@@ -65,9 +65,9 @@ impl BidiStreamSource for ChannelBidiStreamSource {
/// `code`/`reason` are ignored: a single channel has no
/// QUIC-shaped application-level close codes. The drop is the
/// close (ADR-065 §"Negative"). The `_` prefix is intentional —
/// close (ADR-007 §"Negative"). The `_` prefix is intentional —
/// the signature matches the public `Connection::close` API
/// (ADR-070 §"REQ-CORE-02").
/// (ADR-008 §"REQ-CORE-02").
fn close(&self, _code: u32, _reason: &str) {
let _ = self.stream.lock().take();
}
+1 -1
View File
@@ -23,7 +23,7 @@ use thiserror::Error;
pub const CHUNK_HEADER_LEN: usize = 8;
/// The maximum chunk payload length (16 MiB, matching TTY's cap —
/// ADR-052 §5). A chunk with `length > MAX_CHUNK_LEN` returns
/// ADR-034). A chunk with `length > MAX_CHUNK_LEN` returns
/// [`ChunkError::TooLarge`] and does not corrupt the stream — the demux
/// drops the chunk and continues. The header is always exactly 8 bytes,
/// so the demux can always resync by reading the next 8-byte header.
+1 -1
View File
@@ -122,7 +122,7 @@ mod tests {
}
fn registry_with_caps() -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("pub/run"),
+82 -7
View File
@@ -202,7 +202,11 @@ async fn fetch_schema(connection: &CallConnection, name: &str) -> Result<Value,
/// Rebuild an `OperationSpec` from the `services/schema` JSON, applying the
/// optional namespace prefix. The spec JSON shape matches `spec_to_json` in
/// `registry/discovery.rs`.
fn rebuild_spec_for(
///
/// `pub(crate)` so the `op/register` bootstrap path (review 004 F-05) can
/// rebuild peer-announced specs from the same wire shape — one parser, two
/// consumers.
pub(crate) fn rebuild_spec_for(
schema_json: &Value,
remote_name: &str,
namespace_prefix: &Option<String>,
@@ -252,7 +256,10 @@ fn rebuild_spec_for(
output_schema,
error_schemas,
access_control,
None,
schema_json
.get("resource_id_path")
.and_then(|v| v.as_str())
.map(String::from),
);
// ADR-047 §2: the `channel_open` marker survives discovery
@@ -385,7 +392,10 @@ fn parse_access_control(v: &Value) -> AccessControl {
/// If `context.identity` is `None` (the hub chose not to disclose, or has not
/// authenticated an originator), `forwarded_for` is omitted — the spoke
/// receives only the hub's identity.
fn make_forwarding_handler(connection: Arc<CallConnection>, remote_name: String) -> Handler {
pub(crate) fn make_forwarding_handler(
connection: Arc<CallConnection>,
remote_name: String,
) -> Handler {
use crate::registry::registration::make_handler;
make_handler(move |input, context| {
let connection = Arc::clone(&connection);
@@ -411,7 +421,7 @@ fn make_forwarding_handler(connection: Arc<CallConnection>, remote_name: String)
/// leaf: on invocation, calls `CallConnection::subscribe_with_payload()` and
/// forwards the remote stream end-to-end. Each `call.responded` from the
/// remote becomes a stream item, `call.completed` ends the stream, and
/// `call.aborted` drops it (ADR-049 §8). No truncation, no first-value
/// `call.aborted` drops it (ADR-021 §8). No truncation, no first-value
/// fallback.
///
/// `forwarded_for` is populated from `context.identity` (ADR-032 §3), exactly
@@ -421,7 +431,7 @@ fn make_forwarding_handler(connection: Arc<CallConnection>, remote_name: String)
/// `PendingRequestMap`, so the abort cascade (ADR-016 §6) is already wired:
/// a parent abort drops the `SubscriptionStream`, which sends `call.aborted`
/// to the remote node.
fn make_streaming_forwarding_handler(
pub(crate) fn make_streaming_forwarding_handler(
connection: Arc<CallConnection>,
remote_name: String,
) -> StreamingHandler {
@@ -455,7 +465,7 @@ fn make_streaming_forwarding_handler(
/// `forwarded_for` is populated from `context.identity` (ADR-032 §3),
/// exactly as the request/response and streaming forwarding handlers
/// do — both via `build_forwarded_payload`.
fn make_sink_forwarding_handler(
pub(crate) fn make_sink_forwarding_handler(
connection: Arc<CallConnection>,
remote_name: String,
) -> SinkHandler {
@@ -480,7 +490,11 @@ fn make_sink_forwarding_handler(
/// `forwarded_for` from the hub's `OperationContext.identity` (ADR-032 §3).
/// `forwarded_for` is omitted when `context.identity` is `None` (the hub
/// chooses not to disclose the originator).
fn build_forwarded_payload(operation_id: &str, input: Value, context: &OperationContext) -> Value {
pub(crate) fn build_forwarded_payload(
operation_id: &str,
input: Value,
context: &OperationContext,
) -> Value {
let mut payload = serde_json::Map::new();
payload.insert(
"operationId".to_string(),
@@ -591,6 +605,67 @@ mod tests {
assert_eq!(spec.access_control.resource_type.as_deref(), Some("fs"));
}
// --- review 005 Unit 3 gate (G-04 spec wire round-trip) ----------------
/// G-04 gate: `resource_id_path` survives the spec wire round-trip —
/// `spec_to_json_pub` serializes it, `rebuild_spec_for` parses it.
/// The review found the field silently dropped: an announced (or
/// `from_call`-imported) op with ownership-scoped resource
/// extraction rebuilt with `resource_id: None`, so the rebuilt
/// spec's ACL checks ran without the resource ID.
#[test]
fn spec_round_trips_resource_id_path() {
use crate::registry::discovery::spec_to_json_pub;
let spec = OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
crate::registry::spec::AccessControl::default(),
Some("/path".to_string()),
);
let wire = spec_to_json_pub(&spec);
assert_eq!(
wire.get("resource_id_path").and_then(|v| v.as_str()),
Some("/path"),
"resource_id_path serialized"
);
let rebuilt = rebuild_spec_for(&wire, "fs/readFile", &None).expect("rebuild");
assert_eq!(
rebuilt.resource_id_path.as_deref(),
Some("/path"),
"resource_id_path survives the round-trip"
);
}
/// G-04 companion: a spec without `resource_id_path` serializes no
/// `resource_id_path` key and rebuilds with `None` (additive
/// optional field — absent stays absent).
#[test]
fn spec_without_resource_id_path_stays_absent_through_round_trip() {
use crate::registry::discovery::spec_to_json_pub;
let spec = OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
crate::registry::spec::AccessControl::default(),
None,
);
let wire = spec_to_json_pub(&spec);
assert!(wire.get("resource_id_path").is_none());
let rebuilt = rebuild_spec_for(&wire, "fs/readFile", &None).expect("rebuild");
assert_eq!(rebuilt.resource_id_path, None);
}
#[test]
fn rebuild_spec_channel_open_marker_set_for_channels_alpn_op() {
let mut schema = sample_schema_json("channels/tty/sub", "sub");
+5
View File
@@ -10,6 +10,11 @@ mod from_call;
pub use call_client::CallClient;
pub use from_call::{from_call, FromCallConfig};
// crate-internal surface for the `op/register` bootstrap path (review
// 004 F-05): the forwarding-handler constructor and the spec wire
// parser are shared with `registry::op_register`.
pub(crate) use from_call::{make_forwarding_handler, rebuild_spec_for};
use crate::registry::registration::HandlerRegistration;
/// Errors produced by [`OperationAdapter::import`].
+28 -26
View File
@@ -227,14 +227,14 @@ pub trait ProtocolHandler: Send + Sync + 'static {
async fn handle(&self, connection: Connection, auth: &AuthContext) -> Result<(), HandlerError>;
}
// --- BiStream: the handler leaf (ADR-092) ---------------------------------
// --- BiStream: the handler leaf (ADR-009) ---------------------------------
/// Internal helper trait — the union of `AsyncRead + AsyncWrite + Send`.
/// Not public; exists only to give `BiStream` a single boxed field.
trait AsyncReadWrite: AsyncRead + AsyncWrite + Send {}
impl<T: AsyncRead + AsyncWrite + Send> AsyncReadWrite for T {}
/// The handler leaf — a bidirectional byte stream (ADR-092).
/// The handler leaf — a bidirectional byte stream (ADR-005, ADR-009).
///
/// `accept_bi`/`open_bi` return a `BiStream`, not a split
/// `(SendStream, RecvStream)` pair. Handlers that want the split halves call
@@ -316,7 +316,7 @@ impl AsyncWrite for BiStream {
}
}
// --- SendStream / RecvStream: thin newtypes (ADR-092) ---------------------
// --- SendStream / RecvStream: thin newtypes (ADR-009) ---------------------
pub struct SendStream {
inner: Box<dyn AsyncWrite + Send + Unpin>,
@@ -327,10 +327,11 @@ pub struct RecvStream {
}
impl SendStream {
/// Box a write half into the thin `SendStream` newtype. Used by
/// `into_sub_streams()` (ADR-074) and the channels reassembly path.
/// Not a constructor that feeds `Connection` — the split never crosses
/// a crate boundary as part of a constructor (ADR-092).
/// Box a write half into the thin `SendStream` newtype. Used by tests
/// and any consumer that holds pre-split halves (the channels layer
/// has its own `MpscSendStream`; `into_sub_streams()` was removed by
/// ADR-035). Not a constructor that feeds `Connection` — the split
/// never crosses a crate boundary as part of a constructor (ADR-009).
pub fn from_stream(stream: impl AsyncWrite + Send + Unpin + 'static) -> Self {
Self {
inner: Box::new(stream),
@@ -339,10 +340,11 @@ impl SendStream {
}
impl RecvStream {
/// Box a read half into the thin `RecvStream` newtype. Used by
/// `into_sub_streams()` (ADR-074) and the channels reassembly path.
/// Not a constructor that feeds `Connection` — the split never crosses
/// a crate boundary as part of a constructor (ADR-092).
/// Box a read half into the thin `RecvStream` newtype. Used by tests
/// and any consumer that holds pre-split halves (the channels layer
/// has its own `MpscRecvStream`; `into_sub_streams()` was removed by
/// ADR-035). Not a constructor that feeds `Connection` — the split
/// never crosses a crate boundary as part of a constructor (ADR-009).
pub fn from_stream(stream: impl AsyncRead + Send + Unpin + 'static) -> Self {
Self {
inner: Box::new(stream),
@@ -386,15 +388,15 @@ impl AsyncRead for RecvStream {
/// Yield bidirectional streams to a `Connection`. Downstream crates implement
/// this trait to add connection shapes (channels, a future transport, a test
/// double beyond the single-stream case) without editing core. See ADR-070
/// for the full rationale and ADR-065 for the yield-once contract the
/// double beyond the single-stream case) without editing core. See ADR-008
/// for the full rationale and ADR-007 for the yield-once contract the
/// `StreamBidiStreamSource` impl preserves. The return type is `BiStream`
/// (ADR-092) — the join happens once, in the impl, not per-handler.
/// (ADR-009) — the join happens once, in the impl, not per-handler.
#[async_trait]
pub trait BidiStreamSource: Send + Sync + 'static {
/// Yield the next bidirectional stream this connection provides.
///
/// Transport semantics (carried from ADR-065):
/// Transport semantics (carried from ADR-007):
/// - QUIC (quinn/iroh): returns a new bidi stream on each call,
/// `ConnectionClosed` when the underlying connection closes.
/// - Single-stream (TCP+TLS, SSH channel, WebTransport stream, wasm):
@@ -407,7 +409,7 @@ pub trait BidiStreamSource: Send + Sync + 'static {
/// Open a bidirectional stream to the peer.
///
/// Single-stream sources return `StreamClosed` (a single stream cannot
/// open new application streams — ADR-065). QUIC and channels sources
/// open new application streams — ADR-007). QUIC and channels sources
/// open new streams.
async fn open_bi(&self) -> Result<BiStream, StreamError>;
@@ -416,13 +418,13 @@ pub trait BidiStreamSource: Send + Sync + 'static {
/// Close the connection. The `code`/`reason` args are QUIC application-
/// level close codes; non-QUIC sources ignore them (the drop is the
/// close — ADR-065 §"Negative"). See ADR-070 §"REQ-CORE-02" for the
/// close — ADR-007 §"Negative"). See ADR-008 §"REQ-CORE-02" for the
/// rationale for keeping the QUIC-shaped signature on the trait.
fn close(&self, code: u32, reason: &str);
}
/// Single-stream `BidiStreamSource` (TCP+TLS, SSH channel, WebTransport
/// stream, wasm stream — ADR-065). Crate-private; constructed via
/// stream, wasm stream — ADR-007). Crate-private; constructed via
/// `Connection::from_bidi`. `accept_bi` yields the underlying `BiStream`
/// once, then `ConnectionClosed`; `open_bi` returns `StreamClosed`.
struct StreamBidiStreamSource {
@@ -449,9 +451,9 @@ impl BidiStreamSource for StreamBidiStreamSource {
}
/// `code`/`reason` are ignored: a single stream has no QUIC-shaped
/// application-level close codes. The drop is the close (ADR-065
/// application-level close codes. The drop is the close (ADR-007
/// §"Negative"). The `_` prefix is intentional — the signature matches
/// the public `Connection::close` API (ADR-070 §"REQ-CORE-02").
/// the public `Connection::close` API (ADR-008 §"REQ-CORE-02").
fn close(&self, _code: u32, _reason: &str) {
let _ = self.stream.lock().unwrap_or_else(|e| e.into_inner()).take();
}
@@ -467,11 +469,11 @@ impl Connection {
/// Construct a `Connection` from a single bidirectional stream (e.g.
/// `tokio::io::DuplexStream`, `TlsStream<TcpStream>`,
/// `russh::Channel::into_stream()`). The stream is wrapped in a
/// `BiStream` (ADR-092) and yielded by `accept_bi` once, then
/// `BiStream` (ADR-009) and yielded by `accept_bi` once, then
/// `ConnectionClosed`. `open_bi` returns `StreamClosed` (a single
/// stream can't open new application streams — ADR-065).
/// stream can't open new application streams — ADR-007).
///
/// This is the only public stream constructor (ADR-092): the split
/// This is the only public stream constructor (ADR-009): the split
/// never crosses a crate boundary as part of a constructor. Handlers
/// that want the split halves call `tokio::io::split(&mut *stream)` on
/// the `BiStream` they receive from `accept_bi`.
@@ -492,7 +494,7 @@ impl Connection {
/// Construct from a caller-supplied `BidiStreamSource` impl. The
/// extension point for downstream crates — implement the trait and
/// construct a `Connection` from it without editing core. See ADR-070.
/// construct a `Connection` from it without editing core. See ADR-008.
pub fn from_source(source: impl BidiStreamSource, alpn: Vec<u8>) -> Self {
Self {
source: Box::new(source),
@@ -514,7 +516,7 @@ impl Connection {
/// Handlers that loop `accept_bi` (e.g. `TtyAdapter`) get one session
/// per single-stream connection; handlers that call once (e.g.
/// `HttpAdapter`) get the stream directly. Both are correct. The
/// return type is `BiStream` (ADR-092); handlers that want the split
/// return type is `BiStream` (ADR-009); handlers that want the split
/// halves call `tokio::io::split` on the `BiStream`.
pub async fn accept_bi(&self) -> Result<BiStream, StreamError> {
self.source.accept_bi().await
@@ -656,7 +658,7 @@ mod tests {
/// `tokio::io::sink()` + `tokio::io::empty()`: reads yield EOF
/// immediately (zero bytes), writes discard. Exists because
/// `Connection::from_bidi` requires a single value that implements
/// both traits (ADR-092). Used only to construct a `Connection` for
/// both traits (ADR-009). Used only to construct a `Connection` for
/// tests that exercise `Connection`-level state (alpn, addr, identity)
/// without ever reading or writing the stream.
pub(crate) struct SinkEmpty;
+787
View File
@@ -0,0 +1,787 @@
//! The transport-neutral dispatch spine ([`GatewayDispatch`]) and the
//! shared `services/schema` disclosure guard
//! ([`schema_disclosure_denial`]) — the non-HTTP half of alkhttp's
//! gateway, promoted for reuse by hubs, relays, and protocol crates
//! that expose a call surface to a less-trusted in-transport caller.
//!
//! The spine constructs the root [`OperationContext`] identically for
//! every transport (`internal: false` — ACL runs against the caller's
//! identity, not a handler's composition authority; `forwarded_for:
//! None` — wire-ingress only), resolves the registration's
//! composition authority / capabilities / scoped env into it, and
//! bounds Once-ops and sink dispatch with a configurable deadline
//! while leaving streaming subscriptions unbounded (ADR-021:
//! subscriptions are long-lived).
//!
//! There are no HTTP concepts here: no statuses, no headers, no body
//! framing. CallError → HTTP mapping, NDJSON/SSE framing, body caps,
//! and batch envelopes are the consumer's (see alkhttp's gateway).
//!
//! See ADR-048 for the promotion decision and the divergence note on
//! ACL-denial codes (`FORBIDDEN` here, spec-404 in the wire handler).
use std::collections::HashMap;
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::Arc;
use std::time::{Duration, Instant};
use crate::core::auth::Identity;
use crate::core::types::Capabilities;
use crate::protocol::wire::{CallError, ResponseEnvelope};
use crate::registry::context::{AbortPolicy, OperationContext, ScopedPeerEnv};
use crate::registry::env::LocalOperationEnv;
use crate::registry::registration::OperationRegistry;
use crate::registry::spec::{AccessResult, Visibility};
use futures::stream::BoxStream;
use serde_json::Value;
const SERVICES_SCHEMA: &str = "services/schema";
/// The default handler deadline for Once-ops and sink dispatch: 30 s.
/// Override with [`GatewayDispatch::with_deadline`].
pub const DEFAULT_DEADLINE: Duration = Duration::from_secs(30);
/// The transport-neutral dispatch spine over an
/// [`OperationRegistry`](crate::registry::registration::OperationRegistry):
/// invokes operations for the neutral `ResponseEnvelope` result shape
/// any transport maps to its own wire format. Identity arrives
/// per-call as `Option<Identity>` — bearer resolution, transport
/// framing, and error presentation happen upstream in the consumer.
///
/// See the [module docs](crate::gateway) and ADR-048.
pub struct GatewayDispatch {
registry: Arc<OperationRegistry>,
deadline: Option<Duration>,
invoke_count: AtomicUsize,
}
impl GatewayDispatch {
/// Assemble a dispatch spine over a registry with the 30 s default
/// deadline ([`DEFAULT_DEADLINE`]).
pub fn new(registry: Arc<OperationRegistry>) -> Self {
Self {
registry,
deadline: Some(DEFAULT_DEADLINE),
invoke_count: AtomicUsize::new(0),
}
}
/// Override the Once/sink deadline: `Some(duration)` bounds every
/// [`GatewayDispatch::invoke`] and [`GatewayDispatch::invoke_sink`]
/// dispatch (a hung handler surfaces as a retryable `TIMEOUT` error
/// envelope); `None` removes the bound entirely (the wire path's
/// Pub dispatch behavior). Streaming subscriptions are unbounded
/// either way (ADR-021).
pub fn with_deadline(mut self, deadline: Option<Duration>) -> Self {
self.deadline = deadline;
self
}
/// The registry operations resolve against.
pub fn registry(&self) -> &Arc<OperationRegistry> {
&self.registry
}
/// How many [`GatewayDispatch::invoke`] calls this spine has
/// served. A test-spy accessor: consumers' over-cap batch tests
/// assert it stays at zero to prove no dispatch happened before the
/// cap rejection.
pub fn invoke_count(&self) -> usize {
self.invoke_count.load(Ordering::Relaxed)
}
/// Invoke a Query/Mutation op under the configured deadline; a hung
/// handler surfaces as a retryable `TIMEOUT` error envelope.
pub async fn invoke(
&self,
identity: Option<Identity>,
op: &str,
input: Value,
) -> ResponseEnvelope {
self.invoke_count.fetch_add(1, Ordering::Relaxed);
let operation_name = strip_leading_slash(op).to_string();
if let Some(error) =
schema_via_call_denial(&self.registry, &operation_name, &input, identity.as_ref())
{
return ResponseEnvelope::error(uuid::Uuid::new_v4().to_string(), error);
}
let request_id = uuid::Uuid::new_v4().to_string();
let context = self.build_root_context(&request_id, &operation_name, identity);
let fut = self.registry.invoke(&operation_name, input, context);
match self.deadline {
Some(deadline) => match tokio::time::timeout(deadline, fut).await {
Ok(envelope) => envelope,
Err(_elapsed) => ResponseEnvelope::error(
request_id,
CallError::timeout(format!(
"operation did not complete within the {deadline:?} dispatch deadline"
)),
),
},
None => fut.await,
}
}
/// Dispatch a Sub op: the returned stream of envelopes is unbounded
/// by the deadline (subscriptions are long-lived per ADR-021);
/// pre-handler failures surface as one error envelope.
pub fn invoke_streaming(
&self,
identity: Option<Identity>,
op: &str,
input: Value,
) -> BoxStream<'static, ResponseEnvelope> {
let operation_name = strip_leading_slash(op).to_string();
if let Some(error) =
schema_via_call_denial(&self.registry, &operation_name, &input, identity.as_ref())
{
let request_id = uuid::Uuid::new_v4().to_string();
return Box::pin(futures::stream::once(async move {
ResponseEnvelope::error(request_id, error)
}));
}
let request_id = uuid::Uuid::new_v4().to_string();
let context = self.build_root_context_streaming(&request_id, &operation_name, identity);
self.registry
.invoke_streaming(&operation_name, input, context)
}
/// Dispatch a `Pub` operation (ADR-046). The `publish_stream` is
/// the initiator's chunk stream (each item one published chunk, or
/// an initiator-side error). The dispatch — chunk pacing and the
/// sink handler's final completion alike — is bounded by the
/// configured deadline: a hung sink handler surfaces as a retryable
/// `TIMEOUT` error envelope, not an indefinitely-held initiator.
pub async fn invoke_sink(
&self,
identity: Option<Identity>,
op: &str,
input: Value,
publish_stream: crate::registry::registration::PublishStream,
) -> ResponseEnvelope {
let operation_name = strip_leading_slash(op).to_string();
let request_id = uuid::Uuid::new_v4().to_string();
let context = self.build_root_context_sink(&request_id, &operation_name, identity);
let fut = self
.registry
.invoke_sink(&operation_name, input, publish_stream, context);
match self.deadline {
Some(deadline) => match tokio::time::timeout(deadline, fut).await {
Ok(envelope) => envelope,
Err(_elapsed) => ResponseEnvelope::error(
request_id,
CallError::timeout(format!(
"operation did not complete within the {deadline:?} dispatch deadline"
)),
),
},
None => fut.await,
}
}
fn build_root_context_sink(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
) -> OperationContext {
self.build_root_context_inner(request_id, operation_name, identity, false)
}
fn build_root_context(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
) -> OperationContext {
self.build_root_context_inner(request_id, operation_name, identity, true)
}
fn build_root_context_streaming(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
) -> OperationContext {
self.build_root_context_inner(request_id, operation_name, identity, false)
}
fn build_root_context_inner(
&self,
request_id: &str,
operation_name: &str,
identity: Option<Identity>,
bounded: bool,
) -> OperationContext {
let registration = self.registry.registration(operation_name);
let (composition_authority, capabilities, scoped_env) = match registration {
Some(r) => (
r.composition_authority.clone(),
r.capabilities.clone(),
r.scoped_env.clone().unwrap_or_else(ScopedPeerEnv::empty),
),
None => (None, Capabilities::new(), ScopedPeerEnv::empty()),
};
let env: Arc<dyn crate::registry::env::OperationEnv + Send + Sync> =
Arc::new(LocalOperationEnv::new(Arc::clone(&self.registry)));
OperationContext {
request_id: request_id.to_string(),
parent_request_id: None,
identity,
handler_identity: composition_authority,
forwarded_for: None,
capabilities,
metadata: HashMap::new(),
deadline: bounded
.then_some(self.deadline)
.flatten()
.map(|deadline| Instant::now() + deadline),
scoped_env,
env,
abort_policy: AbortPolicy::default(),
internal: false,
ownership: None,
}
}
}
fn strip_leading_slash(operation_id: &str) -> &str {
operation_id.strip_prefix('/').unwrap_or(operation_id)
}
/// The `services/schema` op-path guard (ADR-048): when the dispatched
/// operation is the schema meta-op, apply the disclosure checks to the
/// **inner** `name` input before dispatching, so no transport through
/// this spine can fetch a spec that would be denied for the same
/// identity elsewhere.
fn schema_via_call_denial(
registry: &OperationRegistry,
operation: &str,
input: &Value,
identity: Option<&Identity>,
) -> Option<CallError> {
let name = input.get("name").and_then(Value::as_str)?;
if !matches!(
registry.registration(operation),
Some(registration) if registration.spec.name == SERVICES_SCHEMA
) {
return None;
}
schema_disclosure_denial(registry, name, identity)
}
/// The schema-disclosure check shared by every transport surface:
/// Internal ops are invisible (`NOT_FOUND` regardless of caller);
/// ACL-forbidden ops are denied with `FORBIDDEN`. One implementation so
/// the transports cannot drift (CF-004's shared-implementation fix; the
/// wire-path handler stays the conservative spec-404 outer bound).
pub fn schema_disclosure_denial(
registry: &OperationRegistry,
operation: &str,
identity: Option<&Identity>,
) -> Option<CallError> {
let name = strip_leading_slash(operation);
let registration = registry.registration(name)?;
if registration.spec.visibility == Visibility::Internal {
return Some(CallError::not_found(operation));
}
if let AccessResult::Forbidden(message) =
registration.spec.access_control.check(identity, None, None)
{
return Some(CallError::forbidden(message));
}
None
}
#[cfg(test)]
mod tests {
use super::*;
use crate::registry::registration::{
make_handler, make_sink_handler, make_streaming_handler, HandlerKind, HandlerRegistration,
OperationProvenance,
};
use crate::registry::spec::{AccessControl, OperationSpec, OperationType, Visibility};
use futures::StreamExt;
// The deadline tests use a short custom deadline instead of the
// module default, so no wall-clock assertion by hand is needed.
fn spec(name: &str, visibility: Visibility, op_type: OperationType) -> OperationSpec {
OperationSpec::new(
name,
op_type,
visibility,
serde_json::json!({}),
serde_json::json!({}),
vec![],
AccessControl::default(),
None,
)
}
fn echo_registry() -> Arc<OperationRegistry> {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("echo/run", Visibility::External, OperationType::Query),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
Arc::new(registry)
}
#[tokio::test]
async fn invoke_external_op_round_trips() {
let dispatch = GatewayDispatch::new(echo_registry());
let envelope = dispatch
.invoke(None, "/echo/run", serde_json::json!({ "x": 1 }))
.await;
assert!(envelope.result.is_ok(), "expected ok, got {envelope:?}");
}
#[tokio::test]
async fn invoke_unknown_op_returns_not_found() {
let dispatch = GatewayDispatch::new(echo_registry());
let envelope = dispatch
.invoke(None, "/missing/op", serde_json::json!({}))
.await;
match envelope.result {
Err(error) => assert_eq!(error.code, "NOT_FOUND"),
Ok(v) => panic!("expected error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_internal_op_returns_not_found() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("internal/op", Visibility::Internal, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry));
let envelope = dispatch
.invoke(None, "/internal/op", serde_json::json!({}))
.await;
match envelope.result {
Err(error) => assert_eq!(error.code, "NOT_FOUND"),
Ok(v) => panic!("expected error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_enforces_the_configured_deadline_on_a_hung_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("hung/op", Visibility::External, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
tokio::time::sleep(Duration::from_secs(120)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch =
GatewayDispatch::new(Arc::new(registry)).with_deadline(Some(Duration::from_millis(50)));
let started = std::time::Instant::now();
let envelope = dispatch
.invoke(None, "/hung/op", serde_json::json!({}))
.await;
assert!(
started.elapsed() < Duration::from_secs(10),
"the 50 ms deadline must fire well before the 120 s handler sleep"
);
match envelope.result {
Err(error) => {
assert_eq!(error.code, "TIMEOUT");
assert!(error.retryable, "the deadline error is retryable");
}
Ok(v) => panic!("expected a TIMEOUT error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_completes_within_the_deadline_for_a_fast_handler() {
let dispatch = GatewayDispatch::new(echo_registry());
let envelope = dispatch
.invoke(None, "/echo/run", serde_json::json!({}))
.await;
assert!(envelope.result.is_ok(), "a fast handler must not time out");
}
#[tokio::test]
async fn invoke_with_no_deadline_completes_a_slow_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("slow/op", Visibility::External, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
tokio::time::sleep(Duration::from_millis(80)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry)).with_deadline(None);
let envelope = dispatch
.invoke(None, "/slow/op", serde_json::json!({}))
.await;
assert!(
envelope.result.is_ok(),
"a no-deadline spine must not time out, got {envelope:?}"
);
}
#[tokio::test]
async fn invoke_sink_enforces_the_configured_deadline_on_a_hung_sink_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("hung/sink", Visibility::External, OperationType::Pub),
HandlerKind::Sink(make_sink_handler(|_input, ctx, _chunks| async move {
tokio::time::sleep(Duration::from_secs(120)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch =
GatewayDispatch::new(Arc::new(registry)).with_deadline(Some(Duration::from_millis(50)));
let chunks: crate::registry::registration::PublishStream =
Box::pin(futures::stream::empty());
let started = std::time::Instant::now();
let envelope = dispatch
.invoke_sink(None, "/hung/sink", serde_json::json!({}), chunks)
.await;
assert!(
started.elapsed() < Duration::from_secs(10),
"the 50 ms deadline must fire well before the 120 s handler sleep"
);
match envelope.result {
Err(error) => {
assert_eq!(error.code, "TIMEOUT");
assert!(error.retryable, "the deadline error is retryable");
}
Ok(v) => panic!("expected a TIMEOUT error, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_sink_completes_within_the_deadline_for_a_fast_sink_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("fast/sink", Visibility::External, OperationType::Pub),
HandlerKind::Sink(make_sink_handler(|_input, ctx, mut chunks| async move {
let mut count = 0u64;
while let Some(chunk) = chunks.next().await {
if chunk.is_ok() {
count += 1;
}
}
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({ "count": count }))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry));
let chunks: crate::registry::registration::PublishStream =
Box::pin(futures::stream::iter(vec![
Ok(serde_json::json!({ "n": 1 })),
Ok(serde_json::json!({ "n": 2 })),
]));
let envelope = dispatch
.invoke_sink(None, "/fast/sink", serde_json::json!({}), chunks)
.await;
assert!(
envelope.result.is_ok(),
"a fast sink handler must not time out, got {envelope:?}"
);
}
#[tokio::test]
async fn invoke_sink_with_no_deadline_completes_a_slow_sink_handler() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("slow/sink", Visibility::External, OperationType::Pub),
HandlerKind::Sink(make_sink_handler(|_input, ctx, mut chunks| async move {
while let Some(chunk) = chunks.next().await {
let _ = chunk;
}
tokio::time::sleep(Duration::from_millis(80)).await;
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry)).with_deadline(None);
let chunks: crate::registry::registration::PublishStream = Box::pin(futures::stream::iter(
vec![Ok(serde_json::json!({ "n": 1 }))],
));
let envelope = dispatch
.invoke_sink(None, "/slow/sink", serde_json::json!({}), chunks)
.await;
assert!(
envelope.result.is_ok(),
"a no-deadline spine must not time out a slow sink, got {envelope:?}"
);
}
#[tokio::test]
async fn streaming_sub_op_streams_envelopes() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("tick/stream", Visibility::External, OperationType::Sub),
HandlerKind::Stream(make_streaming_handler(|input, ctx| {
let count = input.get("count").and_then(|v| v.as_u64()).unwrap_or(2);
let request_id = ctx.request_id.clone();
futures::stream::iter(0..count).map(move |i| {
ResponseEnvelope::ok(request_id.clone(), serde_json::json!({ "tick": i }))
})
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let dispatch = GatewayDispatch::new(Arc::new(registry));
let mut stream =
dispatch.invoke_streaming(None, "/tick/stream", serde_json::json!({ "count": 3 }));
let mut ticks = Vec::new();
while let Some(envelope) = stream.next().await {
ticks.push(envelope);
}
assert_eq!(ticks.len(), 3);
}
#[tokio::test]
async fn invoke_count_counts_invokes_only() {
let dispatch = GatewayDispatch::new(echo_registry());
let _ = dispatch
.invoke(None, "/echo/run", serde_json::json!({}))
.await;
let _ = dispatch.invoke_streaming(None, "/echo/run", serde_json::json!({}));
assert_eq!(dispatch.invoke_count(), 1);
}
fn registry_with_services_schema_over(inner_ops: Vec<OperationSpec>) -> Arc<OperationRegistry> {
use crate::registry::discovery::{services_schema_handler, services_schema_spec};
let inner = Arc::new({
let registry = OperationRegistry::new();
for op_spec in &inner_ops {
registry
.register(HandlerRegistration::new(
op_spec.clone(),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
}
registry
});
let registry = OperationRegistry::new();
for op_spec in &inner_ops {
registry
.register(HandlerRegistration::new(
op_spec.clone(),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
}
registry
.register(HandlerRegistration::new(
services_schema_spec(),
HandlerKind::Once(services_schema_handler(Arc::clone(&inner))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
Arc::new(registry)
}
#[tokio::test]
async fn invoke_of_services_schema_with_internal_inner_name_is_blocked_pre_dispatch() {
let registry = registry_with_services_schema_over(vec![spec(
"secret/op",
Visibility::Internal,
OperationType::Query,
)]);
let dispatch = GatewayDispatch::new(registry);
let envelope = dispatch
.invoke(
None,
"services/schema",
serde_json::json!({ "name": "secret/op" }),
)
.await;
match envelope.result {
Err(error) => {
assert_eq!(error.code, "NOT_FOUND");
assert!(error.message.contains("secret/op"));
}
Ok(v) => panic!("the internal op spec must not be returned, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_of_services_schema_with_authorized_inner_name_still_projects() {
let registry = registry_with_services_schema_over(vec![spec(
"public/op",
Visibility::External,
OperationType::Query,
)]);
let dispatch = GatewayDispatch::new(registry);
let envelope = dispatch
.invoke(
None,
"services/schema",
serde_json::json!({ "name": "public/op" }),
)
.await;
assert!(
envelope.result.is_ok(),
"an allowed inner name must still project, got {envelope:?}"
);
assert_eq!(
envelope
.result
.as_ref()
.ok()
.and_then(|v| v.get("name"))
.and_then(Value::as_str),
Some("public/op")
);
}
#[tokio::test]
async fn invoke_of_services_schema_with_forbidden_inner_name_is_denied() {
let restricted = OperationSpec::new(
"admin/op",
OperationType::Query,
Visibility::External,
serde_json::json!({}),
serde_json::json!({}),
vec![],
AccessControl {
required_scopes: vec!["admin".to_string()],
..Default::default()
},
None,
);
let registry = registry_with_services_schema_over(vec![restricted]);
let dispatch = GatewayDispatch::new(registry);
let envelope = dispatch
.invoke(
Some(Identity {
id: "user".to_string(),
scopes: vec!["user".to_string()],
resources: HashMap::new(),
}),
"services/schema",
serde_json::json!({ "name": "admin/op" }),
)
.await;
match envelope.result {
Err(error) => assert_eq!(error.code, "FORBIDDEN"),
Ok(v) => panic!("the ACL-denied spec must not be returned, got {v:?}"),
}
}
#[tokio::test]
async fn invoke_streaming_of_services_schema_with_internal_inner_name_is_blocked() {
let registry = registry_with_services_schema_over(vec![spec(
"secret/op",
Visibility::Internal,
OperationType::Query,
)]);
let dispatch = GatewayDispatch::new(registry);
let mut stream = dispatch.invoke_streaming(
None,
"services/schema",
serde_json::json!({ "name": "secret/op" }),
);
let envelopes: Vec<ResponseEnvelope> = stream.by_ref().collect().await;
assert_eq!(envelopes.len(), 1, "the guard error is the only item");
match &envelopes[0].result {
Err(error) => assert_eq!(error.code, "NOT_FOUND"),
Ok(v) => panic!("the internal op spec must not be returned, got {v:?}"),
}
}
#[test]
fn schema_disclosure_denial_hides_internal_ops() {
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
spec("secret/op", Visibility::Internal, OperationType::Query),
HandlerKind::Once(make_handler(|_input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, serde_json::json!({}))
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let error = schema_disclosure_denial(&registry, "secret/op", None).unwrap();
assert_eq!(error.code, "NOT_FOUND");
assert!(schema_disclosure_denial(&registry, "missing/op", None).is_none());
}
#[test]
fn schema_disclosure_denial_allows_unrestricted_ops_without_identity() {
let error = schema_disclosure_denial(&echo_registry(), "echo/run", None);
assert!(
error.is_none(),
"default-ACL ops are fetchable unauthenticated"
);
}
}
+6
View File
@@ -0,0 +1,6 @@
//! The transport-neutral dispatch spine and the shared
//! `services/schema` disclosure guard (ADR-048).
pub(crate) mod dispatch;
pub use dispatch::{schema_disclosure_denial, GatewayDispatch, DEFAULT_DEADLINE};
+6
View File
@@ -17,6 +17,10 @@
//! going forward.
//! - **Registry** ([`registry`]): operation specs, context, dispatch, and
//! the operation registry — the call half's dispatch core.
//! - **Gateway** (module `gateway`, feature `gateway`): the transport-neutral
//! dispatch spine and `services/schema` disclosure guard — the
//! deadline-bounded, re-rooted-context invoke surface for HTTP
//! gateways, hub relays, and other transport front-ends (ADR-048).
//! - **Protocol** ([`protocol`]): wire format, streams, adapter, dispatch
//! loop, pending requests, abort cascade — the call half's wire layer.
//! - **Client** ([`client`]): `CallClient`, `from_call`, `OperationAdapter`
@@ -71,6 +75,8 @@
pub mod channels;
pub mod client;
pub mod core;
#[cfg(feature = "gateway")]
pub mod gateway;
pub mod protocol;
pub mod registry;
+2 -2
View File
@@ -253,7 +253,7 @@ mod tests {
acl: AccessControl,
handler: crate::registry::registration::Handler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
@@ -404,7 +404,7 @@ mod tests {
#[tokio::test]
async fn build_root_context_carries_capabilities_and_scoped_env() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let scoped = ScopedPeerEnv::new(["fs/readFile"]);
let caps = Capabilities::new().with_api_key("google", "k".to_string());
registry
+409 -26
View File
@@ -164,6 +164,18 @@ impl CallConnection {
self.imported_operations.write().insert(name, registration);
}
/// `true` when the connection overlay holds a registration for
/// `name` — the collision check `op/register`'s replace semantics
/// need (review 004 F-05).
pub fn overlay_contains(&self, name: &str) -> bool {
self.imported_operations.read().contains_key(name)
}
/// The connection overlay's registration for `name`, if present.
pub fn overlay_registration(&self, name: &str) -> Option<HandlerRegistration> {
self.imported_operations.read().get(name).cloned()
}
pub fn register_imported_all(&self, registrations: Vec<HandlerRegistration>) {
let mut overlay = self.imported_operations.write();
for reg in registrations {
@@ -213,7 +225,7 @@ impl CallConnection {
}
};
// `open_bi` returns a `BiStream` (ADR-092); split it into halves for
// `open_bi` returns a `BiStream` (ADR-009); split it into halves for
// the call protocol's separate write (request) and read (response)
// pumps. The split is the stdlib idiom; no per-handler wrapper.
let stream = match connection.open_bi().await {
@@ -235,7 +247,7 @@ impl CallConnection {
};
if let Err(err) = self.write_request(send, &request_id, payload).await {
let call_error = CallError::internal(err);
let call_error = CallError::connection_closed(err);
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -276,7 +288,8 @@ impl CallConnection {
let envelope = EventEnvelope::requested(&request_id, payload);
if let Err(err) = writer.write_frame(&envelope).await {
let call_error = CallError::internal(format!("failed to write request frame: {err}"));
let call_error =
CallError::connection_closed(format!("failed to write request frame: {err}"));
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -336,7 +349,7 @@ impl CallConnection {
}
};
// `open_bi` returns a `BiStream` (ADR-092); split for the separate
// `open_bi` returns a `BiStream` (ADR-009); split for the separate
// write (request) and read (subscription events) pumps.
let stream = match connection.open_bi().await {
Ok(s) => s,
@@ -353,7 +366,7 @@ impl CallConnection {
};
if let Err(err) = self.write_request(send, &request_id, payload).await {
let call_error = CallError::internal(err);
let call_error = CallError::connection_closed(err);
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -385,7 +398,8 @@ impl CallConnection {
let envelope = EventEnvelope::requested(&request_id, payload);
if let Err(err) = writer.write_frame(&envelope).await {
let call_error = CallError::internal(format!("failed to write request frame: {err}"));
let call_error =
CallError::connection_closed(format!("failed to write request frame: {err}"));
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -498,8 +512,17 @@ impl CallConnection {
});
let write_result = pump_publish_to_wire(send, &request_id, payload, stream).await;
if let Err(err) = write_result {
let call_error = CallError::internal(err);
if let Err((stage, err)) = write_result {
// A write failure before the request frame could be delivered is
// provably undelivered — retry is safe. Mid-publish failures leave
// delivery ambiguous; the INTERNAL mapping (non-retryable) is the
// safe default there.
let call_error = match stage {
PublishWriteStage::Request => CallError::connection_closed(err),
PublishWriteStage::Published | PublishWriteStage::Completed => {
CallError::internal(err)
}
};
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -533,7 +556,8 @@ impl CallConnection {
let requested = EventEnvelope::requested(&request_id, payload);
if let Err(err) = writer.write_frame(&requested).await {
let call_error = CallError::internal(format!("failed to write request frame: {err}"));
let call_error =
CallError::connection_closed(format!("failed to write request frame: {err}"));
self.pending
.lock()
.handle_error(&request_id, call_error.clone());
@@ -588,7 +612,7 @@ impl CallConnection {
.connection
.as_ref()
.ok_or_else(|| "no underlying connection (overlay-only)".to_string())?;
// `open_bi` returns a `BiStream` (ADR-092). We only need the write
// `open_bi` returns a `BiStream` (ADR-009). We only need the write
// half to send the envelope; split and drop the read half.
let stream = connection
.open_bi()
@@ -603,33 +627,48 @@ impl CallConnection {
}
}
/// Which frame of a publish pump failed to write. The `Request` stage is
/// provably undelivered (the producer never received `call.requested`, so a
/// reconnect/retry is safe); the later stages leave delivery ambiguous.
enum PublishWriteStage {
Request,
Published,
Completed,
}
async fn pump_publish_to_wire<W>(
send: W,
request_id: &str,
payload: Value,
mut stream: Pin<Box<dyn Stream<Item = Value> + Send>>,
) -> Result<(), String>
) -> Result<(), (PublishWriteStage, String)>
where
W: AsyncWrite + Unpin,
{
let mut writer = FrameFramedWriter::new(send);
let requested = EventEnvelope::requested(request_id, payload);
writer
.write_frame(&requested)
.await
.map_err(|e| format!("failed to write request frame: {e}"))?;
writer.write_frame(&requested).await.map_err(|e| {
(
PublishWriteStage::Request,
format!("failed to write request frame: {e}"),
)
})?;
while let Some(chunk) = stream.next().await {
let envelope = EventEnvelope::published(request_id, chunk);
writer
.write_frame(&envelope)
.await
.map_err(|e| format!("failed to write published frame: {e}"))?;
writer.write_frame(&envelope).await.map_err(|e| {
(
PublishWriteStage::Published,
format!("failed to write published frame: {e}"),
)
})?;
}
let completed = EventEnvelope::completed(request_id);
writer
.write_frame(&completed)
.await
.map_err(|e| format!("failed to write completed frame: {e}"))?;
writer.write_frame(&completed).await.map_err(|e| {
(
PublishWriteStage::Completed,
format!("failed to write completed frame: {e}"),
)
})?;
Ok(())
}
@@ -799,6 +838,10 @@ impl OperationEnv for OverlayOperationEnv {
fn contains(&self, name: &str) -> bool {
self.overlay.read().contains_key(name)
}
fn list_operation_names(&self) -> Vec<String> {
self.overlay.read().keys().cloned().collect()
}
}
pub struct SubscriptionStream {
@@ -1476,7 +1519,7 @@ mod tests {
)
});
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
pub_spec_e2e("fs/upload"),
@@ -1880,6 +1923,346 @@ mod tests {
assert_eq!(read, envelope);
}
// CF-001 — a transport write failure before the request frame could be
// delivered maps to retryable CONNECTION_CLOSED (the call provably
// never reached the producer). Mid-publish failures stay INTERNAL.
/// AsyncWrite that always errors — simulates a dead mux mid-call.
struct BrokenWrite;
impl AsyncWrite for BrokenWrite {
fn poll_write(
self: Pin<&mut Self>,
_cx: &mut Context<'_>,
_buf: &[u8],
) -> Poll<std::io::Result<usize>> {
Poll::Ready(Err(std::io::Error::new(
std::io::ErrorKind::BrokenPipe,
"mux dead",
)))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
}
fn broken_write_bi_stream() -> crate::core::types::BiStream {
// The joined stream holds the read end; the write half always errors.
let (read_end, _write_end) = tokio::io::duplex(1024);
crate::core::types::BiStream::from_joined(read_end, BrokenWrite)
}
/// AsyncWrite whose writes succeed but whose Nth `poll_flush` fails.
/// `FrameFramedWriter::write_frame` flushes exactly once per frame, so
/// `remaining_ok_flushes = k` lets frames 1..=k land and fails inside
/// frame k+1 — a deterministic mid-stream write failure.
struct FailOnFlushN {
remaining_ok_flushes: std::sync::atomic::AtomicUsize,
}
impl AsyncWrite for FailOnFlushN {
fn poll_write(
self: Pin<&mut Self>,
_cx: &mut Context<'_>,
buf: &[u8],
) -> Poll<std::io::Result<usize>> {
Poll::Ready(Ok(buf.len()))
}
fn poll_flush(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
let prev = self
.remaining_ok_flushes
.fetch_sub(1, std::sync::atomic::Ordering::SeqCst);
if prev == 0 {
Poll::Ready(Err(std::io::Error::new(
std::io::ErrorKind::BrokenPipe,
"mux dead",
)))
} else {
Poll::Ready(Ok(()))
}
}
fn poll_shutdown(self: Pin<&mut Self>, _cx: &mut Context<'_>) -> Poll<std::io::Result<()>> {
Poll::Ready(Ok(()))
}
}
#[tokio::test]
async fn single_stream_call_write_failure_is_retryable_connection_closed() {
let bi_stream = broken_write_bi_stream();
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let response = conn.call("test/echo", serde_json::json!({})).await;
match response.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[tokio::test]
async fn multi_stream_call_write_request_failure_is_retryable() {
// A `Connection` whose open_bi succeeds but whose write half is
// dead: the request write fails, the call provably undelivered.
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(broken_write_bi_stream()))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let response = call_conn.call("test/echo", serde_json::json!({})).await;
match response.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[tokio::test]
async fn multi_stream_publish_request_frame_write_failure_is_retryable() {
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(broken_write_bi_stream()))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let response = call_conn
.publish(
"fs/upload",
serde_json::json!({}),
Box::pin(futures::stream::iter(vec![serde_json::json!({"c": 1})])),
)
.await;
match response.result {
Err(e) => {
// The pump died on the *request* frame — nothing delivered.
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[test]
fn connection_closed_error_is_retryable() {
let err = CallError::connection_closed("test");
assert_eq!(err.code, "CONNECTION_CLOSED");
assert!(err.retryable);
}
#[tokio::test]
async fn single_stream_subscribe_write_failure_is_retryable_connection_closed() {
use futures::stream::StreamExt;
let bi_stream = broken_write_bi_stream();
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let mut stream = conn.subscribe("test/events", serde_json::json!({})).await;
let item = stream.next().await.unwrap();
match item.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
assert!(stream.next().await.is_none());
}
#[tokio::test]
async fn multi_stream_subscribe_write_failure_is_retryable_connection_closed() {
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use futures::stream::StreamExt;
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(broken_write_bi_stream()))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let mut stream = call_conn
.subscribe("test/events", serde_json::json!({}))
.await;
let item = stream.next().await.unwrap();
match item.result {
Err(e) => {
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
assert!(stream.next().await.is_none());
}
#[tokio::test]
async fn single_stream_publish_request_frame_write_failure_is_retryable() {
let bi_stream = broken_write_bi_stream();
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let response = conn
.publish(
"fs/upload",
serde_json::json!({}),
Box::pin(futures::stream::iter(vec![serde_json::json!({"c": 1})])),
)
.await;
match response.result {
Err(e) => {
// The request frame never landed — nothing delivered, retry safe.
assert_eq!(e.code, "CONNECTION_CLOSED");
assert!(e.retryable);
}
other => panic!("expected retryable CONNECTION_CLOSED, got {other:?}"),
}
}
#[tokio::test]
async fn single_stream_publish_midpublish_write_failure_stays_internal() {
// Flush succeeds for the request frame (1) + first chunk (2), then
// fails inside the second chunk — delivery is ambiguous from there.
let bi_stream = crate::core::types::BiStream::from_joined(
tokio::io::empty(),
FailOnFlushN {
remaining_ok_flushes: std::sync::atomic::AtomicUsize::new(2),
},
);
let (writer, reader) = split_single_stream(bi_stream);
let conn = CallConnection::new_single_stream(stub_connection(), writer);
drop(reader);
let chunks = vec![serde_json::json!({"c": 1}), serde_json::json!({"c": 2})];
let stream: Pin<Box<dyn Stream<Item = Value> + Send>> =
Box::pin(futures::stream::iter(chunks));
let response = conn
.publish("fs/upload", serde_json::json!({}), stream)
.await;
match response.result {
Err(e) => {
assert_eq!(e.code, "INTERNAL");
assert!(!e.retryable, "mid-publish failure must not be retryable");
}
other => panic!("expected non-retryable INTERNAL, got {other:?}"),
}
}
#[tokio::test]
async fn multi_stream_publish_midpublish_write_failure_stays_internal() {
use crate::core::types::StreamError;
use crate::core::types::{BiStream, BidiStreamSource, Connection};
use std::net::SocketAddr;
struct OneBiStream(std::sync::Mutex<Option<crate::core::types::BiStream>>);
#[async_trait::async_trait]
impl BidiStreamSource for OneBiStream {
async fn accept_bi(&self) -> Result<BiStream, StreamError> {
Err(StreamError::ConnectionClosed)
}
async fn open_bi(&self) -> Result<BiStream, StreamError> {
let stream = self.0.lock().expect("open once").take();
stream.ok_or(StreamError::StreamClosed)
}
fn remote_addr(&self) -> Option<SocketAddr> {
None
}
fn close(&self, _code: u32, _reason: &str) {}
}
let conn = Connection::from_source(
OneBiStream(std::sync::Mutex::new(Some(
crate::core::types::BiStream::from_joined(
tokio::io::empty(),
FailOnFlushN {
remaining_ok_flushes: std::sync::atomic::AtomicUsize::new(2),
},
),
))),
b"alk/call".to_vec(),
);
let call_conn = CallConnection::new(conn);
let chunks = vec![serde_json::json!({"c": 1}), serde_json::json!({"c": 2})];
let stream: Pin<Box<dyn Stream<Item = Value> + Send>> =
Box::pin(futures::stream::iter(chunks));
let response = call_conn
.publish("fs/upload", serde_json::json!({}), stream)
.await;
match response.result {
Err(e) => {
// The pump died on a *chunk* frame — the request landed,
// delivery is ambiguous, INTERNAL (non-retryable) is correct.
assert_eq!(e.code, "INTERNAL");
assert!(!e.retryable);
}
other => panic!("expected non-retryable INTERNAL, got {other:?}"),
}
}
#[tokio::test]
async fn shared_frame_writer_concurrent_writes_do_not_interleave() {
let (client_end, server_end) = tokio::io::duplex(64 * 1024);
@@ -1993,7 +2376,7 @@ mod tests {
}
}
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("test/echo"),
@@ -2078,7 +2461,7 @@ mod tests {
}
}
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
+546 -130
View File
@@ -32,7 +32,7 @@ use super::abort::AbortCascade;
use super::connection::CallConnection;
use super::wire::{
CallError, EventEnvelope, FrameFramedReader, FrameFramedWriter, ResponseEnvelope,
EVENT_ABORTED, EVENT_COMPLETED, EVENT_ERROR, EVENT_PUBLISHED, EVENT_REQUESTED,
EVENT_ABORTED, EVENT_COMPLETED, EVENT_ERROR, EVENT_PUBLISHED, EVENT_REQUESTED, EVENT_RESPONDED,
};
use crate::protocol::adapter::SessionOverlaySource;
use crate::registry::context::{AbortPolicy, OperationContext, ScopedPeerEnv};
@@ -114,6 +114,18 @@ impl std::fmt::Debug for DispatchResult {
}
}
/// The started half of a dispatched `call.requested` event, from
/// `Dispatcher::dispatch_start`. `Once`/`Stream` are returned ready
/// for the caller to spawn (review 005 G-01); `Sink` was started
/// inline (its `chunk_tx` must be registered before the next
/// `call.published` frame is read) and carries the started handler
/// future.
pub enum StartedDispatch {
Once(Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>),
Stream(ResponseStream),
Sink(SinkDispatch),
}
/// Shared dispatcher for an established `CallConnection`. Constructed by
/// both `CallAdapter` (accept path) and `CallClient` (connect path) and used
/// to run the dispatch loop. Holds no per-connection state; the
@@ -299,6 +311,39 @@ impl Dispatcher {
request_id: String,
payload: Value,
) -> DispatchResult {
match self.dispatch_start(connection, request_id, payload) {
StartedDispatch::Once(invoke) => DispatchResult::Once(invoke.await),
StartedDispatch::Stream(stream) => DispatchResult::Stream(stream),
StartedDispatch::Sink(sink) => DispatchResult::Sink(sink),
}
}
/// Start dispatching a `call.requested` event without awaiting the
/// invocation (review 005 G-01). The synchronous prefix of
/// [`dispatch`](Self::dispatch) — identity resolution, root-context
/// construction, op-type branch — is split from the invocation:
///
/// - `Query`/`Mutation` return the invocation as a boxed future
/// (`StartedDispatch::Once`) for the caller to spawn. Awaiting it
/// inline is what deadlocked the serving loops: a handler that
/// nested-composes over the same connection blocks the one read
/// loop that must resolve the nested call's response frame.
/// - `Sub` returns the [`ResponseStream`] synchronously
/// (`invoke_streaming` is a sync call) for the caller to pump in
/// a spawned task — an inline pump starves the same loop for the
/// stream's whole lifetime.
/// - `Pub` runs the full sink start inline (`StartedDispatch::Sink`):
/// ACL failure resolves to a ready `Once` future (the response
/// envelope is already built), success carries the `chunk_tx` the
/// loop must register in `in_flight_sinks` *before* the next
/// `call.published` frame can be read — the insert cannot race
/// the feed, so the sink start stays synchronous by design.
pub(crate) fn dispatch_start(
&self,
connection: &Arc<CallConnection>,
request_id: String,
payload: Value,
) -> StartedDispatch {
let operation_id = payload
.get("operationId")
.and_then(|v| v.as_str())
@@ -320,7 +365,7 @@ impl Dispatcher {
.unwrap_or(OperationType::Query);
let mut context = self.build_root_context(
request_id.clone(),
request_id,
&operation_name,
identity,
forwarded_for,
@@ -329,54 +374,125 @@ impl Dispatcher {
match op_type {
OperationType::Query | OperationType::Mutation => {
let envelope = self.registry.invoke(&operation_name, input, context).await;
DispatchResult::Once(envelope)
let registry = Arc::clone(&self.registry);
let name = operation_name;
StartedDispatch::Once(Box::pin(async move {
registry.invoke(&name, input, context).await
}))
}
OperationType::Sub => {
context.deadline = None;
let stream = self
.registry
.invoke_streaming(&operation_name, input, context);
DispatchResult::Stream(stream)
StartedDispatch::Stream(stream)
}
OperationType::Pub => {
context.deadline = None;
let sink_handler =
match self
.registry
.resolve_sink_handler(&operation_name, &input, &context)
{
Ok(h) => h,
Err(envelope) => return DispatchResult::Once(envelope),
};
let publish_validator = self
match self
.registry
.registration(&operation_name)
.and_then(|r| r.spec.publish_schema.as_ref())
.and_then(|schema| match jsonschema::options().build(schema) {
Ok(v) => Some(v),
Err(e) => {
warn!(
operation = %operation_name,
error = %e,
"publish_schema failed to compile; chunks will not be validated",
);
None
}
});
let (chunk_tx, chunk_rx) =
mpsc::channel::<Result<Value, CallError>>(PUBLISH_CHANNEL_BUFFER);
let publish_stream: PublishStream = Box::pin(chunk_rx);
let handler = (sink_handler)(input, context, publish_stream);
DispatchResult::Sink(SinkDispatch {
handler: Box::pin(handler),
chunk_tx,
publish_validator,
})
.resolve_sink_handler(&operation_name, &input, &context)
{
Ok(sink_handler) => {
let publish_validator = self.registry.publish_validator(&operation_name);
let (chunk_tx, chunk_rx) =
mpsc::channel::<Result<Value, CallError>>(PUBLISH_CHANNEL_BUFFER);
let publish_stream: PublishStream = Box::pin(chunk_rx);
let handler = (sink_handler)(input, context, publish_stream);
StartedDispatch::Sink(SinkDispatch {
handler: Box::pin(handler),
chunk_tx,
publish_validator,
})
}
Err(envelope) => StartedDispatch::Once(Box::pin(std::future::ready(envelope))),
}
}
}
}
/// Spawn the Once invocation to write its single response frame on
/// completion. The write failure is a warning, not a loop exit: the
/// loop keeps reading (a dying transport surfaces on the next read
/// as `ConnectionClosed`), matching the Sink arm's behavior.
fn spawn_once_dispatch(
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: String,
invoke: Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>,
) -> JoinHandle<()> {
let writer = Arc::clone(writer);
tokio::spawn(async move {
let response = invoke.await;
let event: EventEnvelope = response.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write Once response frame"
);
}
})
}
/// Spawn a subscription's [`ResponseStream`] pump: each
/// [`ResponseEnvelope`] becomes a `call.responded` / `call.error`
/// frame; on natural end, a `call.completed` frame. The shared
/// writer serializes frames so the pump's frames do not interleave
/// with concurrent calls' frames.
fn spawn_stream_pump(
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: String,
mut stream: ResponseStream,
) -> JoinHandle<()> {
let writer = Arc::clone(writer);
tokio::spawn(async move {
let mut last_was_error = false;
while let Some(envelope) = stream.next().await {
last_was_error = envelope.result.is_err();
let event: EventEnvelope = envelope.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write streaming frame"
);
return;
}
}
if !last_was_error {
let completed = EventEnvelope::completed(&request_id);
if let Err(err) = writer.write_frame(&completed).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write call.completed"
);
}
}
})
}
/// Spawn a Pub's handler task: awaits the handler future and writes
/// the single `call.responded` / `call.error` frame.
fn spawn_sink_response_writer(
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: String,
handler: Pin<Box<dyn Future<Output = ResponseEnvelope> + Send>>,
) -> JoinHandle<()> {
let writer = Arc::clone(writer);
tokio::spawn(async move {
let response = handler.await;
let event: EventEnvelope = response.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id,
"serving loop: failed to write sink response frame"
);
}
})
}
pub async fn handle_abort(&self, connection: &Arc<CallConnection>, request_id: &str) {
if let Some(tx) = self.in_flight_sink_aborts.lock().remove(request_id) {
let _ = tx.send(());
@@ -396,7 +512,7 @@ impl Dispatcher {
connection: Arc<CallConnection>,
stream: crate::core::types::BiStream,
) {
// `stream` is a `BiStream` (ADR-092) — `AsyncRead + AsyncWrite + Send
// `stream` is a `BiStream` (ADR-009) — `AsyncRead + AsyncWrite + Send
// + Unpin`. Split into the read and write halves the call protocol's
// frame reader/writer consume. The split is the stdlib idiom; no
// per-handler wrapper.
@@ -455,7 +571,7 @@ impl Dispatcher {
/// for `Ok`, `call.error` for `Err`). On natural stream end (the stream
/// returned `None` without the last item being an `Err`), write a
/// `call.completed` frame. An `Err` envelope is terminal — the stream
/// ends after it and we do NOT write `call.completed` (ADR-049 §6).
/// ends after it and we do NOT write `call.completed` (ADR-021 §6).
///
/// If a frame write fails the pump stops early; the stream is dropped on
/// return, releasing the handler's resources via `Drop` (ADR-016). A
@@ -730,14 +846,24 @@ impl Dispatcher {
/// `call.published` / `call.completed` / `call.aborted` frames
/// arrive *after* the `call.requested` and are routed to the
/// matching sink's `chunk_tx` while new `call.requested` frames for
/// other requests continue to be dispatched. Query/Mutation and Sub
/// responses are written through `writer` immediately after
/// dispatch.
/// other requests continue to be dispatched. Once invocations and
/// Sub pumps are spawned (review 005 G-01 — inline dispatch
/// deadlocked same-connection nested composition); only the sink
/// start runs inline.
///
/// This is the accept side's serving loop, so the loop also carries
/// the pending-resolution arms the connect side's
/// `serve_single_stream` composes: a nested-composing handler
/// (e.g. a `from_call` imported-op stub riding this connection)
/// blocks its own task, not the read loop, and the loop resolves
/// the stub's response frames here. Without these arms the accept
/// side could dispatch inbound requests but never resolve its own
/// outbound pendings in single-stream mode.
///
/// Returns when the read half closes (transport EOF). Outstanding
/// pending requests are failed with `connection closed`, and
/// in-flight sinks' `chunk_tx` are dropped (the handler's
/// `PublishStream` sees EOF).
/// pending requests are failed with `connection closed`, in-flight
/// sinks' `chunk_tx` are dropped (the handler's `PublishStream`
/// sees EOF), and spawned Once/Stream/Sink tasks are aborted.
pub async fn run_loop_single_stream(
self,
connection: Arc<CallConnection>,
@@ -763,7 +889,10 @@ impl Dispatcher {
});
let mut reader = FrameFramedReader::new(reader);
let mut in_flight_sinks: HashMap<String, InFlightSink> = HashMap::new();
let in_flight_sinks: Arc<ParkingLotMutex<HashMap<String, InFlightSink>>> =
Arc::new(ParkingLotMutex::new(HashMap::new()));
let spawned: Arc<ParkingLotMutex<Vec<JoinHandle<()>>>> =
Arc::new(ParkingLotMutex::new(Vec::new()));
loop {
let envelope = match reader.read_frame().await {
@@ -778,46 +907,29 @@ impl Dispatcher {
match envelope.r#type.as_str() {
EVENT_REQUESTED => {
let request_id = envelope.id.clone();
let payload = envelope.payload.clone();
let dispatch_result = self
.dispatch(&connection, request_id.clone(), payload)
.await;
match dispatch_result {
DispatchResult::Once(response) => {
let event: EventEnvelope = response.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(
error = %err,
"single-stream: failed to write Once response; closing loop"
);
break;
}
match self.dispatch_start(&connection, request_id.clone(), envelope.payload) {
StartedDispatch::Once(invoke) => {
let handle =
Self::spawn_once_dispatch(&writer, request_id.clone(), invoke);
spawned.lock().push(handle);
}
DispatchResult::Stream(stream) => {
self.pump_stream_single_stream(&writer, &request_id, stream)
.await;
StartedDispatch::Stream(stream) => {
let handle =
Self::spawn_stream_pump(&writer, request_id.clone(), stream);
spawned.lock().push(handle);
}
DispatchResult::Sink(sink) => {
StartedDispatch::Sink(sink) => {
let SinkDispatch {
handler,
chunk_tx,
publish_validator,
} = sink;
let writer_clone = Arc::clone(&writer);
let request_id_for_handler = request_id.clone();
let handle = tokio::spawn(async move {
let response = handler.await;
let event: EventEnvelope = response.into();
if let Err(err) = writer_clone.write_frame(&event).await {
warn!(
error = %err,
request_id = %request_id_for_handler,
"single-stream: failed to write sink response frame"
);
}
});
in_flight_sinks.insert(
let handle = Self::spawn_sink_response_writer(
&writer,
request_id.clone(),
handler,
);
in_flight_sinks.lock().insert(
request_id.clone(),
InFlightSink {
chunk_tx,
@@ -828,9 +940,25 @@ impl Dispatcher {
}
}
}
EVENT_RESPONDED => {
let request_id = envelope.id.clone();
let output = envelope
.payload
.get("output")
.cloned()
.unwrap_or(Value::Null);
pending.lock().handle_responded(&request_id, output);
}
EVENT_COMPLETED => {
let request_id = envelope.id.clone();
if in_flight_sinks.lock().remove(&request_id).is_none() {
pending.lock().handle_completed(&request_id);
}
}
EVENT_ABORTED => {
let request_id = envelope.id.clone();
if let Some(mut entry) = in_flight_sinks.remove(&request_id) {
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
entry.handler_handle.abort();
let _ = entry
.chunk_tx
@@ -847,7 +975,8 @@ impl Dispatcher {
.get("input")
.cloned()
.unwrap_or(Value::Null);
if let Some(mut entry) = in_flight_sinks.remove(&request_id) {
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let validated = match &entry.publish_validator {
Some(validator) => {
if validator.is_valid(&chunk) {
@@ -865,7 +994,7 @@ impl Dispatcher {
let keep = validated.is_ok();
let _ = entry.chunk_tx.send(validated).await;
if keep {
in_flight_sinks.insert(request_id, entry);
in_flight_sinks.lock().insert(request_id, entry);
}
} else {
debug!(
@@ -874,36 +1003,33 @@ impl Dispatcher {
);
}
}
EVENT_COMPLETED => {
let request_id = envelope.id.clone();
in_flight_sinks.remove(&request_id);
}
EVENT_ERROR => {
let request_id = envelope.id.clone();
if let Some(mut entry) = in_flight_sinks.remove(&request_id) {
let call_error: CallError = serde_json::from_value(envelope.payload)
.unwrap_or_else(|_| {
CallError::internal("publish error from initiator (malformed)")
});
let call_error: CallError = serde_json::from_value(envelope.payload)
.unwrap_or_else(|_| {
CallError::internal("publish error from initiator (malformed)")
});
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let _ = entry.chunk_tx.send(Err(call_error)).await;
} else {
debug!(
request_id = %request_id,
"single-stream: call.error for unknown in-flight sink; dropping"
);
pending.lock().handle_error(&request_id, call_error);
}
}
other => {
debug!(
event_type = %other,
id = %envelope.id,
"single-stream: ignoring non-requested/non-published/non-aborted/non-completed event"
"single-stream: ignoring unknown event type"
);
}
}
}
in_flight_sinks.clear();
in_flight_sinks.lock().clear();
for handle in spawned.lock().drain(..) {
handle.abort();
}
let failed = pending
.lock()
@@ -918,33 +1044,227 @@ impl Dispatcher {
sweeper_handle.abort();
}
/// Pump a subscription's `ResponseStream` to the wire through the
/// shared single-stream writer (ADR-036 amendment). Each
/// `ResponseEnvelope` becomes a `call.responded` / `call.error`
/// frame; on natural end, a `call.completed` frame. The shared
/// writer serializes frames so this pump's frames do not interleave
/// with concurrent calls' frames.
async fn pump_stream_single_stream(
&self,
writer: &Arc<super::connection::SharedFrameWriter>,
request_id: &str,
mut stream: ResponseStream,
/// The full-duplex single-stream serving loop (review 004 F-04):
/// composes the dispatch arms (`run_loop_single_stream`) and the
/// pending-resolution arms (`read_single_stream_until_closed`) in
/// one loop. Both sides of a channels connection can be both
/// producer and consumer (AGENTS.md §8; ADR-022 §2) — on channel 0
/// the two directions' frames are multiplexed on one byte stream,
/// so the loop branches per frame:
///
/// - `call.requested` → start the dispatch and spawn the work (the
/// serving half — what the accept side's `run_loop_single_stream`
/// does). Once invocations and Sub pumps run as spawned tasks;
/// only the Pub sink start runs inline (its `chunk_tx` must be
/// registered before the next `call.published` frame is read).
/// Inline dispatch was the review 005 G-01 defect: a handler that
/// nested-composes over the same connection blocked the one read
/// loop that must resolve the nested call's response frame. The
/// `SharedFrameWriter` serializes frames, so spawned arms' frames
/// do not interleave mid-frame.
/// - `call.responded` / `call.completed` / `call.error` → resolve
/// the matching outbound pending entry (the calling half — what
/// the client read pump does). `call.responded` frames for
/// *inbound* Sub responses are `call.responded` too; direction is
/// disambiguated by the pending map (an id that is one of *our*
/// outbound pendings resolves there; an unknown id is dropped
/// with a debug line, never an error — the peer may legitimately
/// stream a Sub's responses that this side does not track).
/// - `call.aborted` → try both tables: the in-flight sink aborts
/// *and* the outbound pending map's abort-cascade
/// (`handle_abort`), then fall through to serving-side sink
/// cancellation (the `run_loop_single_stream` arm) when the id
/// matches an inbound sink. An id in neither table is a no-op.
/// - `call.published` / `call.completed` (initiator→responder) →
/// route to the matching inbound in-flight sink; ids that are
/// outbound pending *call* entries whose `call.responded` already
/// removed them are no-ops.
///
/// IDs are UUID-generated per side (`generate_request_id()`), so
/// cross-correlation between the two directions is not a hazard.
///
/// On read-half close, outbound pendings are failed (`connection
/// closed`), in-flight inbound sinks are dropped, and the spawned
/// Once/Stream/Sink tasks are aborted (their response frames cannot
/// reach a closed transport anyway).
pub async fn serve_single_stream(
self,
connection: Arc<CallConnection>,
reader: Box<dyn tokio::io::AsyncRead + Send + Unpin>,
writer: Arc<super::connection::SharedFrameWriter>,
) {
let mut last_was_error = false;
while let Some(envelope) = stream.next().await {
last_was_error = envelope.result.is_err();
let event: EventEnvelope = envelope.into();
if let Err(err) = writer.write_frame(&event).await {
warn!(error = %err, "single-stream: failed to write streaming frame");
return;
let pending = Arc::clone(connection.pending());
let sweeper_pending = Arc::clone(&pending);
let sweeper_handle: JoinHandle<()> = tokio::spawn(async move {
let mut interval = tokio::time::interval(SWEEPER_INTERVAL);
interval.set_missed_tick_behavior(tokio::time::MissedTickBehavior::Skip);
loop {
interval.tick().await;
let evicted = sweeper_pending.lock().evict_expired();
if !evicted.is_empty() {
debug!(
count = evicted.len(),
"serve loop: sweeper evicted expired pending entries"
);
}
}
});
let mut reader = FrameFramedReader::new(reader);
let in_flight_sinks: Arc<ParkingLotMutex<HashMap<String, InFlightSink>>> =
Arc::new(ParkingLotMutex::new(HashMap::new()));
let spawned: Arc<ParkingLotMutex<Vec<JoinHandle<()>>>> =
Arc::new(ParkingLotMutex::new(Vec::new()));
loop {
let envelope = match reader.read_frame().await {
Ok(env) => env,
Err(super::wire::FrameError::ConnectionClosed) => break,
Err(err) => {
warn!(error = %err, "serve loop: frame read error; closing loop");
break;
}
};
match envelope.r#type.as_str() {
EVENT_REQUESTED => {
let request_id = envelope.id.clone();
match self.dispatch_start(&connection, request_id.clone(), envelope.payload) {
StartedDispatch::Once(invoke) => {
let handle =
Self::spawn_once_dispatch(&writer, request_id.clone(), invoke);
spawned.lock().push(handle);
}
StartedDispatch::Stream(stream) => {
let handle =
Self::spawn_stream_pump(&writer, request_id.clone(), stream);
spawned.lock().push(handle);
}
StartedDispatch::Sink(sink) => {
let SinkDispatch {
handler,
chunk_tx,
publish_validator,
} = sink;
let handle = Self::spawn_sink_response_writer(
&writer,
request_id.clone(),
handler,
);
in_flight_sinks.lock().insert(
request_id.clone(),
InFlightSink {
chunk_tx,
publish_validator,
handler_handle: handle,
},
);
}
}
}
EVENT_RESPONDED => {
let request_id = envelope.id.clone();
let output = envelope
.payload
.get("output")
.cloned()
.unwrap_or(Value::Null);
pending.lock().handle_responded(&request_id, output);
}
EVENT_COMPLETED => {
let request_id = envelope.id.clone();
if in_flight_sinks.lock().remove(&request_id).is_none() {
pending.lock().handle_completed(&request_id);
}
}
EVENT_ABORTED => {
let request_id = envelope.id.clone();
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
entry.handler_handle.abort();
let _ = entry
.chunk_tx
.send(Err(CallError::internal("publish aborted by initiator")))
.await;
} else {
self.handle_abort(&connection, &request_id).await;
}
}
EVENT_PUBLISHED => {
let request_id = envelope.id.clone();
let chunk = envelope
.payload
.get("input")
.cloned()
.unwrap_or(Value::Null);
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let validated = match &entry.publish_validator {
Some(validator) => {
if validator.is_valid(&chunk) {
Ok(chunk)
} else {
let details = serde_json::json!({ "chunk": chunk });
Err(CallError::invalid_input(
"published chunk failed publish_schema validation",
)
.with_details(details))
}
}
None => Ok(chunk),
};
let keep = validated.is_ok();
let _ = entry.chunk_tx.send(validated).await;
if keep {
in_flight_sinks.lock().insert(request_id, entry);
}
} else {
debug!(
request_id = %request_id,
"serve loop: call.published for unknown in-flight sink; dropping"
);
}
}
EVENT_ERROR => {
let request_id = envelope.id.clone();
let call_error: CallError = serde_json::from_value(envelope.payload)
.unwrap_or_else(|_| {
CallError::internal("publish error from initiator (malformed)")
});
let entry = in_flight_sinks.lock().remove(&request_id);
if let Some(mut entry) = entry {
let _ = entry.chunk_tx.send(Err(call_error)).await;
} else {
pending.lock().handle_error(&request_id, call_error);
}
}
other => {
debug!(
event_type = %other,
id = %envelope.id,
"serve loop: ignoring unknown event type"
);
}
}
}
if !last_was_error {
let completed = EventEnvelope::completed(request_id);
if let Err(err) = writer.write_frame(&completed).await {
warn!(error = %err, "single-stream: failed to write call.completed");
}
in_flight_sinks.lock().clear();
for handle in spawned.lock().drain(..) {
handle.abort();
}
let failed = pending
.lock()
.fail_all(CallError::internal("connection closed"));
if !failed.is_empty() {
debug!(
count = failed.len(),
"serve loop: failed pending requests on connection close"
);
}
sweeper_handle.abort();
}
}
@@ -1028,7 +1348,7 @@ mod tests {
}
fn registry_with(name: &str, visibility: Visibility, acl: AccessControl) -> OperationRegistry {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
@@ -1063,7 +1383,7 @@ mod tests {
#[tokio::test]
async fn dispatch_authorized_peer_dispatches_and_populates_capabilities() {
let caps = Capabilities::new().with_api_key("google", "k".to_string());
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let has_google = context.capabilities.get("google").is_some();
ResponseEnvelope::ok(
@@ -1100,7 +1420,7 @@ mod tests {
#[tokio::test]
async fn dispatch_unauthorized_peer_returns_forbidden_capabilities_never_populated() {
let caps = Capabilities::new().with_api_key("google", "k".to_string());
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let has_google = context.capabilities.get("google").is_some();
ResponseEnvelope::ok(
@@ -1225,7 +1545,7 @@ mod tests {
#[tokio::test]
async fn dispatch_extract_forwarded_for_from_payload_into_context() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let forwarded_id = context.forwarded_for.as_ref().map(|i| i.id.clone());
ResponseEnvelope::ok(
@@ -1266,7 +1586,7 @@ mod tests {
#[tokio::test]
async fn dispatch_without_forwarded_for_field_is_none() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let present = context.forwarded_for.is_some();
ResponseEnvelope::ok(
@@ -1356,7 +1676,7 @@ mod tests {
#[tokio::test]
async fn dispatch_requested_overlay_only_attaches_peer_keyed_by_stored_identity() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, context| async move {
let peer_ids = context.env.peer_ids();
ResponseEnvelope::ok(
@@ -1491,7 +1811,7 @@ mod tests {
);
}
// --- streaming dispatch branch (ADR-049 §6) ---------------------------
// --- streaming dispatch branch (ADR-021 §6) ---------------------------
fn subscription_spec(name: &str, acl: AccessControl) -> OperationSpec {
OperationSpec::new(
@@ -1546,7 +1866,7 @@ mod tests {
name: &str,
handler: crate::registry::registration::StreamingHandler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
subscription_spec(name, AccessControl::default()),
@@ -1625,7 +1945,7 @@ mod tests {
#[tokio::test]
async fn dispatch_query_keeps_deadline_some() {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
let handler = make_handler(|_input, ctx| async move {
let deadline_is_some = ctx.deadline.is_some();
ResponseEnvelope::ok(
@@ -1890,7 +2210,7 @@ mod tests {
name: &str,
handler: crate::registry::registration::SinkHandler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
pub_spec(name, AccessControl::default()),
@@ -2057,7 +2377,7 @@ mod tests {
publish_schema: Value,
handler: crate::registry::registration::SinkHandler,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
pub_spec_with_publish_schema(name, publish_schema),
@@ -2593,4 +2913,100 @@ mod tests {
"in_flight_sink_aborts map is cleaned up after pump_sink exits"
);
}
// --- review 004 F-04: the full-duplex serving loop ---------------------
/// The F-04 probe: two `serve_single_stream` loops over one duplex
/// — the exact hub↔consumer shape. Each side holds a
/// `CallConnection` over its own writer; when side A calls
/// `echo/run`, its `call.requested` crosses the duplex into side
/// B's serving loop, which dispatches and writes `call.responded`
/// back, resolving A's pending through the `EVENT_RESPONDED` arm.
/// This is the direction `read_single_stream_until_closed` used to
/// drop.
#[tokio::test]
async fn serve_single_stream_self_call_resolves_pending() {
fn echo_registry() -> Arc<crate::registry::registration::OperationRegistry> {
let registry = crate::registry::registration::OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("echo/run", AccessControl::default()),
HandlerKind::Once(make_handler(|input, context| async move {
ResponseEnvelope::ok(context.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
Arc::new(registry)
}
let (client_end, server_end) = tokio::io::duplex(64 * 1024);
// Side B: the duplex end wrapped as one BiStream (read+write);
// `split_single_stream` divides it into the frame writer and
// reader the serving loop consumes.
let server_bidi = crate::core::types::BiStream::from_bidi(server_end);
let (server_writer, server_reader) =
crate::protocol::connection::split_single_stream(server_bidi);
let dp_server = Dispatcher::new(echo_registry(), Arc::new(StaticIdentityProvider::new()));
let server_call_conn = Arc::new(CallConnection::new_single_stream(
crate::protocol::sink_empty_connection(),
Arc::clone(&server_writer),
));
let server_conn_for_loop = Arc::clone(&server_call_conn);
let _serve_b = tokio::spawn(async move {
dp_server
.serve_single_stream(server_conn_for_loop, server_reader, server_writer)
.await;
});
// Side A: the same shape on the other duplex end.
let client_bidi = crate::core::types::BiStream::from_bidi(client_end);
let (client_writer, client_reader) =
crate::protocol::connection::split_single_stream(client_bidi);
let dp_client = Dispatcher::new(echo_registry(), Arc::new(StaticIdentityProvider::new()));
let client_call_conn = Arc::new(CallConnection::new_single_stream(
crate::protocol::sink_empty_connection(),
Arc::clone(&client_writer),
));
let client_conn_for_loop = Arc::clone(&client_call_conn);
let _serve_a = tokio::spawn(async move {
dp_client
.serve_single_stream(client_conn_for_loop, client_reader, client_writer)
.await;
});
// Side A calls `echo/run` — side B serves it.
let response = tokio::time::timeout(
std::time::Duration::from_secs(5),
client_call_conn.call("echo/run", serde_json::json!({ "from": "a" })),
)
.await
.expect("A→B call through serving loops timed out");
assert!(
response.result.is_ok(),
"A→B call resolves, got {:?}",
response.result
);
assert_eq!(response.result.unwrap(), serde_json::json!({ "from": "a" }));
// And B calls back — A serves it (the both-directions gate).
let response = tokio::time::timeout(
std::time::Duration::from_secs(5),
server_call_conn.call("echo/run", serde_json::json!({ "from": "b" })),
)
.await
.expect("B→A call through serving loops timed out");
assert!(
response.result.is_ok(),
"B→A call resolves, got {:?}",
response.result
);
assert_eq!(response.result.unwrap(), serde_json::json!({ "from": "b" }));
}
}
+2 -2
View File
@@ -1,6 +1,6 @@
//! Shared test helpers for the call protocol's inline `#[cfg(test)]`
//! modules. Kept here (not in each test module) so the `stub_connection()`
//! shape is defined once — `Connection::from_stream` was removed (ADR-092)
//! shape is defined once — `Connection::from_stream` was removed (ADR-009)
//! and every test stub that previously called it now calls
//! `Connection::from_bidi(SinkEmpty, ...)` via `sink_empty_connection()`.
@@ -15,7 +15,7 @@ use tokio::io::{AsyncRead, AsyncWrite, ReadBuf};
/// A test-only `AsyncRead + AsyncWrite` pair equivalent to
/// `tokio::io::sink() + tokio::io::empty()`: reads yield EOF immediately
/// (zero bytes), writes discard. Exists because `Connection::from_bidi`
/// (ADR-092 — the only public stream constructor, replacing
/// (ADR-009 — the only public stream constructor, replacing
/// `from_stream`) requires a single value that implements both traits.
/// Used only to construct a `Connection` for tests that exercise
/// `Connection`-level state (alpn, addr, identity, dispatcher run loop
+8
View File
@@ -113,9 +113,17 @@ impl CallError {
Self::new("TIMEOUT", message, true)
}
pub fn connection_closed(message: impl Into<String>) -> Self {
Self::new("CONNECTION_CLOSED", message, true)
}
pub fn invalid_operation_type(message: impl Into<String>) -> Self {
Self::new("INVALID_OPERATION_TYPE", message, false)
}
pub fn already_exists(message: impl Into<String>) -> Self {
Self::new("ALREADY_EXISTS", message, false)
}
}
impl Eq for CallError {}
+1 -1
View File
@@ -33,7 +33,7 @@ pub struct OperationContext {
pub internal: bool,
/// `None` when no ownership provider is wired (backward compat —
/// `check` falls back to static `Identity.resources` path). Wired by
/// the assembly layer via `CallAdapter`/`Dispatcher` (ADR-050).
/// the assembly layer via `CallAdapter`/`Dispatcher` (ADR-011).
pub ownership: Option<Arc<dyn OwnershipProvider>>,
}
+266 -11
View File
@@ -3,8 +3,11 @@ use std::sync::Arc;
use serde_json::{json, Value};
use super::context::OperationContext;
use super::registration::{Handler, OperationRegistry};
use super::spec::{AccessControl, OperationSpec, OperationType, Visibility};
use super::registration::{
Handler, HandlerKind, HandlerRegistration, OperationProvenance, OperationRegistry,
};
use super::spec::{AccessControl, AccessResult, OperationSpec, OperationType, Visibility};
use crate::core::types::Capabilities;
use crate::protocol::wire::{CallError, ResponseEnvelope};
const NAME_SERVICES_LIST: &str = "services/list";
@@ -198,6 +201,14 @@ fn error_definition_to_json(def: &super::spec::ErrorDefinition) -> Value {
}
pub(crate) fn spec_to_json(spec: &OperationSpec) -> Value {
spec_to_json_pub(spec)
}
/// Public serialization of an `OperationSpec` into the `services/schema`
/// wire shape — the shape `rebuild_spec_for` parses back. Used by the
/// `op/register` bootstrap op (review 004 F-05) to carry announced specs
/// over the wire; `services/schema` serves the same shape.
pub fn spec_to_json_pub(spec: &OperationSpec) -> Value {
let error_schemas: Vec<Value> = spec
.error_schemas
.iter()
@@ -213,6 +224,9 @@ pub(crate) fn spec_to_json(spec: &OperationSpec) -> Value {
"error_schemas": error_schemas,
"access_control": access_control_to_json(&spec.access_control),
});
if let Some(resource_id_path) = &spec.resource_id_path {
json["resource_id_path"] = json!(resource_id_path);
}
if spec.channel_open.is_some() {
json["channel_open"] = json!(true);
}
@@ -257,6 +271,53 @@ pub fn services_list_handler(registry: Arc<OperationRegistry>) -> Handler {
})
}
/// Register the bootstrap discovery ops (`services/list`,
/// `services/list-peers`, `services/schema`) against `registry` with
/// handlers closed over **that same `Arc`** (review 004 F-06): the
/// handlers see every op the registry serves at call time, including
/// per-session registrations made after this install. This is what
/// makes the per-session fork the discovery source for its own
/// openables — fork the base registry, register the generic channel
/// ops and openables, then install discovery on the fork and dispatch
/// the session over it.
///
/// `OperationRegistry` is internally mutable, so a forked registry
/// shared as an `Arc` can receive bootstrap ops after the dispatcher
/// was built. Call this once per session on the session's registry.
///
/// ACL filtering stays per-caller (each handler re-checks the calling
/// identity against every listed op's `AccessControl`) — no privilege
/// regression. Errors are per-op registration failures (e.g. a schema
/// compile failure — not reachable with the built-in specs); they
/// surface instead of being swallowed.
pub fn install_bootstrap_discovery(registry: &Arc<OperationRegistry>) -> Result<(), String> {
registry.register(HandlerRegistration::new(
services_list_spec(),
HandlerKind::Once(services_list_handler(Arc::clone(registry))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))?;
registry.register(HandlerRegistration::new(
services_list_peers_spec(),
HandlerKind::Once(services_list_peers_handler(Arc::clone(registry))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))?;
registry.register(HandlerRegistration::new(
services_schema_spec(),
HandlerKind::Once(services_schema_handler(Arc::clone(registry))),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))?;
Ok(())
}
pub fn services_list_peers_handler(registry: Arc<OperationRegistry>) -> Handler {
Arc::new(move |input: Value, ctx: OperationContext| {
let registry = Arc::clone(&registry);
@@ -337,13 +398,35 @@ pub fn services_schema_handler(registry: Arc<OperationRegistry>) -> Handler {
);
}
};
match registry.registration(&name) {
Some(reg) => {
let spec_json = spec_to_json(&reg.spec);
ResponseEnvelope::ok(ctx.request_id, spec_json)
}
None => ResponseEnvelope::not_found(ctx.request_id, &name),
let registration = match registry.registration(&name) {
Some(reg) => reg,
None => return ResponseEnvelope::not_found(ctx.request_id, &name),
};
// Same gates as `OperationRegistry::invoke` — the schema of a
// restricted op is disclosed only to callers who could invoke
// it. Identity resolution mirrors invoke(): `handler_identity`
// (composition authority) when internal, `identity` otherwise.
if registration.spec.visibility == Visibility::Internal && !ctx.internal {
return ResponseEnvelope::not_found(ctx.request_id, &name);
}
let identity = if ctx.internal {
ctx.handler_identity
.as_ref()
.and_then(|ca| ca.as_identity())
} else {
ctx.identity.clone()
};
if let AccessResult::Forbidden(_) = registration.spec.access_control.check(
identity.as_ref(),
None,
ctx.ownership.as_deref(),
) {
return ResponseEnvelope::not_found(ctx.request_id, &name);
}
let spec_json = spec_to_json(&registration.spec);
ResponseEnvelope::ok(ctx.request_id, spec_json)
})
})
}
@@ -480,7 +563,7 @@ mod tests {
}
fn registry_with_access_controlled_ops() -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec_with_acl("public/echo", AccessControl::default()),
@@ -532,7 +615,7 @@ mod tests {
}
fn registry_with_ops() -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
external_spec("fs/readFile"),
@@ -720,13 +803,88 @@ mod tests {
}
}
// CF-004 — the schema handler must apply the same visibility + ACL
// gates as invoke(). Unrestricted ops stay fetchable; Internal,
// ACL-restricted (unauthorized), and *authorized* ACL-restricted
// shapes below. Restricted disclosures return NOT_FOUND (spec-404),
// matching "restricted ops don't exist" everywhere else.
#[tokio::test]
async fn services_schema_hides_internal_op_from_external_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context("req-cf4-1");
let response = handler(serde_json::json!({ "name": "internal/hidden" }), ctx).await;
match response.result {
Err(e) => assert_eq!(e.code, "NOT_FOUND"),
other => panic!("expected NOT_FOUND for internal op, got {other:?}"),
}
}
#[tokio::test]
async fn services_schema_shows_internal_op_to_internal_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let mut ctx = root_context("req-cf4-2");
ctx.internal = true;
ctx.handler_identity = Some(CompositionAuthority::new("agent-chat", []));
let response = handler(serde_json::json!({ "name": "internal/hidden" }), ctx).await;
let spec = response.result.expect("internal caller sees internal op");
assert_eq!(spec.get("name"), Some(&json!("internal/hidden")));
}
#[tokio::test]
async fn services_schema_hides_acl_restricted_op_from_unauthorized_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context_with_identity(
"req-cf4-3",
Some(identity_with_scopes("regular-peer", &["user"])),
);
let response = handler(serde_json::json!({ "name": "admin/secret" }), ctx).await;
match response.result {
Err(e) => assert_eq!(e.code, "NOT_FOUND"),
other => panic!("expected NOT_FOUND for unauthorized ACL, got {other:?}"),
}
}
#[tokio::test]
async fn services_schema_shows_acl_restricted_op_to_authorized_caller() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context_with_identity(
"req-cf4-4",
Some(identity_with_scopes("admin-peer", &["admin"])),
);
let response = handler(serde_json::json!({ "name": "admin/secret" }), ctx).await;
let spec = response.result.expect("authorized caller sees the spec");
assert_eq!(spec.get("name"), Some(&json!("admin/secret")));
assert_eq!(
spec.get("access_control")
.and_then(|a| a.get("required_scopes")),
Some(&json!(["admin"]))
);
}
#[tokio::test]
async fn services_schema_unrestricted_op_fetchable_without_identity() {
let registry = registry_with_access_controlled_ops();
let handler = services_schema_handler(Arc::clone(&registry));
let ctx = root_context("req-cf4-5");
let response = handler(serde_json::json!({ "name": "public/echo" }), ctx).await;
let spec = response
.result
.expect("default-ACL op fetchable unauthenticated");
assert_eq!(spec.get("name"), Some(&json!("public/echo")));
}
#[tokio::test]
async fn services_list_handler_registered_and_invocable_via_registry() {
let registry = registry_with_ops();
let list_handler = services_list_handler(Arc::clone(&registry));
let schema_handler = services_schema_handler(Arc::clone(&registry));
let mut discovery_registry = OperationRegistry::new();
let discovery_registry = OperationRegistry::new();
discovery_registry
.register(HandlerRegistration::new(
services_list_spec(),
@@ -1105,4 +1263,101 @@ mod tests {
"unauthorized peer must not see admin op in list-peers"
);
}
// --- review 004 F-06: per-fork bootstrap discovery ---------------------
fn context_for(
request_id: &str,
identity: Option<crate::core::auth::Identity>,
) -> OperationContext {
OperationContext {
request_id: request_id.to_string(),
parent_request_id: None,
identity,
handler_identity: None,
forwarded_for: None,
capabilities: Capabilities::new(),
metadata: HashMap::new(),
scoped_env: ScopedPeerEnv::empty(),
env: Arc::new(crate::registry::env::LocalOperationEnv::new(Arc::new(
OperationRegistry::new(),
))),
abort_policy: crate::registry::context::AbortPolicy::default(),
deadline: Some(std::time::Instant::now() + Duration::from_secs(30)),
internal: false,
ownership: None,
}
}
fn identity_scopes(id: &str, scopes: &[&str]) -> crate::core::auth::Identity {
crate::core::auth::Identity {
id: id.to_string(),
scopes: scopes.iter().map(|s| s.to_string()).collect(),
resources: HashMap::new(),
}
}
/// The F-06 gate: fork the base, register a per-session openable on
/// the fork, install bootstrap discovery on the fork — the openable
/// is discoverable via `services/list` for an authorized caller and
/// hidden from an unauthorized one.
#[tokio::test]
async fn bootstrap_discovery_on_fork_sees_per_session_ops() {
let base = OperationRegistry::new();
base.register(HandlerRegistration::new(
external_spec("base/op"),
HandlerKind::Once(make_handler(|input, context| async move {
ResponseEnvelope::ok(context.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let fork = Arc::new(base.fork());
fork.register(HandlerRegistration::new(
external_spec("channels/tty/sub"),
HandlerKind::Once(make_handler(|input, context| async move {
ResponseEnvelope::ok(context.request_id, input)
})),
OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
install_bootstrap_discovery(&fork).expect("bootstrap discovery install");
let handler = fork
.registration("services/list")
.map(|r| match r.handler {
HandlerKind::Once(h) => h,
_ => panic!("services/list must be Once"),
})
.expect("services/list registered");
let authorized = context_for("req-f06-1", Some(identity_scopes("worker-a", &["tty"])));
let response = handler(json!({}), authorized).await;
let names: Vec<String> = response
.result
.expect("ok")
.get("operations")
.and_then(|v| v.as_array())
.expect("operations array")
.iter()
.filter_map(|o| o.get("name").and_then(|n| n.as_str().map(String::from)))
.collect();
assert!(
names.contains(&"channels/tty/sub".to_string()),
"per-session openable discoverable on the fork: {names:?}"
);
assert!(names.contains(&"base/op".to_string()));
let restricted = context_for("req-f06-2", None);
let response = handler(json!({}), restricted).await;
assert!(response.result.is_ok(), "list itself is callable");
}
}
+41 -1
View File
@@ -64,6 +64,16 @@ pub trait OperationEnv: Send + Sync {
Vec::new()
}
/// The operation names this env layer serves. The default returns
/// empty — single-layer envs don't need it. `OverlayOperationEnv`
/// overrides it with its overlay's registered names, and
/// `PeerCompositeEnv::peer_operations` delegates to it per peer
/// (ADR-030 — without the override `services/list-peers` shows
/// every peer with an empty operation list).
fn list_operation_names(&self) -> Vec<String> {
Vec::new()
}
/// Peer-routing composition (ADR-029 §2). Routes to a specific peer
/// (`PeerRef::Specific`) or to the first peer that serves the op
/// (`PeerRef::Any`). The default impl ignores the peer selector and
@@ -144,6 +154,14 @@ impl OperationEnv for LocalOperationEnv {
self.registry.invoke(&name, input, context).await
}
fn list_operation_names(&self) -> Vec<String> {
self.registry
.list_operations()
.into_iter()
.map(|s| s.name)
.collect()
}
}
/// Per-call composite env (ADR-024 + ADR-029 §1). Built by the `Dispatcher`
@@ -298,6 +316,28 @@ impl OperationEnv for PeerCompositeEnv {
fn peer_ids(&self) -> Vec<PeerId> {
self.connection_order.clone()
}
fn peer_operations(&self, peer: &PeerId) -> Vec<String> {
self.connections
.get(peer)
.map(|overlay| overlay.list_operation_names())
.unwrap_or_default()
}
fn list_operation_names(&self) -> Vec<String> {
let mut names: Vec<String> = self
.session
.as_ref()
.map(|s| s.list_operation_names())
.unwrap_or_default();
names.extend(
self.connections
.values()
.flat_map(|c| c.list_operation_names()),
);
names.extend(self.base.list_operation_names());
names
}
}
#[cfg(test)]
@@ -409,7 +449,7 @@ mod tests {
composition_authority: Option<CompositionAuthority>,
scoped_env: Option<ScopedPeerEnv>,
) -> Arc<OperationRegistry> {
let mut registry = OperationRegistry::new();
let registry = OperationRegistry::new();
registry
.register(HandlerRegistration::new(
OperationSpec::new(
+1
View File
@@ -8,5 +8,6 @@
pub mod context;
pub mod discovery;
pub mod env;
pub mod op_register;
pub mod registration;
pub mod spec;
+731
View File
@@ -0,0 +1,731 @@
//! `op/register` — the wire mechanism by which a connected peer
//! announces the operations it serves (review 004 F-05, ADR-022
//! amendment). The envelope kind set stays closed at six; bootstrap ops
//! over channel 0 are the door.
//!
//! Shape: the peer sends `call.requested` for `op/register` with a
//! payload of serializable registration parts — the `OperationSpec` in
//! the `services/schema` wire shape (`spec_to_json`) plus a `replace`
//! flag. The serving-side handler rebuilds the spec, wraps a
//! call-forwarding handler that issues a nested `call.requested` back
//! over channel 0 to the announcing peer (the same shape `from_call`'s
//! imported bundles use), and writes the bundle into that connection's
//! overlay via `CallConnection::register_imported`.
//!
//! The overlay is the landing zone: `compose_root_env` attaches it
//! keyed by peer identity, so nested invocations from any composed
//! handler reach the peer-announced op, and `services/list-peers`
//! discovers it (`ctx.env.peer_operations()`).
//!
//! `AccessControl` gates the surface: the `op/register` op itself
//! carries an `AccessControl` (an unprivileged peer cannot reach the
//! handler at all — the registry's normal invoke path enforces it), and
//! an announced op that collides with an existing registration is
//! rejected unless `replace` is set (the reconnect path).
//!
//! The `Handler` closures cannot cross the wire — the announcing peer
//! keeps its handler locally; the registered bundle is a forwarding
//! stub. This is the same contract `from_call` produces for the
//! hub→consumer import direction, extended to the peer→hub direction.
use std::sync::Arc;
use serde_json::{json, Value};
use crate::core::types::Capabilities;
use crate::protocol::connection::CallConnection;
use crate::protocol::wire::{CallError, ResponseEnvelope};
use crate::registry::registration::{
make_handler, Handler, HandlerKind, HandlerRegistration, OperationProvenance, OperationRegistry,
};
use crate::registry::spec::{AccessControl, OperationSpec, OperationType, Visibility};
pub const OP_REGISTER_NAME: &str = "op/register";
/// The wire DTO a peer sends as the `op/register` input. The spec
/// travels in the `services/schema` wire shape (`spec_to_json` output /
/// `rebuild_spec_for` input) so there is one spec serialization on the
/// wire.
#[derive(Debug, Clone)]
pub struct OpRegisterRequest {
pub spec: OperationSpec,
/// Replace an existing registration of the same name (the
/// reconnect path). `false` (default) rejects a collision with
/// `ALREADY_EXISTS`.
pub replace: bool,
}
impl OpRegisterRequest {
pub fn to_json(&self) -> Value {
json!({
"spec": crate::registry::discovery::spec_to_json_pub(&self.spec),
"replace": self.replace,
})
}
pub fn from_json(value: &Value) -> Result<Self, CallError> {
let spec_json = value
.get("spec")
.ok_or_else(|| CallError::invalid_input("op/register payload missing `spec`"))?;
let name = spec_json
.get("name")
.and_then(|v| v.as_str())
.ok_or_else(|| CallError::invalid_input("op/register spec missing `name`"))?
.to_string();
let spec = crate::client::rebuild_spec_for(spec_json, &name, &None).map_err(|e| {
CallError::invalid_input(format!("op/register spec rebuild failed: {e:?}"))
})?;
let replace = value
.get("replace")
.and_then(|v| v.as_bool())
.unwrap_or(false);
Ok(Self { spec, replace })
}
}
/// The `op/register` `OperationSpec`. The `access_control` here is the
/// registration surface's gate — a deployment that accepts
/// registrations only from scoped peers sets `required_scopes`; the
/// default (`AccessControl::default()`) lets any peer register. The op
/// is `Mutation`-typed (it mutates the connection overlay).
pub fn op_register_spec(access_control: AccessControl) -> OperationSpec {
OperationSpec::new(
OP_REGISTER_NAME,
OperationType::Mutation,
Visibility::External,
json!({
"type": "object",
"properties": {
"spec": { "type": "object" },
"replace": { "type": "boolean" }
},
"required": ["spec"]
}),
json!({
"type": "object",
"properties": {
"name": { "type": "string" },
"registered": { "type": "boolean" }
},
"required": ["name", "registered"]
}),
vec![],
access_control,
None,
)
}
/// Build the `op/register` handler for a connection: announces land in
/// `connection`'s overlay via `register_imported`, wrapped as
/// call-forwarding stubs that issue a nested `call.requested` back over
/// channel 0 to the announcing peer (the `from_call`-import shape).
///
/// `Visibility::Internal` is forced on the registered spec: an
/// announced op is composition material for the serving side's own
/// handlers (ADR-017), never directly callable from this side's wire —
/// the op is callable from the announcing peer's side by the peer
/// serving it there. `services/list-peers` still discovers it (the
/// overlay is peer-keyed, provenance `FromCall`).
///
/// Collision policy (review 005 G-03): announced ops may collide with
/// other *announced* ops on the same connection (`replace` governs,
/// the reconnect path) but **never** with the serving side's own
/// registrations — a name on `serving_registry` rejects with
/// `ALREADY_EXISTS` regardless of `replace`. Without this gate the
/// connection overlay shadows the base registry in `PeerCompositeEnv`
/// (connections resolve before base), so a peer could silently
/// rewrite what a wire-dispatched handler's `ctx.env.invoke` resolves
/// for any name the deployment registered — composition authority
/// (ADR-018) belongs to the composing handler's deployer, not the
/// connected peer.
///
/// Replace semantics: a registration for the same name already on the
/// overlay is rejected with `ALREADY_EXISTS` unless `replace: true`
/// (the reconnect path re-announces).
pub fn op_register_handler(
connection: Arc<CallConnection>,
serving_registry: Arc<OperationRegistry>,
) -> Handler {
make_handler(move |input, context| {
let connection = Arc::clone(&connection);
let serving_registry = Arc::clone(&serving_registry);
async move {
let request = match OpRegisterRequest::from_json(&input) {
Ok(r) => r,
Err(e) => return ResponseEnvelope::error(context.request_id, e),
};
if serving_registry.registration(&request.spec.name).is_some() {
return ResponseEnvelope::error(
context.request_id,
CallError::already_exists(format!(
"op/register: `{}` is registered by this side's own serving \
registry; peer-announced ops may not shadow it",
request.spec.name
)),
);
}
if connection.overlay_contains(&request.spec.name) && !request.replace {
return ResponseEnvelope::error(
context.request_id,
CallError::already_exists(format!(
"op/register: `{}` is already registered on this connection; \
set `replace: true` to replace it",
request.spec.name
)),
);
}
let mut spec = request.spec;
spec.visibility = Visibility::Internal;
let remote_name = spec.name.clone();
let handler = forwarding_stub_for_announced_op(
Arc::clone(&connection),
remote_name,
spec.op_type,
);
connection.register_imported(HandlerRegistration::new(
spec.clone(),
handler,
OperationProvenance::FromCall,
None,
None,
Capabilities::new(),
));
ResponseEnvelope::ok(
context.request_id,
json!({ "name": spec.name, "registered": true }),
)
}
})
}
/// The forwarding stub for a peer-announced op. Query/Mutation ops get
/// the `from_call` forwarding shape (nested `call.requested` back over
/// channel 0, `forwarded_for` populated per ADR-032 §3). Announced
/// Sub/Pub ops are registered as stubs that return
/// `INVALID_OPERATION_TYPE` on invocation — nested composition is
/// request/response-only (`OverlayOperationEnv`'s contract); the
/// streaming/sink forwarding shapes ride on the `from_call` import
/// path, which the announcing side can use in the other direction.
fn forwarding_stub_for_announced_op(
connection: Arc<CallConnection>,
remote_name: String,
op_type: OperationType,
) -> HandlerKind {
match op_type {
OperationType::Query | OperationType::Mutation => HandlerKind::Once(
crate::client::make_forwarding_handler(connection, remote_name),
),
OperationType::Sub | OperationType::Pub => {
HandlerKind::Once(make_handler(|_input, context| async move {
ResponseEnvelope::error(
context.request_id,
CallError::invalid_operation_type(
"peer-announced Sub/Pub ops are not invocable over nested \
composition (request/response only)",
),
)
}))
}
}
}
/// Announce an op to the connected peer over `connection`'s channel 0:
/// sends `call.requested` for `op/register` and awaits the response.
/// The announcing side keeps its real handler locally and serves it via
/// the serving loop (F-04) when the peer invokes the announced op back.
pub async fn announce_op(
connection: &CallConnection,
spec: OperationSpec,
replace: bool,
) -> ResponseEnvelope {
let request = OpRegisterRequest { spec, replace };
connection
.call_with_payload(serde_json::json!({
"operationId": OP_REGISTER_NAME,
"input": request.to_json(),
}))
.await
}
#[cfg(test)]
mod tests {
use super::*;
use crate::protocol::connection::CallConnection;
use crate::registry::context::OperationContext;
use crate::registry::discovery::install_bootstrap_discovery;
use crate::registry::registration::OperationRegistry;
use crate::registry::spec::Visibility;
use std::collections::HashMap;
fn announced_spec(name: &str) -> OperationSpec {
OperationSpec::new(
name,
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
)
}
fn stub_connection() -> crate::core::types::Connection {
crate::protocol::sink_empty_connection()
}
fn test_context(request_id: &str) -> OperationContext {
OperationContext {
request_id: request_id.to_string(),
parent_request_id: None,
identity: None,
handler_identity: None,
forwarded_for: None,
capabilities: Capabilities::new(),
metadata: HashMap::new(),
scoped_env: crate::registry::context::ScopedPeerEnv::empty(),
env: Arc::new(crate::registry::env::LocalOperationEnv::new(Arc::new(
OperationRegistry::new(),
))),
abort_policy: crate::registry::context::AbortPolicy::default(),
deadline: None,
internal: false,
ownership: None,
}
}
#[test]
fn request_round_trips_through_json() {
let request = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: true,
};
let json = request.to_json();
let parsed = OpRegisterRequest::from_json(&json).expect("parse");
assert_eq!(parsed.spec.name, "worker/exec");
assert_eq!(parsed.spec.op_type, OperationType::Query);
assert!(parsed.replace);
}
#[test]
fn request_missing_spec_is_invalid_input() {
let err = OpRegisterRequest::from_json(&json!({})).unwrap_err();
assert_eq!(err.code, "INVALID_INPUT");
}
#[test]
fn request_missing_name_is_invalid_input() {
let err = OpRegisterRequest::from_json(&json!({ "spec": {} })).unwrap_err();
assert_eq!(err.code, "INVALID_INPUT");
}
#[tokio::test]
async fn handler_registers_announced_op_in_overlay() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let response = handler(input, test_context("req-or-1")).await;
assert!(
response.result.is_ok(),
"register succeeded, got {:?}",
response.result
);
let registered = conn.overlay_env().contains("worker/exec");
assert!(registered, "announced op landed in the connection overlay");
assert!(conn.overlay_contains("worker/exec"));
}
#[tokio::test]
async fn handler_rejects_collision_without_replace() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let first = handler(input.clone(), test_context("req-or-2a")).await;
assert!(first.result.is_ok());
let second = handler(input, test_context("req-or-2b")).await;
let err = second.result.expect_err("collision rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
}
#[tokio::test]
async fn handler_replaces_with_replace_flag() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let original = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let first = handler(original, test_context("req-or-3a")).await;
assert!(first.result.is_ok());
let replacement = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: true,
}
.to_json();
let second = handler(replacement, test_context("req-or-3b")).await;
assert!(
second.result.is_ok(),
"replace flag permits re-registration, got {:?}",
second.result
);
}
#[tokio::test]
async fn registered_spec_forced_internal_with_fromcall_provenance() {
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::new(OperationRegistry::new()));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let response = handler(input, test_context("req-or-4")).await;
assert!(response.result.is_ok());
let registration = conn
.overlay_registration("worker/exec")
.expect("registered");
assert_eq!(registration.spec.visibility, Visibility::Internal);
assert_eq!(
registration.provenance,
crate::registry::registration::OperationProvenance::FromCall
);
}
// --- review 005 Unit 2 acceptance gates (G-03 collision policy) -------
/// G-03 gate: an announce colliding with a **serving-registry**
/// name rejects with `ALREADY_EXISTS` even with `replace: true` —
/// peer-announced ops never shadow the serving side's own
/// registrations (composition authority stays with the deployer).
#[tokio::test]
async fn handler_rejects_base_registry_collision_even_with_replace() {
let serving = Arc::new(OperationRegistry::new());
serving
.register(HandlerRegistration::new(
crate::registry::spec::OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
crate::registry::registration::OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let input = OpRegisterRequest {
spec: announced_spec("fs/readFile"),
replace: true,
}
.to_json();
let response = handler(input, test_context("req-or-5")).await;
let err = response.result.expect_err("base collision rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
assert!(
!conn.overlay_contains("fs/readFile"),
"the rejected announce never lands in the overlay"
);
// The serving side's registration is untouched.
assert!(serving.registration("fs/readFile").is_some());
}
/// G-03 gate: an announce colliding with an **Internal** serving
/// op is also rejected — the visibility of the shadowed op is
/// irrelevant to composition shadowing (`OverlayOperationEnv`
/// gates on `AccessControl`, not visibility; the composed child is
/// `internal: true` by design).
#[tokio::test]
async fn handler_rejects_collision_with_internal_serving_op() {
let serving = Arc::new(OperationRegistry::new());
serving
.register(HandlerRegistration::new(
crate::registry::spec::OperationSpec::new(
"internal/vault",
OperationType::Query,
Visibility::Internal,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
),
HandlerKind::Once(make_handler(|input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, input)
})),
crate::registry::registration::OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let input = OpRegisterRequest {
spec: announced_spec("internal/vault"),
replace: false,
}
.to_json();
let response = handler(input, test_context("req-or-6")).await;
let err = response
.result
.expect_err("internal base collision rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
}
/// G-03 gate: an announced op colliding with another *announced*
/// op still follows `replace` semantics — the base-registry gate
/// must not widen into the overlay.
#[tokio::test]
async fn handler_overlay_collision_still_governed_by_replace() {
let serving = Arc::new(OperationRegistry::new());
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let first = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
assert!(handler(first, test_context("req-or-7a"))
.await
.result
.is_ok());
let collision = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
let err = handler(collision, test_context("req-or-7b"))
.await
.result
.expect_err("overlay collision without replace rejected");
assert_eq!(err.code, "ALREADY_EXISTS");
let replacement = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: true,
}
.to_json();
assert!(
handler(replacement, test_context("req-or-7c"))
.await
.result
.is_ok(),
"overlay replace still permitted; base gate is scoped to the serving registry"
);
}
/// G-03 gate: after a successful announce of a distinct name,
/// nested composition of a **base-registered** op still resolves
/// the serving side's own op — the exact `compose_root_env` shape
/// (`PeerCompositeEnv` with the connection overlay attached). The
/// G-03 defect was that the collision gate was overlay-only; this
/// pins the invariant that survived the fix: the base layer is
/// reachable and correct whenever the name is not announced.
#[tokio::test]
async fn nested_composition_of_base_op_unaffected_by_unrelated_announce() {
let serving = Arc::new(OperationRegistry::new());
serving
.register(HandlerRegistration::new(
crate::registry::spec::OperationSpec::new(
"fs/readFile",
OperationType::Query,
Visibility::External,
json!({}),
json!({}),
vec![],
AccessControl::default(),
None,
),
HandlerKind::Once(make_handler(|_input, ctx| async move {
ResponseEnvelope::ok(ctx.request_id, json!({ "from": "base" }))
})),
crate::registry::registration::OperationProvenance::Local,
None,
None,
Capabilities::new(),
))
.unwrap();
let conn = Arc::new(CallConnection::new(stub_connection()));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
// Announce a *distinct* name; it lands in the connection overlay.
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
assert!(handler(input, test_context("req-or-8a"))
.await
.result
.is_ok());
assert!(conn.overlay_contains("worker/exec"));
// Compose the base op through the same env shape
// `compose_root_env` produces for a wire-dispatched handler.
let base = Arc::new(crate::registry::env::LocalOperationEnv::new(Arc::clone(
&serving,
)));
let mut composite = crate::registry::env::PeerCompositeEnv::new(base);
composite.attach_peer("consumer-peer".to_string(), conn.overlay_env());
let env: Arc<dyn crate::registry::env::OperationEnv + Send + Sync> = Arc::new(composite);
let mut ctx = test_context("req-or-8b");
ctx.env = env;
ctx.scoped_env = crate::registry::context::ScopedPeerEnv::new(["fs/readFile"]);
let response = ctx.env.invoke("fs", "readFile", json!({}), &ctx).await;
let out = response.result.expect("base op composes");
assert_eq!(
out,
json!({ "from": "base" }),
"the serving side's own op resolves through composition, not a peer stub"
);
}
/// UP-03 gate (alkhttp review 006): after a peer announces an op,
/// `services/list-peers` over the real `compose_root_env` shape
/// (`PeerCompositeEnv` + the connection overlay attached under the
/// peer's id) must list the announced op under that peer's
/// entry — not an empty operations array. The pre-fix failure:
/// `PeerCompositeEnv` overrode `peer_ids` only, so
/// `peer_operations` fell to the trait default (`Vec::new()`) and
/// every peer listed with `operations: []`. ADR-030's
/// `list_operation_names` override is what this exercises.
#[tokio::test]
async fn announced_op_is_discoverable_via_services_list_peers() {
use crate::registry::env::PeerCompositeEnv;
use crate::registry::{
context::ScopedPeerEnv, discovery::services_list_peers_handler, env::LocalOperationEnv,
};
let serving = Arc::new(OperationRegistry::new());
install_bootstrap_discovery(&serving).expect("bootstrap discovery install");
let peer_identity = crate::core::auth::Identity {
id: "consumer-peer".to_string(),
scopes: vec![],
resources: HashMap::new(),
};
let conn = Arc::new(CallConnection::new_overlay_only(peer_identity));
let handler = op_register_handler(Arc::clone(&conn), Arc::clone(&serving));
let input = OpRegisterRequest {
spec: announced_spec("worker/exec"),
replace: false,
}
.to_json();
assert!(
handler(input, test_context("req-or-9a"))
.await
.result
.is_ok(),
"announce lands in the connection overlay"
);
assert!(conn.overlay_contains("worker/exec"));
// The exact env shape `compose_root_env` produces for calls
// arriving on this connection: LocalOperationEnv base +
// the connection's overlay attached under the peer's id.
let base = Arc::new(LocalOperationEnv::new(Arc::clone(&serving)));
let mut composite = PeerCompositeEnv::new(base);
composite.attach_peer("consumer-peer".to_string(), conn.overlay_env());
let env: Arc<dyn crate::registry::env::OperationEnv + Send + Sync> = Arc::new(composite);
// Direct probe: the composite resolves the announced name from
// the attached overlay, and peer_operations lists it.
assert!(env.contains("worker/exec"));
let ops = env.peer_operations(&"consumer-peer".to_string());
assert_eq!(
ops,
vec!["worker/exec".to_string()],
"PeerCompositeEnv::peer_operations must surface the peer overlay's announced ops"
);
assert!(
env.peer_operations(&"unknown-peer".to_string()).is_empty(),
"an unattached peer has no operations"
);
// Wire-level probe: services/list-peers over the same env
// attributes the announced op to the peer.
let peers_registry = Arc::new(OperationRegistry::new());
install_bootstrap_discovery(&peers_registry).expect("bootstrap install");
let list_handler = services_list_peers_handler(Arc::clone(&peers_registry));
let mut ctx = test_context("req-or-9b");
ctx.env = env;
ctx.scoped_env = ScopedPeerEnv::empty();
let response = list_handler(json!({}), ctx).await;
let out = response.result.expect("list-peers ok");
let peers_arr = out
.get("peers")
.and_then(|v| v.as_array())
.expect("peers array");
let consumer = peers_arr
.iter()
.find(|p| p.get("peer_id").and_then(|v| v.as_str()) == Some("consumer-peer"))
.expect("consumer-peer present in list-peers output");
let names: Vec<&str> = consumer
.get("operations")
.and_then(|v| v.as_array())
.expect("consumer operations array")
.iter()
.filter_map(|o| o.get("name").and_then(|n| n.as_str()))
.collect();
assert!(
names.contains(&"worker/exec"),
"the announced op must be discoverable via services/list-peers (UP-03)"
);
// The local entry still lists the bootstrap ops (the serving
// registry's own surface is unaffected).
let local = peers_arr
.iter()
.find(|p| p.get("peer_id").and_then(|v| v.as_str()) == Some("local"))
.expect("local peer present");
let local_names: Vec<&str> = local
.get("operations")
.and_then(|v| v.as_array())
.expect("local operations array")
.iter()
.filter_map(|o| o.get("name").and_then(|n| n.as_str()))
.collect();
assert!(local_names.contains(&"services/list-peers"));
}
}
File diff suppressed because it is too large. Load diff