--- id: review-001-ws-eof-signal name: Lossless EOF signal + pending-map sweep for from_wss (WS-02, CON-02) status: completed depends_on: [] scope: narrow risk: high impact: component level: implementation tags: [websocket, adapters, review-001, from-wss] --- ## Description Review 001 findings WS-02 + CON-02 — one mechanism, verified end-to-end: `src/websocket/byte_adapter.rs:138-139` (and the tungstenite twin at `:319`) fires `read_eof.notify_waiters()`, which wakes only *already-registered* waiters and stores no permit. `from_wss` spawns its drop-monitor *after* session setup (`from_wss.rs:156-166`); if the read task hits EOF before the monitor first polls `Notified`, the signal is lost. `import()` does `std::mem::forget(session)` (`:193`), so the `close_rx` fallback never fires either — the monitor never runs `fail_all`, and in-flight imported-op calls hang (Once-calls recover only at the 30 s sweeper *if* a sweeper runs; CON-02 establishes it doesn't on this path — `Dispatcher::run_loop`'s sweeper is never taken; `Sub`/`Pub` pendings hang forever). The module doc at `from_wss.rs:111-113` promises the opposite of the behavior. Fix both halves: - Replace `notify_waiters` with a permit-storing signal: `tokio::sync::watch`, `CancellationToken`, or a checked `AtomicBool` — anything a late subscriber observes. Apply to both the axum and tungstenite pump paths. - Extend the from_wss monitor to sweep the pending map periodically while the session lives (or otherwise ensure post-`fail_all` registrations still resolve), since calls registered after the one-shot `fail_all` are currently never resolved. Acceptance gates from the review: (1) a from_wss test that drops the connection **while a call is being registered** — the CON-02 race — with no hang; (2) the module doc's promise ("no hang") becomes true. ## Acceptance Criteria - [x] EOF-notify is stored (late subscriber observes it) — race test: drop during session setup/first call registration resolves all in-flight calls as retryable - [x] Post-`fail_all`-registered pendings also resolve (sweep or equivalent), not hang forever - [x] The `connection_drop_fails_in_flight_calls_retryable_no_hang` test remains green; add the racing-drop variant (COV gap 10) - [x] Module doc at `from_wss.rs:111-113` matches implemented behavior - [x] `cargo test` and `cargo clippy --all-targets -- -D warnings` pass ## References - docs/reviews/001-initial-implementation-review.md (Part B, WS-02; Part G, CON-02; COV-03) - docs/architecture/decisions/070-from-wss-consumer-adapter.md ## Notes > Agent fills during implementation. Highest-priority WS fix — lossy > notification hangs calls; everything else in the WS subsystem can > follow. ## Summary Replaced the `Notify`-based read-EOF signal in `WsPumps` (both axum and tungstenite pump paths in `byte_adapter.rs`) with a retained `tokio::sync::watch` channel (`Sender`/`subscribe()` receiver) — the lossless EOF property WS-02 required. The `from_wss` drop monitor now selects on `eof_rx.changed()` / `close_rx` / a 1 s sweep tick: on EOF it fails all pendings with retryable `CONNECTION_CLOSED` and keeps sweeping every tick so calls registered *after* the initial `fail_all` (the `std::mem::forget` fire-and-forget import path) are also failed — the CON-02 no-hang guarantee, independent of registration-vs-EOF ordering. `std::mem::forget` semantics were left untouched per the task scope. Tests (in `from_wss.rs` `mod tests`): a killable producer harness (`drop_on_signal_producer` — the upgrade handler stashes the server side's `WsPumps` in a slot the test can `abort()`, forcing consumer-side EOF deterministically); `forget_session_drop_during_call_registration…` and `held_session_drop_during_call_registration…` cover the CON-02 race (EOF before/during registration → call resolves Err, no hang); the `call_registered_after_eof…` test registers a call well after EOF and asserts the sweep resolves it retryable; the original `connection_drop_fails_in_flight_calls_retryable_no_hang` stays green. One tolerated outcome note: the held-session race can also resolve via alkcall's write-failure path (`failed to write request frame`, `INTERNAL` — the mux dies between registration and write), which is prompt-but-not-retryable; that path is accepted in the race tests (the retryable assertion lives with the sweep and pre-drop tests where the call is guaranteed in-flight). Verification: cargo test 219 ok; cargo test --features wss 231 ok (3x flake check); cargo clippy --all-targets -- -D warnings (default; all-features blocked only by an unrelated in-flight `src/server/adapter.rs` edit from a parallel agent); cargo fmt --check.