Files
alkhttp/tasks/adapters/review-002-con18-wss-sweep-exit.md
T

2.9 KiB

id, name, status, depends_on, scope, risk, impact, level, tags
id name status depends_on scope risk impact level tags
review-002-con18-wss-sweep-exit Dead-connection fast-fail + bounded from_wss sweep exit (CON-18) completed
narrow low component implementation
adapters
review-002
from-wss

Description

Review 002 CON-18 [minor]. The review-001 CON-02 fix added a 1 s pending-map sweep to from_wss's drop monitor — but it never exits: once eof_observed holds, the sweep branch runs fail_all every second forever (from_wss.rs:254-272). Each fire-and-forget import() whose peer dies leaves a spawned task + pending map + watch receiver alive for the process lifetime (a true task/allocator leak scaling with import count). Note the naive fix is wrong: the sweep test proves post-EOF registrations race fail_all and hang without the sweep — so exit requires making post-EOF registration fail fast first.

Acceptance Criteria

  • Post-EOF (dead connection) registration fails fast: calls registered after the observed EOF resolve immediately with CONNECTION_CLOSED (retryable, matching the established mapping) instead of relying on the next sweep tick
  • The sweep then exits after a bounded grace period following EOF (long enough that the fast-fail path is exercised; not forever) — no per-dead-session permanent task remains
  • Re-verify the existing race tests still pass (pre-drop, forgotten/held drop-during-registration, post-EOF sweep — from_wss.rs:809-1001) and tighten assertions where the fast-fail makes timing deterministic
  • A task-lifecycle test (or task-count assertion) proving the monitor ends after teardown
  • cargo test --features wss, cargo clippy --features wss --all-targets -- -D warnings, cargo fmt --check pass

References

  • docs/reviews/002-post-remediation-review.md (Part B', CON-18 / CON-16-withdrawn merge note)
  • src/adapters/from_wss.rs:244-273 (the monitor loop), :75 (interval), :305 (fire-and-forget import)
  • tasks/adapters/review-001-ws-eof-signal.md (the WS-02/CON-02 fix this bounds)
  • alkcall/docs/reviews/consumer-findings-ledger.md CF-001 (the retryability classification — keep consistent)

Notes

Implementation shape suggestion: an AtomicBool/watch-dead flag on the call connection consulted in the registration path (fail-fast) plus a sweep-generation counter (exit after N idle sweeps post-EOF). The alkcall-side CF-001 note matters: if/when the dead-mux write mapping changes, the fail-fast class here should match. Keep the WS-02 losslessness invariant intact — fast-fail is an additional resolution path, not a replacement for the watch signal.

Summary

50ms drain interval + monitor return after 8 consecutive empty post-EOF drains (~400ms grace); post-EOF registrations fail fast with retryable CONNECTION_CLOSED; fixed a busy-spin in the old select's eof branch. ADR-070 updated. No public API change.