Critical: - C1: Set server-side idle/keep-alive timeouts on both hyper builders (http2 keep_alive_interval=15s + keep_alive_timeout, http1 header_read_timeout). Both builders now set TokioTimer (required to avoid runtime panic). Prevents FD exhaustion from abandoned TLS connections — the root cause of the 2026-07-24 outage. - C2: Add Semaphore(max_connections) gating the accept loop. Provides backpressure via OS TCP backlog when all permits are taken. Warnings: - W1: Add SIGUSR1 log-reopen handler. New ReopenableFileWriter (Arc<ArcSwap<File>> via custom MakeWriter) atomically swaps the log file. Enables postrotate logrotate without copytruncate, which caused the 1.15GB sparse file that wedged fail2ban. - W2: Set pool_max_idle_per_host(10) on both upstream clients, bounding idle upstream connections per host. - W3: Add connection_idle_timeout_secs to StaticConfig (default 60). - W4: Add max_connections to StaticConfig (default 1024). Both new fields are validated (> 0) and included in static config drift detection on reload. Docs (config.md, README, ADR-009) updated.
2.7 KiB
2.7 KiB
ADR-009: Signal Handling Strategy
Status
Accepted
Context
The proxy needs to handle Unix signals for:
- Graceful shutdown: SIGTERM and SIGINT should stop accepting new connections, drain in-flight requests, then exit.
- Config reload: SIGHUP should trigger a DynamicConfig reload from disk.
- Log reopen: SIGUSR1 should close and reopen the log file, enabling
postrotatelogrotate configs withoutcopytruncate(see review #007 W1).
Two approaches for signal handling:
tokio::signal: Built into tokio. Handles SIGTERM and SIGINT viactrl_c(). Does not directly handle SIGHUP.signal-hook: External crate. Handles all Unix signals including SIGHUP. More flexible but adds a dependency.
Decision
Use signal-hook for all signal handling. Specifically:
signal-hook::flagto set termination flags on SIGTERM/SIGINTsignal-hookto register a SIGHUP handler that triggers config reloadsignal-hookto register a SIGUSR1 handler that reopens the log file
tokio::signal::ctrl_c() is registered as a secondary shutdown trigger; both
mechanisms converge on the same shutdown path. This is a belt-and-suspenders
approach: signal-hook handles all signals including SIGHUP, while
ctrl_c() provides a fallback for environments where signal handling may not
be fully wired (e.g., container runtimes).
The shutdown sequence:
- On SIGTERM or SIGINT: stop accepting new connections, wait up to 30 seconds for in-flight requests to complete, then exit with code 0.
- On SIGHUP: re-read config file, validate, and swap DynamicConfig if valid. Log the result.
- On SIGUSR1: close the current log file handle and open a new one at the
same path. Enables standard
postrotatelogrotate configs (rename + signal) withoutcopytruncate, which creates sparse files when the FD offset is high. See review #007 W1.
Rationale
- SIGHUP handling is required for config reload —
tokio::signaldoesn't support SIGHUP. signal-hookis well-maintained, widely used, and handles all Unix signals.- Using one signal handling mechanism (rather than mixing
tokio::signalandsignal-hook) is simpler and avoids edge cases. signal-hook::flagis a minimal, safe API for signal-triggered flags.
Consequences
Positive:
- SIGHUP for config reload is simple and well-understood
- Single signal handling mechanism for all signals
- Compatible with systemd (SIGTERM for shutdown) and standard Unix conventions
Negative:
signal-hookis an additional dependency (but a well-established one)- Signal handling requires careful coordination with the tokio runtime (async signal receivers must be properly integrated)