Add TLS handshake timeout (review #010 C3)

tls_handshake_timeout_secs (default 10s, must be > 0) wraps
tls_acceptor.accept() in tokio::time::timeout so stalled handshakes
(slowloris / crawler slow-handshake vector) can no longer hold an FD +
semaphore permit indefinitely. Config plumbing through FullConfig /
StaticConfig / validation / reload diff + docs.
This commit is contained in:
glm-5.3-flash committed 2026-09-13 07:16:43 +00:00
1 parent 978a362536
commit 9e26295cf7
11 files changed
+352 -19

No files matched your search

+6
View File
@@ -91,6 +91,7 @@ Immutable after startup. Changes require a process restart.
| `admin_key_path` | `String` | Path to file containing the admin Bearer token (default: `/etc/reverse-proxy/admin-key`; empty string to disable admin endpoints; see ADR-028) |
| `shutdown_timeout_secs` | `u64` | Maximum seconds to wait for in-flight requests during graceful shutdown (default: `30`) |
| `connection_idle_timeout_secs` | `u64` | Server-side idle timeout for client TLS connections. Idle HTTP/2 connections are closed after this duration (with keep-alive pings at 15s intervals to detect dead peers). HTTP/1.1 connections are closed if the client doesn't send a complete request header within this duration. Prevents FD exhaustion from abandoned connections (default: `60`; must be > 0; see review #007 C1) |
| `tls_handshake_timeout_secs` | `u64` | Maximum seconds a client may take to complete the TLS handshake. Stalled handshakes (e.g. crawlers/slowloris clients that open a TCP connection but never send a ClientHello) are closed after this duration, releasing the FD and connection slot. Without it, a stalled handshake holds an FD + connection semaphore permit indefinitely (default: `10`; must be > 0; see review #010 C3) |
| `max_connections` | `usize` | Maximum number of concurrent client TLS connections. When the limit is reached, new connections wait in the OS TCP backlog until a slot frees (default: `1024`; must be > 0; see review #007 C2) |
| `logging` | `LoggingConfig` | Logging configuration (see below) |
@@ -185,6 +186,7 @@ Phase 2.
| `admin_key_path` | `String` | `/etc/reverse-proxy/admin-key` | No |
| `shutdown_timeout_secs` | `u64` | `30` | No |
| `connection_idle_timeout_secs` | `u64` | `60` | No |
| `tls_handshake_timeout_secs` | `u64` | `10` | No |
| `max_connections` | `usize` | `1024` | No |
| `logging.level` | `String` | `"info"` | No |
| `logging.format` | `String` | `"text"` | No |
@@ -313,6 +315,7 @@ certificate:
health_check_port = 9900 # Local health check (0 to disable)
admin_key_path = "/etc/reverse-proxy/admin-key" # Empty string to disable
# connection_idle_timeout_secs = 60 # Server-side idle timeout (default: 60)
# tls_handshake_timeout_secs = 10 # TLS handshake timeout (default: 10)
# max_connections = 1024 # Max concurrent TLS connections (default: 1024)
[logging]
@@ -461,6 +464,9 @@ On startup, the config is validated:
review #007 C1.
22. `max_connections` must be > 0. A zero value would deadlock the connection
semaphore, preventing any client connection from being accepted.
23. `tls_handshake_timeout_secs` must be > 0. A zero value would immediately
kill every TLS handshake, preventing any client connection from
completing (review #010 C3).
On SIGHUP reload, the same validation applies. If the new config fails
validation, the reload is rejected and the old config remains active. An error
+20 -17
View File
@@ -20,8 +20,9 @@ fixes:
C1: FIXED 2026-09-13 — accept-loop error classification + backoff +
signature-keyed log de-duplication (src/server.rs).
- >-
C3: OPEN — next priority. No TLS handshake timeout; stalled handshakes
hold an FD + a semaphore permit indefinitely (src/server.rs:370).
C3: FIXED 2026-09-13 — TLS handshake timeout (tls_handshake_timeout_secs,
default 10s) wraps tls_acceptor.accept() in src/server.rs; stalled
handshakes release FD + permit.
- >-
C4: OPEN — connection semaphore is per-listener (src/server.rs:338), so
the effective cap is max_connections × listeners; sequence before C2 so
@@ -177,14 +178,16 @@ Notes from implementation:
accept errors (axum 0.8.9 `src/serve/listener.rs` →
`handle_accept_error`). No change needed there; the busy-spin existed
only in the custom HTTPS loop.
- Two follow-up findings surfaced while fixing C1 (still open, tracked
in "Recommended next steps"):
- **C3 (new)**: `tls_acceptor.accept()` has no timeout — a stalled TLS
handshake holds an FD *and* a semaphore permit indefinitely (the idle
watchdog only starts after the handshake completes). This is the
likely actual FD-exhaustion vector for slow/held crawler
- Two follow-up findings surfaced while fixing C1 (tracked in
"Recommended next steps"):
- ~~**C3 (new)**~~ — **FIXED 2026-09-13**: `tls_acceptor.accept()` had
no timeout, so a stalled TLS handshake held an FD *and* a semaphore
permit indefinitely (the idle watchdog only starts after the
handshake completes). Fixed via `tls_handshake_timeout_secs`
(default 10s) wrapping the accept in `tokio::time::timeout`. This
was the likely actual FD-exhaustion vector for slow/held crawler
handshakes, and a slowloris amplifier.
- **C4 (new)**: `conn_sem` is per-listener (`main.rs` spawns one
- **C4 (open)**: `conn_sem` is per-listener (`main.rs` spawns one
`serve_https_listener` per listener, each creating its own
semaphore), so the effective connection cap is
`max_connections × listeners`. Any C2 RLIMIT cross-check must
@@ -262,7 +265,7 @@ observed limit).
Residual risk after mitigation: none identified for FD exhaustion at
current traffic (peak concurrent connections observed ≪ 800); the code
findings C3/C4/C2 remain the durable fix (C1 landed 2026-09-13).
findings C4/C2 remain the durable fix (C1 + C3 landed 2026-09-13).
## Traffic-analysis side note (from the same investigation)
@@ -288,13 +291,13 @@ behavior, which is exactly what made the EMFILE state reachable.
1. ~~Land C1 (accept-loop error backoff + log de-duplication)~~ — DONE
2026-09-13 (see Finding C1).
2. Land C3 (TLS handshake timeout) — bound stalled handshakes so they
cannot hold FD + permit indefinitely; directly closes the crawler
slow-handshake vector described under "Trigger conditions". **Next
priority**: it is the likely actual trigger of the incident (stalled
crawler handshakes under the pre-mitigation 1024 cap), it is the only
finding that defends against a live attacker (slowloris amplifier),
and it does not interact with C2/C4 design decisions.
2. ~~Land C3 (TLS handshake timeout)~~ — DONE 2026-09-13. New static config
`tls_handshake_timeout_secs` (default 10s, must be > 0) wraps
`tls_acceptor.accept()` in `tokio::time::timeout`
(`accept_tls_with_timeout()`, src/server.rs). On timeout the handshake
future is dropped, releasing the TCP FD and the semaphore permit; the
idle watchdog never needs to run for a stalled handshake. Closes the
crawler slow-handshake vector described under "Trigger conditions".
3. Land C4 (shared connection semaphore across listeners) so
`max_connections` is a global cap rather than per-listener. Sequenced
before C2 because the RLIMIT cross-check's FD budget depends on the