Add TLS handshake timeout (review #010 C3)
tls_handshake_timeout_secs (default 10s, must be > 0) wraps tls_acceptor.accept() in tokio::time::timeout so stalled handshakes (slowloris / crawler slow-handshake vector) can no longer hold an FD + semaphore permit indefinitely. Config plumbing through FullConfig / StaticConfig / validation / reload diff + docs.
This commit is contained in:
1 parent
978a362536
commit
9e26295cf7
11 files changed
+352
-19
No files matched your search
@@ -91,6 +91,7 @@ Immutable after startup. Changes require a process restart.
|
||||
| `admin_key_path` | `String` | Path to file containing the admin Bearer token (default: `/etc/reverse-proxy/admin-key`; empty string to disable admin endpoints; see ADR-028) |
|
||||
| `shutdown_timeout_secs` | `u64` | Maximum seconds to wait for in-flight requests during graceful shutdown (default: `30`) |
|
||||
| `connection_idle_timeout_secs` | `u64` | Server-side idle timeout for client TLS connections. Idle HTTP/2 connections are closed after this duration (with keep-alive pings at 15s intervals to detect dead peers). HTTP/1.1 connections are closed if the client doesn't send a complete request header within this duration. Prevents FD exhaustion from abandoned connections (default: `60`; must be > 0; see review #007 C1) |
|
||||
| `tls_handshake_timeout_secs` | `u64` | Maximum seconds a client may take to complete the TLS handshake. Stalled handshakes (e.g. crawlers/slowloris clients that open a TCP connection but never send a ClientHello) are closed after this duration, releasing the FD and connection slot. Without it, a stalled handshake holds an FD + connection semaphore permit indefinitely (default: `10`; must be > 0; see review #010 C3) |
|
||||
| `max_connections` | `usize` | Maximum number of concurrent client TLS connections. When the limit is reached, new connections wait in the OS TCP backlog until a slot frees (default: `1024`; must be > 0; see review #007 C2) |
|
||||
| `logging` | `LoggingConfig` | Logging configuration (see below) |
|
||||
|
||||
@@ -185,6 +186,7 @@ Phase 2.
|
||||
| `admin_key_path` | `String` | `/etc/reverse-proxy/admin-key` | No |
|
||||
| `shutdown_timeout_secs` | `u64` | `30` | No |
|
||||
| `connection_idle_timeout_secs` | `u64` | `60` | No |
|
||||
| `tls_handshake_timeout_secs` | `u64` | `10` | No |
|
||||
| `max_connections` | `usize` | `1024` | No |
|
||||
| `logging.level` | `String` | `"info"` | No |
|
||||
| `logging.format` | `String` | `"text"` | No |
|
||||
@@ -313,6 +315,7 @@ certificate:
|
||||
health_check_port = 9900 # Local health check (0 to disable)
|
||||
admin_key_path = "/etc/reverse-proxy/admin-key" # Empty string to disable
|
||||
# connection_idle_timeout_secs = 60 # Server-side idle timeout (default: 60)
|
||||
# tls_handshake_timeout_secs = 10 # TLS handshake timeout (default: 10)
|
||||
# max_connections = 1024 # Max concurrent TLS connections (default: 1024)
|
||||
|
||||
[logging]
|
||||
@@ -461,6 +464,9 @@ On startup, the config is validated:
|
||||
review #007 C1.
|
||||
22. `max_connections` must be > 0. A zero value would deadlock the connection
|
||||
semaphore, preventing any client connection from being accepted.
|
||||
23. `tls_handshake_timeout_secs` must be > 0. A zero value would immediately
|
||||
kill every TLS handshake, preventing any client connection from
|
||||
completing (review #010 C3).
|
||||
|
||||
On SIGHUP reload, the same validation applies. If the new config fails
|
||||
validation, the reload is rejected and the old config remains active. An error
|
||||
|
||||
@@ -20,8 +20,9 @@ fixes:
|
||||
C1: FIXED 2026-09-13 — accept-loop error classification + backoff +
|
||||
signature-keyed log de-duplication (src/server.rs).
|
||||
- >-
|
||||
C3: OPEN — next priority. No TLS handshake timeout; stalled handshakes
|
||||
hold an FD + a semaphore permit indefinitely (src/server.rs:370).
|
||||
C3: FIXED 2026-09-13 — TLS handshake timeout (tls_handshake_timeout_secs,
|
||||
default 10s) wraps tls_acceptor.accept() in src/server.rs; stalled
|
||||
handshakes release FD + permit.
|
||||
- >-
|
||||
C4: OPEN — connection semaphore is per-listener (src/server.rs:338), so
|
||||
the effective cap is max_connections × listeners; sequence before C2 so
|
||||
@@ -177,14 +178,16 @@ Notes from implementation:
|
||||
accept errors (axum 0.8.9 `src/serve/listener.rs` →
|
||||
`handle_accept_error`). No change needed there; the busy-spin existed
|
||||
only in the custom HTTPS loop.
|
||||
- Two follow-up findings surfaced while fixing C1 (still open, tracked
|
||||
in "Recommended next steps"):
|
||||
- **C3 (new)**: `tls_acceptor.accept()` has no timeout — a stalled TLS
|
||||
handshake holds an FD *and* a semaphore permit indefinitely (the idle
|
||||
watchdog only starts after the handshake completes). This is the
|
||||
likely actual FD-exhaustion vector for slow/held crawler
|
||||
- Two follow-up findings surfaced while fixing C1 (tracked in
|
||||
"Recommended next steps"):
|
||||
- ~~**C3 (new)**~~ — **FIXED 2026-09-13**: `tls_acceptor.accept()` had
|
||||
no timeout, so a stalled TLS handshake held an FD *and* a semaphore
|
||||
permit indefinitely (the idle watchdog only starts after the
|
||||
handshake completes). Fixed via `tls_handshake_timeout_secs`
|
||||
(default 10s) wrapping the accept in `tokio::time::timeout`. This
|
||||
was the likely actual FD-exhaustion vector for slow/held crawler
|
||||
handshakes, and a slowloris amplifier.
|
||||
- **C4 (new)**: `conn_sem` is per-listener (`main.rs` spawns one
|
||||
- **C4 (open)**: `conn_sem` is per-listener (`main.rs` spawns one
|
||||
`serve_https_listener` per listener, each creating its own
|
||||
semaphore), so the effective connection cap is
|
||||
`max_connections × listeners`. Any C2 RLIMIT cross-check must
|
||||
@@ -262,7 +265,7 @@ observed limit).
|
||||
|
||||
Residual risk after mitigation: none identified for FD exhaustion at
|
||||
current traffic (peak concurrent connections observed ≪ 800); the code
|
||||
findings C3/C4/C2 remain the durable fix (C1 landed 2026-09-13).
|
||||
findings C4/C2 remain the durable fix (C1 + C3 landed 2026-09-13).
|
||||
|
||||
## Traffic-analysis side note (from the same investigation)
|
||||
|
||||
@@ -288,13 +291,13 @@ behavior, which is exactly what made the EMFILE state reachable.
|
||||
|
||||
1. ~~Land C1 (accept-loop error backoff + log de-duplication)~~ — DONE
|
||||
2026-09-13 (see Finding C1).
|
||||
2. Land C3 (TLS handshake timeout) — bound stalled handshakes so they
|
||||
cannot hold FD + permit indefinitely; directly closes the crawler
|
||||
slow-handshake vector described under "Trigger conditions". **Next
|
||||
priority**: it is the likely actual trigger of the incident (stalled
|
||||
crawler handshakes under the pre-mitigation 1024 cap), it is the only
|
||||
finding that defends against a live attacker (slowloris amplifier),
|
||||
and it does not interact with C2/C4 design decisions.
|
||||
2. ~~Land C3 (TLS handshake timeout)~~ — DONE 2026-09-13. New static config
|
||||
`tls_handshake_timeout_secs` (default 10s, must be > 0) wraps
|
||||
`tls_acceptor.accept()` in `tokio::time::timeout`
|
||||
(`accept_tls_with_timeout()`, src/server.rs). On timeout the handshake
|
||||
future is dropped, releasing the TCP FD and the semaphore permit; the
|
||||
idle watchdog never needs to run for a stalled handshake. Closes the
|
||||
crawler slow-handshake vector described under "Trigger conditions".
|
||||
3. Land C4 (shared connection semaphore across listeners) so
|
||||
`max_connections` is a global cap rather than per-listener. Sequenced
|
||||
before C2 because the RLIMIT cross-check's FD budget depends on the
|
||||
|
||||
Reference in new issue
Block a user