Share connection semaphore across listeners (review #010 C4)
This commit is contained in:
1 parent
9e26295cf7
commit
1bbb9c237c
6 files changed
+183
-25
No files matched your search
@@ -92,7 +92,7 @@ Immutable after startup. Changes require a process restart.
|
||||
| `shutdown_timeout_secs` | `u64` | Maximum seconds to wait for in-flight requests during graceful shutdown (default: `30`) |
|
||||
| `connection_idle_timeout_secs` | `u64` | Server-side idle timeout for client TLS connections. Idle HTTP/2 connections are closed after this duration (with keep-alive pings at 15s intervals to detect dead peers). HTTP/1.1 connections are closed if the client doesn't send a complete request header within this duration. Prevents FD exhaustion from abandoned connections (default: `60`; must be > 0; see review #007 C1) |
|
||||
| `tls_handshake_timeout_secs` | `u64` | Maximum seconds a client may take to complete the TLS handshake. Stalled handshakes (e.g. crawlers/slowloris clients that open a TCP connection but never send a ClientHello) are closed after this duration, releasing the FD and connection slot. Without it, a stalled handshake holds an FD + connection semaphore permit indefinitely (default: `10`; must be > 0; see review #010 C3) |
|
||||
| `max_connections` | `usize` | Maximum number of concurrent client TLS connections. When the limit is reached, new connections wait in the OS TCP backlog until a slot frees (default: `1024`; must be > 0; see review #007 C2) |
|
||||
| `max_connections` | `usize` | Maximum number of concurrent client TLS connections **process-wide**. The connection semaphore is shared across all HTTPS listeners, so this is a global cap — not a per-listener cap. When the limit is reached, new connections wait in the OS TCP backlog until a slot frees (default: `1024`; must be > 0; see review #007 C2, review #010 C4) |
|
||||
| `logging` | `LoggingConfig` | Logging configuration (see below) |
|
||||
|
||||
**LoggingConfig** (nested in `[logging]` TOML section):
|
||||
@@ -316,7 +316,7 @@ health_check_port = 9900 # Local health check (0 to disable)
|
||||
admin_key_path = "/etc/reverse-proxy/admin-key" # Empty string to disable
|
||||
# connection_idle_timeout_secs = 60 # Server-side idle timeout (default: 60)
|
||||
# tls_handshake_timeout_secs = 10 # TLS handshake timeout (default: 10)
|
||||
# max_connections = 1024 # Max concurrent TLS connections (default: 1024)
|
||||
# max_connections = 1024 # Max concurrent TLS connections, process-wide across all listeners (default: 1024)
|
||||
|
||||
[logging]
|
||||
level = "info"
|
||||
@@ -463,7 +463,9 @@ On startup, the config is validated:
|
||||
the server-side idle timeout, reintroducing the FD exhaustion bug from
|
||||
review #007 C1.
|
||||
22. `max_connections` must be > 0. A zero value would deadlock the connection
|
||||
semaphore, preventing any client connection from being accepted.
|
||||
semaphore, preventing any client connection from being accepted. The
|
||||
semaphore is shared across all listeners (review #010 C4), so
|
||||
`max_connections` is the process-wide cap rather than per-listener.
|
||||
23. `tls_handshake_timeout_secs` must be > 0. A zero value would immediately
|
||||
kill every TLS handshake, preventing any client connection from
|
||||
completing (review #010 C3).
|
||||
|
||||
@@ -24,9 +24,9 @@ fixes:
|
||||
default 10s) wraps tls_acceptor.accept() in src/server.rs; stalled
|
||||
handshakes release FD + permit.
|
||||
- >-
|
||||
C4: OPEN — connection semaphore is per-listener (src/server.rs:338), so
|
||||
the effective cap is max_connections × listeners; sequence before C2 so
|
||||
the RLIMIT cross-check uses the real FD budget.
|
||||
C4: FIXED 2026-09-13 — connection semaphore is shared across all
|
||||
listeners (ConnectionSemaphore, src/server.rs); max_connections is now a
|
||||
process-wide cap, not max_connections × listeners.
|
||||
- >-
|
||||
C2: OPEN — no startup cross-check that max_connections fits under
|
||||
RLIMIT_NOFILE with headroom. Sequenced after C3/C4.
|
||||
@@ -187,11 +187,13 @@ Notes from implementation:
|
||||
(default 10s) wrapping the accept in `tokio::time::timeout`. This
|
||||
was the likely actual FD-exhaustion vector for slow/held crawler
|
||||
handshakes, and a slowloris amplifier.
|
||||
- **C4 (open)**: `conn_sem` is per-listener (`main.rs` spawns one
|
||||
`serve_https_listener` per listener, each creating its own
|
||||
semaphore), so the effective connection cap is
|
||||
`max_connections × listeners`. Any C2 RLIMIT cross-check must
|
||||
account for this (or the semaphore should be shared).
|
||||
- ~~**C4 (open)**~~ — **FIXED 2026-09-13**: `conn_sem` was per-listener
|
||||
(`main.rs` spawns one `serve_https_listener` per listener, each creating
|
||||
its own semaphore), so the effective connection cap was
|
||||
`max_connections × listeners`. Fixed via a single `ConnectionSemaphore`
|
||||
(`Arc<Semaphore>` wrapper, src/server.rs) created in `main.rs` and shared
|
||||
by every listener; `max_connections` is now a process-wide cap. This is
|
||||
the topology the C2 RLIMIT cross-check should assume.
|
||||
|
||||
### Original finding (pre-fix, preserved for context)
|
||||
|
||||
@@ -265,7 +267,7 @@ observed limit).
|
||||
|
||||
Residual risk after mitigation: none identified for FD exhaustion at
|
||||
current traffic (peak concurrent connections observed ≪ 800); the code
|
||||
findings C4/C2 remain the durable fix (C1 + C3 landed 2026-09-13).
|
||||
findings C2 remains the durable fix (C1 + C3 + C4 landed 2026-09-13).
|
||||
|
||||
## Traffic-analysis side note (from the same investigation)
|
||||
|
||||
@@ -298,13 +300,18 @@ behavior, which is exactly what made the EMFILE state reachable.
|
||||
future is dropped, releasing the TCP FD and the semaphore permit; the
|
||||
idle watchdog never needs to run for a stalled handshake. Closes the
|
||||
crawler slow-handshake vector described under "Trigger conditions".
|
||||
3. Land C4 (shared connection semaphore across listeners) so
|
||||
`max_connections` is a global cap rather than per-listener. Sequenced
|
||||
before C2 because the RLIMIT cross-check's FD budget depends on the
|
||||
final semaphore topology (shared vs per-listener).
|
||||
3. ~~Land C4 (shared connection semaphore across listeners) so
|
||||
`max_connections` is a global cap rather than per-listener~~ — DONE
|
||||
2026-09-13. `serve_https_listener()` no longer creates its own semaphore;
|
||||
it receives an `Arc<ConnectionSemaphore>` (new public wrapper in
|
||||
src/server.rs) built once in `main.rs` and cloned into every listener
|
||||
task. `max_connections` is now the process-wide concurrent TLS connection
|
||||
cap; the effective cap no longer scales with listener count. This is the
|
||||
topology C2's RLIMIT cross-check should assume.
|
||||
4. Land C2 (RLIMIT cross-check at startup) with a prominent warning or
|
||||
hard validation error. Must account for C4 (per-listener semaphore
|
||||
multiplication → shared after C4 lands) when computing the FD budget.
|
||||
hard validation error. The FD budget assumes a shared semaphore
|
||||
(landed, see step 3): connection FDs are capped at `max_connections`
|
||||
process-wide.
|
||||
Lower urgency after M1: with the deploy baseline (`nofile 8192`,
|
||||
`max_connections 800`) the ceiling is ~10% of the limit, so C2 is a
|
||||
validation guard, not an active exposure.
|
||||
|
||||
Reference in new issue
Block a user