Share connection semaphore across listeners (review #010 C4)

This commit is contained in:
glm-5.3-flash committed 2026-09-13 07:46:58 +00:00
1 parent 9e26295cf7
commit 1bbb9c237c
6 files changed
+183 -25

No files matched your search

+5 -3
View File
@@ -92,7 +92,7 @@ Immutable after startup. Changes require a process restart.
| `shutdown_timeout_secs` | `u64` | Maximum seconds to wait for in-flight requests during graceful shutdown (default: `30`) |
| `connection_idle_timeout_secs` | `u64` | Server-side idle timeout for client TLS connections. Idle HTTP/2 connections are closed after this duration (with keep-alive pings at 15s intervals to detect dead peers). HTTP/1.1 connections are closed if the client doesn't send a complete request header within this duration. Prevents FD exhaustion from abandoned connections (default: `60`; must be > 0; see review #007 C1) |
| `tls_handshake_timeout_secs` | `u64` | Maximum seconds a client may take to complete the TLS handshake. Stalled handshakes (e.g. crawlers/slowloris clients that open a TCP connection but never send a ClientHello) are closed after this duration, releasing the FD and connection slot. Without it, a stalled handshake holds an FD + connection semaphore permit indefinitely (default: `10`; must be > 0; see review #010 C3) |
| `max_connections` | `usize` | Maximum number of concurrent client TLS connections. When the limit is reached, new connections wait in the OS TCP backlog until a slot frees (default: `1024`; must be > 0; see review #007 C2) |
| `max_connections` | `usize` | Maximum number of concurrent client TLS connections **process-wide**. The connection semaphore is shared across all HTTPS listeners, so this is a global cap — not a per-listener cap. When the limit is reached, new connections wait in the OS TCP backlog until a slot frees (default: `1024`; must be > 0; see review #007 C2, review #010 C4) |
| `logging` | `LoggingConfig` | Logging configuration (see below) |
**LoggingConfig** (nested in `[logging]` TOML section):
@@ -316,7 +316,7 @@ health_check_port = 9900 # Local health check (0 to disable)
admin_key_path = "/etc/reverse-proxy/admin-key" # Empty string to disable
# connection_idle_timeout_secs = 60 # Server-side idle timeout (default: 60)
# tls_handshake_timeout_secs = 10 # TLS handshake timeout (default: 10)
# max_connections = 1024 # Max concurrent TLS connections (default: 1024)
# max_connections = 1024 # Max concurrent TLS connections, process-wide across all listeners (default: 1024)
[logging]
level = "info"
@@ -463,7 +463,9 @@ On startup, the config is validated:
the server-side idle timeout, reintroducing the FD exhaustion bug from
review #007 C1.
22. `max_connections` must be > 0. A zero value would deadlock the connection
semaphore, preventing any client connection from being accepted.
semaphore, preventing any client connection from being accepted. The
semaphore is shared across all listeners (review #010 C4), so
`max_connections` is the process-wide cap rather than per-listener.
23. `tls_handshake_timeout_secs` must be > 0. A zero value would immediately
kill every TLS handshake, preventing any client connection from
completing (review #010 C3).
+22 -15
View File
@@ -24,9 +24,9 @@ fixes:
default 10s) wraps tls_acceptor.accept() in src/server.rs; stalled
handshakes release FD + permit.
- >-
C4: OPEN — connection semaphore is per-listener (src/server.rs:338), so
the effective cap is max_connections × listeners; sequence before C2 so
the RLIMIT cross-check uses the real FD budget.
C4: FIXED 2026-09-13 — connection semaphore is shared across all
listeners (ConnectionSemaphore, src/server.rs); max_connections is now a
process-wide cap, not max_connections × listeners.
- >-
C2: OPEN — no startup cross-check that max_connections fits under
RLIMIT_NOFILE with headroom. Sequenced after C3/C4.
@@ -187,11 +187,13 @@ Notes from implementation:
(default 10s) wrapping the accept in `tokio::time::timeout`. This
was the likely actual FD-exhaustion vector for slow/held crawler
handshakes, and a slowloris amplifier.
- **C4 (open)**: `conn_sem` is per-listener (`main.rs` spawns one
`serve_https_listener` per listener, each creating its own
semaphore), so the effective connection cap is
`max_connections × listeners`. Any C2 RLIMIT cross-check must
account for this (or the semaphore should be shared).
- ~~**C4 (open)**~~ — **FIXED 2026-09-13**: `conn_sem` was per-listener
(`main.rs` spawns one `serve_https_listener` per listener, each creating
its own semaphore), so the effective connection cap was
`max_connections × listeners`. Fixed via a single `ConnectionSemaphore`
(`Arc<Semaphore>` wrapper, src/server.rs) created in `main.rs` and shared
by every listener; `max_connections` is now a process-wide cap. This is
the topology the C2 RLIMIT cross-check should assume.
### Original finding (pre-fix, preserved for context)
@@ -265,7 +267,7 @@ observed limit).
Residual risk after mitigation: none identified for FD exhaustion at
current traffic (peak concurrent connections observed ≪ 800); the code
findings C4/C2 remain the durable fix (C1 + C3 landed 2026-09-13).
findings C2 remains the durable fix (C1 + C3 + C4 landed 2026-09-13).
## Traffic-analysis side note (from the same investigation)
@@ -298,13 +300,18 @@ behavior, which is exactly what made the EMFILE state reachable.
future is dropped, releasing the TCP FD and the semaphore permit; the
idle watchdog never needs to run for a stalled handshake. Closes the
crawler slow-handshake vector described under "Trigger conditions".
3. Land C4 (shared connection semaphore across listeners) so
`max_connections` is a global cap rather than per-listener. Sequenced
before C2 because the RLIMIT cross-check's FD budget depends on the
final semaphore topology (shared vs per-listener).
3. ~~Land C4 (shared connection semaphore across listeners) so
`max_connections` is a global cap rather than per-listener~~ — DONE
2026-09-13. `serve_https_listener()` no longer creates its own semaphore;
it receives an `Arc<ConnectionSemaphore>` (new public wrapper in
src/server.rs) built once in `main.rs` and cloned into every listener
task. `max_connections` is now the process-wide concurrent TLS connection
cap; the effective cap no longer scales with listener count. This is the
topology C2's RLIMIT cross-check should assume.
4. Land C2 (RLIMIT cross-check at startup) with a prominent warning or
hard validation error. Must account for C4 (per-listener semaphore
multiplication → shared after C4 lands) when computing the FD budget.
hard validation error. The FD budget assumes a shared semaphore
(landed, see step 3): connection FDs are capped at `max_connections`
process-wide.
Lower urgency after M1: with the deploy baseline (`nofile 8192`,
`max_connections 800`) the ceiling is ~10% of the limit, so C2 is a
validation guard, not an active exposure.