Commit Graph
5 Commits
Author SHA1 Message Date
glm-5.2 71ad3c2905 Fix streaming-body watchdog bug (review #009 C2) + dead-upstream test wiring (C3)
C2: IdleTrackingService decremented in_flight when the handler returned
the Response, but for streaming responses (e.g. git clone) the body is
still in flight, so the watchdog could kill the connection mid-stream.
Fix: wrap the response body in IdleTrackingBody which owns the decrement,
releasing it on body EOF/error/drop (guarded against double-decrement by
an AtomicBool) and calling touch() on each data frame. The handler-level
decrement now only happens on the Err branch. Verified by a reproducer
test (streaming_body_not_killed_by_idle_watchdog) that streams 15 chunks
over 3s with a 500ms idle timeout — red on 4ab8c51, green with this fix.

C3: the two pre-existing idle-timeout tests pointed at 127.0.0.1:18080
where nothing listens, so they validated watchdog behaviour against an
immediate 504 error response (HTTP/1.1 prefix + in_flight==1 hold for
a 504 just as for a 200) and never exercised a real upstream round-trip.
Rewired them to spawn a real TestUpstream and assert HTTP/1.1 200 OK.
Removed the dead no-arg make_proxy_router/start_test_https_server wrappers.

Adds http-body as a direct dep (for the Frame/Body trait) and
tokio-stream as a dev-dep (for the slow-stream test helper).

cargo test: 226 unit + 40 integration green; cargo clippy --all-targets clean.

C1 (deploy to dev1) remains open and out of scope for this repo session.
2026-08-19 09:48:13 +00:00
glm-5.2 ff8819e950 Add review #009: undeployed #008 fix + streaming-body watchdog bug
Documents two issues discovered 2026-08-19:
- C1: review #008 fix (commit 4ab8c51) never deployed to dev1; the
  running binary still leaks FDs (433/1024 after 8 days, ~9 FDs/min).
- C2: the watchdog in IdleTrackingService decrements in_flight when
  the handler returns the Response, but for streaming responses (e.g.
  git clone) the body is still in flight, so the watchdog can kill
  the connection mid-stream. Requires an IdleTrackingBody wrapper
  that owns the decrement until the body reaches EOF or is dropped.
2026-08-19 09:29:42 +00:00
glm-5.2 4ab8c516d8 Fix HTTP/2 idle timeout defeated by keep-alive pings (review #008)
The C1 fix from review #007 added keep_alive_interval(15s) +
keep_alive_timeout(60s) on the HTTP/2 builders. However,
keep_alive_timeout is the timeout for receiving a PONG response to a
PING, not an idle timeout. Well-behaved HTTP/2 clients (including
crawlers) respond to PINGs, resetting the timeout indefinitely. After
6 days of uptime, the proxy had 930+ idle connections (some 145 hours
old) that were never reaped.

Fix: add a custom idle timeout that tracks real request activity and
closes connections with no in-flight requests for longer than the
configured timeout, regardless of PING/PONG activity.

- IdleState: per-connection last_activity + in_flight counter
- IdleTrackingService: tower Service wrapper that updates last_activity
  and in_flight on call() and on response completion (prevents killing
  long-running requests like large git clones)
- idle_watchdog: races against serve_connection in tokio::select!,
  fires only when in_flight == 0 and idle_for >= timeout
- HTTP/1.1 path: header_read_timeout already covers between-request
  idle (verified in hyper source); watchdog is defense-in-depth
- HTTP/2 path: watchdog is the primary fix (hyper has no native
  request-activity-based idle timeout)

Also refactored build_manual_server_config to extract
build_manual_server_config_from_certs for testability (in-memory
certs/keys for integration tests with rcgen).

Tests: 8 unit tests for IdleState/watchdog logic, 2 integration tests
verifying idle connections are closed and active ones are not.

Closes review #008.
2026-08-10 09:35:21 +00:00
glm-5.2 0885486028 Fix connection lifecycle, pool bounds, and log rotation (review #007)
Critical:
- C1: Set server-side idle/keep-alive timeouts on both hyper builders
  (http2 keep_alive_interval=15s + keep_alive_timeout, http1
  header_read_timeout). Both builders now set TokioTimer (required to
  avoid runtime panic). Prevents FD exhaustion from abandoned TLS
  connections — the root cause of the 2026-07-24 outage.
- C2: Add Semaphore(max_connections) gating the accept loop. Provides
  backpressure via OS TCP backlog when all permits are taken.

Warnings:
- W1: Add SIGUSR1 log-reopen handler. New ReopenableFileWriter
  (Arc<ArcSwap<File>> via custom MakeWriter) atomically swaps the log
  file. Enables postrotate logrotate without copytruncate, which caused
  the 1.15GB sparse file that wedged fail2ban.
- W2: Set pool_max_idle_per_host(10) on both upstream clients, bounding
  idle upstream connections per host.
- W3: Add connection_idle_timeout_secs to StaticConfig (default 60).
- W4: Add max_connections to StaticConfig (default 1024).

Both new fields are validated (> 0) and included in static config drift
detection on reload. Docs (config.md, README, ADR-009) updated.
2026-07-28 10:16:26 +00:00
glm-5.2 e803817350 Add fail2ban 4xx/badbots filters, jail backend fix, and review #007
- Add reverse-proxy-4xx and reverse-proxy-badbots fail2ban filters
- Set backend=auto and ignoreip on all jails (fixes silent no-match
  when defaults-debian.conf inherits systemd backend)
- Document three-jail setup and REQUEST log format in README
- Add review #007 covering connection lifecycle, logging, and deployment
  drift triggered by the 2026-07-24 FD exhaustion incident
2026-07-28 10:02:19 +00:00