Add ADR-029/030, implementation tasks, and spec updates for admin socket removal

Security review #005 identified critical vulnerabilities in the Unix domain
socket admin API (C1 symlink race, C2 no auth, C3 info leak, W1-W7, S1-S6).
ADR-028 (already accepted) replaces the socket with an authenticated HTTP
admin API on the health check port. This commit adds the remaining spec work:

- ADR-029: Config file TOCTOU mitigation (mtime check on reload)
- ADR-030: Store cli_allow_wildcard_bind in ConfigReloadHandle for consistent
  reload validation
- Implementation tasks for the admin HTTP migration (fix/admin-http-api),
  TOCTOU fix (fix/config-reload-toctou), and wildcard flag fix
  (fix/wildcard-flag-reload)
- Updated review #005 status to resolved with per-finding disposition
- Resolved OQ-16: POST for state-changing admin endpoints, GET for read-only
- Updated all architecture docs to reference new ADRs, use admin_key_path
  instead of admin_socket_path, and reflect POST method for /admin/reload
This commit is contained in:
glm-5.1 committed 2026-06-15 05:19:42 +00:00
1 parent 9096ec5873
commit 161049a17d
21 files changed
+1210 -119

No files matched your search

+10 -5
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-06-12
last_updated: 2026-06-14
---
# Reverse Proxy — Architecture
@@ -48,7 +48,7 @@ connections remain HTTP/1.1.
| [011](decisions/011-multi-domain-tls.md) | Multi-Domain TLS Configuration | Accepted |
| [012](decisions/012-cipher-suite-restriction.md) | Restrict Cipher Suites to nginx Scope | Accepted |
| [013](decisions/013-health-check-port.md) | Health Check on Separate Local Port | Accepted |
| [014](decisions/014-unix-socket-reload.md) | Unix Domain Socket Config Reload API | Accepted |
| [014](decisions/014-unix-socket-reload.md) | ~~Unix Domain Socket Config Reload API~~ | Superseded by ADR-028 |
| [015](decisions/015-per-site-timeouts.md) | Per-Site Upstream Timeouts with Defaults | Accepted |
| [016](decisions/016-explicit-bind-address.md) | Explicit Bind Address Requirement | Accepted |
| [017](decisions/017-upstream-connection-defaults.md) | Upstream Connection Defaults | Accepted |
@@ -56,12 +56,15 @@ connections remain HTTP/1.1.
| [019](decisions/019-multi-config-listeners.md) | Multi-Config Listener Support | Accepted |
| [020](decisions/020-container-deployment.md) | Container Deployment Model | Accepted |
| [021](decisions/021-x-forwarded-for-edge-proxy.md) | X-Forwarded-For Edge Proxy Model | Accepted |
| [022](decisions/022-health-check-scope.md) | Health Check Scope — Local Port and Admin Socket Only | Accepted |
| [022](decisions/022-health-check-scope.md) | Health Check Scope — Local Port and Admin HTTP Only | Accepted |
| [023](decisions/023-http2-client-facing.md) | HTTP/2 Client-Facing Support | Accepted |
| [024](decisions/024-ansi-disabled-logging.md) | ANSI-Disabled Logging for Container Deployments | Accepted |
| [025](decisions/025-rate-limiter-ip-source.md) | Rate Limiter IP Source — ConnectInfo Only | Accepted |
| [026](decisions/026-connector-timeout-ceiling.md) | Connector Timeout Ceiling for Per-Site Timeouts | Accepted |
| [027](decisions/027-admin-socket-resource-limits.md) | Admin Socket Resource Limits | Accepted |
| [027](decisions/027-admin-socket-resource-limits.md) | ~~Admin Socket Resource Limits~~ | Deprecated (ADR-028) |
| [028](decisions/028-admin-http-api.md) | Authenticated HTTP Admin API | Accepted |
| [029](decisions/029-config-reload-toctou.md) | Config File TOCTOU Mitigation on Reload | Accepted |
| [030](decisions/030-wildcard-flag-consistency.md) | Store cli_allow_wildcard_bind in ConfigReloadHandle | Accepted |
## Open Questions
@@ -72,7 +75,7 @@ See [open-questions.md](open-questions.md) for the full tracker.
| ~~OQ-01~~ | ~~Should cipher suites be restricted beyond rustls defaults?~~ | ~~medium~~ | **resolved** (ADR-012) |
| ~~OQ-02~~ | ~~What log format should fail2ban consume?~~ | ~~high~~ | **resolved** (ADR-007) |
| ~~OQ-03~~ | ~~Should the health check endpoint be on a separate port?~~ | ~~low~~ | **resolved** (ADR-013) |
| ~~OQ-04~~ | ~~Config reload: SIGHUP only or also Unix socket API?~~ | ~~low~~ | **resolved** (ADR-014) |
| ~~OQ-04~~ | ~~Config reload: SIGHUP only or also Unix socket API?~~ | ~~low~~ | **resolved** (ADR-014, superseded by ADR-028: authenticated HTTP admin API) |
| ~~OQ-05~~ | ~~Should the proxy bind to multiple addresses?~~ | ~~low~~ | **resolved** (single bind_addr sufficient) |
| ~~OQ-06~~ | ~~Should upstream timeouts be configurable per-site?~~ | ~~low~~ | **resolved** (ADR-015) |
| ~~OQ-07~~ | ~~Should per-site TLS overrides be supported for mixed ACME/manual domains?~~ | ~~low~~ | **resolved** (ADR-019) |
@@ -83,6 +86,8 @@ See [open-questions.md](open-questions.md) for the full tracker.
| ~~OQ-12~~ | ~~Should request access logging be mandatory or optional?~~ | ~~high~~ | **resolved** (mandatory, always-on per operations.md) |
| OQ-13 | Should `acme_contact` support multiple email addresses? | low | open |
| OQ-14 | Should rate limiter eviction interval and max age be configurable? | low | open |
| OQ-15 | Should admin key rotation persist across restarts? | medium | open |
| ~~OQ-16~~ | ~~Should /admin/reload use POST instead of GET?~~ | ~~low~~ | **resolved** (ADR-028: POST for state-changing, GET for read-only) |
## Document Lifecycle
+43 -16
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-06-12
last_updated: 2026-06-14
---
# Configuration
@@ -36,7 +36,7 @@ config.toml
│ (immutable) │
│ │
│ health_check_port │
│ admin_socket_path │
│ admin_key_path │
│ log_level │
│ log_format │
│ │
@@ -88,7 +88,7 @@ Immutable after startup. Changes require a process restart.
| `listeners` | `Vec<ListenerConfig>` | Independent TLS endpoints, each with its own bind address and TLS config (see ADR-019) |
| `allow_wildcard_bind` | `bool` | Allow `0.0.0.0` as a bind address. Required for container deployments. Default: `false` (see ADR-016, ADR-020) |
| `health_check_port` | `u16` | Port for local health check endpoint (default: `9900`; set to `0` to disable; bound to `127.0.0.1` only; see ADR-013, ADR-022) |
| `admin_socket_path` | `String` | Unix domain socket path for admin API (default: `/run/reverse-proxy/admin.sock`; empty string to disable; see ADR-014) |
| `admin_key_path` | `String` | Path to file containing the admin Bearer token (default: `/etc/reverse-proxy/admin-key`; empty string to disable admin endpoints; see ADR-028) |
| `shutdown_timeout_secs` | `u64` | Maximum seconds to wait for in-flight requests during graceful shutdown (default: `30`) |
| `logging` | `LoggingConfig` | Logging configuration (see below) |
@@ -178,7 +178,7 @@ Phase 2.
|-------|------|---------|----------|
| `allow_wildcard_bind` | `bool` | `false` | No |
| `health_check_port` | `u16` | `9900` | No |
| `admin_socket_path` | `String` | `/run/reverse-proxy/admin.sock` | No |
| `admin_key_path` | `String` | `/etc/reverse-proxy/admin-key` | No |
| `shutdown_timeout_secs` | `u64` | `30` | No |
| `logging.level` | `String` | `"info"` | No |
| `logging.format` | `String` | `"text"` | No |
@@ -220,10 +220,10 @@ Config reload is triggered by two mechanisms:
valid. Simple and well-understood, but provides no feedback on success or
failure.
2. **Admin socket**: The `reload` command via the admin Unix domain socket
performs the same action as SIGHUP but returns a structured response
indicating success or failure with an error message. See ADR-014 for
details.
2. **Admin HTTP endpoint**: The `POST /admin/reload` endpoint on the health
check listener (`127.0.0.1:9900`) performs the same action as SIGHUP but
returns a structured response indicating success or failure. Requires a
Bearer token from the admin key file. See ADR-028 for details.
Both mechanisms converge on the same code path:
1. Re-read the config file from disk
@@ -233,7 +233,7 @@ Both mechanisms converge on the same code path:
### Static Config Changes During Reload
When the config file is reloaded (via SIGHUP or admin socket), the entire file
When the config file is reloaded (via SIGHUP or admin HTTP endpoint), the entire file
is read and validated — both static and dynamic portions. This provides early
error detection for misconfigurations that would prevent a restart from
succeeding.
@@ -259,16 +259,36 @@ reloads, the operator should see the warning only once, not on every reload.
### Reload Serialization
Reload operations are serialized using a `tokio::sync::Mutex` on the reload
code path. If a reload is in progress (triggered by SIGHUP or admin socket) and
code path. If a reload is in progress (triggered by SIGHUP or admin HTTP) and
a second reload is requested, the second request waits for the first to
complete, then re-reads the config file (getting the latest version) and
proceeds. This prevents race conditions where two concurrent reloads could apply
an older config over a newer one.
### Config File TOCTOU Protection
Both SIGHUP and admin HTTP reload paths compare the config file's modification
timestamp before and after reading. If the file changed during the read, the
reload is rejected with a "please retry" message. This detects the common case
of a config management tool mid-write. See ADR-029.
Operators should use atomic file replacement for config rotation: write to a
temporary file in the same directory, then `rename()` over the target. This
is the standard safe pattern and is what tools like Ansible do with their
`copy` module's `validate` parameter.
### Reload Validation Consistency
The `cli_allow_wildcard_bind` flag is stored in `ConfigReloadHandle` at startup
and used for reload validation, ensuring that a config accepted at startup
will also be accepted on reload. If the proxy was started with
`--allow-wildcard-bind`, reload will also allow `0.0.0.0` bind addresses. See
ADR-030.
### Out of Scope: File Watching
Automatic file watching (inotify, fsnotify, etc.) is out of scope for Phase 1.
Config reload is triggered explicitly by SIGHUP or admin socket command. File
Config reload is triggered explicitly by SIGHUP or admin HTTP endpoint. File
watching adds complexity (debouncing, handling atomic renames, handling editor
swap files) that is not justified for a single-instance proxy with infrequent
config changes.
@@ -285,7 +305,7 @@ certificate:
# Global settings
health_check_port = 9900 # Local health check (0 to disable)
admin_socket_path = "/run/reverse-proxy/admin.sock" # Empty string to disable
admin_key_path = "/etc/reverse-proxy/admin-key" # Empty string to disable
[logging]
level = "info"
@@ -343,7 +363,7 @@ A single listener serving multiple domains with one SAN certificate:
```toml
# Global settings
health_check_port = 9900
admin_socket_path = "/run/reverse-proxy/admin.sock"
admin_key_path = "/etc/reverse-proxy/admin-key"
[logging]
level = "info"
@@ -425,6 +445,9 @@ On startup, the config is validated:
(e.g., `"mailto:admin@example.com"`). Values like `"mailto:"` (empty
email) or `"mailto:user"` (no `@`) are rejected. Let's Encrypt requires
a contact email for production certificate requests.
20. `admin_key_path` must be either an empty string (disabled) or an absolute
path. Relative paths and paths containing `..` are rejected. This prevents
path traversal attacks on the admin key file.
On SIGHUP reload, the same validation applies. If the new config fails
validation, the reload is rejected and the old config remains active. An error
@@ -445,13 +468,16 @@ All design decisions are documented as ADRs in [decisions/](decisions/).
| [010](decisions/010-multi-site-phase1.md) | Multi-site in Phase 1 | Multiple domains from initial release |
| [011](decisions/011-multi-domain-tls.md) | Multi-domain TLS config | Single SAN certificate covering all domains |
| [013](decisions/013-health-check-port.md) | Health check on separate local port | Localhost-only HTTP health check, configurable port |
| [014](decisions/014-unix-socket-reload.md) | Unix domain socket config reload API | Programmatic reload with success/failure feedback |
| [014](decisions/014-unix-socket-reload.md) | ~~Unix domain socket config reload API~~ | ~~Programmatic reload with success/failure feedback~~ (Superseded by ADR-028) |
| [015](decisions/015-per-site-timeouts.md) | Per-site upstream timeouts with defaults | 5s connect / 60s request defaults, per-site overrides |
| [016](decisions/016-explicit-bind-address.md) | Explicit bind address required | Rejects `0.0.0.0` to prevent accidental exposure |
| [019](decisions/019-multi-config-listeners.md) | Multi-config listeners | `[[listeners]]` supporting both dedicated-IP and shared-IP deployment models |
| [020](decisions/020-container-deployment.md) | Container deployment model | Flexible upstream addressing; `allow_wildcard_bind` override for containers |
| [026](decisions/026-connector-timeout-ceiling.md) | Connector timeout ceiling | 30s ceiling on connector, per-site timeout via tokio::time::timeout |
| [027](decisions/027-admin-socket-resource-limits.md) | Admin socket resource limits | 5s read timeout, 4096 byte line length limit |
| [027](decisions/027-admin-socket-resource-limits.md) | ~~Admin socket resource limits~~ | ~~5s read timeout, 4096 byte line length limit~~ (Deprecated by ADR-028) |
| [028](decisions/028-admin-http-api.md) | Authenticated HTTP admin API | Bearer token auth on health check port; replaces Unix domain socket |
| [029](decisions/029-config-reload-toctou.md) | Config file TOCTOU mitigation | mtime check before and after config read; reject reload if file changed |
| [030](decisions/030-wildcard-flag-consistency.md) | Store cli_allow_wildcard_bind in ConfigReloadHandle | Consistent validation between startup and reload |
## Open Questions
@@ -459,7 +485,8 @@ Open questions are tracked in [open-questions.md](open-questions.md). Key
questions affecting this document:
- ~~**OQ-04**: Should config reload support a Unix domain socket API in addition
to SIGHUP?~~ (resolved — ADR-014: Unix domain socket admin API added)
to SIGHUP?~~ (resolved — ADR-014, now superseded by ADR-028: authenticated
HTTP admin API)
- ~~**OQ-07**: Should per-site TLS overrides be supported for mixed ACME/manual
domains?~~ (resolved — ADR-019: `[[listeners]]` with per-listener TLS config)
- **OQ-13**: Should `acme_contact` support multiple email addresses? (see
@@ -48,8 +48,8 @@ listener (see ADR-022).
- Configurable port allows different deployment scenarios (some monitoring runs
on different ports)
- Disabling via `health_check_port = 0` removes the health check entirely —
the admin socket's `status` command remains available as an alternative
health/status mechanism
the admin HTTP endpoint's `/admin/status` (with Bearer token) remains available
as an alternative health/status mechanism (ADR-028)
- When this project is folded into alknet, the health check will use alknet's
existing patterns, making the separate port unnecessary in that context
@@ -69,4 +69,5 @@ listener (see ADR-022).
- [operations.md](../operations.md)
- [ADR-022](022-health-check-scope.md) — Health check scope (no `/health` on main listener)
- [ADR-028](028-admin-http-api.md) — Authenticated HTTP admin API (admin endpoints on health check port)
- OQ-03 (now resolved)
@@ -2,7 +2,7 @@
## Status
Accepted
Superseded by [ADR-028](028-admin-http-api.md)
## Context
@@ -99,7 +99,7 @@ Example configuration:
```toml
# Global settings
health_check_port = 9900
admin_socket_path = "/run/reverse-proxy/admin.sock"
admin_key_path = "/etc/reverse-proxy/admin-key"
[logging]
level = "info"
@@ -56,10 +56,11 @@ to consume directly from the host filesystem.
and `journalctl`). File logging is the authoritative source for fail2ban
because it avoids the fragility of Docker log driver parsing.
5. **ACME state and admin socket are volume-mounted.** The ACME cache directory
(`/var/lib/reverse-proxy/acme-cache/`) and admin socket
(`/run/reverse-proxy/admin.sock`) are mounted as volumes so state persists
across container restarts and the host can send reload commands.
5. **ACME state and admin key are volume-mounted.** The ACME cache directory
(`/var/lib/reverse-proxy/acme-cache/`) and admin key file
(`/etc/reverse-proxy/admin-key`) are mounted as volumes so state persists
across container restarts and the host can send authenticated reload commands
(ADR-028).
6. **Health checks use Docker's native mechanism.** The health check endpoint
on port 9900 (localhost only) is used directly by Docker's `HEALTHCHECK`
@@ -1,4 +1,4 @@
# ADR-022: Health Check Scope — Local Port and Admin Socket Only
# ADR-022: Health Check Scope — Local Port and Admin HTTP Endpoint Only
## Status
@@ -25,8 +25,8 @@ handled exclusively by:
1. **Local health check port** (default: 9900, bound to `127.0.0.1`) — serves
`GET /health → 200 OK`. This is the primary health check mechanism for
container orchestration, load balancers, and monitoring systems.
2. **Admin socket** (`status` command) — returns process information including
uptime and site count.
2. **Admin HTTP endpoint** (`GET /admin/status` with Bearer token) — returns
process information including uptime and site count. See ADR-028.
The `/health` route is removed from the main listener entirely. No configurable
path is needed because the route simply does not exist on the public listener.
@@ -52,5 +52,6 @@ path is needed because the route simply does not exist on the public listener.
## References
- ADR-013: Health check on separate local port
- ADR-028: Authenticated HTTP admin API (admin socket replaced by HTTP endpoint)
- OQ-08: Resolved by this ADR
- Implementation review finding W5 (hardcoded `/health` path)
@@ -2,7 +2,9 @@
## Status
Accepted
Deprecated — the Unix domain socket admin API has been replaced by an
authenticated HTTP admin endpoint (ADR-028). Socket resource limits are no
longer needed.
## Context
@@ -0,0 +1,228 @@
# ADR-028: Authenticated HTTP Admin API (Replacing Unix Domain Socket)
## Status
Accepted
## Context
The proxy has a Unix domain socket admin API (ADR-014) that provides two
commands: `reload` (trigger config reload with success/failure feedback) and
`status` (return uptime and site count). Security review #005 identified three
critical and seven warning-level vulnerabilities in the socket implementation:
- **C1**: Symlink race in stale socket cleanup enables arbitrary file deletion
- **C2**: No authentication — any local user can trigger config reload
- **C3**: Error responses leak filesystem paths and config structure details
- **W1**: No connection concurrency limit
- **W3**: Socket path not validated or sanitized
- **W4**: `is_socket_active` side-effect on other processes
- **W5**: Reload validation uses different `cli_allow_wildcard_bind` flag than
startup
These vulnerabilities stem from the fundamental design choice of using a Unix
domain socket. The socket introduces an entire class of filesystem-based attack
surface that does not exist with an HTTP endpoint: symlink races, stale socket
cleanup, path traversal, permission management, and directory mount issues in
containers.
Additionally, the socket requires `socat` for interaction — a non-standard tool
that must be installed separately, complicating container images and CI/CD
pipelines.
The proxy already has a localhost-only HTTP listener (`src/health.rs`) bound to
`127.0.0.1:9900` that serves `/health`. This listener is axum-based, supports
middleware layers, and has integration tests. Co-locating admin endpoints on
this listener is the natural replacement.
## Decision
Replace the Unix domain socket admin API with authenticated HTTP endpoints on
the existing health check listener. Authentication uses a Bearer token verified
against a SHA-256 hash stored in memory.
### Admin Key Management
The admin key is stored in a file on disk (specified by `admin_key_path` in
StaticConfig). The proxy reads this file once at startup, hashes its contents
with SHA-256, and stores only the hash in memory. The plaintext key is never
held in memory after startup initialization.
Key file setup:
```bash
openssl rand -hex 32 > /etc/reverse-proxy/admin-key
chmod 600 /etc/reverse-proxy/admin-key
```
Setting `admin_key_path` to an empty string disables admin endpoints entirely.
### Authentication
Admin endpoints require a Bearer token in the `Authorization` header:
```
Authorization: Bearer <key>
```
The provided token is SHA-256 hashed and compared against the stored hash using
constant-time comparison (`subtle::ConstantTimeEq`) to prevent timing attacks.
Error behavior by auth state:
| Scenario | Response |
|----------|----------|
| Admin disabled (`admin_key_path` empty) | 404 (endpoint does not exist) |
| Missing `Authorization` header | 401 |
| Wrong token | 401 |
| Correct token | Proceed to handler |
Returning 404 when admin is disabled prevents discovery of the endpoint's
existence. Returning 401 for wrong tokens (rather than 404) allows operators to
confirm the endpoint is available without revealing information to attackers
who lack any valid token.
### Endpoints
| Method | Path | Auth | Description |
|--------|------|------|-------------|
| GET | `/health` | None | Health check (unchanged) |
| POST | `/admin/reload` | Bearer token | Trigger config reload |
| GET | `/admin/status` | Bearer token | Return uptime and site count |
| POST | `/admin/rotate-key` | Bearer token | Generate and return a new random admin key |
**`/admin/reload`** — Triggers the same config reload as SIGHUP. Returns
structured JSON:
```json
{"status": "ok"}
```
On error:
```json
{"status": "error", "message": "reload failed"}
```
Error messages are generic — no filesystem paths, no config structure details.
Full error information is logged server-side only.
**`/admin/status`** — Returns process information:
```json
{"status": "ok", "uptime_secs": 1234, "sites": 2}
```
**`/admin/rotate-key`** — Generates a new 256-bit random key using
`rand::RngCore`, returns it in the response, and replaces the stored hash in
memory with the SHA-256 hash of the new key:
```json
{"status": "ok", "key": "<new-plaintext-key-hex>"}
```
The operator should capture this key and update the key file on disk for
subsequent restarts. In-memory rotation does **not** persist across restarts —
on restart, the proxy re-reads the key file. This is by design: the file on
disk is the source of truth for the admin key, and runtime rotation is a
temporary override.
If an attacker has the current admin key, they could rotate it to lock out the
legitimate operator. But if an attacker has the admin key, they can already
trigger config reloads — the most dangerous operation. Rotation is strictly
less damaging than what they could already do.
### Config Change
Replace `admin_socket_path` (StaticConfig) with `admin_key_path` (StaticConfig):
```toml
# Before (ADR-014)
admin_socket_path = "/run/reverse-proxy/admin.sock" # empty = disabled
# After (ADR-028)
admin_key_path = "/etc/reverse-proxy/admin-key" # empty = disabled
```
Default: `"/etc/reverse-proxy/admin-key"`.
### Comparison with Previous Design
| Aspect | Unix Socket (ADR-014) | HTTP Admin (ADR-028) |
|--------|-----------------------|----------------------|
| Authentication | None (filesystem permissions only) | Bearer token with constant-time comparison |
| Attack surface | Filesystem: symlinks, stale cleanup, path traversal | File: read once at startup, no management |
| Client tool | `socat` (non-standard) | `curl` (universal) |
| Error leakage | Paths and config details in responses | Generic messages, details logged server-side |
| Container setup | Volume mount for socket directory | Volume mount for key file (single file, `:ro`) |
| Feedback | Structured JSON responses | Structured JSON responses (same) |
| SIGHUP fallback | Yes (both work) | Yes (both work) |
### What This Eliminates from Review #005
| Finding | Eliminated? | Reason |
|---------|------------|--------|
| C1 (symlink race) | Yes | No socket file management at all |
| C2 (no authentication) | Yes | Bearer token with constant-time comparison |
| C3 (info leak) | Yes | Generic error messages, no paths |
| W1 (no conn limit) | Yes | axum/TCP backlog handles this naturally |
| W3 (path validation) | Yes | No socket path to validate; key file path is read-only, no creation/cleanup |
| W4 (is_socket_active) | Yes | No stale socket detection needed |
| W5 (wildcard flag) | No | Still exists (separate fix) |
| W2 (config TOCTOU) | No | Still exists (separate fix) |
### Remaining Findings
W2 (config file TOCTOU on reload) and W5 (reload validation uses different
`cli_allow_wildcard_bind` flag) still apply to both the SIGHUP and HTTP admin
reload paths. These are independent of the admin interface choice and require
separate fixes.
## Rationale
- **Eliminates an attack surface class**: Every critical finding in review #005
stems from the socket being a filesystem object. Removing the socket removes
the class.
- **Read-once semantics**: The proxy reads the key file once at startup and
never manages it — no creation, no cleanup, no stale detection. This is
fundamentally different from the socket, which required bind, listen, accept,
cleanup-on-startup, cleanup-on-shutdown, and stale detection.
- **Standard tooling**: `curl` is available everywhere. `socat` requires
separate installation in container images and CI environments.
- **Authentication**: Bearer tokens are the standard pattern for HTTP APIs.
Constant-time comparison prevents timing attacks. SHA-256 hashing means the
plaintext key is never held in memory after startup.
- **Key file is lower-risk than socket**: Reading a file is a single syscall.
The socket required managing a filesystem object across the entire process
lifecycle. If an attacker can read the key file, they can also read the
config file — the key file does not expand the trust boundary.
## Consequences
**Positive:**
- Eliminates C1, C2, C3, W1, W3, W4 from security review #005
- Authentication for admin operations (the socket had none)
- Universal client tooling (`curl` instead of `socat`)
- Simpler container setup (single file mount vs. directory mount)
- No socket lifecycle management (startup cleanup, shutdown cleanup, stale
detection)
- Generic error responses prevent information disclosure
**Negative:**
- Key file must exist on disk for admin endpoints to work
- Key file must be readable by the proxy process
- In-memory key rotation does not persist across restarts (operator must
update the key file separately)
- Adds `subtle` and `sha2` crate dependencies
- Admin endpoints share the health check port (operational port serves both
authenticated and unauthenticated routes)
## References
- [operations.md](../operations.md)
- [config.md](../config.md)
- [overview.md](../overview.md)
- [ADR-014](014-unix-socket-reload.md) — Superseded by this ADR
- [ADR-027](027-admin-socket-resource-limits.md) — Deprecated (no longer needed)
- [Review #005](../../reviews/005-admin-socket-security-review.md)
- [Review #006](../../reviews/006-attack-surface-review.md)
@@ -0,0 +1,90 @@
# ADR-029: Config File TOCTOU Mitigation on Reload
## Status
Accepted
## Context
Both the SIGHUP reload path (`src/shutdown.rs:handle_sighup_reload`) and the
admin HTTP reload path (`src/admin/socket.rs:handle_reload`, soon
`src/admin/handler.rs`) read the config file from disk with
`tokio::fs::read_to_string()`, then parse and apply it. If another process is
writing to the config file at the same time (e.g., a configuration management
tool like Ansible writing a partial file), the proxy could read a partially
written config and either fail to parse it (resulting in a reload error) or,
in an unlikely worst case, parse a structurally valid but semantically wrong
config.
This is a filesystem-level time-of-check/time-of-use (TOCTOU) issue. The
window is small but the impact of applying a partial config is significant.
Security review #005 identified this as finding W2.
## Decision
Detect mid-write file changes by comparing file metadata before and after
reading. If the modification timestamp changes between the two `stat` calls,
reject the reload and return a retry message.
```rust
let metadata_before = tokio::fs::metadata(&config_path).await?;
let config_content = tokio::fs::read_to_string(&config_path).await?;
let metadata_after = tokio::fs::metadata(&config_path).await?;
if metadata_before.modified()? != metadata_after.modified()? {
return Err("config file changed during read, please retry");
}
```
This applies to **both** the SIGHUP reload path and the admin HTTP reload path.
For operators, the documentation will recommend the atomic replacement pattern
(write to a temp file in the same directory, then `rename()` over the target).
This is the standard safe pattern for config file rotation and is what tools
like Ansible already do with `copy` module's `validate` parameter.
## Rationale
- **Simple and effective**: The mtime check catches the common case of a
config management tool mid-write. It requires no changes to the config
file format or directory layout.
- **No false negatives**: If mtime changed, the file definitely changed. If
mtime did not change within the typical filesystem timestamp granularity
(1 second on most Linux filesystems), the window is so small that a partial
read is extremely unlikely.
- **Atomic rename is the gold standard**: Recommending it in documentation is
better than trying to enforce it in code. The proxy can't control how
operators write config files, but it can detect when a file might be
inconsistent and ask for a retry.
- **Same pattern in both reload paths**: SIGHUP and admin HTTP share the same
file-reading logic (or should — currently they duplicate it). This ADR
ensures both paths are protected.
## Consequences
**Positive:**
- Config reload will reject a file that changed during the read, preventing
partial or inconsistent configs from being applied.
- Clear error message ("config file changed during read, please retry") tells
operators exactly what happened.
- Documenting the atomic replacement pattern gives operators a clear
recommendation for safe config rotation.
**Negative:**
- In very rare cases, a legitimate config change that happens to land within
the same filesystem timestamp granularity as the read could be falsely
rejected. The operator would need to retry the reload, which is an
acceptable trade-off for safety.
- The mtime check does not protect against all TOCTOU scenarios (e.g., a
write that starts before the first `stat` and completes before the read).
However, combined with the atomic replacement recommendation, this is a
defense-in-depth measure, not a complete solution. A complete solution would
require file locking or checksum verification, which adds complexity for
marginal benefit.
## References
- [operations.md](../operations.md) — Config reload, admin HTTP endpoint
- [config.md](../config.md) — Config reload behavior
- [Review #005](../../reviews/005-admin-socket-security-review.md) — W2 finding
@@ -0,0 +1,99 @@
# ADR-030: Store cli_allow_wildcard_bind Flag in ConfigReloadHandle
## Status
Accepted
## Context
When the proxy starts, `cli_allow_wildcard_bind` can be set to `true` via
the `--allow-wildcard-bind` CLI flag or the `allow_wildcard_bind = true` config
option. The startup validation uses this flag to decide whether `0.0.0.0` bind
addresses are allowed.
However, when a config reload is triggered (via SIGHUP or admin HTTP), the
`validate()` call is invoked with `cli_allow_wildcard_bind: false` — hardcoded
in `ConfigReloadHandle::reload()`. This means that a config that was accepted
at startup (because the CLI flag was set) will be rejected on reload, even
though the running process has `allow_wildcard_bind = true` in effect.
Security review #005 identified this as finding W5. The consequence is that
an operator who started the proxy with `--allow-wildcard-bind` cannot reload
the config without getting a validation error about `0.0.0.0` bind addresses
— even though those bind addresses are currently active and working.
## Decision
Store the `cli_allow_wildcard_bind` flag in `ConfigReloadHandle` at
construction time, and use the stored value during reload validation instead
of hardcoding `false`.
```rust
pub struct ConfigReloadHandle {
config: Arc<ArcSwap<DynamicConfig>>,
static_config: ArcSwap<StaticConfig>,
reload_mutex: Mutex<()>,
cli_allow_wildcard_bind: bool,
}
impl ConfigReloadHandle {
pub fn new(
config: Arc<ArcSwap<DynamicConfig>>,
static_config: StaticConfig,
cli_allow_wildcard_bind: bool,
) -> Self {
Self {
config,
static_config: ArcSwap::from_pointee(static_config),
reload_mutex: Mutex::new(()),
cli_allow_wildcard_bind,
}
}
}
```
In `reload()`, pass `self.cli_allow_wildcard_bind` to `validate()` instead of
`false`:
```rust
validate(&new_static, &new_dynamic, self.cli_allow_wildcard_bind)?;
```
This ensures reload validation uses the same flag as startup validation. The
flag is immutable — it's set once at startup and never changed — so storing it
in `ConfigReloadHandle` is safe.
## Rationale
- **Consistency**: Startup and reload should apply the same validation rules.
If `0.0.0.0` was allowed at startup, it should be allowed on reload.
- **The flag is immutable**: `cli_allow_wildcard_bind` is set once from CLI
args and never changes. Storing it in `ConfigReloadHandle` is a simple,
correct solution.
- **No config file change needed**: The `allow_wildcard_bind` config option is
already in `StaticConfig`. The CLI flag is a separate override. The fix is
purely in how the reload path uses the flag.
- **OR logic preserved**: The validation uses OR logic (`config_flag ||
cli_flag`). If either is true, wildcard binds are allowed. This is unchanged.
## Consequences
**Positive:**
- Config reload will no longer reject valid configs that were accepted at
startup due to the `--allow-wildcard-bind` CLI flag.
- Consistent validation between startup and reload paths.
**Negative:**
- `ConfigReloadHandle::new()` gains an additional parameter. This is a minor
API change but affects all construction sites.
- The flag cannot be changed at runtime. If an operator wants to remove
`--allow-wildcard-bind`, they must restart the process. This is correct
behavior — wildcard bind is a security-sensitive setting that should
require a restart.
## References
- [config.md](../config.md) — Validation rules, allow_wildcard_bind
- [Review #005](../../reviews/005-admin-socket-security-review.md) — W5 finding
- `src/config/dynamic_config.rs` — ConfigReloadHandle, reload()
- `src/config/validation.rs` — validate()
+41 -12
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-06-12
last_updated: 2026-06-14
---
# Open Questions
@@ -51,8 +51,8 @@ last_updated: 2026-06-12
- **Resolution**: Add a configurable local health check port (default: 9900)
bound to `127.0.0.1` only. Health checks work even when TLS is misconfigured.
There is no `/health` route on the main HTTPS listener — health checking is
handled exclusively by the local port and admin socket. See ADR-013 and
ADR-022.
handled exclusively by the local port and admin HTTP endpoint. See ADR-013,
ADR-022, and ADR-028.
- **Cross-references**: ADR-013, ADR-022
## Configuration
@@ -62,10 +62,13 @@ last_updated: 2026-06-12
- **Origin**: [config.md](config.md)
- **Status**: resolved
- **Priority**: low
- **Resolution**: Yes. Add a Unix domain socket admin API alongside SIGHUP.
The socket accepts a `reload` command and returns structured success/failure
responses. SIGHUP is retained as a fallback. See ADR-014.
- **Cross-references**: ADR-014
- **Resolution**: Yes, via ADR-014 (Unix domain socket). The socket has since
been replaced by an authenticated HTTP admin API (ADR-028) due to security
vulnerabilities identified in review #005. SIGHUP is retained as a fallback.
Admin HTTP endpoints (`/admin/reload`, `/admin/status`, `/admin/rotate-key`)
are served on the health check listener at `127.0.0.1:9900` behind Bearer
token authentication.
- **Cross-references**: ADR-014 (superseded), ADR-028
## Deployment
@@ -100,14 +103,14 @@ last_updated: 2026-06-12
- **Priority**: medium
- **Resolution**: The `/health` route does not belong on the main listener at
all. Health checking is an operational concern served by the dedicated local
port (9900) and the admin socket's `status` command — not by intercepting
port (9900) and the admin HTTP endpoint's `/admin/status` — not by intercepting
traffic on the public-facing proxy. Serving `/health` on the main listener
creates collision with upstream applications, requires special-case routing
logic before host-based matching, and is architecturally wrong: the main
listener's job is to proxy requests, not to serve operational endpoints. The
local health check port (bound to `127.0.0.1:9900`) and the admin socket are
the sole health/status mechanisms. See ADR-022.
- **Cross-references**: ADR-013, ADR-022
local health check port (bound to `127.0.0.1:9900`) and the admin HTTP endpoint
are the sole health/status mechanisms. See ADR-022 and ADR-028.
- **Cross-references**: ADR-013, ADR-022, ADR-028
### ~~OQ-09: How should `upstream_connect_timeout_secs` be enforced?~~
@@ -200,4 +203,30 @@ last_updated: 2026-06-12
longer max age might be desirable. These would be dynamic config fields
(hot-reloadable via ArcSwap) if added. For Phase 1, the hardcoded values
are reasonable defaults.
- **Cross-references**: ADR-006
- **Cross-references**: ADR-006
## Security
### OQ-15: Should admin key rotation persist across restarts?
- **Origin**: [operations.md](operations.md), ADR-028
- **Status**: open
- **Priority**: medium
- **Details**: The `/admin/rotate-key` endpoint generates a new random key and
replaces the stored hash in memory, but this change does not persist across
restarts. On restart, the proxy re-reads the admin key file. To make rotation
permanent, the operator must also update the key file on disk separately.
Options for persisting rotation: (1) write the new hash to the key file on
rotation (breaks the "read-only key file" model), (2) write a separate state
file with the rotated hash, (3) document the current behavior and accept it
(operator updates the file manually after rotation). Option 3 is simplest
and maintains the read-only semantics of the key file mount.
- **Cross-references**: ADR-028
### ~~OQ-16: Should the admin HTTP API use POST for /admin/reload instead of GET?~~
- **Origin**: ADR-28 specifies GET for all admin endpoints, but HTTP semantics suggest POST for state-changing operations (reload).
- **Status**: resolved
- **Priority**: low
- **Resolution**: State-changing admin endpoints (`/admin/reload`, `/admin/rotate-key`) use POST. Read-only endpoints (`/admin/status`, `/health`) use GET. This follows standard HTTP semantics — GET requests should be safe and idempotent, and triggering a config reload is neither. The convenience argument (shorter curl command) does not justify violating HTTP method semantics. See ADR-028.
- **Cross-references**: ADR-028
+78 -45
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-06-12
last_updated: 2026-06-14
---
# Operations
@@ -205,13 +205,18 @@ GET http://127.0.0.1:9900/health → 200 OK (empty body)
The port is configurable via `health_check_port` in StaticConfig. Setting it
to `0` disables the health check listener entirely.
The admin socket's `status` command provides an additional health/status
The admin HTTP endpoint's `status` command provides an additional health/status
mechanism that returns process information:
```
{"status": "ok", "uptime_secs": 1234, "sites": 2}
GET http://127.0.0.1:9900/admin/status
Authorization: Bearer <admin-key>
→ {"status": "ok", "uptime_secs": 1234, "sites": 2}
```
Requires Bearer token authentication. See ADR-028 for details.
### What It Checks
- Process is running and the tokio runtime is responsive
@@ -276,45 +281,68 @@ The proxy handles three signals via `signal-hook` (see [ADR-009](decisions/009-s
for in-flight requests to complete (up to a configurable timeout), then exit.
- **SIGHUP**: Config reload. Re-read the config file, validate, and swap
DynamicConfig if valid. No feedback on success or failure.
- **Admin socket reload**: Send `reload` command via the Unix domain socket
(default: `/run/reverse-proxy/admin.sock`). Returns structured response
indicating success or failure. See ADR-014 for details.
- **Admin HTTP endpoint**: Send `POST /admin/reload` with a Bearer token to the
health check listener (`http://127.0.0.1:9900/admin/reload`). Returns
structured response indicating success or failure. See ADR-028 for details.
### SIGHUP for Config Reload
SIGHUP triggers config reload (see [config.md](config.md) for details). The
process does not exit on SIGHUP.
### Admin Socket for Config Reload
### Admin HTTP Endpoint for Config Reload
The admin Unix domain socket provides programmatic config reload with feedback.
This is useful for CI/CD pipelines and automation tools. See ADR-014 for the
rationale.
The admin HTTP endpoint provides programmatic config reload with feedback
and authentication. This is useful for CI/CD pipelines and automation tools.
See ADR-028 for the rationale and security model.
**Protocol:**
**Authentication**: All `/admin/*` endpoints require a Bearer token in the
`Authorization` header. The token is compared against a SHA-256 hash of the
admin key file contents using constant-time comparison (`subtle::ConstantTimeEq`)
to prevent timing attacks.
- **Connection lifecycle**: One command per connection. Client connects, sends
one newline-terminated command, receives one newline-terminated JSON
response, then the server closes the connection.
- **Message framing**: Newline-delimited (`\n`). Responses end with `\n`.
- **Resource limits** (see ADR-027):
- Read timeout: 5 seconds. Connections that send no complete command within
5 seconds are closed. The timeout is logged at `debug` level.
- Line length limit: 4096 bytes. Connections that send more than 4096 bytes
without a newline are closed. The event is logged at `warn` level.
- **Commands**:
- `reload` — Re-read config file, validate, and swap DynamicConfig. Returns
`{"status": "ok"}` or `{"status": "error", "message": "..."}`.
- `status` — Return basic process info. Returns
`{"status": "ok", "uptime_secs": 1234, "sites": 2}`.
- **Error responses**: Unrecognized commands return
`{"status": "error", "message": "unknown command: <cmd>"}`. Invalid or empty
input returns `{"status": "error", "message": "invalid input"}`.
- **Concurrency**: Multiple clients can connect simultaneously, but reload
operations are serialized (see Config Reload section in config.md).
- **Socket cleanup**: The proxy removes any existing socket file at startup
before binding. If the file exists and another process is listening, a warning
is logged and the admin socket is disabled (but the proxy continues starting).
**Key management**: The admin key is stored in a file on disk (specified by
`admin_key_path` in StaticConfig, default: `/etc/reverse-proxy/admin-key`).
The proxy reads this file once at startup, hashes its contents with SHA-256,
and stores only the hash in memory. The plaintext key is never held in memory
after startup. Setting `admin_key_path` to an empty string disables all admin
endpoints.
**Endpoints**:
| Method | Path | Description |
|--------|------|-------------|
| POST | `/admin/reload` | Trigger config reload. Returns `{"status": "ok"}` or `{"status": "error", "message": "..."}` |
| GET | `/admin/status` | Return process info: `{"status": "ok", "uptime_secs": N, "sites": N}` |
| POST | `/admin/rotate-key` | Generate new random key, return it, replace stored hash |
**Error responses**: Admin endpoint errors return generic messages — no
filesystem paths, no config structure details. Full error information is
logged server-side only. When admin endpoints are disabled (empty
`admin_key_path`), `/admin/*` returns 404 (endpoint appears not to exist).
Wrong or missing Bearer token returns 401.
**Key rotation**: `/admin/rotate-key` generates a new 256-bit random key,
returns the plaintext key in the response, and replaces the stored SHA-256
hash in memory. In-memory rotation does **not** persist across restarts — on
restart, the proxy re-reads the key file. To make rotation permanent, the
operator must also update the key file on disk.
**Concurrency**: Reload operations are serialized (see Config Reload section
in config.md). Multiple concurrent `/admin/reload` requests are handled
correctly.
**Config file TOCTOU protection**: Both SIGHUP and admin HTTP reload compare
the config file's modification timestamp before and after reading. If the
file changed during the read, the reload is rejected with a "please retry"
message. Operators should use atomic file replacement (write to a temp file,
then `rename()` over the target) for safe config rotation. See ADR-029.
**Reload validation consistency**: The `cli_allow_wildcard_bind` flag is
stored in `ConfigReloadHandle` at startup, ensuring that reload validation
uses the same flag as startup validation. If the proxy was started with
`--allow-wildcard-bind`, reload will also allow wildcard bind addresses.
See ADR-030.
### Shutdown Sequence
@@ -339,7 +367,7 @@ On SIGTERM or SIGINT, the proxy performs a graceful shutdown:
4. **Force-close remaining connections** — After the timeout, any remaining
connections are forcefully closed via TCP RST.
5. **Cancel background tasks** — ACME renewal tasks, rate limiter eviction task,
and admin socket listener are all cancelled.
and rate limiter eviction task are all cancelled.
6. **Exit with code 0**.
The `shutdown_timeout_secs` is configurable in StaticConfig (default: 30
@@ -427,7 +455,7 @@ override is required for this configuration (see ADR-016, ADR-020).
| `/etc/reverse-proxy/config.toml` | Config file (read-only) | Proxy configuration |
| `/var/lib/reverse-proxy/acme-cache/` | ACME state directory | Certificate persistence across restarts |
| `/var/log/reverse-proxy/` | Log directory | fail2ban reads from host |
| `/run/reverse-proxy/admin.sock` | Admin socket | Host-side config reload commands |
| `/etc/reverse-proxy/admin-key` | Admin key file | Bearer token for admin endpoints (read-only) |
### Docker Compose Example
@@ -448,7 +476,7 @@ services:
- /etc/reverse-proxy/config.toml:/etc/reverse-proxy/config.toml:ro
- /var/lib/reverse-proxy/acme-cache:/var/lib/reverse-proxy/acme-cache
- /var/log/reverse-proxy:/var/log/reverse-proxy
- /run/reverse-proxy:/run/reverse-proxy
- /etc/reverse-proxy/admin-key:/etc/reverse-proxy/admin-key:ro
networks:
- proxy-net
healthcheck:
@@ -495,7 +523,7 @@ Corresponding proxy config (inside the container):
```toml
allow_wildcard_bind = true
health_check_port = 9900
admin_socket_path = "/run/reverse-proxy/admin.sock"
admin_key_path = "/etc/reverse-proxy/admin-key"
[logging]
level = "info"
@@ -576,10 +604,10 @@ and correct dependency initialization:
4. **Bind health check port** (if enabled) — Bind `127.0.0.1:{health_check_port}`.
Fail-fast if bind fails.
5. **Bind admin socket** (if enabled) — Remove any stale socket file first, then
bind the Unix domain socket. If the socket file exists and another process is
listening, log a warning and fail the admin socket (but continue starting —
the admin socket is non-critical).
5. **Read admin key** (if enabled) — Read the admin key file, compute SHA-256
hash, and store the hash in memory for Bearer token authentication. If the
file is not readable, log a warning and disable admin endpoints (but
continue starting — admin endpoints are non-critical).
6. **Bind all listener ports** — For each listener: bind HTTP port (if enabled)
and HTTPS port. If any bind fails, fail-fast and exit. All ports are bound
@@ -593,7 +621,7 @@ and correct dependency initialization:
8. **Start TCP listeners** — Begin accepting connections on all bound ports.
9. **Start background tasks** — ACME renewal tasks (per listener in ACME mode),
rate limiter eviction task, signal handler task, admin socket handler task.
rate limiter eviction task, signal handler task.
10. **Signal readiness** — Send `sd_notify("READY=1")` to systemd (if running
under systemd).
@@ -613,11 +641,14 @@ All design decisions are documented as ADRs in [decisions/](decisions/).
| [007](decisions/007-custom-log-format.md) | Custom structured log format | key=value pairs with RATE_LIMIT prefix for fail2ban |
| [009](decisions/009-signal-handling.md) | Signal handling strategy | signal-hook for SIGTERM/SIGINT/SIGHUP |
| [013](decisions/013-health-check-port.md) | Health check on separate local port | Localhost-only HTTP health check, configurable port |
| [014](decisions/014-unix-socket-reload.md) | Unix domain socket config reload API | Programmatic reload with success/failure feedback |
| [014](decisions/014-unix-socket-reload.md) | ~~Unix domain socket config reload API~~ | ~~Programmatic reload with success/failure feedback~~ (Superseded by ADR-028) |
| [020](decisions/020-container-deployment.md) | Container deployment model | Defense-in-depth via container isolation; file-primary logging |
| [024](decisions/024-ansi-disabled-logging.md) | ANSI-disabled logging | All log output uses `with_ansi(false)` for fail2ban and Docker compatibility |
| [025](decisions/025-rate-limiter-ip-source.md) | Rate limiter IP source | ConnectInfo only, never client-supplied X-Forwarded-For |
| [027](decisions/027-admin-socket-resource-limits.md) | Admin socket resource limits | 5s read timeout, 4096 byte line length limit |
| [027](decisions/027-admin-socket-resource-limits.md) | ~~Admin socket resource limits~~ | ~~5s read timeout, 4096 byte line length limit~~ (Deprecated by ADR-028) |
| [028](decisions/028-admin-http-api.md) | Authenticated HTTP admin API | Bearer token auth on health check port; replaces Unix domain socket |
| [029](decisions/029-config-reload-toctou.md) | Config file TOCTOU mitigation | mtime check before and after config read; reject reload if file changed |
| [030](decisions/030-wildcard-flag-consistency.md) | Store cli_allow_wildcard_bind in ConfigReloadHandle | Consistent validation between startup and reload |
## Open Questions
@@ -628,9 +659,11 @@ questions affecting this document:
— ADR-013: separate local port, default 9900, localhost only)
- ~~**OQ-08**: Should `/health` use a less common path?~~ (resolved — ADR-022:
no `/health` route on the main listener at all; health checking is via port
9900 and admin socket only)
9900 and admin HTTP endpoint only)
- ~~**OQ-12**: Should request access logging be mandatory or optional?~~ (resolved
— access logging is mandatory and always-on at `info` level; no configuration
option to disable it)
- **OQ-14**: Should rate limiter eviction interval and max age be configurable?
(see [open-questions.md](open-questions.md))
- **OQ-15**: Should admin key rotation persist across restarts?
(see [open-questions.md](open-questions.md))
+25 -20
View File
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-06-12
last_updated: 2026-06-14
---
# Overview
@@ -57,7 +57,7 @@ details.
- 100 MB body size limit (global)
- Configurable bind addresses (must be explicit, no `0.0.0.0`)
- Local health check endpoint on separate port (default: 9900, localhost only)
- Unix domain socket admin API for config reload with feedback
- Authenticated HTTP admin API for config reload with feedback (ADR-028)
- Graceful shutdown (SIGTERM handling with in-flight request drain)
- Systemd unit file
- Dual licensing: MIT OR Apache-2.0
@@ -115,17 +115,17 @@ details.
│ │ └─ Rate limiting, headers │ │
│ └────────────────────────────────┘ │
│ │
│ /health → 200 OK (port 9900) │
│ Admin socket (Unix domain) │
└────────────────────────────────────┘
│ │
┌──────┘ └──────┐
│ │
Docker network Volume mounts:
(upstream DNS) ├─ config (ro)
├─ gitea:3000 ├─ ACME cache (rw)
├─ app:8080 ├─ log dir (rw, fail2ban)
└─ admin socket (rw)
│ /health → 200 OK (port 9900) │
│ /admin/* → Bearer token auth │
└────────────────────────────────────┘
│
┌──────┘
│
Docker network Volume mounts:
(upstream DNS) ├─ config (ro)
├─ gitea:3000 ├─ ACME cache (rw)
├─ app:8080 ├─ log dir (rw, fail2ban)
└─ admin key (ro)
```
Each listener has its own `axum::Router` instance with its own middleware stack,
@@ -133,8 +133,7 @@ but all routers share `Arc<ArcSwap<DynamicConfig>>` and
`Arc<Mutex<HashMap<IpAddr, TokenBucket>>>` via axum State. Site routing is
global: the `Host` header is matched against a single routing table collected
from all listeners' site definitions. Hostnames must be unique across all
listeners. Hostnames must be unique across all listeners — see Security & Bug
Review #003, finding C1, resolved by ADR-025.
listeners — see Security & Bug Review #003, finding C1, resolved by ADR-025.
In container deployments (ADR-020), the proxy runs in a minimal container with
`0.0.0.0` bind address and Docker port publishing. Upstream addresses use Docker
@@ -171,6 +170,8 @@ loopback, LAN, and tunnel endpoints for multi-host deployments.
| `rustls-pki-types` | 1 | TLS types | CertificateDer, PrivateKeyDer |
| `clap` | 4 | CLI arguments | Server startup options |
| `signal-hook` | 0.3 | Signal handling | SIGTERM/SIGINT for shutdown, SIGHUP for config reload |
| `subtle` | 2 | Constant-time comparison | Bearer token auth for admin endpoints |
| `sha2` | 0.10 | SHA-256 hashing | Admin key hashing at startup |
Versions listed are minimum major versions. Implementation should pin exact
versions in `Cargo.toml` per standard Rust practice.
@@ -207,7 +208,7 @@ All design decisions are documented as ADRs in [decisions/](decisions/).
| [011](decisions/011-multi-domain-tls.md) | Multi-domain TLS config | Single SAN certificate covering all domains via rustls-acme |
| [012](decisions/012-cipher-suite-restriction.md) | Restrict cipher suites | Match nginx scope: ECDHE-AES-GCM for TLS 1.2, all TLS 1.3 |
| [013](decisions/013-health-check-port.md) | Health check on separate local port | Localhost-only HTTP health check, configurable port |
| [014](decisions/014-unix-socket-reload.md) | Unix domain socket config reload API | Programmatic reload with success/failure feedback |
| [014](decisions/014-unix-socket-reload.md) | ~~Unix domain socket config reload API~~ | ~~Programmatic reload with success/failure feedback~~ (Superseded by ADR-028) |
| [015](decisions/015-per-site-timeouts.md) | Per-site upstream timeouts with defaults | 5s connect / 60s request defaults, per-site overrides |
| [016](decisions/016-explicit-bind-address.md) | Explicit bind address required | Rejects `0.0.0.0` to prevent accidental exposure |
| [017](decisions/017-upstream-connection-defaults.md) | Upstream connection defaults | HTTP/1.1, no redirects, connection pooling |
@@ -215,12 +216,15 @@ All design decisions are documented as ADRs in [decisions/](decisions/).
| [019](decisions/019-multi-config-listeners.md) | Multi-config listeners | `[[listeners]]` supporting both dedicated-IP and shared-IP deployment models |
| [020](decisions/020-container-deployment.md) | Container deployment model | Defense-in-depth via container isolation; file-primary logging; flexible upstream addressing |
| [021](decisions/021-x-forwarded-for-edge-proxy.md) | X-Forwarded-For edge proxy model | Replace, don't append — proxy is the edge, no trusted upstream proxies |
| [022](decisions/022-health-check-scope.md) | Health check scope — local port and admin socket only | No `/health` route on main listener; health check is port 9900/admin socket only |
| [022](decisions/022-health-check-scope.md) | Health check scope — local port and admin HTTP only | No `/health` route on main listener; health check is port 9900/admin HTTP only |
| [023](decisions/023-http2-client-facing.md) | HTTP/2 client-facing support | ALPN-based protocol detection; HTTP/2 to clients, HTTP/1.1 to upstreams |
| [024](decisions/024-ansi-disabled-logging.md) | ANSI-disabled logging | All log output uses `with_ansi(false)` for fail2ban and Docker compatibility |
| [025](decisions/025-rate-limiter-ip-source.md) | Rate limiter IP source | ConnectInfo only, never client-supplied X-Forwarded-For |
| [026](decisions/026-connector-timeout-ceiling.md) | Connector timeout ceiling | 30s ceiling on connector, per-site timeout via tokio::time::timeout |
| [027](decisions/027-admin-socket-resource-limits.md) | Admin socket resource limits | 5s read timeout, 4096 byte line length limit |
| [027](decisions/027-admin-socket-resource-limits.md) | ~~Admin socket resource limits~~ | ~~5s read timeout, 4096 byte line length limit~~ (Deprecated — socket removed by ADR-028) |
| [028](decisions/028-admin-http-api.md) | Authenticated HTTP admin API | Bearer token auth on health check port; replaces Unix domain socket |
| [029](decisions/029-config-reload-toctou.md) | Config file TOCTOU mitigation | mtime check before and after config read; reject reload if file changed |
| [030](decisions/030-wildcard-flag-consistency.md) | Store cli_allow_wildcard_bind in ConfigReloadHandle | Consistent validation between startup and reload |
## Open Questions
@@ -231,6 +235,7 @@ questions affecting this document have been resolved:
- ~~**OQ-03**: Should the health check endpoint be on a separate port?~~ (resolved — ADR-013)
- ~~**OQ-05**: Should the proxy bind to multiple addresses?~~ (resolved — single `bind_addr` per listener)
- ~~**OQ-07**: Should per-site TLS overrides be supported for mixed ACME/manual domains?~~ (resolved — ADR-019: `[[listeners]]` with per-listener TLS config)
- ~~**OQ-08**: Should `/health` use a less common path?~~ (resolved — ADR-022: no `/health` route on main listener; health check is port 9900/admin socket only)
- ~~**OQ-08**: Should `/health` use a less common path?~~ (resolved — ADR-022: no `/health` route on main listener; health check is port 9900/admin HTTP only)
- **OQ-13**: Should `acme_contact` support multiple email addresses? (see [open-questions.md](open-questions.md))
- **OQ-14**: Should rate limiter eviction interval and max age be configurable? (see [open-questions.md](open-questions.md))
- **OQ-14**: Should rate limiter eviction interval and max age be configurable? (see [open-questions.md](open-questions.md))
- **OQ-15**: Should admin key rotation persist across restarts? (see [open-questions.md](open-questions.md))
+4 -4
View File
@@ -112,9 +112,9 @@ known host are forwarded to the upstream without modification.
The proxy does **not** serve a `/health` route on the main listener. Health
checking is an operational concern handled by the dedicated local health check
port (default: 9900, bound to `127.0.0.1` only) and the admin socket's `status`
command — not by intercepting traffic on the public-facing proxy. See ADR-013
and ADR-022.
port (default: 9900, bound to `127.0.0.1` only) and the admin HTTP endpoint's
`/admin/status` (with Bearer token) — not by intercepting traffic on the
public-facing proxy. See ADR-013, ADR-022, and ADR-028.
### 2. Rate Limiter IP Source
@@ -340,7 +340,7 @@ questions affecting this document:
ADR-015: per-site timeout overrides with defaults)
- ~~**OQ-08**: Should the `/health` path use a less common endpoint to avoid
upstream collision?~~ (resolved — ADR-022: no `/health` route on the main
listener; health checking is via port 9900 and admin socket only)
listener; health checking is via port 9900 and admin HTTP endpoint only)
- ~~**OQ-09**: How should `upstream_connect_timeout_secs` be enforced?~~
(resolved — ADR-026: 30s connector ceiling, per-site timeout via
`tokio::time::timeout`)
@@ -1,6 +1,6 @@
---
status: draft
last_updated: 2026-06-14
status: resolved
last_updated: 2026-06-15
reviewed_code:
- src/admin/socket.rs
- src/admin/mod.rs
@@ -20,6 +20,29 @@ based_on: docs/reviews/004-post-fix-review.md
# Admin Socket Security Review #005
## Resolution
This review's architectural recommendation (replace Unix domain socket with
authenticated HTTP admin endpoint) has been accepted and implemented as
ADR-028. Each finding is resolved as follows:
| Finding | Resolution |
|---------|------------|
| C1 (symlink race) | **Resolved by ADR-028** — no socket file management at all |
| C2 (no authentication) | **Resolved by ADR-028** — Bearer token with constant-time comparison |
| C3 (info leak) | **Resolved by ADR-028** — generic error messages, details logged server-side |
| W1 (no conn limit) | **Resolved by ADR-028** — axum/TCP backlog handles this naturally |
| W2 (config TOCTOU) | **Tracked separately** — ADR-029, task `fix/config-reload-toctou` |
| W3 (path validation) | **Resolved by ADR-028** — no socket path to validate; `admin_key_path` validation added (config.md rule 20) |
| W4 (is_socket_active side effect) | **Resolved by ADR-028** — no stale socket detection needed |
| W5 (wildcard flag) | **Tracked separately** — ADR-030, task `fix/wildcard-flag-reload` |
| W6 (changed_fields in response) | **Will be addressed** — `fix/admin-http-api` task includes `changed_fields` in `/admin/reload` response |
| W7 (health check port recon) | **Accepted risk** — localhost-only, minimal information. Admin endpoints add authentication |
| S1–S6 (suggestions) | **Resolved by ADR-028** — all suggestions relate to the socket, which is removed |
Implementation tasks: `fix/admin-http-api`, `fix/config-reload-toctou`,
`fix/wildcard-flag-reload`.
## Purpose
Focused security review of `src/admin/socket.rs` and related code paths,
@@ -661,7 +684,7 @@ Replace the Unix domain socket admin API with authenticated HTTP endpoints on
the existing health check listener:
```
GET http://127.0.0.1:9900/admin/reload → triggers config reload
POST http://127.0.0.1:9900/admin/reload → triggers config reload
GET http://127.0.0.1:9900/admin/status → returns uptime + site count
```
@@ -745,7 +768,7 @@ ADMIN_KEY=your-secure-random-token-here
**Added:**
- `src/admin/auth.rs` — Bearer token middleware with `subtle::ConstantTimeEq`
- `src/admin/handler.rs` — HTTP handlers for `/admin/reload` and `/admin/status`
- `src/admin/handler.rs` — HTTP handlers for `/admin/reload` (POST), `/admin/status` (GET), `/admin/rotate-key` (POST)
- `admin_key` config field (or env var)
- Admin route registration on the health check listener
@@ -770,7 +793,7 @@ echo "reload" | socat - UNIX-CONNECT:/run/reverse-proxy/admin.sock
echo "status" | socat - UNIX-CONNECT:/run/reverse-proxy/admin.sock
# After (HTTP with Bearer token)
curl -H "Authorization: Bearer $ADMIN_KEY" http://127.0.0.1:9900/admin/reload
curl -X POST -H "Authorization: Bearer $ADMIN_KEY" http://127.0.0.1:9900/admin/reload
curl -H "Authorization: Bearer $ADMIN_KEY" http://127.0.0.1:9900/admin/status
```
+184
View File
@@ -0,0 +1,184 @@
---
id: fix/admin-http-api
name: Replace Unix domain socket admin API with authenticated HTTP admin API (ADR-028)
status: open
depends_on: []
scope: broad
risk: high
impact: component
level: implementation
review_findings: [C1, C2, C3, W1, W3, W4, S1, S2, S3, S4, S5, S6]
adr: [028]
---
## Description
Replace the Unix domain socket admin API (`src/admin/socket.rs`) with
authenticated HTTP endpoints on the existing health check listener. This
eliminates the entire class of filesystem-based vulnerabilities identified in
security review #005 (C1 symlink race, C2 no authentication, C3 info leak, W1
no concurrency limit, W3 path validation, W4 is_socket_active side effect, and
S1–S6 suggestions).
ADR-028 defines the replacement design. The health check listener on
`127.0.0.1:9900` already runs an axum router. Admin endpoints are added behind
Bearer token authentication middleware.
### Changes Required
**Remove:**
- `src/admin/socket.rs` — entire file (826 lines of Unix socket code)
- `src/admin/mod.rs` — current re-exports (`AdminSocket`, `AdminSocketError`,
`start_admin_socket`)
**Add:**
- `src/admin/auth.rs` — Bearer token middleware:
- `AdminAuthConfig` struct holding `Option<String>` for the SHA-256 hash of
the admin key (or `None` to disable admin endpoints)
- `admin_auth_middleware` axum middleware that validates `Authorization:
Bearer <token>` against the stored hash using `subtle::ConstantTimeEq`
- Returns 404 when admin is disabled (empty `admin_key_path`), 401 on
missing/wrong token, passes through on valid token
- `load_admin_key(path: &str) -> Result<Option<[u8; 32]>, AdminKeyError>`
function that reads the key file, hashes it with SHA-256, and returns
the hash. Returns `None` if path is empty (disabled). Logs a warning
and returns `None` if the file doesn't exist or is unreadable (admin
endpoints disabled, process continues starting).
- `src/admin/handler.rs` — HTTP handlers:
- `reload_handler(State<...>) -> Json<ReloadResponse>` — **POST** `/admin/reload`.
Triggers `ConfigReloadHandle::reload()`, returns `{"status": "ok"}` or
`{"status": "error", "message": "reload failed"}`. Generic error
messages only; details logged server-side.
- `status_handler(State<...>) -> Json<StatusResponse>` — **GET**
`/admin/status`. Returns
`{"status": "ok", "uptime_secs": N, "sites": N}`
- `rotate_key_handler(State<...>) -> Json<RotateKeyResponse>` — **POST**
`/admin/rotate-key`. Generates new 256-bit random key, returns plaintext
in response, replaces stored hash in memory. Returns
`{"status": "ok", "key": "<hex>"}`.
**Modify:**
- `src/admin/mod.rs` — re-export `AdminAuthConfig`, `AdminKeyError`,
`admin_auth_middleware`, `load_admin_key`, and the handler functions
- `src/health.rs` — expand `health_router()` to `admin_router()` that nests
admin routes under `/admin` with auth middleware. Merge into the health
check listener. The full router becomes:
```
/health → health_handler (GET, no auth)
/admin/* → auth middleware → admin handlers (POST for state-changing, GET for read-only)
```
The `start_health_check_listener` function signature changes to accept
`Option<Arc<AdminAuthConfig>>` and `Arc<ConfigReloadHandle>` and
`Arc<ArcSwap<[u8; 32]>>` for key rotation. If `AdminAuthConfig` is `None`,
`/admin/*` routes return 404.
- `src/main.rs` — remove admin socket initialization entirely (lines 102-127).
Add admin key loading step after config parsing:
```rust
let admin_auth = if !static_config.admin_key_path.is_empty() {
match admin::load_admin_key(&static_config.admin_key_path) {
Ok(Some(hash)) => Some(Arc::new(AdminAuthConfig { admin_key_hash: hash })),
Ok(None) => None, // disabled
Err(e) => {
warn!("admin key load failed, disabling admin endpoints: {}", e);
None
}
}
} else {
None
};
```
Pass `admin_auth`, `reload_handle`, and `start_time` to
`start_health_check_listener`.
- `src/config/static_config.rs` — replace `admin_socket_path` field with
`admin_key_path`:
```rust
#[serde(default = "default_admin_key_path")]
pub admin_key_path: String,
```
Default: `"/etc/reverse-proxy/admin-key"`. Empty string disables admin
endpoints.
- `src/config/dynamic_config.rs` — `ConfigReloadHandle` gains
`cli_allow_wildcard_bind: bool` field (see task `fix/wildcard-flag-reload`).
No other changes needed — `reload()` method stays the same.
- `src/config/validation.rs` — add validation that `admin_key_path` is empty
or an absolute path (no `..` traversal, no relative paths). This is a new
validation rule.
- `Cargo.toml` — add `subtle` and `sha2` dependencies (already in overview.md)
**Tests:**
- Replace all `src/admin/socket.rs` tests with HTTP-based tests using
`reqwest` (already a dev dependency). Test:
- POST `/admin/reload` with valid Bearer token returns `{"status": "ok"}`
- POST `/admin/reload` with wrong token returns 401
- POST `/admin/reload` with no token returns 401
- POST `/admin/reload` when admin disabled returns 404
- GET `/admin/status` with valid token returns uptime and site count
- POST `/admin/rotate-key` with valid token returns new key and updates stored
hash
- POST `/admin/rotate-key` subsequent requests use the new key (old key returns
401)
- GET `/health` always returns 200 regardless of auth state
**Deployment:**
- `deploy/docker-compose.yml` — remove `/run/reverse-proxy` socket volume,
add `/etc/reverse-proxy/admin-key:/etc/reverse-proxy/admin-key:ro` volume
- `deploy/reverse-proxy.service` — remove any socket directory setup
- `deploy/README.md` — replace `socat` commands with `curl` examples
## Acceptance Criteria
- [ ] `src/admin/socket.rs` is deleted entirely
- [ ] `src/admin/auth.rs` implements Bearer token auth with constant-time
comparison and SHA-256 hashing
- [ ] `src/admin/handler.rs` implements `/admin/reload` (POST),
`/admin/status` (GET), `/admin/rotate-key` (POST)
- [ ] `src/health.rs` serves both `/health` (no auth) and `/admin/*`
(auth required) on port 9900
- [ ] `src/config/static_config.rs` uses `admin_key_path` (not
`admin_socket_path`)
- [ ] `src/main.rs` loads admin key at startup, passes auth config to
health check listener
- [ ] Admin disabled (`admin_key_path` empty or file missing) → `/admin/*`
returns 404
- [ ] Wrong/missing Bearer token → 401
- [ ] Error responses are generic (no filesystem paths, no config details)
- [ ] Full error details logged server-side only
- [ ] Key rotation works in-memory (new key replaces stored hash, old key
rejected)
- [ ] Key rotation does not persist across restarts (documented behavior)
- [ ] SIGHUP reload continues to work unchanged
- [ ] All existing tests pass (minus deleted socket tests)
- [ ] New HTTP-based admin tests pass
- [ ] `cargo clippy` passes with no warnings
- [ ] Deployment files updated (docker-compose, systemd, README)
## References
- docs/architecture/decisions/028-admin-http-api.md — ADR-028
- docs/architecture/decisions/014-unix-socket-reload.md — superseded ADR
- docs/architecture/decisions/027-admin-socket-resource-limits.md — deprecated
- docs/architecture/operations.md — admin HTTP endpoint, key management
- docs/architecture/config.md — admin_key_path, StaticConfig
- docs/architecture/overview.md — crate dependencies, architecture diagram
- docs/reviews/005-admin-socket-security-review.md — C1, C2, C3, W1, W3, W4
- src/admin/socket.rs — code to remove
- src/health.rs — code to extend
- src/main.rs — admin socket init to remove/replace
- src/config/static_config.rs — field rename
## Notes
> This is the primary implementation task for the admin socket → HTTP API
> migration. It directly implements ADR-028 and resolves findings C1, C2, C3,
> W1, W3, W4, S1–S6 from security review #005.
>
> W2 (config TOCTOU) and W5 (wildcard flag) are independent fixes tracked in
> separate tasks.
>
> The `subtle` and `sha2` crates are already listed in the architecture spec
> (overview.md crate dependencies). Add them to `Cargo.toml` with appropriate
> versions.
## Summary
> To be filled on completion
+80
View File
@@ -0,0 +1,80 @@
---
id: fix/agents-md-project-structure
name: Update AGENTS.md project structure and common modifications after admin refactor
status: open
depends_on: [fix/admin-http-api]
scope: narrow
risk: low
impact: docs
level: documentation
review_findings: []
adr: [028]
---
## Description
After the admin socket → HTTP API migration, `AGENTS.md` needs updates to
reflect the new project structure, config format, and operational procedures.
### Changes Required
**Project Structure section** — Update to reflect new admin module layout:
```
src/
├── admin/
│ ├── auth.rs # Bearer token auth middleware (subtle, SHA-256)
│ ├── handler.rs # HTTP handlers for /admin/reload, /status, /rotate-key
│ └── mod.rs # Re-exports
```
Remove:
```
│ ├── socket.rs # REMOVED — was Unix domain socket admin API
```
**Key Architecture Concepts section** — Update the admin socket description:
- Replace "Unix domain socket (`admin_socket_path`)" with "Authenticated HTTP
admin API (`admin_key_path`) on health check port"
- Note that admin endpoints require Bearer token auth
- Note that `admin_key_path` empty string = disabled (returns 404)
**Config Format section** — Update:
- Replace `admin_socket_path` references with `admin_key_path`
- Note that `admin_key_path` default is `/etc/reverse-proxy/admin-key`
- Add key file format info (plaintext, one line, read once at startup)
**Common Modifications section** — Replace:
```bash
# Before (Unix socket)
echo "reload" | socat - UNIX-CONNECT:/run/reverse-proxy/admin.sock
# After (HTTP with Bearer token)
curl -H "Authorization: Bearer $ADMIN_KEY" http://127.0.0.1:9900/admin/reload
curl -H "Authorization: Bearer $ADMIN_KEY" http://127.0.0.1:9900/admin/status
```
**Build & Run section** — No changes needed (build commands unchanged).
**Testing section** — Note that admin tests now use HTTP (reqwest) instead of
Unix socket (tokio::net::UnixStream).
## Acceptance Criteria
- [ ] Project structure shows `auth.rs` and `handler.rs`, not `socket.rs`
- [ ] Key architecture concepts mention `admin_key_path` and Bearer token auth
- [ ] Config format section mentions `admin_key_path`
- [ ] Common modifications section uses `curl` examples, not `socat`
- [ ] No references to `admin_socket_path` remain in AGENTS.md
## References
- AGENTS.md — current project structure and common modifications
- docs/architecture/decisions/028-admin-http-api.md — ADR-028
## Notes
> Depends on `fix/admin-http-api` being complete so the new file names are
> accurate.
## Summary
> To be filled on completion
+100
View File
@@ -0,0 +1,100 @@
---
id: fix/config-reload-toctou
name: Add mtime check to config reload to detect mid-write file changes (ADR-029)
status: open
depends_on: []
scope: narrow
risk: low
impact: component
level: implementation
review_findings: [W2]
adr: [029]
---
## Description
Both the SIGHUP reload path (`src/shutdown.rs:handle_sighup_reload`) and the
admin HTTP reload path (`src/admin/socket.rs:handle_reload`, soon
`src/admin/handler.rs`) read the config file from disk with
`tokio::fs::read_to_string()`, then parse and apply it. If another process is
writing to the config file at the same time, the proxy could read a partially
written config.
ADR-029 specifies a simple mitigation: compare file modification timestamps
before and after reading. If mtime changed, reject the reload and return a
"please retry" message.
### Changes Required
**Shared reload function** — Extract the common file-read-and-validate logic
from `src/shutdown.rs:handle_sighup_reload()` and
`src/admin/socket.rs:handle_reload()` into a shared function (e.g.,
`src/config/dynamic_config.rs` or a new `src/config/reload.rs`):
```rust
pub async fn read_and_validate_config(
config_path: &str,
cli_allow_wildcard_bind: bool,
) -> Result<(StaticConfig, DynamicConfig), ReloadError> {
let metadata_before = tokio::fs::metadata(config_path).await
.map_err(ReloadError::Io)?;
let config_content = tokio::fs::read_to_string(config_path).await
.map_err(ReloadError::Io)?;
let metadata_after = tokio::fs::metadata(config_path).await
.map_err(ReloadError::Io)?;
if metadata_before.modified().ok() != metadata_after.modified().ok() {
return Err(ReloadError::FileChangedDuringRead);
}
let full_config = FullConfig::parse(&config_content)?;
let (new_static, new_dynamic) = full_config.into_static_and_dynamic();
validate(&new_static, &new_dynamic, cli_allow_wildcard_bind)?;
Ok((new_static, new_dynamic))
}
```
**`src/shutdown.rs`** — Replace inline file read + parse + validate with a
call to `read_and_validate_config()`. On `ReloadError::FileChangedDuringRead`,
log a warning: "config file changed during read, please retry SIGHUP".
**`src/admin/handler.rs`** (after admin-http-api task) — Same call. On
`ReloadError::FileChangedDuringRead`, return
`{"status": "error", "message": "config file changed during read, please retry"}`.
**Error type** — Define `ReloadError` enum with variants:
- `Io(std::io::Error)`
- `Parse(toml::de::Error)`
- `Validation(String)`
- `FileChangedDuringRead`
## Acceptance Criteria
- [ ] Both SIGHUP and admin HTTP reload paths use the same file-reading logic
- [ ] mtime is checked before and after reading the config file
- [ ] If mtime changed, reload is rejected with a clear error message
- [ ] Error message in admin HTTP response is generic ("config file changed
during read, please retry") — no filesystem paths leaked
- [ ] Full error details are logged server-side (path, mtime values)
- [ ] SIGHUP path logs the same error at warn level
- [ ] `cargo test` passes
- [ ] `cargo clippy` passes with no warnings
## References
- docs/architecture/decisions/029-config-reload-toctou.md — ADR-029
- docs/reviews/005-admin-socket-security-review.md — W2 finding
- src/shutdown.rs — handle_sighup_reload
- src/admin/socket.rs — handle_reload (to be replaced by admin/handler.rs)
## Notes
> This fix is independent of the admin socket → HTTP migration. It applies to
> both reload paths (SIGHUP and admin). The implementation should be done
> after or alongside the admin-http-api task since that task replaces
> socket.rs with handler.rs.
## Summary
> To be filled on completion
+71
View File
@@ -0,0 +1,71 @@
---
id: fix/review-005-status-update
name: Update security review #005 status to reflect ADR-028 decision
status: open
depends_on: []
scope: narrow
risk: low
impact: docs
level: documentation
review_findings: [C1, C2, C3, W1, W3, W4, S1, S2, S3, S4, S5, S6]
adr: [028]
---
## Description
Security review #005 (`docs/reviews/005-admin-socket-security-review.md`) is
currently marked as `status: draft`. The review's architectural recommendation
to replace the Unix domain socket with an authenticated HTTP admin endpoint has
been accepted as ADR-028. The review findings should be annotated with their
resolution status.
### Changes Required
**`docs/reviews/005-admin-socket-security-review.md`**:
- Update frontmatter `status` from `draft` to the appropriate post-decision
status (e.g., `accepted` or `resolved`)
- Add a resolution section at the top of the document noting:
- C1, C2, C3, W1, W3, W4, S1–S6: **Resolved by ADR-028** (replacing Unix
domain socket with authenticated HTTP admin API)
- W2 (config file TOCTOU): **Tracked separately** — ADR-029, task
`fix/config-reload-toctou`
- W5 (wildcard flag inconsistency): **Tracked separately** — ADR-030, task
`fix/wildcard-flag-reload`
- W6 (changed_fields in reload response): **Tracked** — will be implemented
as part of `fix/admin-http-api` (the new `/admin/reload` endpoint will
include changed_fields in its response per operations.md)
- W7 (health check port recon): **Accepted risk** — health check is
localhost-only, returns minimal information. The admin HTTP endpoint adds
authentication for `/admin/*` routes.
**`docs/reviews/006-attack-surface-review.md`**:
- Update Category 5 (Admin Socket) references from `src/admin/socket.rs` to
`src/admin/auth.rs` and `src/admin/handler.rs` (after admin-http-api task
is complete)
- Update entry 4.3 (admin reload config file) to reference the shared
`read_and_validate_config()` function with mtime check
- Remove or update entries that are eliminated by the socket removal (e.g.,
Category 4: Unix Domain Socket entries)
## Acceptance Criteria
- [ ] Review #005 frontmatter status updated
- [ ] Review #005 has a resolution section annotating each finding with its
disposition (resolved by ADR-028, tracked separately, accepted risk)
- [ ] Review #006 admin socket references updated (after admin-http-api task)
- [ ] No inline content removed — findings are annotated, not deleted
## References
- docs/reviews/005-admin-socket-security-review.md
- docs/reviews/006-attack-surface-review.md
- docs/architecture/decisions/028-admin-http-api.md
## Notes
> This task should be done after the `fix/admin-http-api` task is complete,
> since review #006 references need to point to the new file structure.
## Summary
> To be filled on completion
+112
View File
@@ -0,0 +1,112 @@
---
id: fix/wildcard-flag-reload
name: Store cli_allow_wildcard_bind in ConfigReloadHandle for consistent reload validation (ADR-030)
status: open
depends_on: []
scope: narrow
risk: low
impact: component
level: implementation
review_findings: [W5]
adr: [030]
---
## Description
When the proxy starts with `--allow-wildcard-bind` (or `allow_wildcard_bind =
true` in config), bind addresses using `0.0.0.0` are accepted. But on config
reload, `validate()` is called with `cli_allow_wildcard_bind: false` — a
hardcoded value in `ConfigReloadHandle::reload()`. This means a config that was
valid at startup will be rejected on reload because the flag that enabled
wildcard binding is not preserved.
ADR-030 specifies storing `cli_allow_wildcard_bind` in `ConfigReloadHandle` at
construction time and using the stored value during reload validation.
### Changes Required
**`src/config/dynamic_config.rs`** — `ConfigReloadHandle` struct:
- Add `cli_allow_wildcard_bind: bool` field
- Update `ConfigReloadHandle::new()` to accept and store the flag:
```rust
pub fn new(
config: Arc<ArcSwap<DynamicConfig>>,
static_config: StaticConfig,
cli_allow_wildcard_bind: bool,
) -> Self {
Self {
config,
static_config: ArcSwap::from_pointee(static_config),
reload_mutex: Mutex::new(()),
cli_allow_wildcard_bind,
}
}
```
- In `reload()`, pass `self.cli_allow_wildcard_bind` to `validate()` instead
of `false`:
```rust
validate(&new_static, &new_dynamic, self.cli_allow_wildcard_bind)?;
```
**`src/main.rs`** — Update `ConfigReloadHandle::new()` call to pass
`cli_allow_wildcard_bind` from the loaded config:
```rust
let reload_handle = Arc::new(ConfigReloadHandle::new(
config_arc.clone(),
loaded_config.static_config.clone(),
loaded_config.cli_allow_wildcard_bind, // or args.allow_wildcard_bind
));
```
The `cli_allow_wildcard_bind` value should be the OR of the config flag and
the CLI flag, matching the startup validation logic. Check `src/cli.rs` for
how the flag is currently handled.
**`src/admin/socket.rs`** (or `src/admin/handler.rs` after migration) — Same
change: pass the flag through to `ConfigReloadHandle::new()`.
**`src/config/validation.rs`** — No changes needed; `validate()` already
accepts `cli_allow_wildcard_bind: bool` and uses it correctly.
**Tests** — Update all `ConfigReloadHandle::new()` calls to include the new
parameter. Add a test that verifies:
1. A config with `0.0.0.0` bind address is accepted on reload when
`cli_allow_wildcard_bind: true`
2. A config with `0.0.0.0` bind address is rejected on reload when
`cli_allow_wildcard_bind: false`
## Acceptance Criteria
- [ ] `ConfigReloadHandle` has a `cli_allow_wildcard_bind: bool` field
- [ ] `ConfigReloadHandle::new()` accepts and stores `cli_allow_wildcard_bind`
- [ ] `reload()` passes `self.cli_allow_wildcard_bind` to `validate()`
(not hardcoded `false`)
- [ ] All `ConfigReloadHandle::new()` call sites pass the correct flag
- [ ] Config with `0.0.0.0` bind address is accepted on reload when flag is
true (test)
- [ ] Config with `0.0.0.0` bind address is rejected on reload when flag is
false (test)
- [ ] `cargo test` passes
- [ ] `cargo clippy` passes with no warnings
## References
- docs/architecture/decisions/030-wildcard-flag-consistency.md — ADR-030
- docs/reviews/005-admin-socket-security-review.md — W5 finding
- docs/architecture/config.md — validation rules, allow_wildcard_bind
- src/config/dynamic_config.rs — ConfigReloadHandle
- src/config/validation.rs — validate()
- src/cli.rs — CLI flag handling
## Notes
> This fix is independent of the admin socket → HTTP migration. It should be
> applied to `ConfigReloadHandle` regardless of which admin interface is used.
> The implementation is straightforward: add a field, pass it through.
>
> The flag value should be `allow_wildcard_bind || cli_allow_wildcard_bind`
> (OR logic) matching the startup behavior documented in config.md.
## Summary
> To be filled on completion