Add ADR-029/030, implementation tasks, and spec updates for admin socket removal

Security review #005 identified critical vulnerabilities in the Unix domain
socket admin API (C1 symlink race, C2 no auth, C3 info leak, W1-W7, S1-S6).
ADR-028 (already accepted) replaces the socket with an authenticated HTTP
admin API on the health check port. This commit adds the remaining spec work:

- ADR-029: Config file TOCTOU mitigation (mtime check on reload)
- ADR-030: Store cli_allow_wildcard_bind in ConfigReloadHandle for consistent
  reload validation
- Implementation tasks for the admin HTTP migration (fix/admin-http-api),
  TOCTOU fix (fix/config-reload-toctou), and wildcard flag fix
  (fix/wildcard-flag-reload)
- Updated review #005 status to resolved with per-finding disposition
- Resolved OQ-16: POST for state-changing admin endpoints, GET for read-only
- Updated all architecture docs to reference new ADRs, use admin_key_path
  instead of admin_socket_path, and reflect POST method for /admin/reload
This commit is contained in:
glm-5.1 committed 2026-06-15 05:19:42 +00:00
1 parent 9096ec5873
commit 161049a17d
21 files changed
+1210 -119

No files matched your search

@@ -48,8 +48,8 @@ listener (see ADR-022).
- Configurable port allows different deployment scenarios (some monitoring runs
on different ports)
- Disabling via `health_check_port = 0` removes the health check entirely —
the admin socket's `status` command remains available as an alternative
health/status mechanism
the admin HTTP endpoint's `/admin/status` (with Bearer token) remains available
as an alternative health/status mechanism (ADR-028)
- When this project is folded into alknet, the health check will use alknet's
existing patterns, making the separate port unnecessary in that context
@@ -69,4 +69,5 @@ listener (see ADR-022).
- [operations.md](../operations.md)
- [ADR-022](022-health-check-scope.md) — Health check scope (no `/health` on main listener)
- [ADR-028](028-admin-http-api.md) — Authenticated HTTP admin API (admin endpoints on health check port)
- OQ-03 (now resolved)
@@ -2,7 +2,7 @@
## Status
Accepted
Superseded by [ADR-028](028-admin-http-api.md)
## Context
@@ -99,7 +99,7 @@ Example configuration:
```toml
# Global settings
health_check_port = 9900
admin_socket_path = "/run/reverse-proxy/admin.sock"
admin_key_path = "/etc/reverse-proxy/admin-key"
[logging]
level = "info"
@@ -56,10 +56,11 @@ to consume directly from the host filesystem.
and `journalctl`). File logging is the authoritative source for fail2ban
because it avoids the fragility of Docker log driver parsing.
5. **ACME state and admin socket are volume-mounted.** The ACME cache directory
(`/var/lib/reverse-proxy/acme-cache/`) and admin socket
(`/run/reverse-proxy/admin.sock`) are mounted as volumes so state persists
across container restarts and the host can send reload commands.
5. **ACME state and admin key are volume-mounted.** The ACME cache directory
(`/var/lib/reverse-proxy/acme-cache/`) and admin key file
(`/etc/reverse-proxy/admin-key`) are mounted as volumes so state persists
across container restarts and the host can send authenticated reload commands
(ADR-028).
6. **Health checks use Docker's native mechanism.** The health check endpoint
on port 9900 (localhost only) is used directly by Docker's `HEALTHCHECK`
@@ -1,4 +1,4 @@
# ADR-022: Health Check Scope — Local Port and Admin Socket Only
# ADR-022: Health Check Scope — Local Port and Admin HTTP Endpoint Only
## Status
@@ -25,8 +25,8 @@ handled exclusively by:
1. **Local health check port** (default: 9900, bound to `127.0.0.1`) — serves
`GET /health → 200 OK`. This is the primary health check mechanism for
container orchestration, load balancers, and monitoring systems.
2. **Admin socket** (`status` command) — returns process information including
uptime and site count.
2. **Admin HTTP endpoint** (`GET /admin/status` with Bearer token) — returns
process information including uptime and site count. See ADR-028.
The `/health` route is removed from the main listener entirely. No configurable
path is needed because the route simply does not exist on the public listener.
@@ -52,5 +52,6 @@ path is needed because the route simply does not exist on the public listener.
## References
- ADR-013: Health check on separate local port
- ADR-028: Authenticated HTTP admin API (admin socket replaced by HTTP endpoint)
- OQ-08: Resolved by this ADR
- Implementation review finding W5 (hardcoded `/health` path)
@@ -2,7 +2,9 @@
## Status
Accepted
Deprecated — the Unix domain socket admin API has been replaced by an
authenticated HTTP admin endpoint (ADR-028). Socket resource limits are no
longer needed.
## Context
@@ -0,0 +1,228 @@
# ADR-028: Authenticated HTTP Admin API (Replacing Unix Domain Socket)
## Status
Accepted
## Context
The proxy has a Unix domain socket admin API (ADR-014) that provides two
commands: `reload` (trigger config reload with success/failure feedback) and
`status` (return uptime and site count). Security review #005 identified three
critical and seven warning-level vulnerabilities in the socket implementation:
- **C1**: Symlink race in stale socket cleanup enables arbitrary file deletion
- **C2**: No authentication — any local user can trigger config reload
- **C3**: Error responses leak filesystem paths and config structure details
- **W1**: No connection concurrency limit
- **W3**: Socket path not validated or sanitized
- **W4**: `is_socket_active` side-effect on other processes
- **W5**: Reload validation uses different `cli_allow_wildcard_bind` flag than
startup
These vulnerabilities stem from the fundamental design choice of using a Unix
domain socket. The socket introduces an entire class of filesystem-based attack
surface that does not exist with an HTTP endpoint: symlink races, stale socket
cleanup, path traversal, permission management, and directory mount issues in
containers.
Additionally, the socket requires `socat` for interaction — a non-standard tool
that must be installed separately, complicating container images and CI/CD
pipelines.
The proxy already has a localhost-only HTTP listener (`src/health.rs`) bound to
`127.0.0.1:9900` that serves `/health`. This listener is axum-based, supports
middleware layers, and has integration tests. Co-locating admin endpoints on
this listener is the natural replacement.
## Decision
Replace the Unix domain socket admin API with authenticated HTTP endpoints on
the existing health check listener. Authentication uses a Bearer token verified
against a SHA-256 hash stored in memory.
### Admin Key Management
The admin key is stored in a file on disk (specified by `admin_key_path` in
StaticConfig). The proxy reads this file once at startup, hashes its contents
with SHA-256, and stores only the hash in memory. The plaintext key is never
held in memory after startup initialization.
Key file setup:
```bash
openssl rand -hex 32 > /etc/reverse-proxy/admin-key
chmod 600 /etc/reverse-proxy/admin-key
```
Setting `admin_key_path` to an empty string disables admin endpoints entirely.
### Authentication
Admin endpoints require a Bearer token in the `Authorization` header:
```
Authorization: Bearer <key>
```
The provided token is SHA-256 hashed and compared against the stored hash using
constant-time comparison (`subtle::ConstantTimeEq`) to prevent timing attacks.
Error behavior by auth state:
| Scenario | Response |
|----------|----------|
| Admin disabled (`admin_key_path` empty) | 404 (endpoint does not exist) |
| Missing `Authorization` header | 401 |
| Wrong token | 401 |
| Correct token | Proceed to handler |
Returning 404 when admin is disabled prevents discovery of the endpoint's
existence. Returning 401 for wrong tokens (rather than 404) allows operators to
confirm the endpoint is available without revealing information to attackers
who lack any valid token.
### Endpoints
| Method | Path | Auth | Description |
|--------|------|------|-------------|
| GET | `/health` | None | Health check (unchanged) |
| POST | `/admin/reload` | Bearer token | Trigger config reload |
| GET | `/admin/status` | Bearer token | Return uptime and site count |
| POST | `/admin/rotate-key` | Bearer token | Generate and return a new random admin key |
**`/admin/reload`** — Triggers the same config reload as SIGHUP. Returns
structured JSON:
```json
{"status": "ok"}
```
On error:
```json
{"status": "error", "message": "reload failed"}
```
Error messages are generic — no filesystem paths, no config structure details.
Full error information is logged server-side only.
**`/admin/status`** — Returns process information:
```json
{"status": "ok", "uptime_secs": 1234, "sites": 2}
```
**`/admin/rotate-key`** — Generates a new 256-bit random key using
`rand::RngCore`, returns it in the response, and replaces the stored hash in
memory with the SHA-256 hash of the new key:
```json
{"status": "ok", "key": "<new-plaintext-key-hex>"}
```
The operator should capture this key and update the key file on disk for
subsequent restarts. In-memory rotation does **not** persist across restarts —
on restart, the proxy re-reads the key file. This is by design: the file on
disk is the source of truth for the admin key, and runtime rotation is a
temporary override.
If an attacker has the current admin key, they could rotate it to lock out the
legitimate operator. But if an attacker has the admin key, they can already
trigger config reloads — the most dangerous operation. Rotation is strictly
less damaging than what they could already do.
### Config Change
Replace `admin_socket_path` (StaticConfig) with `admin_key_path` (StaticConfig):
```toml
# Before (ADR-014)
admin_socket_path = "/run/reverse-proxy/admin.sock" # empty = disabled
# After (ADR-028)
admin_key_path = "/etc/reverse-proxy/admin-key" # empty = disabled
```
Default: `"/etc/reverse-proxy/admin-key"`.
### Comparison with Previous Design
| Aspect | Unix Socket (ADR-014) | HTTP Admin (ADR-028) |
|--------|-----------------------|----------------------|
| Authentication | None (filesystem permissions only) | Bearer token with constant-time comparison |
| Attack surface | Filesystem: symlinks, stale cleanup, path traversal | File: read once at startup, no management |
| Client tool | `socat` (non-standard) | `curl` (universal) |
| Error leakage | Paths and config details in responses | Generic messages, details logged server-side |
| Container setup | Volume mount for socket directory | Volume mount for key file (single file, `:ro`) |
| Feedback | Structured JSON responses | Structured JSON responses (same) |
| SIGHUP fallback | Yes (both work) | Yes (both work) |
### What This Eliminates from Review #005
| Finding | Eliminated? | Reason |
|---------|------------|--------|
| C1 (symlink race) | Yes | No socket file management at all |
| C2 (no authentication) | Yes | Bearer token with constant-time comparison |
| C3 (info leak) | Yes | Generic error messages, no paths |
| W1 (no conn limit) | Yes | axum/TCP backlog handles this naturally |
| W3 (path validation) | Yes | No socket path to validate; key file path is read-only, no creation/cleanup |
| W4 (is_socket_active) | Yes | No stale socket detection needed |
| W5 (wildcard flag) | No | Still exists (separate fix) |
| W2 (config TOCTOU) | No | Still exists (separate fix) |
### Remaining Findings
W2 (config file TOCTOU on reload) and W5 (reload validation uses different
`cli_allow_wildcard_bind` flag) still apply to both the SIGHUP and HTTP admin
reload paths. These are independent of the admin interface choice and require
separate fixes.
## Rationale
- **Eliminates an attack surface class**: Every critical finding in review #005
stems from the socket being a filesystem object. Removing the socket removes
the class.
- **Read-once semantics**: The proxy reads the key file once at startup and
never manages it — no creation, no cleanup, no stale detection. This is
fundamentally different from the socket, which required bind, listen, accept,
cleanup-on-startup, cleanup-on-shutdown, and stale detection.
- **Standard tooling**: `curl` is available everywhere. `socat` requires
separate installation in container images and CI environments.
- **Authentication**: Bearer tokens are the standard pattern for HTTP APIs.
Constant-time comparison prevents timing attacks. SHA-256 hashing means the
plaintext key is never held in memory after startup.
- **Key file is lower-risk than socket**: Reading a file is a single syscall.
The socket required managing a filesystem object across the entire process
lifecycle. If an attacker can read the key file, they can also read the
config file — the key file does not expand the trust boundary.
## Consequences
**Positive:**
- Eliminates C1, C2, C3, W1, W3, W4 from security review #005
- Authentication for admin operations (the socket had none)
- Universal client tooling (`curl` instead of `socat`)
- Simpler container setup (single file mount vs. directory mount)
- No socket lifecycle management (startup cleanup, shutdown cleanup, stale
detection)
- Generic error responses prevent information disclosure
**Negative:**
- Key file must exist on disk for admin endpoints to work
- Key file must be readable by the proxy process
- In-memory key rotation does not persist across restarts (operator must
update the key file separately)
- Adds `subtle` and `sha2` crate dependencies
- Admin endpoints share the health check port (operational port serves both
authenticated and unauthenticated routes)
## References
- [operations.md](../operations.md)
- [config.md](../config.md)
- [overview.md](../overview.md)
- [ADR-014](014-unix-socket-reload.md) — Superseded by this ADR
- [ADR-027](027-admin-socket-resource-limits.md) — Deprecated (no longer needed)
- [Review #005](../../reviews/005-admin-socket-security-review.md)
- [Review #006](../../reviews/006-attack-surface-review.md)
@@ -0,0 +1,90 @@
# ADR-029: Config File TOCTOU Mitigation on Reload
## Status
Accepted
## Context
Both the SIGHUP reload path (`src/shutdown.rs:handle_sighup_reload`) and the
admin HTTP reload path (`src/admin/socket.rs:handle_reload`, soon
`src/admin/handler.rs`) read the config file from disk with
`tokio::fs::read_to_string()`, then parse and apply it. If another process is
writing to the config file at the same time (e.g., a configuration management
tool like Ansible writing a partial file), the proxy could read a partially
written config and either fail to parse it (resulting in a reload error) or,
in an unlikely worst case, parse a structurally valid but semantically wrong
config.
This is a filesystem-level time-of-check/time-of-use (TOCTOU) issue. The
window is small but the impact of applying a partial config is significant.
Security review #005 identified this as finding W2.
## Decision
Detect mid-write file changes by comparing file metadata before and after
reading. If the modification timestamp changes between the two `stat` calls,
reject the reload and return a retry message.
```rust
let metadata_before = tokio::fs::metadata(&config_path).await?;
let config_content = tokio::fs::read_to_string(&config_path).await?;
let metadata_after = tokio::fs::metadata(&config_path).await?;
if metadata_before.modified()? != metadata_after.modified()? {
return Err("config file changed during read, please retry");
}
```
This applies to **both** the SIGHUP reload path and the admin HTTP reload path.
For operators, the documentation will recommend the atomic replacement pattern
(write to a temp file in the same directory, then `rename()` over the target).
This is the standard safe pattern for config file rotation and is what tools
like Ansible already do with `copy` module's `validate` parameter.
## Rationale
- **Simple and effective**: The mtime check catches the common case of a
config management tool mid-write. It requires no changes to the config
file format or directory layout.
- **No false negatives**: If mtime changed, the file definitely changed. If
mtime did not change within the typical filesystem timestamp granularity
(1 second on most Linux filesystems), the window is so small that a partial
read is extremely unlikely.
- **Atomic rename is the gold standard**: Recommending it in documentation is
better than trying to enforce it in code. The proxy can't control how
operators write config files, but it can detect when a file might be
inconsistent and ask for a retry.
- **Same pattern in both reload paths**: SIGHUP and admin HTTP share the same
file-reading logic (or should — currently they duplicate it). This ADR
ensures both paths are protected.
## Consequences
**Positive:**
- Config reload will reject a file that changed during the read, preventing
partial or inconsistent configs from being applied.
- Clear error message ("config file changed during read, please retry") tells
operators exactly what happened.
- Documenting the atomic replacement pattern gives operators a clear
recommendation for safe config rotation.
**Negative:**
- In very rare cases, a legitimate config change that happens to land within
the same filesystem timestamp granularity as the read could be falsely
rejected. The operator would need to retry the reload, which is an
acceptable trade-off for safety.
- The mtime check does not protect against all TOCTOU scenarios (e.g., a
write that starts before the first `stat` and completes before the read).
However, combined with the atomic replacement recommendation, this is a
defense-in-depth measure, not a complete solution. A complete solution would
require file locking or checksum verification, which adds complexity for
marginal benefit.
## References
- [operations.md](../operations.md) — Config reload, admin HTTP endpoint
- [config.md](../config.md) — Config reload behavior
- [Review #005](../../reviews/005-admin-socket-security-review.md) — W2 finding
@@ -0,0 +1,99 @@
# ADR-030: Store cli_allow_wildcard_bind Flag in ConfigReloadHandle
## Status
Accepted
## Context
When the proxy starts, `cli_allow_wildcard_bind` can be set to `true` via
the `--allow-wildcard-bind` CLI flag or the `allow_wildcard_bind = true` config
option. The startup validation uses this flag to decide whether `0.0.0.0` bind
addresses are allowed.
However, when a config reload is triggered (via SIGHUP or admin HTTP), the
`validate()` call is invoked with `cli_allow_wildcard_bind: false` — hardcoded
in `ConfigReloadHandle::reload()`. This means that a config that was accepted
at startup (because the CLI flag was set) will be rejected on reload, even
though the running process has `allow_wildcard_bind = true` in effect.
Security review #005 identified this as finding W5. The consequence is that
an operator who started the proxy with `--allow-wildcard-bind` cannot reload
the config without getting a validation error about `0.0.0.0` bind addresses
— even though those bind addresses are currently active and working.
## Decision
Store the `cli_allow_wildcard_bind` flag in `ConfigReloadHandle` at
construction time, and use the stored value during reload validation instead
of hardcoding `false`.
```rust
pub struct ConfigReloadHandle {
config: Arc<ArcSwap<DynamicConfig>>,
static_config: ArcSwap<StaticConfig>,
reload_mutex: Mutex<()>,
cli_allow_wildcard_bind: bool,
}
impl ConfigReloadHandle {
pub fn new(
config: Arc<ArcSwap<DynamicConfig>>,
static_config: StaticConfig,
cli_allow_wildcard_bind: bool,
) -> Self {
Self {
config,
static_config: ArcSwap::from_pointee(static_config),
reload_mutex: Mutex::new(()),
cli_allow_wildcard_bind,
}
}
}
```
In `reload()`, pass `self.cli_allow_wildcard_bind` to `validate()` instead of
`false`:
```rust
validate(&new_static, &new_dynamic, self.cli_allow_wildcard_bind)?;
```
This ensures reload validation uses the same flag as startup validation. The
flag is immutable — it's set once at startup and never changed — so storing it
in `ConfigReloadHandle` is safe.
## Rationale
- **Consistency**: Startup and reload should apply the same validation rules.
If `0.0.0.0` was allowed at startup, it should be allowed on reload.
- **The flag is immutable**: `cli_allow_wildcard_bind` is set once from CLI
args and never changes. Storing it in `ConfigReloadHandle` is a simple,
correct solution.
- **No config file change needed**: The `allow_wildcard_bind` config option is
already in `StaticConfig`. The CLI flag is a separate override. The fix is
purely in how the reload path uses the flag.
- **OR logic preserved**: The validation uses OR logic (`config_flag ||
cli_flag`). If either is true, wildcard binds are allowed. This is unchanged.
## Consequences
**Positive:**
- Config reload will no longer reject valid configs that were accepted at
startup due to the `--allow-wildcard-bind` CLI flag.
- Consistent validation between startup and reload paths.
**Negative:**
- `ConfigReloadHandle::new()` gains an additional parameter. This is a minor
API change but affects all construction sites.
- The flag cannot be changed at runtime. If an operator wants to remove
`--allow-wildcard-bind`, they must restart the process. This is correct
behavior — wildcard bind is a security-sensitive setting that should
require a restart.
## References
- [config.md](../config.md) — Validation rules, allow_wildcard_bind
- [Review #005](../../reviews/005-admin-socket-security-review.md) — W5 finding
- `src/config/dynamic_config.rs` — ConfigReloadHandle, reload()
- `src/config/validation.rs` — validate()