docs(research): fuzzing must not share fate with agent sessions (detached runner)

- OOM in a fuzz target must cost the fuzzer, never the opencode host
- three defense layers: soft rss/malloc limits, fork-mode blast radius,
  setsid+nohup detachment with log-file polling (fuzz/run-detached.sh)
- no ulimit -v with ASAN; corpus replay + CI flag parity notes
- sequencing and summary updated to make the detached runner mandatory
This commit is contained in:
glm-5.3-flash committed 2026-09-27 20:13:35 +00:00
1 parent 791eea7298
commit 7ae0315f19
1 file changed
+101 -3
+101 -3
View File
@@ -372,8 +372,9 @@ comparisons), and a `-dict` of JSON tokens for the envelope targets.
1. `cargo fuzz init` + first two targets (`chunk_header`, `envelope_frame`)
— highest value, lowest setup (both are thin wrappers over existing
sync/`Cursor`-drivable APIs).
2. Run a local 10–30 min campaign per target; fix anything found;
triage §6.2 candidates with targeted corpus entries.
2. Run a local 10–30 min campaign per target **via the detached runner
(§7.6 — never in the foreground of an agent session)**; fix anything
found; triage §6.2 candidates with targeted corpus entries.
3. Add targets 3–5, the smoke CI job, and `.gitignore` entries.
4. Scheduled campaign tier; then OSS-Fuzz application.
@@ -395,6 +396,100 @@ property test for the chunk header (yamux's pattern) if the team wants
in-suite fuzz-adjacent coverage without nightly. Both are optional
nice-to-haves; the core recommendation is the `fuzz/` workspace.
### 7.6 Operational isolation: fuzzing must never share fate with the agent session
This one is an environment constraint, not a nice-to-have. Agent sessions
(opencode) run fuzzer builds/campaigns through a bash tool that spawns the
fuzzer as a **child of the session's own process tree and cgroup**. A fuzz
run is precisely the workload shape that kills its own host:
- **OOM-killer shared fate**: the target's worst case *is* unbounded
allocation (that's what §6.2-1/3 look for). If libFuzzer's own RSS guard
is miscalibrated or races the kernel, the kernel OOM killer fires — and
it picks the largest-RSS process in the *cgroup*, which can be the
opencode server hosting the session, killing the agent mid-run. An
OOM in a fuzz target must cost the fuzzer, never the session.
- **Fork bombs and CPU saturation**: `-fork=N` mode spawns many workers;
a runaway campaign or a buggy target with unbounded task spawning can
starve the session host's CPU/IO.
- **Interactive-shell edge cases**: libFuzzer prints status lines that
confuse non-TTY shells; a crashed fuzzer must never leave the tool's
bash session hanging.
Three layers of defense, all standard, cheapest first:
1. **libFuzzer's own soft limits (always on)**: `-rss_limit_mb=2048` +
`-malloc_limit_mb=2048` make libFuzzer *exit cleanly* (report the
input as OOM-class artifact) when the target exceeds the budget, before
the kernel gets involved. This is the primary defense and it is
already part of the §7.2 flag set. Note the libFuzzer RSS limit is
**soft** (it polls `/proc/self/statm` in a background thread and
exits), not a hard rlimit — a fast single huge allocation can still
beat it; that's what layers 2–3 are for. Do **not** use `ulimit -v`
with ASAN/MSAN (the sanitizer reserves terabytes of virtual address
space; the classic failure mode is immediate `MmapAlloc` death) — the
equivalent knob under ASAN is `malloc_limit_mb`, and libFuzzer's
documented recipe for hard-limiting RSS under sanitizers is `-fork=1`
(see next item).
2. **Fork mode as the blast-radius containment (default for local
campaigns)**: `-fork=1` (or `$(nproc)`) runs each input in a short-
lived child process. A target crash, timeout, OOM, or leak kills only
that one-shot child; the parent harness survives, records the artifact,
and keeps fuzzing. This is libFuzzer's own documented containment
model — and it is mutually exclusive with in-process ASAN crash
reporting (the child still detects, the parent persists the evidence).
Fork mode also solves the *other* shared-fate trap: a target that
deadlocks no longer hangs the campaign (per-input timeout kills the
child, not the session's bash call).
3. **Process detachment from the session tree (mandatory for agent-run
campaigns)**: agent sessions must never run fuzzing as a foreground
child. The runner script below wraps `cargo fuzz run` in
`setsid` + `nohup` + stdin/stdout/stderr redirection to a log file,
which (a) removes the fuzzer from the session's controlling terminal
and signal-relation, (b) makes the run survive the session ending, and
(c) gives the agent a pollable log/artifact tail instead of a blocking
call. Combined with fork mode, an OOM-class finding costs one child
process; combined with the soft RSS limit, the kernel OOM killer should
never be the discovery mechanism at all.
The detached-runner recipe for agent sessions (write as
`fuzz/run-detached.sh` in the implementation step, not inline in a tool
call):
```bash
#!/usr/bin/env bash
# Detached fuzzing runner for agent sessions: the fuzz campaign never
# runs as a foreground child of the session (OOM in a target must not
# take down the agent host), and survives the session ending.
set -euo pipefail
target="${1:?usage: run-detached.sh <target> [extra libfuzzer args...]}"
shift
runtime="${FUZZ_RUNTIME_SECS:-600}"
log="fuzz/artifacts/${target}-$(date -u +%Y%m%d-%H%M%S).log"
mkdir -p fuzz/artifacts
setsid nohup cargo fuzz run "$target" -- \
-fork=1 -rss_limit_mb=2048 -malloc_limit_mb=2048 -timeout=25 \
-max_total_time="$runtime" "$@" \
>"$log" 2>&1 < /dev/null &
echo "pid=$! log=$log"
```
Polling protocol for the agent: `tail -n 50 fuzz/artifacts/<target>-*.log`
and check `fuzz/artifacts/` for `crash-*`/`oom-*`/`timeout-*` files
periodically; `pgrep -f "cargo fuzz run <target>"` to see if it is still
running. Never `wait` on the detached process from a tool call — poll
instead. One session can hold one detached campaign per target; the log
filename carries the UTC timestamp.
**For CI the constraint does not apply** (the runner *is* the job), but
the same flags apply everywhere: `-rss_limit_mb`/`-malloc_limit_mb` and
`-fork=1` should be pinned in the target run configs in both the smoke
tier and the scheduled campaign. For the scheduled campaign tier, also
cap concurrency (`-fork=$(nproc)` is fine on CI runners; locally prefer
an explicit number well below core count) and keep the Actions-job
timeout above the `-max_total_time` budget so a wedged harness is killed
by the runner, not the host.
---
## 8. Answering the "if / how" directly
@@ -409,7 +504,10 @@ nice-to-haves; the core recommendation is the `fuzz/` workspace.
internals exposure where needed, two-tier CI with a 10 s-per-target
smoke on every push (§7.2), nightly confined to the fuzz job, AFL as an
optional secondary engine, bolero deferred, OSS-Fuzz after
stabilization.
stabilization. **Locally (agent sessions), campaigns always run via the
detached runner (§7.6): `setsid` + `nohup` + log file + poll, fork mode
on, soft RSS/malloc limits pinned — an OOM in a fuzz target must cost
the fuzzer, never the agent session.**
## 9. References