docs(research): fuzzing evaluation — adopt cargo-fuzz, five targets, two-tier CI

- web survey of 2025/2026 Rust fuzzing tooling (cargo-fuzz/afl/bolero/libafl)
- practice survey: rustls, quinn, quiche, h2, prost, s2n-quic, yamux/libp2p
- parse-surface inventory of this crate (both wire formats, dispatch, manager)
- pre-fuzz code-review findings: unbounded early-arrival key-space
  (manager.rs early_arrivals), unbounded spawned-JoinHandle growth
  (dispatch.rs), 64 MiB pre-payload allocation window, O(n^2) abort cascade
- recommendation: cargo-fuzz + arbitrary, 5 targets (chunk_header,
  envelope_frame, manager_routing, envelope_semantic, spec_parse),
  committed seeds only, 10s/target smoke CI, OSS-Fuzz after stabilization
This commit is contained in:
glm-5.3-flash committed 2026-09-27 20:01:29 +00:00
1 parent 7bb76cd2b5
commit 791eea7298
1 file changed
+439
+439
View File
@@ -0,0 +1,439 @@
# Fuzzing alkcall: Research and Recommendation
**Status:** recommendation (not yet adopted)
**Date:** 2026-09-27
**Research inputs:** web survey of the 2025–2026 Rust fuzzing ecosystem, survey
of fuzzing practice in comparable Rust protocol crates (rustls, quinn, quiche,
h2, prost, s2n-quic, yamux/libp2p, serde_json), and a first-hand inventory of
alkcall's untrusted-input parse surfaces (§5, verified against the code in
this repo).
---
## 1. Should this crate be fuzzed?
**Yes.** Three independent lines of argument converge:
1. **Position in the dependency graph.** alkcall is the substrate crate:
vendored core types, the call protocol, and the channels demux/mux all
live here, and downstream crates consume the wire formats rather than
re-implementing them (ADR-031 crate decomposition, ADR-007/008/009
vendored types). A parser bug here propagates everywhere; a fuzzed
parser here protects the whole stack. Fuzzing "at this level and working
down" is the right order — it is exactly what rustls did (fuzz the
deframer at the bottom, then the state machines).
2. **Wire formats are stable and attacker-reachable by design.** The
`EventEnvelope` framing (`{ type, id, payload: serde_json::Value }`
behind a 4-byte BE length prefix, ADR-014) and the channels 8-byte chunk
header (`[channel_id:u32 BE][len:u32 BE][payload]`, ADR-034) are
one-way doors. Both are parsed from bytes produced by *arbitrary peers*
— including relayed peers in the ADR-042 hub topology, where the hub
forwards bytes it never validated. The `payload` field is deliberately
schema-free JSON, so serde is on the attacker-controlled path with no
type-level defense.
3. **The failure classes that dominate real-world findings in Rust are
present in this codebase's shape, and several already have concrete
candidate findings.** The rust-fuzz trophy case (250+ entries) shows
that for Rust protocol crates the dominant bug classes are not memory
unsafety but: panics on untrusted input, integer/length arithmetic
errors, OOM via length-prefix-driven allocation, and unbounded state
growth. alkcall is safe Rust end-to-end (zero `unsafe` blocks —
verified), which means **no memory-corruption class at all**, but every
other class maps directly onto code in `src/` (see §5 and §6.2).
The strongest single datapoint: **RUSTSEC-2026-0037 / CVE-2026-31812
(quinn-proto, March 2026, CVSS 8.7)** — a remote DoS via `unwrap()` in
transport-parameter parsing. The maintainers' post-incident statement was
explicit: *"we did not have sufficient fuzzing coverage to find this issue.
We have since added a fuzzing target to cover this code path."* quinn-proto
is structurally very close to alkcall (async Rust RPC-ish protocol crate,
safe Rust, peers on the wire). A panic on attacker wire bytes in this crate
would be the same CVE class — and "the code is well tested" did not save
quinn, because conventional tests do not randomly explore malformed inputs.
**The honest caveat (why this was not obviously required before):** alkcall
has no `unsafe`, uses serde_json (whose 128-depth recursion limit and
non-panicking parser absorb most JSON pathology), bounds both length
prefixes before allocating (64 MiB frame cap, 16 MiB chunk cap — both
verified), and has solid round-trip unit tests. So the *expected* yield is
not memory-safety bugs; it is (a) panic-class bugs in hand-rolled dispatch
parsing, (b) the DoS-class findings in §6.2 which no amount of
example-based testing would have caught, and (c) regression protection for
two stable wire formats. If the campaign's first months find nothing, that
is a successful negative result, not a wasted one — it converts "we think
the parsers are robust" into a demonstrated property.
---
## 2. Tool landscape (2025–2026 snapshot)
| Tool | Version (date) | Engine | Toolchain | Status / verdict for alkcall |
|---|---|---|---|---|
| `cargo-fuzz` | 0.13.2 (2026-06) | libFuzzer | nightly required | **Recommended primary.** Actively maintained (rust-fuzz org), 3 releases since mid-2025, the de-facto standard; every surveyed protocol crate uses it. |
| `libfuzzer-sys` | 0.4.13 (2026-06) | libFuzzer runtime | via cargo-fuzz | Re-exports `arbitrary`; `fuzz_target!` accepts any `Arbitrary` type. |
| `arbitrary` | 1.4.2 (2025-08) | bytes → structured values | stable | **Recommended for structured targets.** `#[derive(Arbitrary)]`; 1.4.x brought notable fuzzing-speed wins. |
| `afl` (afl.rs, wraps AFL++) | 0.18.2 (2026-05) | AFL++ | **stable OK** | Recommended secondary/optional engine — long campaigns, multi-core, no nightly needed. CMPLOG on by default. |
| `honggfuzz` | 0.5.62 (2026-08) | honggfuzz | stable OK | Usable but upstream quiet since mid-2024; no reason to prefer over the above. |
| `bolero` | 0.13.4 (2025-07) | front-end: libfuzzer/AFL/honggfuzz/**Kani**/plain `cargo test` | stable for test/random | Attractive "same test, both engines" model, production-proven at s2n-quic. Considered and **deferred** (§7). |
| `libafl` | 0.16.1 (2025-08) | library for building fuzzers | stable | State of the art but a framework, not a tool; overkill for byte-buffer parser targets. Revisit only for stateful connection fuzzing. |
| OSS-Fuzz | — | hosted service | nightly (provided) | Free continuous fuzzing; expects exactly the cargo-fuzz layout (`libfuzzer` + `address` only). Adopt after targets stabilize. |
| `proptest` / quickcheck | mature | random generation + shrinking | stable | Complement, not replacement: no coverage guidance. Good for round-trip properties in the normal suite. |
Corroborating detail on cargo-fuzz status: published on crates.io, install
via `cargo install cargo-fuzz`; subcommands `init`, `add`, `run`, `fmt`,
`tmin`, `cmin`, `coverage`; recent releases added `--fuzz-engine`,
`--disable-branch-folding`, codegen-units tuning, Windows/MSVC support.
libFuzzer itself is upstream-maintenance-mode (its authors moved to
Centipede), which is irrelevant in practice for Rust app crates —
cargo-fuzz remains the recommended path in the Rust Fuzz Book.
---
## 3. What comparable crates actually do
Full details in the research notes; the load-bearing patterns:
- **rustls** (the reference model): `fuzz/` workspace, 8 targets spanning
byte-parsers *and* full state machines; internals exposed to targets via
`#[cfg(fuzzing)] pub mod fuzzing` (targets get access without widening
the public API); invariants include consumption accounting
(`assert!(processed <= buf.len())`) and re-encode round-trips; fuzz
functions additionally exercised from plain unit tests; corpora in a
dedicated repo; CI runs **10 s per target on every push** + OSS-Fuzz
CIFuzz at 150 s on PRs.
- **quinn**: `#[derive(Arbitrary)]` structured inputs; the `packet.rs`
target asserts the decoder consumed exactly the input
(`assert_eq!(len, decoded.0.len() + rest)`); a stateful target replays
`Vec<Op>` op sequences against `StreamsState`. Post-CVE, this is the
closest structural analogue to alkcall.
- **quiche**: fuzzes the *whole* protocol handler with committed test
certs; 1800+ committed seed corpus files generated by a seed script; Mayhem for continuous fuzzing.
- **h2**: the standout e2e pattern — fuzz input interpreted as a *script*
for mock I/O (length-prefixed chunks drive every socket read/write), so
libFuzzer explores partial-write/partial-read interleavings through the
full tokio state machine.
- **prost**: libFuzzer + AFL side by side; decode-then-re-encode round-trip
invariant; committed seed corpus for the AFL side.
- **yamux / libp2p** (no fuzz targets — the cautionary example): rely on
quickcheck only; their hand-rolled length-prefix handling shows the
defensive patterns fuzzing would police — `MAX_FRAME_BODY_LEN` checked
**before** `vec![0; body_len]` (yamux `frame/io.rs`), "validate RPC
limits by parsing the wire format without allocating" (libp2p gossipsub
`validate_rpc_limits`).
- **s2n-quic**: bolero's flagship consumer — fuzz checks live as ordinary
`#[cfg(test)]` functions, run in normal CI as `cargo test` (corpus
committed as tarball per test), long-run under libFuzzer out-of-band,
and the same properties carry Kani proof attributes.
**Invariant taxonomy observed across all repos** (in ascending strength):
1. **No-panic** — baseline; `drop(result)` suffices because libFuzzer +
ASan + debug assertions catch panics.
2. **Consumption accounting** — decoder consumed exactly the expected
bytes; nothing lost or duplicated (quinn packet, rustls deframer).
3. **Round-trip** — `decode(encode(x)) == x` or `encode(decode(bytes))`
prefix-match (yamux, quiche qpack, s2n-quic, prost).
4. **Resource bounds** — attacker-controlled lengths rejected (or skipped
without allocating) before allocation; bounded map growth.
5. **Stateful op sequences** — structured `Vec<Op>` replay against
internal state machines (quinn streams, h2 MockIo).
CI is universal but budgeted: the norm is a short per-push smoke run
(10–300 s) plus a longer scheduled campaign; crash artifacts uploaded via
`actions/upload-artifact` on failure. Cargo corpora are committed by
quiche/rustls/s2n-quic, gitignored by quinn/h2.
---
## 4. Toolchain decision: cargo-fuzz + arbitrary, AFL optional, bolero deferred
**Primary: `cargo-fuzz` + `arbitrary`.** Rationale: it is what every
surveyed crate uses, so patterns and CI templates transfer directly; the
`fuzz_target!` macro accepts both raw `&[u8]` and `Arbitrary`-derived
types, which maps onto alkcall's two input shapes (malformed-byte hunting
vs. structured logic replay); OSS-Fuzz compatibility for free; nightly
requirement is confined to the fuzz job (a pinned nightly toolchain for
the `fuzz/` workspace only — the crate itself stays stable at MSRV 1.88
and the fuzz crate is excluded from publishing via `exclude` in
`Cargo.toml`).
**Secondary (optional, later): `afl` (AFL++).** Stable-toolchain, good for
long multi-core campaigns and as an independent engine to cross-check
findings (prost runs exactly this pairing). Not needed for the initial
rollout.
**Deferred: bolero.** The s2n-quic model (fuzz check colocated as a normal
unit test, replayed in CI with a committed corpus, long-run out-of-band)
has a real advantage for this repo specifically — alkcall has *no* `fuzz/`
workspace today and its tests are inline `#[cfg(test)]` modules, so
colocated checks fit the house style. But it adds a dependency and an
engine-selection layer between the code and cargo-fuzz, and none of the
RPC-protocol crates surveyed (quinn, rustls, prost, h2) use it. The
deciding factor: alkcall's highest-value targets are *async* (the framing
reader and demux loop are tokio-generic), and bolero's in-process test
form adds friction there (`#[tokio::test]` cannot drive `bolero::check!`
iteration without shims). Recommendation: start with a standard `fuzz/`
workspace; if corpus-replay-as-unit-test proves valuable later, adopting
bolero's pattern (or just hand-rolling a corpus-replay test, which quinn
does in plain `cargo test`) is a cheap follow-up. Note quinn's CI already
does plain corpus replay without bolero — that pattern is available
without any new dependency.
**Not adopted: libafl, honggfuzz, OSS-Fuzz (initially).** libafl is a
framework for bespoke fuzzers — unnecessary for parser targets. honggfuzz
is superseded. OSS-Fuzz is a strong *later* step once targets exist and
have been stable for a while (it is free continuous fuzzing and expects
exactly this target layout; both rustls and prost are on it).
**Sanitizer choice:** default `--sanitizer=address` with debug assertions
(cargo-fuzz's default config, `-O1`). MSAN is not warranted — with zero
`unsafe` there is no uninitialized-memory class; ASAN mainly adds
heap-buffer-overflow detection for slice math in hand-rolled framing and
double-free-style issues in vendored types, plus the debug-assertions
panic tripwire, which is the real detector here.
---
## 5. alkcall's untrusted-input parse surfaces (inventory)
Verified against the code (file:line as of 0.8.0). The crate is safe Rust
end-to-end: **zero `unsafe` blocks** (grep-verified). Two wire formats,
plus JSON-based op payloads:
| # | Surface | Location | Sync? | Notes |
|---|---|---|---|---|
| 1 | `EventEnvelope` framing decode | `src/protocol/wire.rs:204-232` | async (`read_frame`, generic over `AsyncRead`; tests already drive it with `std::io::Cursor`) | 4-byte BE length prefix; rejects len 0 and len > `MAX_FRAME_SIZE` (64 MiB) **before** `vec![0u8; length]` at `:221`; partial prefix/body handled (`ConnectionClosed`); serde_json parse of up to 64 MiB body. |
| 2 | Channels chunk header | `src/channels/wire.rs:96-112` (`parse_header`) | **sync, pure, no allocation** | `len >= 8` check, `length > MAX_CHUNK_LEN` (16 MiB) rejected before payload exists. `write_header` at `:123`. |
| 3 | Demux loop | `src/channels/adapter.rs:126-222` (`run_demux_loop_for_client`, pub) | async | `parse_header` → bounded alloc (`:181`) → `route_payload`; `TooLarge` branch skips in 64 KiB reads with a cumulative 256 MiB budget (`:137`, `:153-170`) instead of allocating; EOF → `clear_all`. |
| 4 | `ChannelManager` routing | `src/channels/manager.rs:439-457` (`route_payload`), `:463-485` (`park_early_arrival`) | async (bounded send await) | Unknown ids park up to `EARLY_ARRIVAL_CAP = 64` chunks per id; **the map's key-space is unbounded** (see §6.2). |
| 5 | Spec rebuild from wire | `src/client/from_call.rs:209-310` (`rebuild_spec_for`, `pub(crate)`, sync pure) | sync | Remote-announced op specs: op_type/visibility/access_control/error_schemas/`publish_schema` — attacker-shaped JSON becomes compiled `jsonschema` validators at registration. |
| 6 | `op/register` DTO | `src/registry/op_register.rs:66-84` (`OpRegisterRequest::from_json`, `pub`, sync pure) | sync | Wraps `rebuild_spec_for` + replace flag. |
| 7 | Pending-request state machine | `src/protocol/pending.rs:111-229` (sync, `pub`) | sync | `handle_responded/completed/aborted/error`, eviction. |
| 8 | Abort cascade | `src/protocol/abort.rs:46` (`cascade_abort`, sync pure) | sync | Tree walk; O(n²) descendant search (§6.2). |
| 9 | Dispatch payload parsing | `src/protocol/dispatch.rs:341-412` (`dispatch_start`), `:646-770` (`pump_sink`), `:897-1027` / `:1120-1250` (single-stream loops) | sync-prefix / async | `payload.get(...)` extraction, `forwarded_for` as `serde_json::from_value::<Identity>`, `call.error` → `CallError`, jsonschema validation of published chunks. |
| 10 | Client-side envelope pump | `src/protocol/connection.rs:693-741` (`read_single_stream_until_closed`, `dispatch_envelope`) | async | pub; decode → pending-resolution pump. |
| 11 | Channel lifecycle ops | `src/channels/operations.rs:182-272` (close/control handlers) | async | `as_u64` → `as u32` cast truncation on `channel_id`. |
| 12 | Relay reply parsing | `src/channels/relay.rs:239-297` (`open_on_producer_leg`), `:308-335` (`map_spoke_error`) | async | reply `channel_id` extraction, `details.reason/message`. |
Also relevant: **`jsonschema` is compiled from schemas that arrive from
the wire** (`from_call.rs:299-303` announced `publish_schema` /
`input_schema`; enforced per chunk at `dispatch.rs:706-718`,
`registration.rs:257-274`) — a compilable-but-pathological remote schema
is a CPU-amplification class the fuzz campaign can probe.
Non-surfaces (deliberately): no TTY/SFTP/varint parsers in this crate; the
channels layer carries opaque `Bytes` and the handler owns sub-stream
framing (ADR-035) — those downstream parsers are downstream crates' fuzz
targets.
---
## 6. Findings the campaign should target (candidate issues in current code)
These are pre-fuzzing code-review findings (mine, verified at 0.8.0) that
define what the fuzz targets must encode as invariants. They are also, in
effect, the first candidates for the fuzzer to confirm or refute.
### 6.1 Framing and parse invariants (encode as assertions in targets)
1. **No-panic on any byte sequence** for both framing paths and all sync
pure parsers — the baseline every target asserts by existing.
2. **Consumption accounting** — `read_frame` consumes exactly
`4 + length` bytes for a valid frame; `parse_header` reads exactly 8.
3. **Round-trip** — envelope encode→decode (structural equality; note
serde_json runs **without** `preserve_order` in this crate's dep tree,
so `Value` object key order is not preserved — byte-identity round-trip
asserts are invalid for envelopes; use structural equality); chunk
header encode→decode identity.
4. **Allocation bounds** — decode paths never allocate more than
`MAX_FRAME_SIZE` / `MAX_CHUNK_LEN` for any input.
5. **Termination** — demux skip loop obeys the 256 MiB budget and always
terminates on a trickle-feed peer.
### 6.2 DoS-class candidate findings (pre-fuzz code review)
1. **Unbounded early-arrival key-space** — `ChannelManager.early_arrivals`
(`src/channels/manager.rs:102`, `park_early_arrival` `:463-485`) caps
parked chunks **per key** at 64, but nothing bounds the number of
distinct never-adopted `channel_id` keys. A peer spraying unknown ids
parks up to 64 chunks × up to 16 MiB *per distinct id* until
`clear_all` at EOF — an OOM class on a long-lived connection. A fuzz
target driving `route_payload` with arbitrary ids and payloads, with a
total-bytes-parked assertion, will find this class immediately.
2. **Unbounded per-connection task accumulation** — `spawned: Vec<JoinHandle>`
(`src/protocol/dispatch.rs:894-895`, pushed at `:914`, drained only at
loop end `:1030-1032`, and the channel-0 twin `:1115-1118`,
`:1253-1255`): one task per inbound `call.requested`, no cap on
concurrent requests per connection. Cheap-request spam grows the
vector unboundedly for the connection's life.
3. **64 MiB pre-payload allocation window** — `read_frame` allocates
`vec![0u8; length]` after validating ≤ 64 MiB but *before reading any
payload byte* (`src/protocol/wire.rs:217-221`): repeated 4-byte
headers force repeated 64-MiB zeroing (allocation-rate pressure).
Bounded per frame, but worth a target asserting decode never
allocates beyond the cap and exploring whether a read-into-capped-
buffer refactor (libp2p gossipsub's validate-without-allocating
pattern) is warranted.
4. **O(n²) abort cascade** — `find_descendants` scans the full pending map
per frontier step (`src/protocol/abort.rs:80-104`); many pendings + an
abort = quadratic CPU. An op-sequence target on `PendingRequestMap` +
`AbortCascade` with many registered requests will surface it.
5. **Remote-schema CPU amplification** — compiled-from-wire jsonschema
validators applied per published chunk (§5); pathological-schema
generation is a semantic fuzz target.
Items 1 and 2 are the strongest arguments that fuzzing (specifically
*stateful*, not just parser-level fuzzing) is warranted here: they are
resource-bound bugs in correct-looking safe Rust that example-based tests
structurally cannot find, in code paths every downstream consumer exposes
to peers.
---
## 7. Recommendation
### 7.1 Adopt fuzzing; scope it to the wire surface
Create `fuzz/` via `cargo fuzz init` (cargo-fuzz ≥ 0.13; keep the fuzz
crate in the main workspace, its default since 0.11.4), add `fuzz/` and
`fuzz/artifacts/` to `.gitignore`, add `"fuzz"` handling to the publish
exclude list if needed (cargo-fuzz's generated layout is already
workspace-compatible). Add five targets, in priority order:
| # | Target | Input style | Drives | Invariants |
|---|---|---|---|---|
| 1 | `chunk_header` | raw `&[u8]` | `parse_header` / `write_header` (`channels/wire.rs`) | no-panic; round-trip identity; `TooLarge` iff `len > MAX_CHUNK_LEN`; consumption accounting (8 bytes) |
| 2 | `envelope_frame` | raw `&[u8]` | `FrameFramedReader::new(Cursor::new(bytes)).read_frame()` under a current-thread runtime's `block_on` (tests already use `Cursor`, `wire.rs:534`) | no-panic; never allocates > `MAX_FRAME_SIZE`; clean `FrameError` on truncation/oversize/zero-length; consumption accounting |
| 3 | `manager_routing` | op-sequence: `#[derive(Arbitrary)]` enum over `{ route_payload(id, bytes), adopt, teardown, clear_all, … }` | `ChannelManager` | **total parked-bytes bound** (§6.2-1); no-panic; id-uniqueness invariants |
| 4 | `envelope_semantic` | `#[derive(Arbitrary)]` envelope-shaped input | constructors → `write_frame` → `read_frame` round-trip | structural round-trip; `CallError` payload parse never panics |
| 5 | `spec_parse` | raw bytes → `serde_json::Value` → `rebuild_spec_for` / `OpRegisterRequest::from_json` | sync pure parsers | no-panic; registration always `Result`; round-trip vs `spec_to_json_pub` |
Second wave (after the first five stabilize, roughly in quinn's
`streams.rs` style): a pending-map/abort op-sequence target (§6.2-4), a
demux-loop byte-sequence target through `run_demux_loop_for_client`
(needs a small current-thread runtime and a spawned mux runner), and a
client-pump target on `read_single_stream_until_closed`.
**Corpus policy: commit hand-made seeds, gitignore grown corpora** (the
quinn/h2 pattern). Seeds per target: valid frame/header, truncation at
every prefix length, `length = 0`, `length = MAX+1`, `length = u32::MAX`,
channel 0, invalid UTF-8 in JSON, deep-ish nesting. libFuzzer runs fine
without seeds but is far more efficient on structured inputs. Seed
generation can be scripted from existing unit tests (quiche's
`gen_fuzz_seeds.sh` pattern).
**Where internals need exposure:** follow quinn/rustls — `#[cfg(fuzzing)]
pub mod fuzzing` modules re-exporting `pub(crate)` items
(`rebuild_spec_for`) rather than widening the public API. The crate
already has `pub` sync parsers (`parse_header`, `from_json`,
`PendingRequestMap`), so exposure needs are small.
### 7.2 CI: two-tier, rustls-style
1. **Per-push smoke** (cheap, catches build rot and shallow bugs): pinned
nightly toolchain + pinned cargo-fuzz; `cargo fuzz build`; then per
target `cargo fuzz run <t> -- -max_total_time=10` (rustls's exact
budget) in a small matrix; upload `fuzz/artifacts` on failure.
2. **Scheduled campaign** (weekly at first; cron-driven Actions workflow):
`-max_total_time=1800` (30 min) per target with a cached corpus
(Actions cache, `restore-keys` prefix matching) and `-fork=$(nproc)`;
grow corpora incrementally. This is the depot.dev pattern; adopt it
only after the smoke tier is green for a while.
3. **No-cargo-fuzz fallback for PRs**: since nightly is a dedicated-job
concern, an even cheaper PR gate is `cargo test --manifest-path
fuzz/Cargo.toml` (replays the committed corpus through the targets as
plain unit tests — quinn's CI does this; corpus replay costs nothing
and prevents seed rot).
libFuzzer flags worth pinning in the targets' run configs:
`-rss_limit_mb=2048` (the OOM tripwire — directly relevant to §6.2-1/3),
`-max_len=65536` on framing targets, `-timeout=25` (well under CI job
limits), `-use_value_profile=1` (helps get past length/JSON-prefix
comparisons), and a `-dict` of JSON tokens for the envelope targets.
### 7.3 What not to do
- **Do not** gate regular development on nightly: fuzz tooling lives
entirely in `fuzz/` + the fuzz CI job; `cargo test` / MSRV / wasm
targets are untouched.
- **Do not** assert byte-identity round-trips for `Value` payloads
(serde_json without `preserve_order` sorts object keys — structural
equality only). This would generate false positives from day one.
- **Do not** commit grown corpora to git at first (quinn/h2 model);
revisit if the scheduled campaign produces inputs worth curating.
- **Do not** reach for libafl/OSS-Fuzz/Mayhem until the five initial
targets are stable; OSS-Fuzz onboarding is the natural next step after
that (free continuous fuzzing, expects exactly this layout).
### 7.4 Sequencing
1. `cargo fuzz init` + first two targets (`chunk_header`, `envelope_frame`)
— highest value, lowest setup (both are thin wrappers over existing
sync/`Cursor`-drivable APIs).
2. Run a local 10–30 min campaign per target; fix anything found;
triage §6.2 candidates with targeted corpus entries.
3. Add targets 3–5, the smoke CI job, and `.gitignore` entries.
4. Scheduled campaign tier; then OSS-Fuzz application.
### 7.5 Relationship to existing tests
The existing suite is strong on *valid-input* round-trips and documented
error paths (framing tests at `wire.rs:262-588`, demux skip tests at
`adapter.rs:279-693`, spec round-trips at `from_call.rs:634-1053`). What
no example-based suite provides, and what fuzzing adds:
- random *malformed* input exploration (the quinn CVE class),
- stateful interleavings (many ids × many ops, the §6.2-1/2 classes),
- resource-bound assertions under unbounded adversarial sequences.
Two zero-cost complements, available without any new toolchain: (a) replay
committed corpora as unit tests in normal CI (quinn's pattern — add it
from day one, it is two lines of CI); (b) a quickcheck-style round-trip
property test for the chunk header (yamux's pattern) if the team wants
in-suite fuzz-adjacent coverage without nightly. Both are optional
nice-to-haves; the core recommendation is the `fuzz/` workspace.
---
## 8. Answering the "if / how" directly
- **If?** Yes — justified by position in the dependency graph, two stable
attacker-reachable wire formats, and two concrete DoS-class findings a
fuzz campaign would have caught (§6.2). The expected bug classes here
are panics and resource exhaustion, not memory unsafety (no `unsafe`
exists), which is precisely the profile where cheap fuzzing pays off.
- **How?** `cargo-fuzz` + `arbitrary`, a standard `fuzz/` workspace with
five targets (§7.1), committed seed corpora only, `#[cfg(fuzzing)]`
internals exposure where needed, two-tier CI with a 10 s-per-target
smoke on every push (§7.2), nightly confined to the fuzz job, AFL as an
optional secondary engine, bolero deferred, OSS-Fuzz after
stabilization.
## 9. References
- RUSTSEC-2026-0037 / CVE-2026-31812 (quinn-proto remote DoS; maintainers'
fuzzing-coverage statement) — https://rustsec.org/advisories/RUSTSEC-2026-0037.html
- cargo-fuzz releases / Rust Fuzz Book — https://github.com/rust-fuzz/cargo-fuzz ·
https://rust-fuzz.github.io/book/
- arbitrary 1.4.2 and the 2025 speed analysis —
https://nnethercote.github.io/2025/08/16/speed-wins-when-fuzzing-rust-code-with-derive-arbitrary.html
- rustls fuzz targets + CI — https://github.com/rustls/rustls/tree/main/fuzz ·
https://github.com/rustls/rustls-fuzzing-corpora
- quinn fuzz targets (packet/streams/params) — https://github.com/quinn-rs/quinn/tree/main/fuzz
- quiche fuzzing + seed generation — https://github.com/cloudflare/quiche/tree/master/fuzz
- h2 MockIo e2e target — https://github.com/hyperium/h2/tree/master/fuzz
- prost fuzz (libfuzzer + AFL, FUZZING.md) — https://github.com/tokio-rs/prost/tree/master/fuzz
- s2n-quic bolero corpus-per-test pattern — https://github.com/aws/s2n-quic
- yamux allocation guard / libp2p validate-without-allocating —
https://github.com/libp2p/rust-yamux ·
https://github.com/libp2p/rust-libp2p (gossipsub `validate_rpc_limits`)
- rust-fuzz trophy case (finding-class taxonomy) — https://github.com/rust-fuzz/trophy-case
- OSS-Fuzz Rust integration — https://google.github.io/oss-fuzz/getting-started/new-project-guide/rust-lang/
- Scheduled CI fuzzing with cached corpora — https://depot.dev/blog/distributed-rust-fuzzing
Internal: ADR-014 (envelope framing), ADR-034 (chunk header), ADR-040
(backpressure/limits), ADR-042 (hub relay), ADR-031 (crate
decomposition); review 006 E-04 (early-arrival cap sizing note,
`src/channels/manager.rs:106-121`).