7 Commits
Author SHA1 Message Date
glm-5.3-flash 3bda29c419 chore(release): 0.8.1 — duplicate adopt/open fix + fuzz harness bookkeeping
- version 0.8.0 -> 0.8.1 (bug-fix release; semver-checks 196 pass vs 0.8.0)
- CHANGELOG 0.8.1: the manager_routing-found duplicate adopt/open fix,
  the fuzz harness, and the verification summary
- publish exclude gains docs/research/ (internal research notes;
  fuzz/ was already excluded — package verified clean of both)
- fuzzing.md §7.9: standing no-hosted-CI policy — corpus replay is the
  release-verification fuzz gate; campaigns manual via the detached
  runner; OSS-Fuzz out; no workflow files in this repo
- AGENTS.md: corpus replay added to the verification checklist
- lockfiles: alkcall 0.8.1 (root + fuzz)

Verification: 684 tests; clippy -D warnings (main, wasm, fuzz/shared);
fmt clean; doc clean; wasm32-unknown-unknown check+clippy clean;
semver-checks 196 pass; publish --dry-run 116 files (no fuzz/research/
reviews/sdd/AGENTS in the tarball); fuzz corpus replay 5/5 green
2026-09-28 06:56:11 +00:00
glm-5.3-flash 77f7006e2e docs(research): fuzzing steps 1-3 complete; 10-min campaigns clean on all five targets
\u00a77.8 campaign results: manager_routing 1.06M execs (exact-counter and
parked-bytes invariants held across every adversarial sequence),
envelope_semantic 2.44M execs (coverage saturated, all event kinds
round-trip), spec_parse 10.8M execs (registry compiled every
attacker-shaped schema or rejected cleanly). All exited 0, zero
artifacts.

Also records the two harness-model defects the fuzzer caught in the
interim (per-receiver EOF-sentinel latch; odd/even Adopt range per
ADR-047 \u00a75). Doc-only + harness fixes; no crate changes.
2026-09-27 23:52:45 +00:00
glm-5.3-flash a1f257757b fix(channels): duplicate adopt/open must not destroy the live channel; fuzz targets 3-5
Found by the manager_routing fuzz target (docs/research/fuzzing.md \u00a77.8):

- ChannelManager::open_channel / adopt_channel used
  HashMap::insert(...).is_some() as a collision check \u2014 insert REPLACES
  the existing entry, so a duplicate adopt/open installed the new state,
  dropped the live channel's demux_sender (spurious EOF to its readers,
  subsequently routed chunks lost) and still returned
  Err(ChannelExists). Fixed with contains_key pre-check; map untouched
  on collision. Regression tests: adopt_channel_duplicate_id_leaves_
  live_channel_intact, open_channel_duplicate_id_leaves_live_channel_
  intact.

New fuzz targets (\u00a77.4 step 3):
- manager_routing: Arbitrary op sequences over ChannelManager; exact
  counter models (parked/dropped must equal the manager's monotonic
  counters), parked-bytes bound per \u00a76.2-1, clear_all ledger-vs-map
  semantics, drainer-byte reconciliation (lossless routing)
- envelope_semantic: constructors -> serde -> write_frame/read_frame
  structural round-trip; event-type constants; call.error parse-back
- spec_parse: OpRegisterRequest::from_json -> rebuild -> registry
  registration (attacker schemas compile at register, CF-003)
- fuzz/shared/src/arbitrary_value.rs: bounded Arbitrary for
  serde_json::Value; 20 spec_parse + 4 manager_routing seeds
- corpus replay for the new targets in fuzz/shared tests

Verification: cargo test 684 passed (682 + 2 regression); clippy
-D warnings clean (main + fuzz/shared); fmt clean (main + fuzz);
cargo fuzz build clean; 20 s smoke on all three new targets clean
(manager_routing 79k, envelope_semantic 141k, spec_parse 517k runs);
crash input replays clean post-fix
2026-09-27 23:19:18 +00:00
glm-5.3-flash c211c163fd docs(research): fuzzing step 1 complete — campaigns ran clean (fuzzing.md \u00a77.7)
chunk_header: saturated (203.9M execs, ~350k exec/s, 33 fork jobs clean);
envelope_frame: 2119 edges/1664 corpus entries, still growing at budget
end, 7.08M execs, 33 fork jobs clean. No crash/OOM/timeout artifacts in
either 10-min detached campaign; exit 0 both. Grown corpus not merged
(per \u00a77.1 corpus policy). Doc-only change.
2026-09-27 21:27:58 +00:00
glm-5.3-flash 80a37e94ab feat(fuzz): cargo-fuzz workspace, chunk_header + envelope_frame targets (fuzzing.md step 1)
- fuzz/ workspace (nightly-pinned via rust-toolchain.toml, excluded from
  the main workspace and the published package): chunk_header and
  envelope_frame targets per fuzzing.md \u00a77.1
- invariant logic in fuzz/shared (stable-toolchain crate): committed
  corpus replay as plain cargo test (quinn CI pattern, \u00a77.2 tier 3)
- envelope target adds FrameError shape-partition asserts, exact
  consumption accounting, structural write_frame round-trip, serde
  key-contract, and a trailing-byte probe (prefix counts body only)
- chunk_header target adds round-trip identity, TooLarge/short error
  shape, is_eof, 8-byte consumption, input-never-mutated
- committed seed corpora: 245 deterministic seeds via
  fuzz/gen_fuzz_seeds.py (truncations, len=0/MAX+1/u32::MAX, channel 0,
  invalid UTF-8, deep nesting); grown corpora + artifacts gitignored
- fuzz/run-detached.sh: \u00a77.6 detached runner (setsid+nohup+log, fork
  mode, rss/malloc limits) \u2014 campaigns never share fate with a session
- fuzz/json.dict; README; research doc \u00a77.7 records step-1 status

Verification: cargo fuzz build clean; stable side green (cargo test 682
passed, clippy -D warnings, fmt --check incl. fuzz/shared); corpus
replay 245 seeds green; detached 10-min campaigns on both targets
completed with zero crashes/OOMs/timeouts
2026-09-27 21:14:35 +00:00
glm-5.3-flash 7ae0315f19 docs(research): fuzzing must not share fate with agent sessions (detached runner)
- OOM in a fuzz target must cost the fuzzer, never the opencode host
- three defense layers: soft rss/malloc limits, fork-mode blast radius,
  setsid+nohup detachment with log-file polling (fuzz/run-detached.sh)
- no ulimit -v with ASAN; corpus replay + CI flag parity notes
- sequencing and summary updated to make the detached runner mandatory
2026-09-27 20:13:35 +00:00
glm-5.3-flash 791eea7298 docs(research): fuzzing evaluation — adopt cargo-fuzz, five targets, two-tier CI
- web survey of 2025/2026 Rust fuzzing tooling (cargo-fuzz/afl/bolero/libafl)
- practice survey: rustls, quinn, quiche, h2, prost, s2n-quic, yamux/libp2p
- parse-surface inventory of this crate (both wire formats, dispatch, manager)
- pre-fuzz code-review findings: unbounded early-arrival key-space
  (manager.rs early_arrivals), unbounded spawned-JoinHandle growth
  (dispatch.rs), 64 MiB pre-payload allocation window, O(n^2) abort cascade
- recommendation: cargo-fuzz + arbitrary, 5 targets (chunk_header,
  envelope_frame, manager_routing, envelope_semantic, spec_parse),
  committed seeds only, 10s/target smoke CI, OSS-Fuzz after stabilization
2026-09-27 20:01:29 +00:00
1465 changed files with 3865 additions and 5 deletions

No files matched your search

+1
View File
@@ -194,6 +194,7 @@ Run these before committing. All must pass.
```bash
cargo test # full suite
cargo test --manifest-path fuzz/shared/Cargo.toml # fuzz corpus replay (stable)
cargo clippy --all-targets -- -D warnings
cargo fmt --check
cargo doc --no-deps # if docs changed
+60
View File
@@ -4,6 +4,65 @@ All notable changes to this crate are documented here. The format is
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this crate adheres to [Semantic Versioning](https://semver.org/).
## [0.8.1] - 2026-09-28
Bug-fix release: a duplicate `adopt_channel`/`open_channel` no longer
destroys the live channel it collided with, found by the crate's new
fuzzing harness (`fuzz/` workspace — five cargo-fuzz targets, committed
seed corpora, three 10-minute campaigns clean). No API or wire-format
change; the error is still `Err(ChannelExists)` — the difference is
entirely in what the collision leaves behind.
### Fixed
- **Duplicate adopt/open no longer replaces the live channel's state
(found by the `manager_routing` fuzz target,
`docs/research/fuzzing.md` §7.8).** `ChannelManager::open_channel`
and `adopt_channel` used `HashMap::insert(...).is_some()` as the
collision check — but `insert` *replaces* an occupied entry and
returns the old value, so a duplicate on an in-use id installed the
new `ChannelState`, dropped the live channel's `demux_sender`
(spurious EOF to its readers; all subsequently routed chunks lost),
and *still* returned `Err(ChannelExists)`. The mux write half kept
framing onto the transport for a channel the demux no longer fed.
Both functions now check `contains_key` and return before any map
mutation — the live channel's routing state is untouched on
collision. Reachable whenever an open-op response is replayed or a
connection-owner race re-announces an id. Regression tests:
`adopt_channel_duplicate_id_leaves_live_channel_intact`,
`open_channel_duplicate_id_leaves_live_channel_intact` (both verify
routing survives a rejected duplicate; the first also verifies the
EOF sentinel remains the only EOF source).
### Added (development)
- **Fuzzing harness (`fuzz/` workspace, `docs/research/fuzzing.md`).**
Five cargo-fuzz targets over the two attacker-reachable wire formats
and the stateful channel-routing surface: `chunk_header` (8-byte
header no-panic/round-trip/accounting), `envelope_frame` (frame
decode error-shape partition + allocation bound), `manager_routing`
(`Arbitrary` op sequences over `ChannelManager` with exact counter
models and the parked-bytes bound), `envelope_semantic` (all six
event kinds: constructors → serde → framing round-trip), and
`spec_parse` (attacker-shaped op-register schemas compile or reject
cleanly at registration). Invariant logic lives in
`fuzz/shared/` — a stable-toolchain crate replayed as plain
`cargo test`, so the committed corpora stay executable without
nightly. Campaigns (10 min per target, detached runner): ~25M
executions total, zero crashes/hangs/OOMs. The `manager_routing`
campaign found the 0.8.1 fix above within minutes — exactly the
stateful interleaving class example-based tests cannot reach.
`fuzz/` is excluded from the workspace and from the published
package.
### Verified
- 684 tests pass; `clippy --all-targets -- -D warnings`, `fmt --check`,
`cargo doc`, wasm32-unknown-unknown check, `cargo semver-checks`
(no API changes vs 0.8.0), and `cargo publish --dry-run` clean;
corpus replay (`cargo test --manifest-path fuzz/shared/Cargo.toml`)
green.
## [0.8.0] - 2026-09-18
Review 008's remediation lands in full — the graduation upstream asks
@@ -678,6 +737,7 @@ Vendored core types (`Connection`, `ProtocolHandler`, `BiStream`,
(ADR-046), the channels protocol with openable-ALPNs-as-operations
(ADR-047), and the `ChannelClient` transport-agnostic client.
[0.8.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.8.1
[0.8.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.8.0
[0.7.1]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.7.1
[0.7.0]: https://git.alk.dev/alkdev/alkcall/releases/tag/v0.7.0
Generated
+1 -1
View File
@@ -27,7 +27,7 @@ dependencies = [
[[package]]
name = "alkcall"
version = "0.8.0"
version = "0.8.1"
dependencies = [
"async-trait",
"bytes",
+6 -2
View File
@@ -1,6 +1,6 @@
[package]
name = "alkcall"
version = "0.8.0"
version = "0.8.1"
edition = "2021"
rust-version = "1.88"
license = "MIT OR Apache-2.0"
@@ -9,7 +9,11 @@ readme = "README.md"
repository = "https://git.alk.dev/alkdev/alkcall"
keywords = ["rpc", "json-rpc", "multiplexing", "wire-format", "alpn"]
categories = ["network-programming", "asynchronous", "encoding"]
exclude = [".opencode/", "AGENTS.md", "docs/reviews/", "docs/sdd_process.md"]
exclude = [".opencode/", "AGENTS.md", "docs/reviews/", "docs/sdd_process.md", "docs/research/", "fuzz/"]
[workspace]
members = ["."]
exclude = ["fuzz"]
[lib]
name = "alkcall"
+814
View File
@@ -0,0 +1,814 @@
# Fuzzing alkcall: Research and Recommendation
**Status:** steps 1–3 of §7.4 adopted — `fuzz/` workspace, five targets,
committed seeds, detached runner, two-tier campaigns run; one real
finding (duplicate adopt/open destroying the live channel) fixed with
regression tests. §7.2/§7.4's CI tiers are **not planned** — this repo
will not run hosted CI (the git host is a minimal gitea that serves git
only; see §7.9). CI's two deliverables are covered instead: the stable
side's corpus replay (`cargo test --manifest-path fuzz/shared/Cargo.toml`)
runs as part of every release's verification checklist (AGENTS.md), and
campaigns run manually via the detached runner (§7.6)
**Date:** 2026-09-27
**Research inputs:** web survey of the 2025–2026 Rust fuzzing ecosystem, survey
of fuzzing practice in comparable Rust protocol crates (rustls, quinn, quiche,
h2, prost, s2n-quic, yamux/libp2p, serde_json), and a first-hand inventory of
alkcall's untrusted-input parse surfaces (§5, verified against the code in
this repo).
---
## 1. Should this crate be fuzzed?
**Yes.** Three independent lines of argument converge:
1. **Position in the dependency graph.** alkcall is the substrate crate:
vendored core types, the call protocol, and the channels demux/mux all
live here, and downstream crates consume the wire formats rather than
re-implementing them (ADR-031 crate decomposition, ADR-007/008/009
vendored types). A parser bug here propagates everywhere; a fuzzed
parser here protects the whole stack. Fuzzing "at this level and working
down" is the right order — it is exactly what rustls did (fuzz the
deframer at the bottom, then the state machines).
2. **Wire formats are stable and attacker-reachable by design.** The
`EventEnvelope` framing (`{ type, id, payload: serde_json::Value }`
behind a 4-byte BE length prefix, ADR-014) and the channels 8-byte chunk
header (`[channel_id:u32 BE][len:u32 BE][payload]`, ADR-034) are
one-way doors. Both are parsed from bytes produced by *arbitrary peers*
— including relayed peers in the ADR-042 hub topology, where the hub
forwards bytes it never validated. The `payload` field is deliberately
schema-free JSON, so serde is on the attacker-controlled path with no
type-level defense.
3. **The failure classes that dominate real-world findings in Rust are
present in this codebase's shape, and several already have concrete
candidate findings.** The rust-fuzz trophy case (250+ entries) shows
that for Rust protocol crates the dominant bug classes are not memory
unsafety but: panics on untrusted input, integer/length arithmetic
errors, OOM via length-prefix-driven allocation, and unbounded state
growth. alkcall is safe Rust end-to-end (zero `unsafe` blocks —
verified), which means **no memory-corruption class at all**, but every
other class maps directly onto code in `src/` (see §5 and §6.2).
The strongest single datapoint: **RUSTSEC-2026-0037 / CVE-2026-31812
(quinn-proto, March 2026, CVSS 8.7)** — a remote DoS via `unwrap()` in
transport-parameter parsing. The maintainers' post-incident statement was
explicit: *"we did not have sufficient fuzzing coverage to find this issue.
We have since added a fuzzing target to cover this code path."* quinn-proto
is structurally very close to alkcall (async Rust RPC-ish protocol crate,
safe Rust, peers on the wire). A panic on attacker wire bytes in this crate
would be the same CVE class — and "the code is well tested" did not save
quinn, because conventional tests do not randomly explore malformed inputs.
**The honest caveat (why this was not obviously required before):** alkcall
has no `unsafe`, uses serde_json (whose 128-depth recursion limit and
non-panicking parser absorb most JSON pathology), bounds both length
prefixes before allocating (64 MiB frame cap, 16 MiB chunk cap — both
verified), and has solid round-trip unit tests. So the *expected* yield is
not memory-safety bugs; it is (a) panic-class bugs in hand-rolled dispatch
parsing, (b) the DoS-class findings in §6.2 which no amount of
example-based testing would have caught, and (c) regression protection for
two stable wire formats. If the campaign's first months find nothing, that
is a successful negative result, not a wasted one — it converts "we think
the parsers are robust" into a demonstrated property.
---
## 2. Tool landscape (2025–2026 snapshot)
| Tool | Version (date) | Engine | Toolchain | Status / verdict for alkcall |
|---|---|---|---|---|
| `cargo-fuzz` | 0.13.2 (2026-06) | libFuzzer | nightly required | **Recommended primary.** Actively maintained (rust-fuzz org), 3 releases since mid-2025, the de-facto standard; every surveyed protocol crate uses it. |
| `libfuzzer-sys` | 0.4.13 (2026-06) | libFuzzer runtime | via cargo-fuzz | Re-exports `arbitrary`; `fuzz_target!` accepts any `Arbitrary` type. |
| `arbitrary` | 1.4.2 (2025-08) | bytes → structured values | stable | **Recommended for structured targets.** `#[derive(Arbitrary)]`; 1.4.x brought notable fuzzing-speed wins. |
| `afl` (afl.rs, wraps AFL++) | 0.18.2 (2026-05) | AFL++ | **stable OK** | Recommended secondary/optional engine — long campaigns, multi-core, no nightly needed. CMPLOG on by default. |
| `honggfuzz` | 0.5.62 (2026-08) | honggfuzz | stable OK | Usable but upstream quiet since mid-2024; no reason to prefer over the above. |
| `bolero` | 0.13.4 (2025-07) | front-end: libfuzzer/AFL/honggfuzz/**Kani**/plain `cargo test` | stable for test/random | Attractive "same test, both engines" model, production-proven at s2n-quic. Considered and **deferred** (§7). |
| `libafl` | 0.16.1 (2025-08) | library for building fuzzers | stable | State of the art but a framework, not a tool; overkill for byte-buffer parser targets. Revisit only for stateful connection fuzzing. |
| OSS-Fuzz | — | hosted service | nightly (provided) | Free continuous fuzzing; expects exactly the cargo-fuzz layout (`libfuzzer` + `address` only). Adopt after targets stabilize. |
| `proptest` / quickcheck | mature | random generation + shrinking | stable | Complement, not replacement: no coverage guidance. Good for round-trip properties in the normal suite. |
Corroborating detail on cargo-fuzz status: published on crates.io, install
via `cargo install cargo-fuzz`; subcommands `init`, `add`, `run`, `fmt`,
`tmin`, `cmin`, `coverage`; recent releases added `--fuzz-engine`,
`--disable-branch-folding`, codegen-units tuning, Windows/MSVC support.
libFuzzer itself is upstream-maintenance-mode (its authors moved to
Centipede), which is irrelevant in practice for Rust app crates —
cargo-fuzz remains the recommended path in the Rust Fuzz Book.
---
## 3. What comparable crates actually do
Full details in the research notes; the load-bearing patterns:
- **rustls** (the reference model): `fuzz/` workspace, 8 targets spanning
byte-parsers *and* full state machines; internals exposed to targets via
`#[cfg(fuzzing)] pub mod fuzzing` (targets get access without widening
the public API); invariants include consumption accounting
(`assert!(processed <= buf.len())`) and re-encode round-trips; fuzz
functions additionally exercised from plain unit tests; corpora in a
dedicated repo; CI runs **10 s per target on every push** + OSS-Fuzz
CIFuzz at 150 s on PRs.
- **quinn**: `#[derive(Arbitrary)]` structured inputs; the `packet.rs`
target asserts the decoder consumed exactly the input
(`assert_eq!(len, decoded.0.len() + rest)`); a stateful target replays
`Vec<Op>` op sequences against `StreamsState`. Post-CVE, this is the
closest structural analogue to alkcall.
- **quiche**: fuzzes the *whole* protocol handler with committed test
certs; 1800+ committed seed corpus files generated by a seed script; Mayhem for continuous fuzzing.
- **h2**: the standout e2e pattern — fuzz input interpreted as a *script*
for mock I/O (length-prefixed chunks drive every socket read/write), so
libFuzzer explores partial-write/partial-read interleavings through the
full tokio state machine.
- **prost**: libFuzzer + AFL side by side; decode-then-re-encode round-trip
invariant; committed seed corpus for the AFL side.
- **yamux / libp2p** (no fuzz targets — the cautionary example): rely on
quickcheck only; their hand-rolled length-prefix handling shows the
defensive patterns fuzzing would police — `MAX_FRAME_BODY_LEN` checked
**before** `vec![0; body_len]` (yamux `frame/io.rs`), "validate RPC
limits by parsing the wire format without allocating" (libp2p gossipsub
`validate_rpc_limits`).
- **s2n-quic**: bolero's flagship consumer — fuzz checks live as ordinary
`#[cfg(test)]` functions, run in normal CI as `cargo test` (corpus
committed as tarball per test), long-run under libFuzzer out-of-band,
and the same properties carry Kani proof attributes.
**Invariant taxonomy observed across all repos** (in ascending strength):
1. **No-panic** — baseline; `drop(result)` suffices because libFuzzer +
ASan + debug assertions catch panics.
2. **Consumption accounting** — decoder consumed exactly the expected
bytes; nothing lost or duplicated (quinn packet, rustls deframer).
3. **Round-trip** — `decode(encode(x)) == x` or `encode(decode(bytes))`
prefix-match (yamux, quiche qpack, s2n-quic, prost).
4. **Resource bounds** — attacker-controlled lengths rejected (or skipped
without allocating) before allocation; bounded map growth.
5. **Stateful op sequences** — structured `Vec<Op>` replay against
internal state machines (quinn streams, h2 MockIo).
CI is universal but budgeted: the norm is a short per-push smoke run
(10–300 s) plus a longer scheduled campaign; crash artifacts uploaded via
`actions/upload-artifact` on failure. Cargo corpora are committed by
quiche/rustls/s2n-quic, gitignored by quinn/h2.
---
## 4. Toolchain decision: cargo-fuzz + arbitrary, AFL optional, bolero deferred
**Primary: `cargo-fuzz` + `arbitrary`.** Rationale: it is what every
surveyed crate uses, so patterns and CI templates transfer directly; the
`fuzz_target!` macro accepts both raw `&[u8]` and `Arbitrary`-derived
types, which maps onto alkcall's two input shapes (malformed-byte hunting
vs. structured logic replay); OSS-Fuzz compatibility for free; nightly
requirement is confined to the fuzz job (a pinned nightly toolchain for
the `fuzz/` workspace only — the crate itself stays stable at MSRV 1.88
and the fuzz crate is excluded from publishing via `exclude` in
`Cargo.toml`).
**Secondary (optional, later): `afl` (AFL++).** Stable-toolchain, good for
long multi-core campaigns and as an independent engine to cross-check
findings (prost runs exactly this pairing). Not needed for the initial
rollout.
**Deferred: bolero.** The s2n-quic model (fuzz check colocated as a normal
unit test, replayed in CI with a committed corpus, long-run out-of-band)
has a real advantage for this repo specifically — alkcall has *no* `fuzz/`
workspace today and its tests are inline `#[cfg(test)]` modules, so
colocated checks fit the house style. But it adds a dependency and an
engine-selection layer between the code and cargo-fuzz, and none of the
RPC-protocol crates surveyed (quinn, rustls, prost, h2) use it. The
deciding factor: alkcall's highest-value targets are *async* (the framing
reader and demux loop are tokio-generic), and bolero's in-process test
form adds friction there (`#[tokio::test]` cannot drive `bolero::check!`
iteration without shims). Recommendation: start with a standard `fuzz/`
workspace; if corpus-replay-as-unit-test proves valuable later, adopting
bolero's pattern (or just hand-rolling a corpus-replay test, which quinn
does in plain `cargo test`) is a cheap follow-up. Note quinn's CI already
does plain corpus replay without bolero — that pattern is available
without any new dependency.
**Not adopted: libafl, honggfuzz, OSS-Fuzz (initially).** libafl is a
framework for bespoke fuzzers — unnecessary for parser targets. honggfuzz
is superseded. OSS-Fuzz is a strong *later* step once targets exist and
have been stable for a while (it is free continuous fuzzing and expects
exactly this target layout; both rustls and prost are on it).
**Sanitizer choice:** default `--sanitizer=address` with debug assertions
(cargo-fuzz's default config, `-O1`). MSAN is not warranted — with zero
`unsafe` there is no uninitialized-memory class; ASAN mainly adds
heap-buffer-overflow detection for slice math in hand-rolled framing and
double-free-style issues in vendored types, plus the debug-assertions
panic tripwire, which is the real detector here.
---
## 5. alkcall's untrusted-input parse surfaces (inventory)
Verified against the code (file:line as of 0.8.0). The crate is safe Rust
end-to-end: **zero `unsafe` blocks** (grep-verified). Two wire formats,
plus JSON-based op payloads:
| # | Surface | Location | Sync? | Notes |
|---|---|---|---|---|
| 1 | `EventEnvelope` framing decode | `src/protocol/wire.rs:204-232` | async (`read_frame`, generic over `AsyncRead`; tests already drive it with `std::io::Cursor`) | 4-byte BE length prefix; rejects len 0 and len > `MAX_FRAME_SIZE` (64 MiB) **before** `vec![0u8; length]` at `:221`; partial prefix/body handled (`ConnectionClosed`); serde_json parse of up to 64 MiB body. |
| 2 | Channels chunk header | `src/channels/wire.rs:96-112` (`parse_header`) | **sync, pure, no allocation** | `len >= 8` check, `length > MAX_CHUNK_LEN` (16 MiB) rejected before payload exists. `write_header` at `:123`. |
| 3 | Demux loop | `src/channels/adapter.rs:126-222` (`run_demux_loop_for_client`, pub) | async | `parse_header` → bounded alloc (`:181`) → `route_payload`; `TooLarge` branch skips in 64 KiB reads with a cumulative 256 MiB budget (`:137`, `:153-170`) instead of allocating; EOF → `clear_all`. |
| 4 | `ChannelManager` routing | `src/channels/manager.rs:439-457` (`route_payload`), `:463-485` (`park_early_arrival`) | async (bounded send await) | Unknown ids park up to `EARLY_ARRIVAL_CAP = 64` chunks per id; **the map's key-space is unbounded** (see §6.2). |
| 5 | Spec rebuild from wire | `src/client/from_call.rs:209-310` (`rebuild_spec_for`, `pub(crate)`, sync pure) | sync | Remote-announced op specs: op_type/visibility/access_control/error_schemas/`publish_schema` — attacker-shaped JSON becomes compiled `jsonschema` validators at registration. |
| 6 | `op/register` DTO | `src/registry/op_register.rs:66-84` (`OpRegisterRequest::from_json`, `pub`, sync pure) | sync | Wraps `rebuild_spec_for` + replace flag. |
| 7 | Pending-request state machine | `src/protocol/pending.rs:111-229` (sync, `pub`) | sync | `handle_responded/completed/aborted/error`, eviction. |
| 8 | Abort cascade | `src/protocol/abort.rs:46` (`cascade_abort`, sync pure) | sync | Tree walk; O(n²) descendant search (§6.2). |
| 9 | Dispatch payload parsing | `src/protocol/dispatch.rs:341-412` (`dispatch_start`), `:646-770` (`pump_sink`), `:897-1027` / `:1120-1250` (single-stream loops) | sync-prefix / async | `payload.get(...)` extraction, `forwarded_for` as `serde_json::from_value::<Identity>`, `call.error` → `CallError`, jsonschema validation of published chunks. |
| 10 | Client-side envelope pump | `src/protocol/connection.rs:693-741` (`read_single_stream_until_closed`, `dispatch_envelope`) | async | pub; decode → pending-resolution pump. |
| 11 | Channel lifecycle ops | `src/channels/operations.rs:182-272` (close/control handlers) | async | `as_u64` → `as u32` cast truncation on `channel_id`. |
| 12 | Relay reply parsing | `src/channels/relay.rs:239-297` (`open_on_producer_leg`), `:308-335` (`map_spoke_error`) | async | reply `channel_id` extraction, `details.reason/message`. |
Also relevant: **`jsonschema` is compiled from schemas that arrive from
the wire** (`from_call.rs:299-303` announced `publish_schema` /
`input_schema`; enforced per chunk at `dispatch.rs:706-718`,
`registration.rs:257-274`) — a compilable-but-pathological remote schema
is a CPU-amplification class the fuzz campaign can probe.
Non-surfaces (deliberately): no TTY/SFTP/varint parsers in this crate; the
channels layer carries opaque `Bytes` and the handler owns sub-stream
framing (ADR-035) — those downstream parsers are downstream crates' fuzz
targets.
---
## 6. Findings the campaign should target (candidate issues in current code)
These are pre-fuzzing code-review findings (mine, verified at 0.8.0) that
define what the fuzz targets must encode as invariants. They are also, in
effect, the first candidates for the fuzzer to confirm or refute.
### 6.1 Framing and parse invariants (encode as assertions in targets)
1. **No-panic on any byte sequence** for both framing paths and all sync
pure parsers — the baseline every target asserts by existing.
2. **Consumption accounting** — `read_frame` consumes exactly
`4 + length` bytes for a valid frame; `parse_header` reads exactly 8.
3. **Round-trip** — envelope encode→decode (structural equality; note
serde_json runs **without** `preserve_order` in this crate's dep tree,
so `Value` object key order is not preserved — byte-identity round-trip
asserts are invalid for envelopes; use structural equality); chunk
header encode→decode identity.
4. **Allocation bounds** — decode paths never allocate more than
`MAX_FRAME_SIZE` / `MAX_CHUNK_LEN` for any input.
5. **Termination** — demux skip loop obeys the 256 MiB budget and always
terminates on a trickle-feed peer.
### 6.2 DoS-class candidate findings (pre-fuzz code review)
1. **Unbounded early-arrival key-space** — `ChannelManager.early_arrivals`
(`src/channels/manager.rs:102`, `park_early_arrival` `:463-485`) caps
parked chunks **per key** at 64, but nothing bounds the number of
distinct never-adopted `channel_id` keys. A peer spraying unknown ids
parks up to 64 chunks × up to 16 MiB *per distinct id* until
`clear_all` at EOF — an OOM class on a long-lived connection. A fuzz
target driving `route_payload` with arbitrary ids and payloads, with a
total-bytes-parked assertion, will find this class immediately.
2. **Unbounded per-connection task accumulation** — `spawned: Vec<JoinHandle>`
(`src/protocol/dispatch.rs:894-895`, pushed at `:914`, drained only at
loop end `:1030-1032`, and the channel-0 twin `:1115-1118`,
`:1253-1255`): one task per inbound `call.requested`, no cap on
concurrent requests per connection. Cheap-request spam grows the
vector unboundedly for the connection's life.
3. **64 MiB pre-payload allocation window** — `read_frame` allocates
`vec![0u8; length]` after validating ≤ 64 MiB but *before reading any
payload byte* (`src/protocol/wire.rs:217-221`): repeated 4-byte
headers force repeated 64-MiB zeroing (allocation-rate pressure).
Bounded per frame, but worth a target asserting decode never
allocates beyond the cap and exploring whether a read-into-capped-
buffer refactor (libp2p gossipsub's validate-without-allocating
pattern) is warranted.
4. **O(n²) abort cascade** — `find_descendants` scans the full pending map
per frontier step (`src/protocol/abort.rs:80-104`); many pendings + an
abort = quadratic CPU. An op-sequence target on `PendingRequestMap` +
`AbortCascade` with many registered requests will surface it.
5. **Remote-schema CPU amplification** — compiled-from-wire jsonschema
validators applied per published chunk (§5); pathological-schema
generation is a semantic fuzz target.
Items 1 and 2 are the strongest arguments that fuzzing (specifically
*stateful*, not just parser-level fuzzing) is warranted here: they are
resource-bound bugs in correct-looking safe Rust that example-based tests
structurally cannot find, in code paths every downstream consumer exposes
to peers.
---
## 7. Recommendation
### 7.1 Adopt fuzzing; scope it to the wire surface
Create `fuzz/` via `cargo fuzz init` (cargo-fuzz ≥ 0.13; keep the fuzz
crate in the main workspace, its default since 0.11.4), add `fuzz/` and
`fuzz/artifacts/` to `.gitignore`, add `"fuzz"` handling to the publish
exclude list if needed (cargo-fuzz's generated layout is already
workspace-compatible). Add five targets, in priority order:
| # | Target | Input style | Drives | Invariants |
|---|---|---|---|---|
| 1 | `chunk_header` | raw `&[u8]` | `parse_header` / `write_header` (`channels/wire.rs`) | no-panic; round-trip identity; `TooLarge` iff `len > MAX_CHUNK_LEN`; consumption accounting (8 bytes) |
| 2 | `envelope_frame` | raw `&[u8]` | `FrameFramedReader::new(Cursor::new(bytes)).read_frame()` under a current-thread runtime's `block_on` (tests already use `Cursor`, `wire.rs:534`) | no-panic; never allocates > `MAX_FRAME_SIZE`; clean `FrameError` on truncation/oversize/zero-length; consumption accounting |
| 3 | `manager_routing` | op-sequence: `#[derive(Arbitrary)]` enum over `{ route_payload(id, bytes), adopt, teardown, clear_all, … }` | `ChannelManager` | **total parked-bytes bound** (§6.2-1); no-panic; id-uniqueness invariants |
| 4 | `envelope_semantic` | `#[derive(Arbitrary)]` envelope-shaped input | constructors → `write_frame` → `read_frame` round-trip | structural round-trip; `CallError` payload parse never panics |
| 5 | `spec_parse` | raw bytes → `serde_json::Value` → `rebuild_spec_for` / `OpRegisterRequest::from_json` | sync pure parsers | no-panic; registration always `Result`; round-trip vs `spec_to_json_pub` |
Second wave (after the first five stabilize, roughly in quinn's
`streams.rs` style): a pending-map/abort op-sequence target (§6.2-4), a
demux-loop byte-sequence target through `run_demux_loop_for_client`
(needs a small current-thread runtime and a spawned mux runner), and a
client-pump target on `read_single_stream_until_closed`.
**Corpus policy: commit hand-made seeds, gitignore grown corpora** (the
quinn/h2 pattern). Seeds per target: valid frame/header, truncation at
every prefix length, `length = 0`, `length = MAX+1`, `length = u32::MAX`,
channel 0, invalid UTF-8 in JSON, deep-ish nesting. libFuzzer runs fine
without seeds but is far more efficient on structured inputs. Seed
generation can be scripted from existing unit tests (quiche's
`gen_fuzz_seeds.sh` pattern).
**Where internals need exposure:** follow quinn/rustls — `#[cfg(fuzzing)]
pub mod fuzzing` modules re-exporting `pub(crate)` items
(`rebuild_spec_for`) rather than widening the public API. The crate
already has `pub` sync parsers (`parse_header`, `from_json`,
`PendingRequestMap`), so exposure needs are small.
### 7.2 CI: two-tier, rustls-style
1. **Per-push smoke** (cheap, catches build rot and shallow bugs): pinned
nightly toolchain + pinned cargo-fuzz; `cargo fuzz build`; then per
target `cargo fuzz run <t> -- -max_total_time=10` (rustls's exact
budget) in a small matrix; upload `fuzz/artifacts` on failure.
2. **Scheduled campaign** (weekly at first; cron-driven Actions workflow):
`-max_total_time=1800` (30 min) per target with a cached corpus
(Actions cache, `restore-keys` prefix matching) and `-fork=$(nproc)`;
grow corpora incrementally. This is the depot.dev pattern; adopt it
only after the smoke tier is green for a while.
3. **No-cargo-fuzz fallback for PRs**: since nightly is a dedicated-job
concern, an even cheaper PR gate is `cargo test --manifest-path
fuzz/Cargo.toml` (replays the committed corpus through the targets as
plain unit tests — quinn's CI does this; corpus replay costs nothing
and prevents seed rot).
libFuzzer flags worth pinning in the targets' run configs:
`-rss_limit_mb=2048` (the OOM tripwire — directly relevant to §6.2-1/3),
`-max_len=65536` on framing targets, `-timeout=25` (well under CI job
limits), `-use_value_profile=1` (helps get past length/JSON-prefix
comparisons), and a `-dict` of JSON tokens for the envelope targets.
### 7.3 What not to do
- **Do not** gate regular development on nightly: fuzz tooling lives
entirely in `fuzz/` + the fuzz CI job; `cargo test` / MSRV / wasm
targets are untouched.
- **Do not** assert byte-identity round-trips for `Value` payloads
(serde_json without `preserve_order` sorts object keys — structural
equality only). This would generate false positives from day one.
- **Do not** commit grown corpora to git at first (quinn/h2 model);
revisit if the scheduled campaign produces inputs worth curating.
- **Do not** reach for libafl/OSS-Fuzz/Mayhem until the five initial
targets are stable; OSS-Fuzz onboarding is the natural next step after
that (free continuous fuzzing, expects exactly this layout).
### 7.4 Sequencing
1. `cargo fuzz init` + first two targets (`chunk_header`, `envelope_frame`)
— highest value, lowest setup (both are thin wrappers over existing
sync/`Cursor`-drivable APIs). **Done — see §7.7.**
2. Run a local 10–30 min campaign per target **via the detached runner
(§7.6 — never in the foreground of an agent session)**; fix anything
found; triage §6.2 candidates with targeted corpus entries.
**Done for targets 1–2 — see §7.7.**
3. Add targets 3–5, the smoke CI job, and `.gitignore` entries.
**Targets 3–5 done, campaigns clean — see §7.8. The CI job will not
exist — see §7.9 (no hosted CI policy); the stable-side corpus
replay `cargo test --manifest-path fuzz/shared/Cargo.toml` is in
the release verification checklist instead (AGENTS.md).**
4. ~~Scheduled campaign tier; then OSS-Fuzz application.~~ **Not
planned — see §7.9.** Long campaigns are run manually via the
detached runner (§7.6) when wanted; OSS-Fuzz is out (it requires
hosted infrastructure and a public repo posture this project does
not have).
### 7.5 Relationship to existing tests
The existing suite is strong on *valid-input* round-trips and documented
error paths (framing tests at `wire.rs:262-588`, demux skip tests at
`adapter.rs:279-693`, spec round-trips at `from_call.rs:634-1053`). What
no example-based suite provides, and what fuzzing adds:
- random *malformed* input exploration (the quinn CVE class),
- stateful interleavings (many ids × many ops, the §6.2-1/2 classes),
- resource-bound assertions under unbounded adversarial sequences.
Two zero-cost complements, available without any new toolchain: (a) replay
committed corpora as unit tests in normal CI (quinn's pattern — add it
from day one, it is two lines of CI); (b) a quickcheck-style round-trip
property test for the chunk header (yamux's pattern) if the team wants
in-suite fuzz-adjacent coverage without nightly. Both are optional
nice-to-haves; the core recommendation is the `fuzz/` workspace.
### 7.6 Operational isolation: fuzzing must never share fate with the agent session
This one is an environment constraint, not a nice-to-have. Agent sessions
(opencode) run fuzzer builds/campaigns through a bash tool that spawns the
fuzzer as a **child of the session's own process tree and cgroup**. A fuzz
run is precisely the workload shape that kills its own host:
- **OOM-killer shared fate**: the target's worst case *is* unbounded
allocation (that's what §6.2-1/3 look for). If libFuzzer's own RSS guard
is miscalibrated or races the kernel, the kernel OOM killer fires — and
it picks the largest-RSS process in the *cgroup*, which can be the
opencode server hosting the session, killing the agent mid-run. An
OOM in a fuzz target must cost the fuzzer, never the session.
- **Fork bombs and CPU saturation**: `-fork=N` mode spawns many workers;
a runaway campaign or a buggy target with unbounded task spawning can
starve the session host's CPU/IO.
- **Interactive-shell edge cases**: libFuzzer prints status lines that
confuse non-TTY shells; a crashed fuzzer must never leave the tool's
bash session hanging.
Three layers of defense, all standard, cheapest first:
1. **libFuzzer's own soft limits (always on)**: `-rss_limit_mb=2048` +
`-malloc_limit_mb=2048` make libFuzzer *exit cleanly* (report the
input as OOM-class artifact) when the target exceeds the budget, before
the kernel gets involved. This is the primary defense and it is
already part of the §7.2 flag set. Note the libFuzzer RSS limit is
**soft** (it polls `/proc/self/statm` in a background thread and
exits), not a hard rlimit — a fast single huge allocation can still
beat it; that's what layers 2–3 are for. Do **not** use `ulimit -v`
with ASAN/MSAN (the sanitizer reserves terabytes of virtual address
space; the classic failure mode is immediate `MmapAlloc` death) — the
equivalent knob under ASAN is `malloc_limit_mb`, and libFuzzer's
documented recipe for hard-limiting RSS under sanitizers is `-fork=1`
(see next item).
2. **Fork mode as the blast-radius containment (default for local
campaigns)**: `-fork=1` (or `$(nproc)`) runs each input in a short-
lived child process. A target crash, timeout, OOM, or leak kills only
that one-shot child; the parent harness survives, records the artifact,
and keeps fuzzing. This is libFuzzer's own documented containment
model — and it is mutually exclusive with in-process ASAN crash
reporting (the child still detects, the parent persists the evidence).
Fork mode also solves the *other* shared-fate trap: a target that
deadlocks no longer hangs the campaign (per-input timeout kills the
child, not the session's bash call).
3. **Process detachment from the session tree (mandatory for agent-run
campaigns)**: agent sessions must never run fuzzing as a foreground
child. The runner script below wraps `cargo fuzz run` in
`setsid` + `nohup` + stdin/stdout/stderr redirection to a log file,
which (a) removes the fuzzer from the session's controlling terminal
and signal-relation, (b) makes the run survive the session ending, and
(c) gives the agent a pollable log/artifact tail instead of a blocking
call. Combined with fork mode, an OOM-class finding costs one child
process; combined with the soft RSS limit, the kernel OOM killer should
never be the discovery mechanism at all.
The detached-runner recipe for agent sessions (write as
`fuzz/run-detached.sh` in the implementation step, not inline in a tool
call):
```bash
#!/usr/bin/env bash
# Detached fuzzing runner for agent sessions: the fuzz campaign never
# runs as a foreground child of the session (OOM in a target must not
# take down the agent host), and survives the session ending.
set -euo pipefail
target="${1:?usage: run-detached.sh <target> [extra libfuzzer args...]}"
shift
runtime="${FUZZ_RUNTIME_SECS:-600}"
log="fuzz/artifacts/${target}-$(date -u +%Y%m%d-%H%M%S).log"
mkdir -p fuzz/artifacts
setsid nohup cargo fuzz run "$target" -- \
-fork=1 -rss_limit_mb=2048 -malloc_limit_mb=2048 -timeout=25 \
-max_total_time="$runtime" "$@" \
>"$log" 2>&1 < /dev/null &
echo "pid=$! log=$log"
```
Polling protocol for the agent: `tail -n 50 fuzz/artifacts/<target>-*.log`
and check `fuzz/artifacts/` for `crash-*`/`oom-*`/`timeout-*` files
periodically; `pgrep -f "cargo fuzz run <target>"` to see if it is still
running. Never `wait` on the detached process from a tool call — poll
instead. One session can hold one detached campaign per target; the log
filename carries the UTC timestamp.
**For CI the constraint does not apply** (the runner *is* the job), but
the same flags apply everywhere: `-rss_limit_mb`/`-malloc_limit_mb` and
`-fork=1` should be pinned in the target run configs in both the smoke
tier and the scheduled campaign. For the scheduled campaign tier, also
cap concurrency (`-fork=$(nproc)` is fine on CI runners; locally prefer
an explicit number well below core count) and keep the Actions-job
timeout above the `-max_total_time` budget so a wedged harness is killed
by the runner, not the host.
---
## 7.7 Implementation status (step 1 of §7.4, 2026-09-27)
**Adopted.** The `fuzz/` workspace exists with targets 1 and 2 from
§7.1, committed seed corpora, the detached runner, and stable-side
corpus replay. Everything below is verified in-tree.
### Layout (as implemented)
```
fuzz/
├── Cargo.toml alkcall-fuzz (nightly-only bins; own [workspace])
├── rust-toolchain.toml pins nightly for this subtree only
├── fuzz_targets/
│ ├── chunk_header.rs thin fuzz_target! wrapper
│ └── envelope_frame.rs thin fuzz_target! wrapper
├── shared/ alkcall-fuzz-shared — STABLE-toolchain library
│ └── src/{chunk_header,envelope_frame}.rs invariant logic + corpus replay
├── corpus/{chunk_header,envelope_frame}/ committed seeds (245 total)
├── gen_fuzz_seeds.py deterministic seed generator (quiche pattern)
├── run-detached.sh §7.6 detached runner (CWD-independent)
├── json.dict JSON token dictionary for envelope targets
└── README.md
```
Deviations from the §7.1 sketch, all load-bearing:
- **The invariant logic lives in `fuzz/shared/` (a separate
stable-toolchain crate), not inline in the target binaries.** This
makes §7.2's tier-3 fallback (corpus replay as plain `cargo test`)
real from day one: `cargo test --manifest-path fuzz/shared/Cargo.toml`
replays every committed seed through the identical invariant functions
the fuzzer runs, on stable, in normal CI. The `fuzz_target!` binaries
are three-line wrappers.
- **The root `Cargo.toml` gained a `[workspace]` table
(`members = ["."]`, `exclude = ["fuzz"]`)** and `"fuzz/"` joined the
publish `exclude` list. Without the explicit workspace table, cargo
auto-discovers `fuzz/shared/` into the main workspace and the MSRV
toolchain tries to build the nightly-consumed dev-deps.
- **`fuzz/rust-toolchain.toml` pins `nightly`** for the fuzz subtree so
`cargo fuzz build` works from any CWD (rustup resolves toolchains per
directory; the initial build attempt from the repo root picked the
stable toolchain and failed on `-Zsanitizer`). Nightly remains
confined to `fuzz/` — the crate itself stays stable at MSRV 1.88.
- **No `#[cfg(fuzzing)]` exposure was needed.** Both target APIs were
already `pub`: `parse_header`/`write_header` (`channels/wire.rs`) and
`FrameFramedReader`/`FrameFramedWriter` (`protocol/wire.rs`). The one
private item, `MAX_FRAME_SIZE` (`wire.rs:20`), is a wire-stable
constant (ADR-014), so the shared crate carries its own copy
(16 MiB/64 MiB are one-way-door values; a drift would be caught by the
shape invariants, which assert the rejection boundary).
- **Seeds are committed** (245 files), gitignored artifacts/corpus-grown
per the generated `fuzz/.gitignore`; the generator script
(`gen_fuzz_seeds.py`) is committed next to them.
### Invariants encoded (beyond the §7.1 table)
`chunk_header`:
- no-panic; round-trip identity (`write_header(parse_header(x)) ==
x[..8]`); `TooLarge` iff `length > MAX_CHUNK_LEN` with the parsed
`length` echoed in the error; `HeaderTooShort` iff `< 8` bytes with
exact `need`/`have`; `is_eof() == (length == 0)`; consumption
accounting (8 bytes); **input buffer never mutated**.
`envelope_frame`:
- no-panic; clean error classes with **shape assertions** on each
`FrameError` variant: `ConnectionClosed` only when the prefix is
truncated or the body is short (a complete body must never yield it);
`InvalidFrame` only for `len == 0` or `len > MAX_FRAME_SIZE`;
`Json` only when the full body was present; `Io` impossible on a
Cursor; Ok only when `len > 0`, `len <= MAX`, body complete.
- allocation bound: the reader never allocates for `len > MAX_FRAME_SIZE`
or `len == 0` (rejection precedes `vec![0u8; length]` — verified by the
`InvalidFrame`/Ok shape partition).
- exact consumption accounting (4 + len bytes read for a valid frame).
- structural round-trip: every decodable envelope re-encodes via
`write_frame` and re-decodes to an equal envelope (structural —
serde_json without `preserve_order` does not preserve key order).
- serde contract: the decoded envelope serializes with `type`/`id`/
`payload` keys intact (`type` survives the rename).
- **trailing-byte probe**: for any input ending in a complete JSON
object body, a copy with one extra byte appended *and counted in the
length prefix* must fail with a `Json` error (serde_json's
from_slice rejects trailing non-whitespace) — catching any future
switch to a streaming/tolerant reader that would silently accept
junk after the JSON document.
Note on the probe: the length prefix counts **only the body**, not the
4-byte prefix — the first probe draft claimed `4 + len + 1` and tripped
its own shape assert (`ConnectionClosed`). The fuzzer would have found
that instantly; corpus replay found it first.
### Verification at adoption
- `cargo fuzz build` clean (nightly-1.95.0-nightly, cargo-fuzz 0.13.2,
libfuzzer-sys 0.4.13).
- Stable side untouched and green: `cargo test` 682 passed / 0 failed;
`cargo clippy --all-targets -- -D warnings` clean; `cargo fmt --check`
clean (main crate, fuzz workspace, and fuzz/shared).
- Corpus replay (`cargo test --manifest-path fuzz/shared/Cargo.toml`):
245 committed seeds through both invariant sets, green.
Commit: 80a37e9 (step-1 implementation, before the campaigns below ran).
### Campaign results (10-min detached runs per target, §7.6 runner)
**No crashes, no hangs, no leak/OOM findings.** Both campaigns ran to
their time budget via `run-detached.sh` and exited 0 (`-fork=1
-rss_limit_mb=2048 -malloc_limit_mb=2048 -timeout=25`;
`-max_len=65536 -dict=json.dict` for the envelope target); artifact
directories empty after both.
- `chunk_header` — saturated its state space in the first seconds and
spent the remaining ~10 min confirming stability: flat
`cov: 51 ft: 54` from job ~5 onward, 203.9M execs at ~350k exec/s,
`oom/timeout/crash: 0/0/0` on all 33 fork jobs. The pure-8-byte-parser
ceiling is fully enumerated.
- `envelope_frame` — still finding coverage when the budget ended:
2119 edges / 7246 features / 1664 in-memory corpus entries
(from the 245 committed seeds), ~7.08M execs at ~12k exec/s,
`oom/timeout/crash: 0/0/0` on all 33 fork jobs. Growth was still
positive (JSON structure exploration), so longer campaigns keep
paying; the 25 s timeout and both memory tripwires never fired.
The grown envelope corpus (1631 new entries) was **not** merged into
the committed seeds — per the §7.1 policy, grown corpora stay
gitignored; worth revisiting if a future scheduled campaign produces
inputs worth curating. §6.2-3 (pre-payload allocation window) remains
bounded as designed. §6.2 items 1/2/4/5 are stateful targets
(§7.4 step 3) and were not exercised by these parser-level campaigns.
---
## 7.8 Step 3: targets 3–5 (2026-09-27)
**Adopted.** Three more targets, all over already-`pub` APIs (the
`#[cfg(fuzzing)]` exposure never materialized — quinn/rustls-style
exposure has not been needed for any of the five targets):
| Target | Drives | Invariants |
|---|---|---|
| `manager_routing` | `ChannelManager` under `#[derive(Arbitrary)]` op sequences (`Route`/`Open`/`Adopt`/`Teardown`/`ClearAll`), 256 ops × 4 KiB payloads, live mux runner, drainer tasks | no-panic; **exact counter models**: harness parked/dropped counters must equal the manager's monotonic `early_arrival_count`/`dropped_unknown_chunks` after every op; parked bytes ≤ distinct-ids × 64 × 16 MiB (the documented per-channel bound, §6.2-1); post-sequence `clear_all` reconciles every routed byte against the drainers' totals (lossless bounded-buffer routing, sentinel semantics included) |
| `envelope_semantic` | `#[derive(Arbitrary)]` `EnvelopeKind` → the six `EventEnvelope` constructors → serde round-trip → `write_frame`/`read_frame` | no-panic; constructors emit exactly the wire-stable event-type constants; any `Value` payload survives a serde round-trip; `call.error` payloads always parse back as `CallError` (ADR-016 closed schema); structural `write_frame`/`read_frame` round-trip |
| `spec_parse` | raw bytes → `serde_json` → `OpRegisterRequest::from_json` (the `pub` wrapper over `rebuild_spec_for`) → `OperationRegistry::register` (attacker-shaped schemas **compile at registration**, CF-003) | no-panic; parse failures are clean `INVALID_INPUT`; `spec_to_json_pub` → `from_json` round-trips the rebuilt spec; registration is always `Result` (uncompilable schema = compile error, never silent); registry bookkeeping survives repeated attacker-shaped inserts |
New supporting pieces: `fuzz/shared/src/arbitrary_value.rs` (a bounded
`Arbitrary` impl for `serde_json::Value` — depth 6, width 6, sized
strings/keys), and 20 `spec_parse` + 4 `manager_routing` committed
seeds (the op-sequence corpus is tiny because structured `Arbitrary`
inputs self-generate; the four committed seeds are semantic fixtures
including the duplicate-adopt sequence that pinned the bug below).
### First real finding: duplicate adopt/open destroyed the live channel
The `manager_routing` campaign found it in its first minutes — the
§6.2 stateful thesis, confirmed: **the bug is invisible to
example-based tests** because it needs an *interleaving* (adopt → drop
write half → duplicate adopt → route) that no unit test thinks to
replay.
**Mechanism** (`src/channels/manager.rs`, `open_channel` and
`adopt_channel`, pre-fix):
```rust
if channels.insert(channel_id, state).is_some() {
return Err(ManagerError::ChannelExists(channel_id));
}
```
`HashMap::insert` **replaces** an occupied entry and returns the old
value — so the collision-detection idiom silently *installed the new
state and dropped the old `ChannelState`*, whose `demux_sender` was
the live channel's read half. A duplicate `adopt_channel`/`open_channel`
on an in-use id returned `Err(ChannelExists)` (correct) **while
destroying the live channel** (incorrect): readers saw a spurious EOF,
subsequently routed chunks were lost, and the mux write half kept
framing onto the transport for a channel the demux no longer fed.
Reproducibility: a failed adopt on a live channel whose mux pump had
finished (the write half dropped — as any caller does) re-registered a
pump, succeeded through `mux.register`, and hit the replacing insert.
The duplicate-adopt path is reachable whenever an open-op response is
replayed or a connection-owner race re-announces an id.
**Fix** (this commit): check `channels.contains_key(&channel_id)` and
return before any insert — the map is never touched on collision in
either `open_channel` or `adopt_channel`. Regression tests:
`adopt_channel_duplicate_id_leaves_live_channel_intact`,
`open_channel_duplicate_id_leaves_live_channel_intact` (both verify
routing survives a rejected duplicate; the first also verifies the
EOF sentinel remains the only EOF source).
**Also fixed in the harness** (fuzzer-vs-harness findings, not crate
bugs): `clear_all` returns the opener-ledger entries (ADR-047 §7),
not channel-map entries — adopted channels carry no ledger entry, so
`drained.len() == open_count()` is *not* an invariant; the monotonic
counters survive `clear_all`; and drainer-byte reconciliation needs
per-id supersession when a fresh stream replaces a drained one.
### Campaign results (10-min detached runs per target, §7.6 runner)
The smoke-tier runs above surfaced two harness-model defects worth
recording (the fuzzer as a harness-checker), fixed before the long
runs: the op-sequence driver's EOF-sentinel latch is **per-receiver,
not per-channel-id** — a fresh adopt installs a fresh reassembled
stream with no latched EOF, while a sentinel inside a *drained
early-arrival queue* latches the new receiver too
(`drain_early_arrivals` delivers it); and `Adopt` ops must respect
ADR-047 §5's odd/even split (an Accept-side manager adopts the peer's
odd ids; its own `Open` allocates even — a cross-range adopt/open
collision is a protocol violation, not a manager bug).
**Final 10-minute detached campaigns, all three targets (post-fixes,
post the `insert`-replace fix): every run exited 0 with an empty
artifacts directory — no crashes, hangs, OOMs, or leaks.**
- `manager_routing` — 1.06M execs (~1.9k exec/s; each exec replays up
to 256 manager ops on a live runtime), 33 fork jobs clean,
`oom/timeout/crash: 0/0/0` throughout. Final coverage 2499 edges /
11332 features / 461 in-memory corpus entries. The exact-counter and
parked-bytes invariants held across every adversarial sequence.
- `envelope_semantic` — 2.44M execs (~4.4k exec/s), 33 fork jobs
clean, coverage saturated at 1669 edges (constructor × payload shape
space is finite), corpus 765. All six event kinds round-tripped
structurally; `CallError` parse-back never failed.
- `spec_parse` — 10.8M execs (~19k exec/s), 33 fork jobs clean,
coverage 978 edges / corpus 1267. The registry compiled every
attacker-shaped schema it was handed (or rejected cleanly); the
`spec_to_json_pub` → `from_json` round-trip never diverged.
§6.2-1 (parked-bytes bound) is now encoded as an exact counter model
and held; §6.2-3 confirmed bounded at the parser level (§7.7). The
stateful coverage §7.4 step 3 called for exists and is clean.
---
## 7.9 No hosted CI — the standing policy (2026-09-28)
§7.2's two CI tiers (per-push smoke, scheduled campaign) and the
OSS-Fuzz step assume a hosted CI platform. This repo has none and will
not get one: the git host is a minimal gitea that serves git and
nothing else, deliberately — gitea/gitlab have had full-compromise CVEs
(app.ini and any plaintext DB tokens exfiltrated in one real incident
elsewhere), so the attack surface stays minimal. All verification and
campaigns are **manual, run from the separate publishing server**:
- **Corpus replay is the fuzz gate, as plain `cargo test`:**
`cargo test --manifest-path fuzz/shared/Cargo.toml` replays every
committed seed through the same invariant functions the fuzzer runs
— on stable, no nightly, no cargo-fuzz. It is part of the release
verification checklist (AGENTS.md). This is CI tier 3's deliverable
(§7.2), minus the automation.
- **Campaigns** run via the detached runner (§7.6) when wanted —
e.g. before a release, or after touching `src/protocol/wire.rs`,
`src/channels/wire.rs`, the demux loop, or `ChannelManager`.
- **OSS-Fuzz is out**: it requires hosted infrastructure and a public
repo posture this project does not have.
- Do not add Actions/Gitea-Actions/workflow files anywhere in this
repo, and do not reintroduce "when CI exists" language into docs —
point at this section instead.
## 8. Answering the "if / how" directly
- **If?** Yes — justified by position in the dependency graph, two stable
attacker-reachable wire formats, and two concrete DoS-class findings a
fuzz campaign would have caught (§6.2). The expected bug classes here
are panics and resource exhaustion, not memory unsafety (no `unsafe`
exists), which is precisely the profile where cheap fuzzing pays off.
- **How?** `cargo-fuzz` + `arbitrary`, a standard `fuzz/` workspace with
five targets (§7.1), committed seed corpora only, `#[cfg(fuzzing)]`
internals exposure where needed, two-tier CI with a 10 s-per-target
smoke on every push (§7.2), nightly confined to the fuzz job, AFL as an
optional secondary engine, bolero deferred, OSS-Fuzz after
stabilization. **Locally (agent sessions), campaigns always run via the
detached runner (§7.6): `setsid` + `nohup` + log file + poll, fork mode
on, soft RSS/malloc limits pinned — an OOM in a fuzz target must cost
the fuzzer, never the agent session.**
## 9. References
- RUSTSEC-2026-0037 / CVE-2026-31812 (quinn-proto remote DoS; maintainers'
fuzzing-coverage statement) — https://rustsec.org/advisories/RUSTSEC-2026-0037.html
- cargo-fuzz releases / Rust Fuzz Book — https://github.com/rust-fuzz/cargo-fuzz ·
https://rust-fuzz.github.io/book/
- arbitrary 1.4.2 and the 2025 speed analysis —
https://nnethercote.github.io/2025/08/16/speed-wins-when-fuzzing-rust-code-with-derive-arbitrary.html
- rustls fuzz targets + CI — https://github.com/rustls/rustls/tree/main/fuzz ·
https://github.com/rustls/rustls-fuzzing-corpora
- quinn fuzz targets (packet/streams/params) — https://github.com/quinn-rs/quinn/tree/main/fuzz
- quiche fuzzing + seed generation — https://github.com/cloudflare/quiche/tree/master/fuzz
- h2 MockIo e2e target — https://github.com/hyperium/h2/tree/master/fuzz
- prost fuzz (libfuzzer + AFL, FUZZING.md) — https://github.com/tokio-rs/prost/tree/master/fuzz
- s2n-quic bolero corpus-per-test pattern — https://github.com/aws/s2n-quic
- yamux allocation guard / libp2p validate-without-allocating —
https://github.com/libp2p/rust-yamux ·
https://github.com/libp2p/rust-libp2p (gossipsub `validate_rpc_limits`)
- rust-fuzz trophy case (finding-class taxonomy) — https://github.com/rust-fuzz/trophy-case
- OSS-Fuzz Rust integration — https://google.github.io/oss-fuzz/getting-started/new-project-guide/rust-lang/
- Scheduled CI fuzzing with cached corpora — https://depot.dev/blog/distributed-rust-fuzzing
Internal: ADR-014 (envelope framing), ADR-034 (chunk header), ADR-040
(backpressure/limits), ADR-042 (hub relay), ADR-031 (crate
decomposition); review 006 E-04 (early-arrival cap sizing note,
`src/channels/manager.rs:106-121`).
+6
View File
@@ -0,0 +1,6 @@
target
corpus/*
!corpus/chunk_header
!corpus/envelope_frame
artifacts
coverage
+1280
View File
File diff suppressed because it is too large. Load diff
+52
View File
@@ -0,0 +1,52 @@
[package]
name = "alkcall-fuzz"
version = "0.0.0"
publish = false
edition = "2021"
[package.metadata]
cargo-fuzz = true
[dependencies]
libfuzzer-sys = "0.4"
alkcall-fuzz-shared = { path = "shared" }
[dependencies.alkcall]
path = ".."
[[bin]]
name = "chunk_header"
path = "fuzz_targets/chunk_header.rs"
test = false
doc = false
bench = false
[[bin]]
name = "envelope_frame"
path = "fuzz_targets/envelope_frame.rs"
test = false
doc = false
bench = false
[[bin]]
name = "manager_routing"
path = "fuzz_targets/manager_routing.rs"
test = false
doc = false
bench = false
[[bin]]
name = "envelope_semantic"
path = "fuzz_targets/envelope_semantic.rs"
test = false
doc = false
bench = false
[[bin]]
name = "spec_parse"
path = "fuzz_targets/spec_parse.rs"
test = false
doc = false
bench = false
[workspace]
+53
View File
@@ -0,0 +1,53 @@
# alkcall fuzzing
libFuzzer targets for alkcall's wire formats (see
`docs/research/fuzzing.md` for the full rationale and campaign plan).
## Layout
- `fuzz_targets/` — one binary per target; thin `fuzz_target!` wrappers.
- `shared/` — the invariant logic, as a plain library so normal
`cargo test` (stable toolchain) can replay the committed corpora
through the same invariants (`fuzz/shared/src/*.rs` `#[cfg(test)]`
modules; quinn's CI pattern). The fuzz binaries are nightly-only;
the shared crate is stable-clean.
- `corpus/<target>/` — committed hand-made seeds (regenerate with
`python3 fuzz/gen_fuzz_seeds.py` from the repo root). Grown corpora
and artifacts are gitignored.
- `run-detached.sh` — mandatory runner for agent sessions: wraps
`cargo fuzz run` in `setsid` + `nohup` + log redirection so an OOM
in a target can never take down the agent host (research doc §7.6).
- `json.dict` — JSON token dictionary for the envelope targets.
## Targets
| Target | Drives | Invariants |
|---|---|---|
| `chunk_header` | `parse_header` / `write_header` (`channels/wire.rs`) | no-panic; round-trip identity; `TooLarge` iff `length > MAX_CHUNK_LEN`; consumption accounting (8 bytes); input never mutated |
| `envelope_frame` | `FrameFramedReader::new(Cursor).read_frame()` on a current-thread runtime | no-panic; never allocates > `MAX_FRAME_SIZE`; clean `FrameError` on truncation/oversize/zero-length; exact consumption; structural `write_frame` round-trip |
## Commands
`cargo fuzz build` and `cargo fuzz run` must be executed with `fuzz/`
(or deeper) as the working directory so rustup selects the pinned
nightly toolchain — the detached runner handles that itself.
```sh
# build (nightly, pinned by rust-toolchain.toml inside fuzz/; run from fuzz/)
cargo fuzz build
# agent sessions: detached campaign (never foreground; CWD-independent)
FUZZ_RUNTIME_SECS=600 fuzz/run-detached.sh chunk_header
FUZZ_RUNTIME_SECS=600 fuzz/run-detached.sh envelope_frame -max_len=65536 -dict=json.dict
# corpus replay through the invariants (stable toolchain, no nightly)
cargo test --manifest-path fuzz/shared/Cargo.toml
# coverage
cargo fuzz coverage <target>
```
The fuzz workspace is excluded from the main workspace and from the
published package (`exclude` in the root `Cargo.toml`); it pins its own
nightly toolchain via `rust-toolchain.toml` and does not affect the
crate's stable MSRV.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+1
View File
@@ -0,0 +1 @@
��������
View File
Whitespace-only changes.
+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@

Binary file not shown.
Loaded 100 of 1465 files, more files were not shown because too many files have changed in this diff. Show more