Targets the bench gaps from the 0.3.0 port review (commit dea96f0):
packet read was ~111-189x hand-rolled, chunk read ~18x.
- ReadPlan gains compile-time fixed_size (cached field-size sum).
Fixed structs skip the cursor size walk entirely (one bounds check
instead); fixed-size union variants skip the plan_walk_variant_size
pre-pass, eliminating the double walk of variant bytes for the
common SFTP-shaped case.
- CompositePlan::Union gains an int_keys dispatch table (pre-parsed
u64 mapping keys); byte-discriminator unions dispatch on the raw
integer instead of stringifying per read. Returned discriminator
String unchanged (public API). String-keyed fallback preserved.
- Additive SequentialReader::read_next_borrowed returns the field
name borrowed from the plan — zero allocs per field for hot loops.
read_next stays the owned-name form (single source of truth).
- plan_walk_struct_size / union shared walk: per-field format! moved
to the error path only.
- materialize: with_capacity for bytes arrays, arrays, and struct
objects.
Benches (1024 chunks/iter, criterion, pre-review baseline vs now):
- read_packet_stream: 600 -> 246 µs (~2.4x; gap to hand 189x -> ~74x)
- read_chunk_stream: 104 -> 67 µs (~1.6x; 18x -> ~11x)
- write/validate groups unchanged (within noise)
- engine_compile +8% (int_keys table + fixed-size precompute), still
one-shot
Verification: 566 tests pass, clippy -D warnings clean, wasm32 build
green. Bench baselines saved as pre-review/post-review.