2 Commits
Author SHA1 Message Date
glm-5.3-flash d4635d28f0 perf: fixed-size struct fast path, integer union dispatch, zero-alloc read_next_borrowed
Targets the bench gaps from the 0.3.0 port review (commit dea96f0):
packet read was ~111-189x hand-rolled, chunk read ~18x.

- ReadPlan gains compile-time fixed_size (cached field-size sum).
  Fixed structs skip the cursor size walk entirely (one bounds check
  instead); fixed-size union variants skip the plan_walk_variant_size
  pre-pass, eliminating the double walk of variant bytes for the
  common SFTP-shaped case.
- CompositePlan::Union gains an int_keys dispatch table (pre-parsed
  u64 mapping keys); byte-discriminator unions dispatch on the raw
  integer instead of stringifying per read. Returned discriminator
  String unchanged (public API). String-keyed fallback preserved.
- Additive SequentialReader::read_next_borrowed returns the field
  name borrowed from the plan — zero allocs per field for hot loops.
  read_next stays the owned-name form (single source of truth).
- plan_walk_struct_size / union shared walk: per-field format! moved
  to the error path only.
- materialize: with_capacity for bytes arrays, arrays, and struct
  objects.

Benches (1024 chunks/iter, criterion, pre-review baseline vs now):
- read_packet_stream: 600 -> 246 µs (~2.4x; gap to hand 189x -> ~74x)
- read_chunk_stream: 104 -> 67 µs (~1.6x; 18x -> ~11x)
- write/validate groups unchanged (within noise)
- engine_compile +8% (int_keys table + fixed-size precompute), still
  one-shot

Verification: 566 tests pass, clippy -D warnings clean, wasm32 build
green. Bench baselines saved as pre-review/post-review.
2026-09-03 17:28:20 +00:00
glm-5.3-flash dea96f0195 bench: port wire_vs_bast from alktty, add union + validate_bytes groups
The wire_vs_bast bench originated in alktty as the uncommitted curiosity
probe that surfaced review #004's 400x read gap (the driver for the 0.3.0
compiled-forms release). It now lives here so alktype owns its perf
story; the alktty-only async roundtrip group (tokio ChunkReader/
ChunkWriter over a duplex pipe) was dropped — that measures alktty's I/O
stack, not this engine. The alktty copy is deleted.

Groups:
- read_chunk_stream / write_chunk_stream — the original ChunkHeader
  shape, byte-identical methodology, so numbers stay comparable with the
  historical series (400x → ~18x on read p64).
- read_packet_stream (new) — SFTP-shaped byte-discriminator union
  (Read/Write variants, Write carries a length-prefixed bytes field):
  exercises CompositePlan::Union dispatch + variant walks + variable
  reads, the case ADR-011's framing argument was about. The alktype
  consumer pattern follows the documented FieldValue::Union contract;
  a pre-measurement parity check locks the pattern (variant walk size
  + disc size == packet size) so the stream loop can't drift silently.
- validate_stream (new) — engine.validate_bytes per buffer (materialize
  + ValidationPlan walk), the read+validate-on-untrusted-stream shape
  alkcall cares about; closes the phase-7/8 bench deferral.
- one-shots — engine_compile, sequential_reader_new, layout_build.

criterion 0.7 dev-dep (default-features off). Benches don't affect the
wasm gate (bench targets never compile under wasm32-unknown-unknown).

Numbers (1024 chunks/iter, Xeon D-1521, shared box — ±10% noise):
- read p64: hand 5.7 µs / alktype 104.8 µs (~18x; parity with the
  phase-2/8 record of 98-99 ns/chunk)
- read p4k: hand 12.1 µs / alktype 107.5 µs
- write p64: hand 14.0 µs / alktype 37.5 µs; p4k: 324/362 µs
- packet read p64: hand 3.3 µs / alktype 622.6 µs (~189x — dominated by
  per-field String allocs + variant reader construction; the read_next
  (String, FieldValue) signature is pinned by the semver contract)
- validate: header 458 ns/chunk, packet p64 2.23 µs, packet p4k 47 µs
- one-shots: compile 615 µs (meta-schema dominated), reader_new 15.7 ns,
  layout_build 343 ns

Verification: cargo test --release (566 tests green), clippy
--all-targets -D warnings, wasm32-unknown-unknown build clean.
2026-09-03 08:45:13 +00:00