70 Commits
Author SHA1 Message Date
glm-5.3-flash b10e980346 Release 0.4.0: fuzzing wave findings, MaxLengthReserved encoding, reservation read/write API
- VariableEncoding gains MaxLengthReserved (public enum, semver-major
  for 0.x; breaks exhaustive matches) — W3-3 root fix
- read_field/write_field correctly treat aligned maxLength reservations
  as raw zero-padded/NUL-trimmed windows (W3-3, a0dd3d2)
- plan_read_array bounds check for truncated fixed-stride arrays (W2-1)
- data_access::read_reservation{,_string}/write_reservation public API
- 0.4.0 changelog entry; fuzz Cargo.lock version sync

Verification: cargo test --release 573 passed/0 failed; clippy
--all-targets -D warnings clean; wasm32-unknown-unknown release build
clean; fuzz corpus replay 30/30; cargo doc --no-deps clean;
cargo publish --dry-run --allow-dirty (67 files, 1.3MiB) verified
2026-09-30 12:51:56 +00:00
glm-5.3-flash f809f6ca7e docs: fuzzing wave-3 status — release-budget campaigns, W3-1/2/3 findings, all five targets in-tree
§Status now covers waves 1-3. §5 gains the release-budget numbers
(bast_compile 818k execs still growing; data_access 11.3M execs with
value-profile lifting the saturated edge count 379 -> 684;
read_opseq/layout_build 63k execs each with the stateful search
active through budget; validate_pair 37k execs clean after the fixes).
§3 target-5 marked implemented; §4 records the 48-seed menu; §6
gains candidate 8 (the W3-3 aligned maxLength reservation bug,
confirmed+fixed) and rewrites the evidence summary; §7 wave 3 is
checked off — corpus replay 30/30 stands as the gate.

Verification: cargo doc clean; 573 main-crate tests + 30 replay tests
green; clippy -D warnings clean (crate + shared).
2026-09-30 08:55:30 +00:00
glm-5.3-flash a0dd3d2de4 fix: W3-3 — read_field/write_field misread aligned maxLength reservations
The running validate_pair campaign found a third crash: in aligned
mode a maxLength reservation (ADR-003 strategy 2, VARCHAR(N)) stores
RAW zero-padded data with no length prefix — materialize and
validate_bytes implement exactly that — but read_field read the entry
through data_access::read_string, i.e. parsed the window's first four
bytes as a u32 length prefix. Raw reservation bytes that look like a
large prefix then fail bounds with Access while validate_bytes says
Ok: the validate⇒read agreement lattice breaks on every aligned
maxLength string/bytes field (any schema declaring maxLength in
aligned mode). write_field had the same mismatch (prefix+data into a
raw window).

Engine fix:
- VariableEncoding gains MaxLengthReserved (additive variant, ADR-003
  strategy 2). OffsetMap::compute records it for maxLength fields with
  the default encoding; maxLength+offset-indirect stays OffsetIndirect
  (the pair read is intentional, the window reserves max_len bytes),
  preserving the W3-1 combination semantics.
- read_field String/Bytes arms dispatch on MaxLengthReserved → new
  data_access::read_reservation_string / read_reservation (raw window
  inside-buffer check + NUL trim — the materializer's exact semantics).
- write_field dispatches → new data_access::write_reservation (zero-
  pads the window, rejects oversized values with Access).
- materialize_aligned reads MaxLengthReserved through the same new
  read_reservation paths (single source of truth; replaces the inline
  trim logic with an identical implementation).
- offset_map compute rejects a MaxLengthReserved encoding reaching the
  walk with a clean Offset error (recorded, never declared).
- builder round-trips: MaxLengthReserved serializes via maxLength (the
  document form), never as an encoding value.
- three engine regression tests: raw-not-prefixed read, zero-pad
  write + oversize rejection, validate⇒read_field agreement.
- fuzz/shared validate_pair invariant updated: the W3-1
  shorter-than-reservation exemption now applies to offset-indirect
  only; reservations assert the full window in-bounds (fixed engine).
- corpus regenerated for generator-consistent numbering (seeds 037-044
  relabeled; W3-1/W3-2 artifacts remain 044/045-047 → now 044, 048-050
  region) — 48 seeds, replay 30/30 green.

Verification: main crate 573 tests pass; clippy -D warnings clean
(crate + shared); wasm clean; cargo fuzz build clean.
2026-09-30 08:20:13 +00:00
glm-5.3-flash b7ead99724 fuzz: W3-2 — serde_json parse-side one-ulp float drift pinned with slack
The restarted validate_pair campaign found a second crash: materialize
produces f64 0x5bffffffffffffff, serde_json emits the shortest repr
1.4536774485912136e+135, and the non-`float_roundtrip` parse side
(lexical concise-float over re-parsed digits) lands one ulp low —
probe-verified upstream of this crate (ryu's own float parser accepts
the same digits exactly; the std parser is exact; only serde_json's
concise reparse drifts). The harness's structural serde round-trip
assertion assumed Value equality holds for every finite f64 — upstream
parse-side drift breaks that assumption on adversarial magnitudes.

- assert_values_agree_with_ulp_slack replaces the bare Value equality:
  keys/shapes exact, numbers equal-or-within-one-ulp (bit diff ≤ 1)
- the exact artifact bytes pinned as corpus seed-047, plus
  deterministic minimal forms as seed-045/seed-046 (45→48 seeds)
- upstream note: enabling serde_json's float_roundtrip feature would
  remove the drift; the crate pins serde_json default features +
  preserve_order by design, so the slack is the honest pin

Verification: corpus replay 30/30 green (48 seeds), fuzz build clean,
clippy -D warnings clean.
2026-09-30 07:43:02 +00:00
glm-5.3-flash 9ca9922fd4 fuzz: W3-1 — pin the indirect-reservation invariant; harness over-assertion fixed, reproducers committed
The running validate_pair campaign found the first wave-3 crash
(artifact crash-e40d...): the harness invariant 'validate_bytes Ok ⇒
every offset-map leaf's range.end ≤ buffer.len()' is WRONG for
offset-indirect entries. In aligned mode a maxLength reservation
contributes its full window to the layout (menu 2's total is 68), while
the {data_offset, data_length} pair is absolute — the data may live
anywhere in the buffer and the all-zero pair {0,0} over a 64-byte
buffer validates and reads fine. The wave-1 data_access bounds
partition is the real contract; the new invariant exempted the read
side but over-asserted the window. Harness-bug, not engine-bug.

- invariant now splits: non-indirect leaves keep the full window
  assertion; offset-indirect leaves assert only the read_ok ⇒ pair
  agreement (the pointed-to window sits inside the buffer)
- the exact artifact bytes pinned as a regression test
  (validated_buffer_may_be_shorter_than_the_indirect_reservation_window)
  and as corpus seed-044; the generator emits the same shape
  deterministically (44→45 seeds)
- PairInput fields made pub for out-of-crate triage probes

Verification: corpus replay 30/30 green; clippy -D warnings clean.
2026-09-30 07:16:53 +00:00
glm-5.3-flash aef8d9f6ab fuzz: wave 3 — validate_pair two-input harness, 44 seeds, release-budget campaigns next
Target 5 (§3): compile an attacker schema (10-lane menu incl. raw JSON
bytes lane) in both modes, then hammer the hostile buffer through
validate_bytes, an independent materialize_packed/materialize_aligned,
read_field over every offset-map leaf, junk field paths, and the packed
sequential walk under the spin bound.

Invariants coded (per §3 target 5): mode agreement (aligned Ok ⇒ packed
Ok; packed-Ok/aligned-Err only for the documented ADR-006/ADR-008/
offset-indirect rejections), the materialize⇄validate_bytes verdict
lattice with verbatim error propagation, unknown-path echo, serde
round-trip of materialized output, non-finite-float Access pinning,
out-of-range enum Validation pinning, and the record-count spin bound.

44 committed seeds (menu/raw lanes × valid/valid, hostile-schema/
valid-bytes, valid-schema/hostile-bytes incl. a per-prefix truncation
sweep, NaN/Inf, enum 99, spin fixtures, mode-agreement pins), hand-
encoded against the pinned arbitrary 1.4.2 derive layout and pinned by
decode tests. The aligned maxLength-reservation offset pin (s@8..72,
tail@72, total 76) caught a fixture assumption error pre-commit.

Verification: corpus replay 29/29 green (44 new seeds decode+replay),
main crate 570 tests pass, clippy -D warnings clean (crate + shared),
wasm clean, cargo fuzz build clean. Hand-run drives (indirect pair
escape, enum-Validation, unknown discriminator, trailing garbage) all
held.
2026-09-30 07:02:35 +00:00
glm-5.3-flash 16b9023f60 fuzz: wave 2 — stateful read_opseq + layout_build targets, seeds, one engine fix
Targets 3-4 of docs/plans/fuzzing.md, per the sibling layout:

- fuzz/shared/src/read_opseq.rs — SequentialReader op sequences
  (Next/NextBorrowed/Field/Reset/End, Arbitrary-derived) over hostile
  buffers under the fixed packed schema menu. Invariants: cursor
  discipline (failed read leaves position untouched, state replay
  deterministic), None sticky at plan end, plan-order full walks with
  a spin bound, read_field leaves a usable reader, ADR-007 reader
  independence (shared Arc, isolated cursors), and the plan §6-1
  record-count ≥4-verified-bytes bound encoded as an explicit End-op
  assertion.
- fuzz/shared/src/layout_build.rs — LayoutBuilder::build with
  adversarial var_sizes over a five-schema menu (string/bytes, nested
  struct, byte-disc union, record+array, fixed control). Invariants:
  Offset-class failures only, position disjointness + total-size
  bounds, variable fields record their 4-byte prefix, failed writes
  leave the buffer byte-identical, write→read pair round trip.
- derive_var_sizes discovers the synthetic keys ('p.__discriminator')
  the builder actually wants by parsing the quoted key from the
  Offset reason.
- 73 committed seeds (58 read_opseq + 15 layout_build) hand-encoded
  against the pinned arbitrary 1.4.2 derive layout (4-byte LE
  multiply-shift variant selectors, keep-going vec elements,
  take-rest last field) and pinned by decode_lands_on_the_intended_variants
  replay tests; gen_fuzz_seeds.py mirrors the encoders.
- Engine fix (finding W2-1): plan_read_array returned Ok for a
  fixed-stride array whose count*stride window extended past the
  buffer — the struct/union arms bounds-check, the array arm did not;
  a truncated array deferred the failure to the next field (wrong
  path) or masked it entirely as an Ok walk. Now an Access error
  naming the array, regression test in sequential_reader.rs.
- Packed-mode 'encoding: offset-indirect' pinned as the documented
  inline-length-prefix no-op (finding W2-2, bast-format.md Default
  strategy selection); open design question recorded as plan §6-7.

Verification: fuzz corpus replay 19/19; main crate 570 tests incl.
the new regression; clippy -D warnings clean (crate + shared); wasm
build clean; cargo fuzz build clean (nightly confined to fuzz/).
Smoke campaigns (10 min detached each): read_opseq 52.1k execs exit 0
empty artifacts, layout_build 42.4k execs exit 0 empty artifacts; no
crash/oom/timeout on any fork job.
2026-09-30 06:26:05 +00:00
glm-5.3-flash 752e36b526 docs: fuzzing wave-1 status — campaigns clean, seeds count, dict fix
- bast_compile smoke: 543k execs, 10830 edges, coverage still growing
  at budget end; data_access smoke: 3.5M execs, saturated at 379 edges
- both exited 0, empty artifact dirs, oom/timeout/crash 0/0/0
- json.dict: libFuzzer's parser rejects \u escapes and unquoted tails
  (caught at campaign launch, not by the fuzzer)
- plan doc §3/§5/§6/§7 updated with wave-1 results
2026-09-30 05:16:11 +00:00
glm-5.3-flash ed41d77e72 fuzz: wave 1 — infra + bast_compile/data_access targets, seeds, corpus-replay gate
- fuzz/ workspace (nightly-pinned subtree, own [workspace]), copied
  from the alkhttp/alkcall pattern: thin fuzz_target wrappers,
  stable-toolchain shared crate holding the invariant logic, detached
  runner, seed generator, json.dict, README
- bast_compile: AlkTypeEngine::compile both modes over attacker BAST
  JSON; meta-schema gate ordering, always-Result, fixed-size leaf
  metadata partition, json_schema lane
- data_access: the hand-rolled decode core over raw bytes at
  attacker-chosen offsets; bool strictness, UTF-8 discipline, bounds
  partitions, indirect {offset,length} pair contract, write-side
  no-touch-on-failure + write/read round trips
- 136 committed seeds (38 + 98) via fuzz/gen_fuzz_seeds.py
- root Cargo.toml: explicit [workspace] exclude=[fuzz]; publish
  exclude gains fuzz/
- AGENTS.md verification checklist gains the corpus-replay gate
- .gitignore: fuzz artifacts + grown-corpus pattern

Verification: cargo test 569 pass; clippy -D warnings clean; corpus
replay 4/4 green (136 seeds); cargo fuzz build clean (nightly
confined to fuzz/)
2026-09-30 05:03:07 +00:00
glm-5.3-flash 8d779e7672 docs: fuzzing plan for alktype (5 targets, 3 waves, sibling pattern) 2026-09-30 04:31:26 +00:00
glm-5.3-flash 9803d3b768 Pre-publish review #008: gate int_keys on canonical keys, restore no-prealloc array rule
Review #008 (docs/reviews/008-pre-publish-review.md) audits the two
post-#007 unreviewed commits (dea96f0 bench port, d4635d2 perf) before
the first crates.io publish of 0.3.0.

- F1: the int_keys integer dispatch accepted non-canonical mapping
  keys ("01", "+1" parse as u64 1) — the reader dispatched disc 1
  while the materializer, validation plan, and tunion rejected the
  same buffer. compile_int_keys now builds the table only when every
  key is canonical (v.to_string() == key); otherwise the string
  fallback applies (agreement restored, perf kept for canonical
  mappings). Tests: r8_non_canonical_mapping_key_disables_int_dispatch,
  r8_canonical_mapping_keys_keep_int_dispatch.
- F2: d4635d2 reintroduced Vec::with_capacity(count) on both array
  materializers. Bounded per array by MAX_ARRAY_ELEMENTS but nesting
  compounds: probe (counting allocator) measured ~477 MB simultaneous
  allocation from a ~1 KB schema + empty buffer (100-level stride-0
  chain, all legal under the caps). H1's layer-1 rule restored:
  Vec::new() + push. Bench unchanged (packet read 220µs vs 246µs
  baseline). Test: r8_deeply_nested_stride0_array_rejects_before_bulk_prealloc.
- N3a disposition (review #007's deferred item): BastField::synthetic
  (pub(crate), zero callers, #[allow(dead_code)]) deleted; the seven
  source() accessors are public API and stay (semver decision —
  removal needs an explicit ask); resolve_typeref_as_def's inline
  struct/union/enum arms probe-verified reachable (inline struct
  union variants are legal) — kept.
- CHANGELOG: 0.3.0 entry (compiled forms, breaking surface, hardening
  fixes, coverage). README: ReadPlan/ValidationPlan roles, union
  conventions, untrusted-schema bounds.

Verification: 569 tests green (491 lib + 17 + 34 + 15 + 12, + 2
ignored doctests), clippy -D warnings clean, cargo doc 0 warnings,
wasm32 build green, cargo publish --dry-run clean.
2026-09-07 10:58:49 +00:00
glm-5.3-flash d4635d28f0 perf: fixed-size struct fast path, integer union dispatch, zero-alloc read_next_borrowed
Targets the bench gaps from the 0.3.0 port review (commit dea96f0):
packet read was ~111-189x hand-rolled, chunk read ~18x.

- ReadPlan gains compile-time fixed_size (cached field-size sum).
  Fixed structs skip the cursor size walk entirely (one bounds check
  instead); fixed-size union variants skip the plan_walk_variant_size
  pre-pass, eliminating the double walk of variant bytes for the
  common SFTP-shaped case.
- CompositePlan::Union gains an int_keys dispatch table (pre-parsed
  u64 mapping keys); byte-discriminator unions dispatch on the raw
  integer instead of stringifying per read. Returned discriminator
  String unchanged (public API). String-keyed fallback preserved.
- Additive SequentialReader::read_next_borrowed returns the field
  name borrowed from the plan — zero allocs per field for hot loops.
  read_next stays the owned-name form (single source of truth).
- plan_walk_struct_size / union shared walk: per-field format! moved
  to the error path only.
- materialize: with_capacity for bytes arrays, arrays, and struct
  objects.

Benches (1024 chunks/iter, criterion, pre-review baseline vs now):
- read_packet_stream: 600 -> 246 µs (~2.4x; gap to hand 189x -> ~74x)
- read_chunk_stream: 104 -> 67 µs (~1.6x; 18x -> ~11x)
- write/validate groups unchanged (within noise)
- engine_compile +8% (int_keys table + fixed-size precompute), still
  one-shot

Verification: 566 tests pass, clippy -D warnings clean, wasm32 build
green. Bench baselines saved as pre-review/post-review.
2026-09-03 17:28:20 +00:00
glm-5.3-flash dea96f0195 bench: port wire_vs_bast from alktty, add union + validate_bytes groups
The wire_vs_bast bench originated in alktty as the uncommitted curiosity
probe that surfaced review #004's 400x read gap (the driver for the 0.3.0
compiled-forms release). It now lives here so alktype owns its perf
story; the alktty-only async roundtrip group (tokio ChunkReader/
ChunkWriter over a duplex pipe) was dropped — that measures alktty's I/O
stack, not this engine. The alktty copy is deleted.

Groups:
- read_chunk_stream / write_chunk_stream — the original ChunkHeader
  shape, byte-identical methodology, so numbers stay comparable with the
  historical series (400x → ~18x on read p64).
- read_packet_stream (new) — SFTP-shaped byte-discriminator union
  (Read/Write variants, Write carries a length-prefixed bytes field):
  exercises CompositePlan::Union dispatch + variant walks + variable
  reads, the case ADR-011's framing argument was about. The alktype
  consumer pattern follows the documented FieldValue::Union contract;
  a pre-measurement parity check locks the pattern (variant walk size
  + disc size == packet size) so the stream loop can't drift silently.
- validate_stream (new) — engine.validate_bytes per buffer (materialize
  + ValidationPlan walk), the read+validate-on-untrusted-stream shape
  alkcall cares about; closes the phase-7/8 bench deferral.
- one-shots — engine_compile, sequential_reader_new, layout_build.

criterion 0.7 dev-dep (default-features off). Benches don't affect the
wasm gate (bench targets never compile under wasm32-unknown-unknown).

Numbers (1024 chunks/iter, Xeon D-1521, shared box — ±10% noise):
- read p64: hand 5.7 µs / alktype 104.8 µs (~18x; parity with the
  phase-2/8 record of 98-99 ns/chunk)
- read p4k: hand 12.1 µs / alktype 107.5 µs
- write p64: hand 14.0 µs / alktype 37.5 µs; p4k: 324/362 µs
- packet read p64: hand 3.3 µs / alktype 622.6 µs (~189x — dominated by
  per-field String allocs + variant reader construction; the read_next
  (String, FieldValue) signature is pinned by the semver contract)
- validate: header 458 ns/chunk, packet p64 2.23 µs, packet p4k 47 µs
- one-shots: compile 615 µs (meta-schema dominated), reader_new 15.7 ns,
  layout_build 343 ns

Verification: cargo test --release (566 tests green), clippy
--all-targets -D warnings, wasm32-unknown-unknown build clean.
2026-09-03 08:45:13 +00:00
glm-5.3-flash 557a0d791e Review #007: fix YAML frontmatter — unquoted colon in reviewer value broke parsing 2026-09-03 06:56:38 +00:00
glm-5.3-flash 120c05cd60 Review #007: expand brace shorthand in reviewed_artifacts frontmatter 2026-09-03 06:54:27 +00:00
glm-5.3-flash cb952c9bf3 Review #007: record post-fix coverage (91.66% lines, +0.99) 2026-09-03 06:48:07 +00:00
glm-5.3-flash 7e5e58aa1b Close review #007 coverage holes C1-C3, L1, L2 (locking tests)
- C1: packed validate_bytes now exercised over all eleven primitive
  kinds (LE battery + BE subset + corrupted-bool rejection) — the plan
  materializer's i16..bool arms had zero public-path executions.
- C2: aligned validate_bytes over the default inline length-prefixed
  encoding (string + bytes; ADR-006 last-position rule honored).
- C3: ReadPlan::compile cycle rejection through a union mapping entry
  (compile_variant's own cycle arm — field-level cycles were already
  covered; this shape reaches the variant path). Arm confirmed
  executed in the post-fix coverage run.
- L1: builder.rs standard JSON-Schema conveniences locked with exact-
  JSON table tests, plus an end-to-end build_validator compile test.
- L2: tunion::read_field_discriminator's enum arm (both endians) —
  the last untested arm of the documented kind set (N1 parity).

docs/reviews/007-coverage-audit.md updated with per-finding
resolution blocks.

Verification: 488 lib + 78 integration tests green, clippy -D
warnings clean, wasm32-unknown-unknown build green.
2026-09-03 06:47:34 +00:00
glm-5.3-flash 844c199fb8 Fix F1/F2 from review #007: zero-progress guard parity + maxLength cap
- F1: materialize_plan_array (validate_bytes' packed path) now carries
  the zero-progress array guard the reader and legacy walker already
  had; validate_bytes no longer accepts an empty buffer against a
  stride-0 empty-struct-element array that SequentialReader rejects.
  Cross-consumer agreement test added (review #007 probe transcript).
- F2: MAX_LENGTH = 2^26 cap on the maxLength annotation — the N2
  dual-layer pattern (clean Schema parse error naming value+maximum,
  meta-schema "maximum": 67108864 so the published contract matches).
  Also closes the silent usize-overflow drop in parse_max_length.
- docs/reviews/007-coverage-audit.md records the full audit: per-file
  numbers, all classifications, and the N3a dead-surface list deferred
  to the pre-release review.

Verification: 477 lib + 78 integration tests green, clippy -D warnings
clean, wasm32-unknown-unknown build green.
2026-09-03 06:39:12 +00:00
glm-5.3-flash bb28ba3006 Resolve N3: maxLength is string/bytes-only (review #006)
- Parse gate in BastField::parse: maxLength on any kind other than
  string/bytes is a clean Schema error (records, arrays, inline
  structs, refs, union shared fields all covered; the choke point
  needs no ref-following since $defs entries are struct/union/enum)
- Meta-schema FieldDef: if kind in {string, bytes} else maxLength
  forbidden — the published alk.dev/bast/v1 contract matches the
  parser (N2 dual-layer pattern)
- M5's compute-side record maxLength arm became unreachable and was
  deleted (offset-indirect arm stays); the two superseded M5
  maxLength tests rewritten as the n3_* parse-rejection family
- ADR-006 remedy message tailored per kind: for records both
  annotated remedies are dead ends, so the error text points at the
  last-position fix only
- Docs aligned: bast-format.md (FieldDef meta-schema + FieldDef/
  Variable-Length Encoding prose), layout-engine.md (Strategy 2 +
  ADR-006 paragraph), schema-layer.md, ADR-003 §2/§3a amended,
  builder .max_length() doc
- Review #006: N3 resolved (all findings now closed), M5 update
  note, test-count bookkeeping note (in-session probes vs static
  counts), status lines flipped to fully resolved

Verified: 547 tests green + 2 ignored doctests in BOTH release and
default profiles (a stale debug artifact from an earlier session
masked one H3 roundtrip test in debug; clean rebuild passes both),
clippy -D warnings clean, cargo doc --no-deps zero warnings, wasm
build green.
2026-09-03 04:30:54 +00:00
glm-5.3-flash 0857ea1c23 Review #006: record L4/N1/M4 closure; add M5/M6 (fixed) and N3 (open)
- Resolution blocks on L4 (BTreeMap index + first-occurrence-wins
  pinned), N1 (tunion extended), M4 items 1-4 (aligned family, reader
  arms, data-access guards, dead-arm verdict with the structural note
  that the legacy packed walker serves only aligned fallbacks now).
- New findings: M5 (aligned record maxLength/offset-indirect silent
  corruption, resolved same-day), M6 (legacy walker field-disc union
  skipped shared fields, resolved same-day), N3 (packed record
  maxLength unenforced in validate_bytes, open - posture decision).
- Stats, resolution log, recommended order, and notes updated.
2026-09-02 20:28:05 +00:00
glm-5.3-flash 2eb086f400 Fix M6, extend M4 coverage over aligned/legacy-walk/reader paths (review #006)
- M6 (new finding, fixed): the legacy packed BAST-walker's field-disc
  union arm materialized only the discriminator field and started the
  variant immediately after it — silently reading the remaining shared
  fields' bytes as variant data whenever the union had any (probe:
  record<union> values produced {"handle": 5} where 5 was seq's value).
  The arm now walks all shared fields in order and starts the variant
  after the whole shared walk, matching H3's convention and the plan
  materializer's object shape (__discriminator + typed disc value +
  shared + variant). Reachable via aligned record/leaf paths only.
- M4 item 1: aligned-materializer test family — nested-struct
  recursion (3-level, previously 0 executions), maxLength trim in
  nested structs, invalid-UTF-8 Access error, offset-indirect
  out-of-bounds + data-after-sibling roundtrip, and records with
  struct/array/byte-disc-union/field-disc-union/wide-primitive values
  driving the legacy walker's previously-dead arms.
- M4 item 2: reader coverage — field-disc uint16/uint32/enum arms,
  byte-disc uint16/uint32 arms, nested-union variant size walk, and
  the public schema()/plan() accessors (all previously 0-execution).
- M4 item 3: data_access indirect-write tests at nonzero pair offset +
  data-region bounds refusal (the u32-truncation guards themselves
  need >4GiB slices and stay documented as defensively unreachable
  on 64-bit).
- M4 item 4 follow-through: OQ-001 rejection for struct elements and
  endian propagation for fixed elements locked with offset-map tests;
  the dead composite arms were removed in the previous commit.

Coverage after: materialize.rs 64.48→85.72% lines, TOTAL 89.59→90.60%.
Verified: 567 tests green (463 crate + 17 + 34 + 15 + 12 + 2 ignored),
clippy -D warnings clean, wasm build green, cargo doc zero warnings.
2026-09-02 20:26:21 +00:00
glm-5.3-flash 5f9793f9c0 Resolve L4/N1, fix M5, close M4 item 4 (review #006)
- L4: BTreeMap path->index for OffsetMap::get and PackedLayout::get;
  the linear scans behind the "random access" doc claim are gone.
  First-occurrence-wins preserved (BastStruct::parse doesn't reject
  duplicate names); locking tests in both modules.
- N1: tunion::read_field_discriminator now accepts uint16/uint32 disc
  fields (matching the reader's plan_discriminator_string_value set);
  one answer to "which field kinds can discriminate a union".
- M5 (new finding, fixed): aligned Record fields accepted maxLength /
  offset-indirect annotations, but the materializer always walks the
  inline count-prefixed form from the entry start — probe-verified
  silent corruption (record data crossed the reservation into the next
  field's bytes; validate_bytes accepted the corrupt buffer).
  Both annotations now rejected at compute with clean Offset errors.
  Parity-preserved from 0.2.0.
- M4 item 4: field_endian_for_element's Struct/Union arms and
  element_alignment were dead — the OQ-001 gate rejects every
  non-fixed-size element kind before either runs, so the phase-5
  "element's own endian is consulted" divergence never existed on any
  reachable path. Dead arms deleted; OQ-001 rejection for struct
  elements and endian propagation for fixed elements locked with tests.

Verified: 546 tests green (446 crate + 17 + 34 + 15 + 12 + 2 ignored),
clippy -D warnings clean, wasm build green.
2026-09-02 20:07:43 +00:00
glm-5.3-flash 0bc5a541ac Resolve N2: bound align annotations at 4096 (review #006)
align: 2^62 compiled and reported total_size = 2^63 — meaningless
layout output the consumer may act on, and the reachable path to the
MAX_ARRAY_BYTES cap used exactly this knob.

- MAX_ALIGN = 4096 (page granularity) in schema.rs, documented with the
  probe arithmetic
- parse_align returns Result and rejects over-cap values with a clean
  Schema error naming path/value/maximum — a silent clamp was rejected
  (it would change layout semantics without telling the consumer);
  both call sites thread the path, so standalone BastDoc::new (which
  never runs the meta-schema) is covered
- Meta-schema: "maximum": 4096 on StructDef.align and FieldDef.align —
  the published alk.dev/bast/v1/schema contract now matches the parser
- H1's byte-cap test retuned to align 4096 x count 2^16 = 2^28 > 2^26
  (the byte cap stays reachable under the new align cap)

Tests: 3 new (struct align above cap, field align above cap, align at
cap accepted). 511 tests green, clippy -D warnings clean, wasm32 build
green, cargo doc zero warnings.
2026-09-02 19:46:51 +00:00
glm-5.3-flash 8739d29550 Resolve L2/L3: real schema threaded through ReadPlan compile; dead locals dropped (review #006)
L2: compile() builds Arc<Value> once and threads &Arc<Value> down the
compile walk; every real ReadPlan carries the document at construction
— the Value::Null-placeholder-then-map-overwrite dance is gone. The one
remaining Null in wrap_leaf is documented as correct-by-construction
(anonymous synthetic wrapper, never escapes).

L3: materialize_plan_field drops its plan parameter (taken solely to
discard) and the field-disc union arm's disc_field/let _ pair is
deleted — the order-walk + by-name capture is the materializer's
correct design, as the finding's parity note described.

508 tests green, clippy -D warnings clean, wasm32 build green,
cargo doc zero warnings.
2026-09-02 19:43:51 +00:00
glm-5.3-flash dcfe9d16ff Fix M1/M2/M3: ADR-006 record gap + aligned record coverage + dead Struct arm (review #006)
M1: field_variable_kind (offset_map) now matches Record — a non-final
inline length-prefixed record field in aligned mode hits the ADR-006
rejection instead of computing silently corrupt offsets (probe-verified
clobber in the review: counts prefix at 0, id at 4).

M2: aligned record path locked with public-path tests —
materialize_aligned roundtrip (record<uint16>, wire arithmetic
asserted) and engine validate_bytes roundtrip + corrupted-buffer
rejection (record<uint32>). Record-as-last-field is the only safe
inline position post-M1.

M3: read_field's unreachable Struct arm (no struct-path entry ever
exists in an OffsetMap) replaced with a documented defensive Offset
error; doc comment states struct paths have no entry and the Offset
miss is the reachable composite failure. FieldValue::Struct (public
API, constructed by the packed reader) untouched.

Tests: 7 new (2 offset_map, 2 materialize, 1 engine M3 lock, plus the
roundtrip pair). 508 tests green, clippy -D warnings clean, wasm32
build green, cargo doc zero warnings.
2026-09-02 19:40:09 +00:00
glm-5.3-flash 5e74b991ac Fix H2: shared reference-graph guard for standalone walkers (review #006)
Cyclic or over-deep $ref graphs stack-overflowed the three standalone
schema walkers (OffsetMap::compute, LayoutBuilder::new,
materialize_aligned) — SIGABRT on probe, parity-preserved from 0.2.0.

- New src/walk_guard.rs: check_ref_graph() — one bounded walk over the
  reachable reference graph (depth cap 128 matching the plan compilers,
  path-scoped cycle set; diamonds allowed, cycles and 201-def chains
  rejected with the plan compilers' error wording)
- All three walkers run the guard at entry, before any recursion;
  materialize_aligned's is defense-in-depth (a cyclic doc can no longer
  produce an OffsetMap, but mismatched doc/map inputs must still fail
  cleanly)
- Behavioral side effect, net-positive: the guard eagerly parses every
  reachable def, so an invalid non-root def now surfaces at
  LayoutBuilder::new instead of build() — four H3 tests updated to
  expect the same Schema error earlier
- Test family: 12 new tests (walk_guard, offset_map, layout_builder,
  materialize) covering self/two-def/composite-carrier cycles, deep
  chains, and diamond non-rejection; no stack-overflow reproducers
  in-tree per the review's Methodology warning
- Stale "walkers have no cycle guard" statements updated in
  validation.md, 030 plan, ADR-012, and the engine gate comment

Verified: 501 tests green (423 + 17 + 34 + 15 + 12 + 2 ignored),
clippy -D warnings clean, wasm32 build green, cargo doc zero warnings.
2026-09-02 19:34:24 +00:00
glm-5.3-flash 05a2a42983 Fix H3: field-disc union wire convention — shared-then-variant (review #006)
Decision (recorded as an ADR-011 addendum): the packed-mode wire layout
for a field-name-discriminator TUnion is shared-then-variant — the
union's declared `fields` (disc + shared fields) first, then the
variant's own fields. Reader and materializer already implemented this;
LayoutBuilder was corrected from variant-only layout.

Enforcement in BastUnion::parse (the choke point every consumer
inherits — union roots at BastDoc::new, referenced unions at
resolve_ref):
- discriminator field must be declared in `fields`
- `fields` must not contain duplicate names
- variants must not re-declare shared fields (checked inline and
  through $ref resolution — parse chain now threads the doc root)
- the discriminator field must be the FIRST entry in `fields` (the
  reader reads the disc at the union start; a later position made it
  dispatch on the wrong bytes — H3 item 2, probe-verified)

Schemas relying on the old variant-only builder convention (variants
re-declaring shared fields) are rejected with a clean Schema error
naming the convention. Breaking for 0.2.0-era re-declaring schemas;
announced with 0.3.x.

- L5: FieldValue::Union::variant_start doc now states per-kind
  semantics (byte-disc: union_start + disc.offset + disc.size;
  field-disc: after the shared walk).
- L6: roundtrip test added (poc_roundtrip.rs) — LayoutBuilder write →
  SequentialReader read → materialize_packed → validate_bytes over a
  field-disc union with a second shared field and non-redeclaring
  variant; pins event.type@0/seq@1/handle@5, total 10.
- ADR-011: Status-block addendum recording the convention decision,
  the no-re-declare rule, and the breaking-constraint note.
- Review #006 updated: H3/L5/L6 resolution blocks, resolution log,
  recommended order.

Verified: 488 tests green (410+17+34+15+12, 2 pre-existing ignored),
clippy -D warnings clean, wasm32 build green, cargo doc zero warnings.
2026-09-02 18:37:51 +00:00
glm-5.3-flash 2d166f567b Fix H1: bound untrusted array counts; resolve L1 (review #006)
- Replace the three Vec::with_capacity(count) sites in materialize.rs
  with Vec::new() — validate_bytes on an adversarial count no longer
  OOM-aborts the process (AGENTS.md §3).
- New compile-time caps in schema.rs: MAX_ARRAY_ELEMENTS (2^16,
  enforced at BastArray::parse — the choke point every consumer
  inherits, bounds the walkers' per-element entry loops) and
  MAX_ARRAY_BYTES (2^26, enforced per walker against the mode-specific
  stride: compile_array, walk_array, compute_array_field).
- L1: fixed_composite_size/fixed_plan_size now return
  Result<Option<usize>>; unwrap_or_default() gone, overflow is a clean
  Schema error instead of silent stride-0.
- Zero-progress guard: stride-0 arrays whose elements consume 0 bytes
  (legal empty-struct elements) now error in plan_walk_variable_array_
  size and materialize_array_packed instead of looping count times.
- Tests: 8 new (parse/build/compile rejections, cap boundary,
  short-buffer clean error) + array_count_large_u64_parses_on_64bit
  rewritten to assert the new cap rejection. In-tree tests assert only
  the safe (compile-time) half per review #006's Methodology warning.
- Review #006 updated: H1/L1 resolution blocks, new finding N2
  (unbounded align annotations, found while re-deriving the cap
  arithmetic), resolution log, recommended order.

Verified: 482 tests green (405+17+34+14+12, 2 pre-existing ignored),
clippy -D warnings clean, wasm32-unknown-unknown build green.
2026-09-02 16:46:46 +00:00
glm-5.3-flash 27be01af93 Add review #006: 0.3.0 post-implementation review
Audits the shipped 0.3.0 code (commits ff85258..9949f91) for
correctness, untrusted-input discipline, 0.2.0 parity, code smells,
and coverage. Gates re-run green in-session (474 tests, clippy -D
warnings, doc, wasm32, publish --dry-run); coverage measured with
cargo-llvm-cov (89.59% lines / 84.80% fn).

Findings:
- H1: huge declared array count OOM-aborts validate_bytes (three
  Vec::with_capacity(count) sites; release-blocking)
- H2: cyclic $ref stack-overflows OffsetMap::compute /
  LayoutBuilder::new / materialize_aligned when driven standalone
  (engine gated, public walkers not; parity-preserved)
- H3: field-disc unions — builder (variant-only layout), reader
  (disc at union start), and materializer (position-correct shared
  walk) disagree on layout and discriminator position; needs a
  convention decision
- M1: ADR-006 check misses non-final inline Record fields (probe-
  confirmed)
- M2: aligned-mode record fields untested end-to-end
- M3: read_field's aligned Struct arm is unreachable dead code
- M4: coverage weak spots (materialize.rs 64.5% lines; aligned
  nested-struct recursion 0 executions through any test)
- L1-L6, N1: stride unwrap_or_default conflation, Null-schema
  placeholder, dead locals, linear-scan get, variant_start doc gap,
  missing field-disc roundtrip test, tunion/reader disc-kind split

Includes an operational warning: the H1/H2 reproducers OOM/stack-
abort the test harness — reproduce in an isolated process only.

Verification: file-only change, no code touched; suite green before
commit.
2026-09-02 15:06:26 +00:00
glm-5.3-flash 9949f914df Release v0.3.0: compiled forms — ReadPlan, owned BastDoc, LeafMeta, ValidationPlan, fingerprinting
Public API bump 0.2.0 -> 0.3.0 (the 030-compiled-forms plan is now
fully implemented; all eight phases landed).

- Cargo.toml: version 0.3.0. lib.rs re-exports complete (ReadPlan +
  sub-types, LeafMeta, OffsetEntry, ValidationPlan + sub-types).
- ADR-007 "Cost" rewritten to the Arc<ReadPlan> cost (15.7 ns) with
  the 0.2.0 "re-parse on demand" framing as a historical note
  (review #004 L2, the last loose end from that review).
- ADR-011/012 status blocks flipped to implemented; architecture
  README ADR table rows updated; layout-engine.md rewritten for the
  0.3.0 surface (engine-factory reader construction, OffsetMap
  OffsetEntry/LeafMeta/fingerprint section, owned BastDoc compute
  signature); SequentialReader module doc points at the engine
  factory. Reviews #004 and #005 flipped to closed.
- Bench re-run (alktty wire_vs_bast, 0.3.0 tree): read p64 98
  ns/chunk (parity with phase 2; hand-rolled 5.7 us/stream),
  layout_build 180 ns (was ~1.2 us — the phase-4 owned-doc cache
  removed the per-build re-parse, ~7x), sequential_reader_new 15.7
  ns, write p64 -3%, engine_compile unchanged (meta-schema
  validation dominates). No dedicated validate_bytes-stream bench:
  the phase-7 spot check (~0.2 us plan-validate vs ~0.6 us
  compile-per-call) stands; a dedicated bench is a follow-up if
  alkcall profiling motivates it.
- Downstream: alktty compiles against the path dep unchanged; alkcall
  has no dependency yet.

Verification (full block, all green): 474 tests; clippy -D warnings
clean; cargo doc zero warnings; wasm32 release build green; cargo
publish --dry-run clean at 0.3.0.
2026-09-02 09:24:30 +00:00
glm-5.3-flash 537a2170fb Fingerprint ReadPlan/OffsetMap: Hash + Eq + fingerprint() (ADR-012 §1/§4, plan phase 6)
- #[derive(Hash, Eq)] on ReadPlan, FieldPlan, CompositePlan, ReadKind,
  DiscriminatorPlan (schema: Arc<Value> hashes via serde_json Value
  Hash + Eq under preserve_order), and on OffsetMap (+ Clone;
  LeafMeta/OffsetEntry/ByteRange payload already Hash from phase 5 /
  this phase).
- fingerprint() -> u64 on both via std DefaultHasher (deferred
  decision 3 resolved: no new dep, not hot, cross-version stability a
  non-goal per ADR-012).
- Contract tests both sides: equal schemas -> equal PartialEq +
  fingerprint; field-kind / field-order / endianness changes each
  break equality and fingerprint; different root names over the same
  document fingerprint differently (ReadPlan).
- ValidationPlan already carries its own Hash/Eq/fingerprint +
  contract test (phase 7 landed early).

Verification: 474 tests pass (9 new fingerprint contract tests);
clippy -D warnings clean; cargo doc zero warnings; wasm32 release
build green.
2026-09-02 09:16:11 +00:00
glm-5.3-flash 255c8c493e OffsetMap carries LeafMeta; read/write_field dispatch on it (ADR-012 §2b, plan phase 5)
Prerequisite (review #005 M2): Hash added to Endian/VariableEncoding
derives (additive; fieldless Eq enums), and to ByteRange.

- New public types LeafMeta { kind, encoding, endian } (Copy + Eq +
  Hash) and OffsetEntry { range, meta } (start()/end() accessors),
  re-exported from lib.rs. Deferred decision 2 resolved: struct —
  get(path) -> Option<&OffsetEntry>, iter() -> (&str, &OffsetEntry).
  Storage: Vec<(String, OffsetEntry)>.
- LeafMeta computed at compute time with effective endian threaded
  through the aligned walk (container default -> field override,
  propagated into nested-struct probes and array elements via the
  referring field, matching the aligned materializer).
- engine read_field/write_field dispatch on the entry's LeafMeta:
  the per-access BastDoc re-parse + lookup_leaf_field walk +
  LeafFieldInfo are gone — the last two review #004 M1 sites.
- Parity note: lookup_leaf_field computed nested-struct defaults from
  the nested struct's own endian annotation; the map now agrees with
  the aligned materializer and packed ReadPlan (referring-field
  propagation). The old divergence (nested struct declaring endian
  under a field that also declares one) is closed; no test pinned it.
- Behavior change: read_field on a map-absent path (whole-struct
  field) errors Offset ("field not found") instead of Access
  ("composite types"); the composite-path test accepted either.
- materialize_aligned's four offset_map.get call sites updated to
  .range.start. alktty/alkcall untouched (bench never uses
  OffsetMap::get; alkcall has no dependency yet).

Verification: 465 tests pass (offset_map tests updated to the
OffsetEntry shape with per-kind LeafMeta expectations; engine test
for the old lookup walk rewritten to assert map entries carry the
LeafMeta); clippy -D warnings clean; cargo doc zero warnings; wasm32
release build green.
2026-09-02 09:13:48 +00:00
glm-5.3-flash b7c7dbe2a1 LayoutBuilder caches the owned BastDoc (ADR-012 §2a, plan phase 4)
The builder stores doc: BastDoc + endian (the doc_value: Value +
root_name: String cache is gone); new parses the typed tree once and
build walks &self.doc — the per-build BastDoc::new re-parse
(layout_builder.rs M1) is retired.

- build's root-is-struct re-check replaces its unreachable!() with a
  clean Schema error (AGENTS.md §3 never-panic; invariant unchanged —
  new already rejects non-struct roots).
- Boxing fallout: the builder now holds the full owned tree, so
  Layout::Packed boxes it (Box<LayoutBuilder>) to keep the engine's
  Layout enum variant sizes balanced (clippy large_enum_variant).
  layout_builder() still returns Option<&LayoutBuilder> via
  auto-deref; public API unchanged.

Verification: 465 tests pass unchanged (layout_builder.rs suites
drive new/build through the public API); clippy -D warnings clean;
wasm32 release build green.
2026-09-02 08:59:13 +00:00
glm-5.3-flash c583762352 Make BastDoc owned: drop Bast* lifetimes (ADR-012 §2a, plan phase 3)
Every Bast* type drops <'a>: &'a str -> String, &'a Value -> Value
(deferred decision 1: plain String/Value — the tree is built once;
Arc<str> name-sharing needs a bench justification that doesn't exist).
BastDoc::new(&Value, &str) still takes references in and clones into
owned storage; the doc gains Clone. resolve_ref/resolve_typeref/
resolve_typeref_as_def return owned types.

- Engine ownership flip: AlkTypeEngine holds the owned BastDoc
  (replacing bast_doc: Value + root_name: String; root_name()
  delegates to the doc), killing its three per-call BastDoc::new
  re-parses (aligned validate_bytes, read_field, write_field — the
  review #004 M1 pattern removed by construction; phase 5 retires the
  lookup_leaf_field walk itself). New public accessor root_name()
  (additive). Engine Send + Sync with the owned doc, asserted in the
  existing thread-share test.
- Bonus cleanup: materialize_typeref_packed's dead _field param
  dropped (phase 2 left it dangling). Under ownership, keeping it
  would force a deep Value clone per array element / record value /
  union variant via dummy_field_for. The param, dummy_field_for, and
  ty_source are gone; no behavior change (the arg was already
  ignored). BastField::synthetic keeps an owned-signature
  #[allow(dead_code)] definition (no remaining callers today).
- Consumers adapted: OffsetMap::compute(&BastDoc),
  materialize_aligned(&BastDoc, ...) (no lifetime), BuildCtx/
  ComputeCtx hold &'d BastDoc, tunion/discriminator name borrows,
  lib.rs module doc. LayoutBuilder's doc_value re-parse cache is
  unchanged pending phase 4.

Verification: 465 tests pass with zero test-logic changes (bast.rs
suites exercise every parser path through the public API); clippy
-D warnings clean; cargo doc zero warnings; wasm32 release build
green.
2026-09-02 08:31:58 +00:00
glm-5.3-flash 1641dab505 Document session-continuity rules in AGENTS.md
opencode ends the turn when an assistant message contains no tool call,
including analysis-only messages. Document the working fix (end bursts
with a tool call or a final report, land work incrementally, resume
without re-deriving) so every session inherits it.
2026-09-02 08:08:27 +00:00
glm-5.3-flash e5f1b9d825 Wire packed read path through ReadPlan (ADR-011 steps 2-4, plan phase 2)
SequentialReader now walks Arc<ReadPlan> instead of re-parsing the
BAST typed tree per field (the 400x read-path gap, review #004 H1);
materialize_packed walks the same plan, unifying the two packed
read-side consumers on one compiled form.

- SequentialReader::new(Arc<ReadPlan>) -> Self, infallible: the
  fallible BastDoc parse moved to ReadPlan::compile (phase 1). The
  reader holds the plan Arc + cursor state only; schema() returns the
  Arc<Value> retained on the plan (review #005 H2 — no
  self-referential struct); new plan() accessor exposes the shared
  plan.
- ReadPlan carries schema: Arc<Value> (set at compile; sub-plans hold
  a Null placeholder — only the root plan is handed out).
- materialize_packed(&ReadPlan, &[u8]): plan-walking packed
  materializer. The aligned path keeps walking BastDoc with the
  retained dummy_field_for/ty_source/materialize_typeref_packed
  helpers (phase 5 Scope Boundary: aligned structure walk is the
  permanent 0.3.0 design).
- Engine: Layout::Packed stores Arc<ReadPlan> alongside the builder;
  sequential_reader() is an Arc::clone (was a full-document Value
  clone); packed validate_bytes calls materialize_packed(&self.plan).
- Stride (deferred decision 4): FieldValue::Array now reports the
  true stride for fixed-size struct/nested-array elements (0.2.0
  returned 0); doc comment documents the behavioral change; no
  existing test asserted the 0, so none needed changing.
- Two parity subtleties found and preserved:
  (a) materialize_plan_composite unwraps the plan's anonymous
      single-field wrapper for primitive array elements/record values
      — without it, materialized records nest each leaf under a
      synthetic object (caught by the record parity test);
  (b) field-disc unions keep 0.2.0's materialized key order
      (__discriminator first), observable under preserve_order.
  Both are now covered by plan-phase tests or construction.

Bench (alktty wire_vs_bast, 1024 chunks/stream): packed read
2.27 us/chunk (review #004) -> 98 ns/chunk p64 / 100 ns/chunk p4k
(~23x; the 400x gap closes to ~17x vs hand-rolled 5.6 ns/chunk).
Residual gap is the per-field String allocation mandated by the
unchanged (String, FieldValue) read_next signature (2 allocs/chunk)
plus data_access bounds checks. sequential_reader() construction:
15.7 ns (was a whole-document clone).

Verification: 465 tests pass unchanged (the existing reader/
materialize/engine suites drive the rewrite through the public API —
only constructor call sites moved to ReadPlan::compile); clippy
-D warnings clean; cargo doc zero warnings; wasm32 release build
green.
2026-09-02 07:52:46 +00:00
glm-5.3-flash ff85258d03 Implement ReadPlan type + compile (ADR-011 step 1, plan phase 1)
Pure addition: the packed read-side compiled form (src/read_plan.rs)
and lib.rs wiring (module + re-exports of ReadPlan, FieldPlan,
CompositePlan, ReadKind, DiscriminatorPlan). No existing engine code
touched — phases 2-5 wire the plan into the reader/materializer/engine.

- Refined union shape (ADR-011 as refined by review #005):
  CompositePlan::Union { disc, shared, variants } with
  shared: Option<Box<ReadPlan>> for field-disc unions and
  variants: Vec<(String, CompositePlan)> — no VariantPlan/VariantKind,
  nested-union variants work by ordinary CompositePlan recursion
  (restores the 0.2.0 capability the POC rejected).
- by_name is BTreeMap (ADR-012 §1 Hash-derive prerequisite).
- True array strides (deferred decision 4): fixed struct/nested-array
  elements compute their real stride via fixed_composite_size;
  variable-length elements stay 0. 0.2.0 returned 0 for fixed struct
  arrays; that behavioral change rides the 0.3.0 bump (phase 2 will
  surface it through SequentialReader).
- Endianness: effective endian baked at every node. Parity lock: the
  plan propagates the referring field's effective endian into nested
  structs/unions — what the 0.2.0 packed reader/materializer actually
  do — and ignores nested containers' own endian annotations (the POC
  baked s.endian() there; latent divergence, never exercised by its
  equivalence tests). Nested-annotation tests lock this in.
- Untrusted input: compile carries its own depth cap (128) +
  definition-level cycle set (mirrors ValidationPlan::compile), so
  standalone compile is safe on adversarial docs: cyclic refs, deep
  chains, dangling refs, non-struct roots, and non-struct/union
  variants all surface as AlkTypeError::Schema, never a panic.
  Overflow-safe stride arithmetic (checked_mul).

Verification: 388 tests pass (355 existing + 33 new: every BastType
arm coverage, field-disc shared/nested-union compile shape, stride
computation, endian parity, cycle/depth/malformed rejection,
Send + Sync static-bound assertion); clippy -D warnings clean;
cargo doc zero warnings; wasm32-unknown-unknown release build green.

Next: phase 2 (SequentialReader + materialize_packed consume the plan).
2026-09-02 07:19:54 +00:00
glm-5.3-flashandopencode e4636e6a44 Implement ValidationPlan (ADR-012 §3, plan phase 7)
The compiled value-domain validation form: replaces the interpretive
BastDoc walk in validate_bytes with a compile-once-walk-many
constraint tree built at engine-compile time. This was the design
session + implementation ADR-012 §3 delegated; the shape decisions
are recorded in new ADR-012 §3a.

- New src/validation_plan.rs: ValidationPlan + ValidNode/ValidField/
  ValidVariant (Debug+Clone+PartialEq+Eq+Hash+Send+Sync),
  compile(&BastDoc) with eager $ref resolution, and a per-buffer walk
  with deferred error-path rendering (zero happy-path allocation,
  byte-identical error messages vs the 0.2.0 walker).
  fingerprint() via DefaultHasher, same as the phase-6 pattern.
- Compile-time graph safety: definition-level cycle set + depth cap
  (128) reject cyclic $ref graphs with AlkTypeError::Schema. The
  interpretive walker resolved refs lazily with no guard (stack-
  overflow hazard); diamond (shared) refs still compile.
- bast_validation.rs: interpretive walker retired (deleted);
  validate_value survives as a one-shot wrapper (compile + validate)
  for callers holding a doc without an engine.
- engine: Arc<ValidationPlan> built at compile in BOTH modes; the
  plan compile runs before the layout build and doubles as the
  engine's cyclic-ref gate (LayoutBuilder/OffsetMap struct recursion
  has no cycle guard; a cyclic doc previously overflowed there).
  validate_bytes walks the plan; new accessor validation_plan().
  validate_bytes signature unchanged.
- lib.rs: pub mod validation_plan + re-exports (ValidationPlan,
  ValidNode, ValidField, ValidVariant).

Verification: cargo test --release (355 pass, incl. parity suite,
fingerprint contract, cycle/depth rejection, Send+Sync + thread-share
assertions); clippy --all-targets -D warnings clean; cargo doc
zero warnings; wasm32-unknown-unknown release build green.

Co-authored-by: opencode <noreply@alk.dev>
2026-08-31 17:45:46 +00:00
glm-5.2 e461f01c97 Resolve review #005: refine 0.3.0 plan + ADR-011/012
Resolve all 11 findings from the 0.3.0 plan review (#005) in one
docs-only pass. No source changes; the crate still builds/tests at
v0.2.0. The one substantive decision change is M3 (per user
direction: ship ValidationPlan in 0.3.0, no more hedging); the rest
are spec corrections or pre-implementation refinements to types that
do not yet exist on main.

- H1: refine ADR-011 CompositePlan::Union to carry
  shared: Option<Box<ReadPlan>> (field-disc shared fields) and
  variants: Vec<(String, CompositePlan)> (drop VariantPlan/
  VariantKind). Plan phase 1 implements the refined shape.
- H2: plan phase 2 specifies ReadPlan stores schema: Arc<Value>
  (not &Value), avoiding the self-referential struct ADR-011
  rejects. Verified serde_json::Value: Hash + Eq holds with
  preserve_order, so phase 6 derives are not blocked.
- M1: nested-union support falls out of the H1 shape refinement
  (a variant can be CompositePlan::Union) — option (a) from the
  review, no behavioral drop vs 0.2.0, no Semver regression row.
- M2: plan phase 5 adds an explicit first sub-step to derive Hash
  on Endian and VariableEncoding in src/schema.rs (additive,
  semver-safe prerequisite the original plan omitted).
- M3: reverse the ValidationPlan deferral. ADR-012's "Deferring
  ValidationPlan" becomes "ValidationPlan — in scope for 0.3.0";
  new ADR-012 §3 commits the decision (compiled form, no per-buffer
  BastDoc walk, Hash + Eq + fingerprint()) and defers only the
  concrete shape to a follow-on design session + the plan's new
  phase 7. Plan gains phase 7 (ValidationPlan); old phase 7 (bump)
  renumbered to phase 8. ADR-011's Out-of-scope and Scope
  Boundaries bullets updated to point at ADR-012 §3. The deferral
  black hole this review's methodology flagged is closed: the work
  is committed with a concrete reactivation trigger, not hedged
  into an unplanned future.
- L1: plan phase 2 corrects the dummy_field_for/ty_source removal
  claim — only packed-side call sites go away; the helpers stay
  for the aligned materialize_leaf_at path.
- L2: plan phase 2 states the packed-vs-aligned
  materialize_typeref_packed split (packed gets a new plan-walking
  function; the existing function stays for aligned).
- L3: plan phase 5 adds a Scope Boundary note — aligned
  materialize's BastDoc structure walk is the permanent 0.3.0
  design; an AlignedPlan is out of scope, tracked as an OQ.
- N1: fix "back-comat" -> "back-compat" typo.
- N2: plan phase 1 verification adds the read_plan_is_send_sync
  static-bound assertion test ADR-011 requires.
- N3: Semver Contract table notes the Result drop on
  SequentialReader::new (Result<Self, AlkTypeError> -> Self)
  alongside the argument-type change.

Also: ADR-012 title -> "Plan Fingerprinting, ValidationPlan, and
Closing the Deferred M1 Sites in 0.3.0"; §3 (Fingerprinting
OffsetMap) renumbered to §4; README ADR table updated; review #005
gets a Resolution section recording how each finding was closed.

Verification (docs-only change, v0.2.0 unchanged):
  cargo test --release                     ok (310 crate + 86 integration + 2 doctests)
  cargo clippy --all-targets -- -D warnings  ok
  cargo doc --no-deps                      ok
2026-08-20 06:46:43 +00:00
glm-5.2 0e7921a02a Add review #005: 0.3.0 plan review
Cross-checks docs/plans/030-compiled-forms.md against the codebase,
ADRs 011/012, the POC on readplan-poc, and review #004.

Findings:
- H1: field-disc union shape is in neither ADR-011 nor the POC
- H2: schema() &Value on Arc<ReadPlan> is the self-referential
  pattern ADR-011 rejects
- M1: nested-union silent behavioral drop (POC rejects what 0.2.0
  accepts); deferral-black-hole pattern
- M2: Endian/VariableEncoding missing Hash derive (phases 5/6 break)
- M3: ValidationPlan deferral flagged for re-evaluation — the
  read+validate-on-untrusted-input case may be hotter than
  ADR-012's 'not a hot loop' dismissal accounts for
- L1/L2/L3: dummy_field_for wording, materialize_packed split,
  materialize_aligned BastDoc walk silence
- N1/N2/N3: typo, Send+Sync assertion test, Result drop on new

Includes a deferral-pattern scan methodology section surfacing
M1/L3/H2 as black-hole instances and confirming the plan's four
explicit deferred decisions are the healthy pattern.

Verification: file-only change, no code touched.
2026-08-20 06:09:57 +00:00
glm-5.2 2310f6cbd8 Propose ADR-012 + 0.3.0 implementation plan
ADR-012 bundles two pieces of work into the 0.3.0 release so the
crate ships one round of breaking changes, not two:

- Fingerprinting: #[derive(Hash, Eq)] + fingerprint() -> u64 on
  ReadPlan and OffsetMap. BTreeMap for ReadPlan.by_name (HashMap
  blocks Hash derive). Fingerprint contract: equal hashes => identical
  reads over identical bytes. Enables cross-run plan caching, alkcall
  hub/spoke schema handshake, schema-version diagnostics.
- Closing the deferred M1 sites via owned BastDoc (lifetime removal,
  scoped to LayoutBuilder/bast_validation/materialize_aligned/
  OffsetMap::compute) + extending OffsetMap with LeafMeta
  {kind, encoding, endian} for the aligned read_field/write_field paths.

Reframes the 'WritePlan' candidate from ADR-011's Future capabilities
section: the packed write-side compiled form is PackedLayout; the
aligned R/W compiled form is OffsetMap; the M1 fixes are 'cache the
parse' and 'extend the compiled form with leaf metadata', not 'add a
third compiled form.' Serves minimal-public-API-changes better than
a literal WritePlan type. ValidationPlan deferred (different shape,
not a hot loop).

The plan (docs/plans/030-compiled-forms.md) is the execution entry
point: seven phases ordered by dependency, each phase a session
boundary. Phase 1-2: ReadPlan (ADR-011). Phase 3: owned BastDoc.
Phase 4: LayoutBuilder M1 fix. Phase 5: OffsetMap LeafMeta. Phase 6:
fingerprinting. Phase 7: version bump + docs + verification. Includes
a semver contract table, deferred decisions, cross-phase invariants,
and the verification block.

ADR-011's Future capabilities section updated to point at ADR-012 for
the items moving into 0.3.0 and record the WritePlan reframe. README
ADR table gets ADR-012 as Proposed.

Verification: docs-only change; cargo test --release, cargo clippy
--all-targets -- -D warnings, cargo doc --no-deps unchanged (no source
touched).
2026-08-19 08:06:55 +00:00
glm-5.2 1037e68091 Accept ADR-011: compiled read plan for packed mode
Tighten framing per pre-acceptance review (no decision changes):

- Be precise about M1 coverage: closes the packed-side site
  (engine.rs:284 validate_bytes); the aligned-side M1 sites
  (engine.rs:334,467, layout_builder.rs:190) are a deliberate
  reversible bet, not a non-issue.
- Replace the 'two type walkers is symmetric with OffsetMap/
  PackedLayout' spin with an honest 'parallel typed tree, permanent
  maintenance tax, justified by ~20-50x composite-dispatch win on
  SFTP-shaped union-with-$ref-variants packets.'
- Move the 'determinism enables fingerprinting/cache/handshake' future
  work out of the positives list into a dedicated 'Future
  capabilities' section — it is a forward reference, not a current win.
- Clarify bast_doc: Value is retained unconditionally (unused in
  packed mode, still needed in aligned mode); add a Send+Sync test
  note for Arc<ReadPlan> shared from the Send+Sync engine.
- Tighten the 'no format! allocations' claim to 'no resolve_typeref,
  no BastDef::parse, no JSON node access on the happy path' (error-
  path format! remains, and is not the cost being removed).
- Add a 'POC coverage' section recording that the readplan-poc branch
  walked every BastType arm and confirmed ReadPlan covers all cases,
  including the union Byte/Field split and the array variable-stride
  (element_stride = 0) case.

Status flipped Proposed -> Accepted. README ADR table updated.

Verification: docs-only change; cargo test --release, cargo clippy
--all-targets -- -D warnings, cargo doc --no-deps unchanged (no source
touched).
2026-08-18 09:32:17 +00:00
glm-5.2 3184818c08 Propose ADR-011: compiled read plan for packed mode
Review #004 measured the packed read path at ~400x slower per chunk
than a hand-rolled codec, root-caused to SequentialReader re-parsing
BastDoc::new on every field read (the 're-parse on demand' framing
from ADR-007). The write path is competitive because it has a compiled
form (PackedLayout); the read path is the only mode/side pair without
one.

ADR-011 proposes ReadPlan — the packed read-side compiled form,
symmetric to OffsetMap (aligned R/W) and PackedLayout (packed W).
Packed positions are data-dependent (variable-length fields shift
subsequent fields), so the compiled form is necessarily a read program
(a pre-resolved tree of read instructions), not a flat lookup table
like OffsetMap. ReadPlan::compile walks BastDoc once at engine
construction, resolves all $ref\s eagerly, computes effective
endianness at every node, and inlines union variants; the read loop
then indexes into a Vec, matches on ReadKind, and calls data_access
with a precomputed Endian — no BastDoc, no resolve_typeref, no JSON
node access at read time.

Scope: SequentialReader and materialize_packed consume the plan (one
walker, not two); bast_validation, LayoutBuilder, and aligned one-shot
paths stay on BastDoc (different concern, not hot loops). Includes a
7-step Recommended Order matching review #004's structure. Breaking
public-API change (0.2.0 -> 0.3.0): SequentialReader::new and
materialize_packed take ReadPlan; BastDoc and the Bast* types are
unchanged (smaller breakage than review #004's Option A).

Cross-links: ADR-007's 'Cost' section gets a Note pointing at review
#004 and ADR-011, marking the 're-parse on demand' framing as the root
cause slated for retirement (the factory decision itself is retained);
the actual Cost/doc-comment rewrite happens in the implementation
commit per ADR-011's recommended order. README ADR table updated.

Verification: docs-only change, no source touched.

Closes review #004 H1, M1 (packed side), L1, L2 (on implementation).
2026-08-18 07:51:45 +00:00
glm-5.2 51cb552715 Add review #004: read-path performance review
Traces the ~400x read-path gap (alktty wire_vs_bast bench) to
SequentialReader::read_field_at re-parsing BastDoc::new on every field
read, with the self-referential lifetime constraint as the root cause.

Findings:
- H1: per-field BastDoc::new re-parse in sequential_reader.rs:262
- M1: same re-parse in four one-shot paths (LayoutBuilder::build,
  validate_bytes, engine read_field/write_field)
- L1: dead _field_schema param + Value clones in SequentialReader
- L2: ADR-007 Cost section + engine doc comment understate the re-parse
- N1: carry-forward of review #003 N2 (no new action)

Lays out fix options: Option A (owned typed tree, recommended, closes
H1+M1, breaking), Option B (read-plan precompute, fallback, H1 only,
non-breaking), Option C (borrow-from-engine, rejected, contradicts
ADR-007).

Verification: docs-only review; alktype source unchanged.

cab4932
2026-08-17 14:18:23 +00:00
glm-5.2 cab493206c Release v0.2.0: BAST pivot
Bump version to 0.2.0, exclude AGENTS.md from the published crate, and
add a CHANGELOG.md covering the breaking BAST pivot (schema format,
compile signature, validation split) plus bug fixes vs v0.1.0.

Verification:
- cargo test --release: all tests pass
- cargo clippy --all-targets -- -D warnings: clean
- cargo publish --dry-run --allow-dirty: packages as v0.2.0, no collision
- AGENTS.md no longer in cargo package --list; CHANGELOG.md included
2026-08-17 05:47:25 +00:00
glm-5.2 82fec45bc6 Sync docs to 18 BAST kinds (Timestamp removal fallout)
- README: 19 -> 18 kinds, drop the timestamp row from the kinds table,
  drop 'timestamp shape' from the validate_bytes constraint list, remove
  the residual 'upcoming alkcall crate' sentence from the crate
  independence section (alkcall exists now), fix two '19 kinds' refs in
  the documentation pointer list and schema-layer link.
- src/data_access.rs: module doc comment still said 'all 19 AlkType
  kinds' -> 18.
- bast-format.md (normative): 19 -> 18 AlkTypeKind enum variants.
- layout-engine.md: cross-reference to 'the 19 AlkType kinds' -> 18.
- ADR-BAST (bast-bast-format.md): the Decision section claimed the post-
  pivot enum has '19 unchanged' variants; now 18, with the wording
  adjusted so it no longer says 'unchanged' across the pivot.
- ADR-VAL-SPLIT: drop 'timestamp shape' from the value-domain constraint
  list and the validator-arm table row (validate_timestamp no longer
  exists).

Left as historically accurate (describe the v0.1.0 pre-pivot state):
ADR-003/004/006 Context mentions of AlkType:Timestamp, ADR-005 'engine
now has 19', and the '19 jsonschema::Keyword factories' references in
the What-is-removed sections of ADR-BAST and ADR-VAL-SPLIT.

Verification: cargo test --release (407 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean), cargo build --target
wasm32-unknown-unknown --release (clean).
2026-08-16 09:45:01 +00:00
deepseek-v4-pro 510553d800 Remove the Timestamp kind
The Timestamp kind was a residual from an early research reference. It
was byte-identical to String everywhere (length-prefixed UTF-8) and its
only distinguishing behavior was a hand-rolled non-strict RFC 3339 check
that the docs admitted was incomplete (Feb 31 passes, seconds range
unchecked, no leap seconds). JSON-level timestamp validation is
jsonschema's job (format: date-time on the validate_json path), not
alktype's.

Removes the AlkTypeKind::Timestamp variant, its to_bast_str/from_bast_str
mapping, the builder's Schema::timestamp() constructor, the
validate_timestamp/is_rfc3339_timestamp validator arms, and the
materializer/reader/engine timestamp arms. Updates the meta-schema
primitive enum (14 -> 13), the spec docs (bast-format.md, schema-layer.md,
builder.md, validation.md, overview.md, data-access.md, README.md), and
the kind-count references (19 -> 18).

Verification: cargo test --release (407 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean), cargo build --target
wasm32-unknown-unknown --release (clean).
2026-08-16 09:12:33 +00:00
deepseek-v4-pro 62ed009281 Resolve review #003 nits N1, N3, N4
- N1: number_from_f64 now returns AlkTypeError::Access for non-finite
  floats instead of silently materializing Value::Null. A NaN/Inf in the
  buffer surfaces as a clear access error rather than a misleading
  'expected a number' validation error.
- N3: drop the dead Value::String arm from check_bytes; the materializer
  only ever emits bytes as an array of u8. Update the validation-model
  docs to match.
- N4: fix stale ADR references in doc comments (ADR-096 -> ADR-002,
  ADR-097 -> ADR-003, ADR-098 -> ADR-004, ADR-101 -> ADR-007).

N2 (BastType::alk_kind returning Struct for ) left as documented;
every current caller resolves the ref first, so forcing Option through
the call sites is churn without benefit.

Verification: cargo test --release (411 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean).
2026-08-16 08:59:41 +00:00
deepseek-v4-pro 230345a867 Validate BAST documents against the meta-schema at compile time
Resolves review #003 finding M3.

- Add bast_meta::validate_bast_doc and call it from AlkTypeEngine::compile
  before parsing. Malformed annotations (unknown endian/encoding strings,
  non-integer align/maxLength, missing required properties) now surface as
  AlkTypeError::Schema instead of being silently tolerated by the parser.
- Align the meta-schema TypeRef with the parser: allow inline struct/union/
  enum as TypeRefs (the parser and builder already accepted them; the
  meta-schema and spec did not). Update bast-format.md TypeRef table and
  the BastType doc comment to the seven-form vocabulary.

Verification: cargo test --release (410 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo doc --no-deps (clean), cargo build --target
wasm32-unknown-unknown --release (clean).
2026-08-16 08:54:57 +00:00
deepseek-v4-pro e5c7cc1ca2 Fix offset-indirect, field-level endian, and aligned materialization
Resolves review #003 findings M1, M2, L1, and L2.

- offset-indirect (L1 + M2 arm): rework read_*_indirect to same-buffer
  absolute offsets (safetensors-style, per the metatensor model) and add
  write_*_indirect. Wire the encoding dispatch into engine.read_field/
  write_field and the aligned materializer. Previously the encoding was
  laid out but never read back.
- field-level endian override (M1): thread field.effective_endian through
  sequential_reader, engine.read_field/write_field, and
  materialize_struct_aligned. Previously a per-field endian override was
  silently ignored, misreading multi-byte values.
- aligned-mode materialization (M2): materialize fixed-size arrays via
  their vals[i] offset-map entries and maxLength reservations as
  zero-padded fixed-size slices (trailing NULs trimmed). Previously both
  read garbage or errored.
- dead endian param (L2): drop the ignored endian argument from
  materialize_packed/materialize_aligned; endianness is read from the
  root struct.

Verification: cargo test --release (404 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo build --target wasm32-unknown-unknown
--release (clean), cargo llvm-cov --release (89.60% lines).
2026-08-16 08:38:26 +00:00
deepseek-v4-pro ec73440c19 Add post-BAST-pivot code review (#003)
Covers the whole crate after the BAST pivot: correctness, code smell,
panic safety, and coverage (cargo-llvm-cov). 3 Medium findings (field-
level endian override ignored, aligned-mode validate_bytes broken for
arrays/maxLength/offset-indirect, meta-schema never applied at compile
time), 2 Low (offset-indirect dead code, dead endian param), 4 Nits.
Timestamp removal recorded as a publisher decision, tracked separately.

Verification: cargo test --release (389 pass), cargo clippy --all-targets
-- -D warnings (clean), cargo llvm-cov --release (90.14% lines / 86.68%
functions).
2026-08-15 14:43:09 +00:00
glm-5.2 562284faf4 Rewrite README for BAST pivot
The README still described the v0.1.0 custom-keyword JSON Schema format
(`AlkType:*` kinds, `compile(&mut schema, mode)`, single jsonschema
validator). Rewrite it for the BAST-era shipped API:

- What-it-is table: two validation specs (bytes via BAST-native, JSON
  via consumer-provided JSON Schema) instead of one.
- Usage example: `Definitions::new().build_doc(ChunkHeader, ...)` +
  `AlkTypeEngine::compile(&doc, ChunkHeader, LayoutMode::Packed,
  None)`; added a json!-literal variant showing the BAST document
  shape.
- The 19 kinds table: lowercase BAST `kind` strings instead of
  `AlkType:*` keywords; note the enum index bounds fix.
- Two layout modes: updated `compile` signature.
- Variable-length handling: BAST field-level `encoding` annotation.
- Union discriminators: `kind: union`, lazy variant $ref
  resolution, D-BAST-005 fields requirement.
- Endianness: struct-level (was 'top-level schema').
- Validation: two validators (ADR-VAL-SPLIT), BAST-native for bytes,
  standard jsonschema for JSON; uniform AlkTypeError::Validation
  payload (D-BAST-009); enum bounds fix noted.
- New 'BAST document shape' section: $defs required, root_name
  parameter, $ref restricted to #/$defs/<name>, BAST_META_SCHEMA
  re-export.
- Untrusted input: BAST document (was 'schema'); BAST parser
  preserves the no-panic invariant.
- Documentation index: updated for the new/rewritten docs and the
  ADR-BAST / ADR-VAL-SPLIT additions.

Verified both code examples compile and pass against the shipped API
(via a throwaway integration test, since removed).
2026-08-15 14:13:34 +00:00
glm-5.2 62270b03ca Sync architecture docs and ADRs to BAST pivot (steps 9-10)
Step 9 (convert tests to BAST format) was a no-op: steps 4-8 converted
the tests as they went. The only remaining  reference in
src/tests was the intentional  rejection test at
src/schema.rs:462 (asserting the old keyword form is rejected). Full
suite passes: 389 tests (312 lib + 77 integration).

Step 10 (sync architecture docs and ADRs):

Descriptive docs rewritten/updated for BAST:
- schema-layer.md: rewritten for the BAST parser (BastDoc/BastDef/
  BastType typed tree, AlkTypeKind enum with to_bast_str/from_bast_str,
  what was removed). Points at bast-format.md for the normative format.
- validation.md: rewritten for the two-validator model
  (bast_validation for validate_bytes, standard jsonschema for
  validate_json). Documents the repurposed build_validator, the
  AlkTypeError::Validation uniform payload (D-BAST-009), and what is
  removed.
- builder.md: updated all output examples to BAST JSON
  (struct_() -> { kind: struct, fields: [...] }; object() -> standard
  JSON Schema). Documents build_doc, count(), and the field-name union
  fields requirement (D-BAST-005).
- overview.md: updated for BAST (what/why, schema-is-the-format table,
  dependencies, architecture pointers, design decisions table).
- README.md (architecture index): updated document table, ADR table
  (new ADR-BAST + ADR-VAL-SPLIT, superseded ADR-001), OQ table
  (OQ-007/OQ-008 resolutions updated for BAST-native validator), and
  key design principles (#1, #2, #7, #10 reworded for BAST).
- data-access.md: updated tunion function signatures to BastUnion and
  the variant resolution to return BastType (resolve_typeref for refs).
- layout-engine.md: updated construct signatures
  (LayoutBuilder::new(bast_doc, root_name), OffsetMap::compute(&doc),
  SequentialReader::new(bast_doc, root_name)), the recursive-walk
  description (BAST typed tree), and composite-kind headings
  (TStruct/TUnion/TArray -> struct/union/array). Added D-BAST-004
  note on array count requirement.

New ADRs:
- ADR-BAST (bast-bast-format.md): the BAST format, meta-schema,
  //kind vocabulary, design principles, what is removed, the
  enum index bounds bug fix. Supersedes ADR-001's format-specific
  content; records D-BAST-001..009.
- ADR-VAL-SPLIT (val-split-two-validator-model.md): the two-validator
  model (BAST-native for validate_bytes, standard jsonschema for
  validate_json), the repurposed build_validator, the uniform
  AlkTypeError::Validation payload. Refines ADR-004's validation
  strategy and ADR-010's validation step; records D-BAST-006/007/009.

Amended ADRs (supersession/amendment notes added; original decision
text preserved as historical record):
- ADR-001: format-specific content superseded by ADR-BAST;
  purpose/scope and schema-is-the-format principle retained.
- ADR-002: unchanged under the pivot; one-line note that the input
  format changed but the modes didn't.
- ADR-003: annotation semantics retained; annotation location moved
  to BAST type-level properties (amended by ADR-BAST).
- ADR-004: AlkTypeError enum retained (D-BAST-009); validation
  strategy section refined by ADR-VAL-SPLIT.
- ADR-009: builder API surface retained; build() output format
  amended to BAST / standard JSON Schema by ADR-BAST (D-BAST-008).
- ADR-010: validate_bytes two-step concept retained; validation step
  amended to the BAST-native validator by ADR-VAL-SPLIT.

Other:
- Cargo.toml description: JSON Schema with AlkType:* custom keywords
  -> BAST document.
- bast-pivot.md research record: status draft -> implemented, with a
  pointer to the ADRs that superseded its decisions.
- bast-implementation.md plan: status draft -> complete, with a note
  that step 9 was a no-op and step 10 is this commit.
- open-questions.md: OQ-006/OQ-007/OQ-008 resolutions updated for the
  BAST-native validator.
- questions/008-unionvalidator-variant-dispatch.md: added a
  post-BAST-pivot note pointing to the current bast_validation
  implementation; v0.1.0 resolution text preserved as historical
  record.

Verification:
- cargo test --release: 389 pass (312 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cross-reference check: every relative link in the new/updated docs
  resolves (verified by script).
2026-08-15 14:03:21 +00:00
glm-5.2 54fd112fde Remove v0.1.0 custom-keyword machinery (step 8)
The BAST parser (step 3) and BAST-native validator (step 5) replaced
the v0.1.0 custom-keyword accessor layer; step 7 moved the builder to
BAST output. This step removes the now-dead code:

Removed from src/schema.rs:
- get_alktype_kind / get_alktype_kind_enum /
  get_alktype_kind_loose / get_alktype_kind_loose_enum
  (replaced by the BAST parser's kind dispatch)
- normalize_refs / inline_union_variant_refs + helpers
  (BAST refs are always #/$defs/<name>; resolution is a single
  hash lookup, variant refs resolve lazily)
- parse_encoding / parse_align / parse_max_length / parse_endian
  (bast.rs has its own BAST-property-form copies)
- parse_discriminator + DiscriminatorKind
  (replaced by bast::BastDiscriminator; builder has its own
  Discriminator enum)
- resolve_ref / resolve_ref_or_inline
  (replaced by BastDoc::lookup_def / resolve_typeref)
- FromStr impl, as_str, Endian::from_schema, ALKTYPE_PREFIX,
  BYTE_DISCRIMINATOR_TYPES, and the associated unit tests

Kept: AlkTypeKind enum + methods (type_size, natural_alignment,
is_fixed_size, needs_endian, is_composite, is_variable_length,
to_bast_str, from_bast_str), Display (now backed by to_bast_str),
Endian, VariableEncoding, U32_SIZE, DISCRIMINATOR_PATH.

src/lib.rs: dropped the 13 schema::* helper re-exports and
DiscriminatorKind from the public surface; kept Endian, AlkTypeKind,
VariableEncoding.

Doc/comment updates: bast.rs, builder.rs, engine.rs, error.rs —
removed references to the deleted functions and the AlkType:*
keyword form.

The jsonschema crate remains a dependency (validate_json path +
BAST meta-schema validation); build_validator was already repurposed
in step 6 (no custom keywords).

Verification:
- cargo test --release: 389 pass (312 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cargo build --target wasm32-unknown-unknown --release: clean
2026-08-15 13:39:05 +00:00
glm-5.2 45f3336201 Builder API produces BAST JSON (step 7)
- Schema internals: flat Map<String, Value> → Repr enum distinguishing
  BAST primitives (bare TypeRef strings), BAST composites (struct/
  union/enum/array/record objects), $ref, standard JSON Schema
  objects, and raw adopted values. Annotations (endian/align/encoding/
  maxLength) stored on the Schema and placed correctly by build()
  (struct-level) or field() (field-level).

- Primitive constructors (uint32, string, bytes, etc.) now produce
  bare BAST kind strings ("uint32") instead of {"AlkType:Uint32":true}.

- struct_().field(...) produces {"kind":"struct","fields":[{name,kind,...annos}]}
  with an ordered fields array (BAST design principle #4 — no reliance
  on preserve_order for field order).

- union_() emits {"kind":"union","discriminator":{...},"mapping":{...}}
  with BAST kind strings in the discriminator ("uint8" not
  "AlkType:Uint8"). Field-name discriminator unions emit the fields
  array; byte-offset unions omit it.

- enum_of() produces {"kind":"enum","values":[...]}.

- array_of(element).count(n) produces {"kind":"array","element":...,"count":n}.
  .count() is an additive method (D-BAST-004 requires count; the
  array_of signature is unchanged per the semver contract).

- record_of(values) produces {"kind":"record","values":...}.

- object()/string_()/etc. unchanged — standard JSON Schema output.

- encoding() now stores the value on the Schema; field() extracts it
  as a field-level property. No more keyword-object duality.

- max_length() on standard types emits the standard keyword; on BAST
  types it is extracted by field() as a field-level constraint.

- Definitions::build_doc(root_name, root) — new additive method
  producing a complete BAST document with the root type inside $defs
  (where BAST requires it). The old build()/merge_into() remain for
  backward compatibility.

- All builder tests updated to expect BAST output shapes. New tests:
  field annotations, field-name union with fields, array with count,
  ref element, encoding emission, Definitions::build_doc round-trip
  through compile(), SFTP-style union compile, array/record compile.

- Module doc comments updated (builder.rs, lib.rs).

Verification:
- cargo test --release: 427 pass (350 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean (no warnings)
- cargo build --target wasm32-unknown-unknown --release: clean
2026-08-15 13:09:18 +00:00
glm-5.2 ba7f8e1bad validate_json against consumer-provided JSON Schema (step 6)
- AlkTypeEngine::compile gains a 4th param json_schema: Option<&Value>.
  When Some, a standard jsonschema::Validator is built from the
  consumer-provided JSON Schema and stored for the JSON-validation
  path. When None, validate_json returns AlkTypeError::Schema and
  is_valid_json returns false (D-BAST-007).

- validate_json / is_valid_json signatures unchanged (per semver
  contract). Behavior: they now validate against the consumer JSON
  Schema, not a custom-keyword validator built from the alktype
  schema. The BAST document is not involved in this path.

- build_validator repurposed (deferred decision #2): same signature,
  now builds a standard jsonschema::Validator with no custom keywords.
  Behavioral break, not a type break. Re-export kept.

- Removed the 19 jsonschema::Keyword implementations and the 4
  define_*_validator! macros (dead on the bytes path since step 5,
  now dead on the JSON path too). The is_rfc3339_timestamp helper
  lives on in bast_validation.rs (already copied there in step 5).

- All compile call sites updated to pass None for json_schema (the
  layout/read/write/validate_bytes tests don't need JSON validation).

- New tests: validate_json accepts/rejects against consumer JSON
  Schema, returns Schema error when no JSON Schema supplied,
  is_valid_json false when no schema, independence from BAST doc,
  malformed JSON Schema build error, nested object JSON Schema.

Verification:
- cargo test --release: 409 pass (332 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo doc --no-deps: clean
2026-08-15 12:53:35 +00:00
glm-5.2 f853dafaf1 Add BAST-native validator for validate_bytes (step 5)
Replace the jsonschema custom-keyword validator on the bytes path with
a recursive walker over the BAST typed tree. The materializer already
guarantees structural correctness; the validator enforces only the
value-domain constraints expressed in the BAST document.

- New `src/bast_validation.rs`: `validate_value` dispatches on
  `BastType`, resolving `$ref`s lazily via `BastDoc::resolve_typeref`.
  Constraint arms: integer ranges (Int8..Uint64), float finiteness,
  string/bytes `maxLength` (from `BastField::max_length`), RFC 3339
  timestamp shape, enum index bounds, union variant dispatch (recurses
  into the variant, recovering OQ-008 per-variant constraints), struct
  field presence, array count, record values.
- `engine::validate_bytes` now calls `bast_validation::validate_value`
  instead of `self.validator.validate`. The `validator` field is still
  built and used by `validate_json`/`is_valid_json` (step 6 reworks
  those).
- Enum index bounds check (`idx < values.len()`) fixes the v0.1.0 dead
  constraint: the built-in `enum` keyword checked string membership, but
  the materializer emits a numeric index that never matched.
- Errors constructed via `jsonschema::ValidationError::custom` so
  `AlkTypeError::Validation` keeps its payload type uniform with the
  `validate_json` path (D-BAST-009).
- `lib.rs`: add `pub mod bast_validation;` (engine-internal, not
  re-exported in the `pub use` block) and update the module doc.

Verification:
- cargo test --release: 446 pass (369 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo doc --no-deps: clean
2026-08-15 12:44:31 +00:00
glm-5.2 04573e1d86 Wire compile() to BAST document + root name (step 4)
Step 4 of the BAST pivot: the layout engines, materializer, tunion
dispatch, and engine now consume the BAST typed tree (BastDoc/
BastStruct/BastField/BastType/...) instead of walking raw JSON with
get_alktype_kind*.

Breaking changes (per the pivot plan's semver contract):
- AlkTypeEngine::compile signature:
    compile(schema: &mut Value, mode)
    -> compile(bast_doc: &Value, root_name: &str, mode)
  Drops &mut (BAST needs no in-place normalize_refs); adds required
  root_name (D-BAST-001); input is a BAST document, not a custom-keyword
  JSON Schema.
- OffsetMap::compute, LayoutBuilder::new, SequentialReader::new now take
  a BAST document (&Value) + root_name (or &BastDoc) instead of a
  v0.1.0 schema.
- tunion::read_byte_discriminator / read_field_discriminator /
  resolve_variant / discriminator_size now take &BastUnion instead of
  &Value.
- materialize::materialize_packed / materialize_aligned now take
  &BastDoc instead of &Value.

Key design points:
- The engine stores a clone of the BAST Value + root_name so
  sequential_reader() and read_field() can re-parse the typed tree on
  demand without lifetime entanglement with the caller's Value.
-  resolution is a single hash lookup via BastDoc::resolve_typeref;
  no normalize_refs, no inline_union_variant_refs.
- bast.rs gains BastField::synthetic() (pub(crate)) for constructing
  synthetic fields wrapping TypeRefs (array elements, record values,
  union variants — these aren't fields and carry no field annotations).
- The v0.1.0 schema.rs helpers and validation.rs custom-keyword
  validators remain defined (step 8 removes them). build_validator still
  runs on the BAST doc — with no AlkType:* keywords present, the custom
  factories don't trigger and jsonschema performs structural validation
  only. The validate_bytes value-constraint enforcement (maxLength, enum
  bounds) is step 5's concern (the BAST-native validator).

Tests:
- All engine, layout, materialize, tunion, and integration tests
  converted to BAST format (kind/fields vocabulary, /
  composition). Expected validation outcomes for the layout path are
  identical; the maxLength/enum-bounds validate_bytes tests are step 5's
  regression target.
- builder.rs::builder_chunk_header_compiles_in_packed_mode uses a
  hand-written BAST doc (the builder still emits v0.1.0 format; step 7
  converts it).

Verification:
- cargo test --release: 425 pass (348 lib + 77 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo doc --no-deps: clean
2026-08-15 12:29:27 +00:00
glm-5.2 f2f9c0326c Add BAST document parser (typed tree, step 3)
- New src/bast.rs: typed surface over a BAST document — BastDoc,
  BastDef, BastDefKind, BastStruct, BastField, BastUnion,
  BastDiscriminator, BastEnum, BastType, BastRef, BastArray,
  BastRecord. Borrows from the source Value (no clone of the tree).
- $ref resolution is a single hash lookup against $defs
  (#/$defs/<name> only); union variant refs resolved lazily via
  BastDoc::resolve_typeref / resolve_ref — replaces
  normalize_refs + inline_union_variant_refs (those stay for now;
  step 8 removes them).
- Untrusted-input safe: every walk returns AlkTypeError::Schema on
  a malformed document, never panic/unwrap (AGENTS.md §3).
  Overflow-safe usize parsing via try_from (AGENTS.md §4).
- D-BAST-005 enforced: fields array only valid with field-name
  discriminators; required for them.
- Additive only: new pub mod bast + re-exports in lib.rs. No
  existing re-exports removed (those go in step 8). 47 new tests.

Verification:
- cargo test --release: 465 pass (379 lib + 86 integration)
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cargo build --target wasm32-unknown-unknown --release: clean
2026-08-15 11:56:03 +00:00
glm-5.2 29134789a9 Embed BAST v1 meta-schema as BAST_META_SCHEMA
Step 2 of the BAST pivot. Adds the normative BAST meta-schema
(Draft 2020-12 JSON Schema) as a public serde_json::Value, embedded
at compile time and available for validating BAST document
well-formedness. Copied verbatim from docs/architecture/bast-format.md
§The Meta-Schema.

- New src/bast_meta.rs: BAST_META_SCHEMA static, built lazily via
  LazyLock (the json! macro allocates, so it can't be a const;
  parsed once, reused as &'static Value thereafter).
- src/lib.rs: pub mod bast_meta + re-export BAST_META_SCHEMA.
  Additive public surface.

The meta-schema validates structure (correct $defs shape, known
kind strings, required properties, no additional properties,
$ref restricted to #/$defs/<name>). Value-domain constraints
(maxLength, enum index bounds) are enforced by the BAST-native
validator (step 5), not this meta-schema.

Verification:
- cargo test --release: 418 tests pass (332 lib + 86 integration);
  16 new unit tests exercise the meta-schema against valid and
  invalid BAST documents (missing $defs, unknown kind, additional
  properties, byte-discriminator union, enum, array+count, record,
  $ref, malformed ref, empty enum).
- cargo clippy --all-targets -- -D warnings: clean.
- cargo build --target wasm32-unknown-unknown --release: clean
  (LazyLock + serde_json::json! macro are wasm-safe).
2026-08-15 11:37:14 +00:00
glm-5.2 66ab9d7d93 Add AlkTypeKind::to_bast_str/from_bast_str for BAST kind strings
Step 1 of the BAST pivot. Adds the lowercase-string mapping
("uint32" <-> AlkTypeKind::Uint32) that the BAST parser and
validator dispatch on (D-BAST-002).

- to_bast_str(self) -> &'static str: returns the lowercase BAST
  string for all 19 variants (14 primitives + struct/union/array/
  record/enum). Boolean -> "bool", distinct from the PascalCase
  variant name.
- from_bast_str(s) -> Result<AlkTypeKind, AlkTypeError>: inverse of
  to_bast_str; returns AlkTypeError::Schema for unknown strings.

These are additive inherent methods on the already-re-exported enum.
The existing FromStr impl (parsing the v0.1.0 "AlkType:Uint32"
keyword form) is unchanged and removed in step 8. The two surfaces
are deliberately distinct: from_bast_str rejects "AlkType:Uint32"
and FromStr rejects "uint32".

Verification:
- cargo test --release: 402 tests pass (316 lib + 86 integration);
  6 new unit tests cover both directions, round-trip, and the
  from_bast_str/from_str distinctness.
- cargo clippy --all-targets -- -D warnings: clean.
2026-08-15 11:29:34 +00:00
glm-5.2 f5f52c61e8 Decompose BAST pivot doc into normative spec + implementation plan
The bast-pivot.md research doc had grown to 1477 lines (~58KB) through
iterative editing, pushing its most actionable content (D-BAST
decisions, POC result, migration steps) past the 50KB Read tool cap.
Agents peeking at the truncated file landed in duplicated/out-of-order
sections. Decompose into three readable-sized files with distinct roles:

- docs/architecture/bast-format.md (28KB, new): the normative BAST
  format spec -- meta-schema, TypeRef, examples, validation model.
  Grounded in the POC and D-BAST-001..009. Stable and safe to write
  now; schema-layer.md/validation.md stay describing current code and
  are rewritten post-implementation (per AGENTS.md ADR-grounding rule).
- docs/plans/bast-implementation.md (31KB, new): the execution entry
  point -- ordered 10-step plan with per-step goal/files/spec-ref/
  verification, the public-API semver contract table up front as a
  scope-creep guardrail, and the ADR-sync checklist at the end. Each
  step links to the specific bast-format.md section and D-BAST anchor.
- docs/research/bast-pivot.md (28KB, trimmed): now the research record
  only -- Summary, Motivation, POC scope/result, Decisions, Risks,
  References. The normative format spec, what-changes tables,
  validator-split details, and migration steps moved to the two new
  docs; pointers added. 1155 lines removed, 216 added.
- docs/architecture/README.md: index updated to list bast-format.md
  and the two in-progress pivot docs, with notes on schema-layer.md
  and validation.md being rewritten when the pivot lands.

All three files are under the 50KB Read cap, so an implementing agent
gets the whole document in one call. Cross-reference anchors verified
to resolve. No code changes; cargo test --release (396 tests) green.

Verification: cargo test --release (310 crate + 86 integration, all pass).
2026-08-15 10:59:14 +00:00
glm-5.2 5796d1c22f Add semver and ADR impact mapping for the BAST pivot
Scope-creep guardrail for the public API during implementation. Maps
every item re-exported from src/lib.rs to a class (breaking / additive /
unchanged) with the specific change, and every ADR (001-010) to an
action (supersede / amend / unchanged) with the reason.

Net breaking: compile (signature), validate_json/is_valid_json
(contract), Schema::build/Definitions::build (output format),
build_validator (signature or removal), and the ~13 schema::* helper
re-exports. Net additive: BAST parser, BAST-native validator,
AlkTypeKind::from_str/to_str. Net unchanged: the entire layout +
data-access + materialize + tunion layer, AlkTypeError (D-BAST-009),
the Discriminator builder, AlkTypeKind variants.

Flags three small decisions deferred to their implementation steps
(validate_json JSON Schema source, build_validator fate, schema::*
re-export retention) so they don't become drive-by semver changes.

This is a living guide — it may shift slightly during implementation,
but capturing the contract now prevents public-surface drift. Doc-only.
2026-08-15 10:29:24 +00:00
glm-5.2 89f05850f2 Resolve OQ-BAST-001: keep Validation(ValidationError<'static>) (D-BAST-009)
Converts the open question into a closed decision. Rationale: consumer
ergonomics on the combined validate_json + validate_bytes path — one
uniform payload type means one match arm downstream. Option 2
(Validation(String)) would force validate_json to flatten its
structured errors to a String, losing information on the richer path to
accommodate the less rich one. The no_std/minimal-build angle that
option 2 was meant to enable is moot: validate_json requires jsonschema
regardless, so a bytes-only no_std build already has to give up
validate_json as a separate larger decision; dropping the type from one
error variant doesn't unlock it.

Updates Phase 1 step 5 to reference D-BAST-009 for the error
construction pattern, and rewrites POC Result observation 4 from
'deferred decision' to 'decided — see D-BAST-009'.

No semver-relevant change to the Validation variant. Doc-only.
2026-08-15 10:13:47 +00:00
glm-5.2 e77268c951 Record BAST validator POC result; track OQ-BAST-001 error-payload decision
Adds the POC Result section (POC on branch bast-validator-poc, commit
f371fe4 — 20/20 tests, full 416-test suite green, clippy/wasm/doc clean).
Hypothesis confirmed: a recursive walker over the BAST type tree fully
replaces the 19 custom keyword validators on the validate_bytes path,
recovers OQ-008 union variant dispatch, and fixes the enum-membership
dead constraint. The POC code is reference scaffolding on the branch
and is not merged to main — it is superseded by Phase 1 step 5.

Marks Phase 1 step 2 and the POC Scope section as done with pointers to
the result section.

Elevates the deferred AlkTypeError::Validation payload-shape question
to OQ-BAST-001: keep jsonschema::ValidationError<'static> (POC choice,
simplest, dependency stays) vs introduce Validation(String) (drops
jsonschema from the error type; semver-relevant public-API change).
Decision belongs to the production refactor.

Verification: doc-only change, no code touched.
2026-08-15 09:38:53 +00:00
glm-5.2 19f8162f1a Refine BAST pivot: validation model, POC scope, resolve open questions
- Rewrite Validator Split around BAST-native validator for
  validate_bytes (walks BAST, checks value-domain constraints, no
  external JSON Schema needed); validate_json uses standard
  jsonschema::Validator from consumer-provided JSON Schema
- Add targeted POC: BAST-native validator replacing 19 custom keyword
  validators on the bytes path, verified via existing test suite
- Fix meta-schema: require count on arrays (variable-element arrays
  deferred per OQ-001), add optional fields array to UnionDef for
  field-name discriminators
- Remove lying no-count array example, replace with deferred note
- Document dead enum constraint (materialized index never matches
  string-membered enum); BAST-native validator fixes it via index
  bounds check
- Resolve all 7 OQs + 3 spec gaps as D-BAST-001 through D-BAST-008
- Update Migration Path, Risks table, Engine internals to reflect
  BAST-native validator
- Clarify 'no pocs needed' was an overcorrection: layout swap needs
  no POC, but the validation model does

Verification: docs-only change, no code affected
2026-08-15 09:05:55 +00:00
deepseek-v4-pro 82bc8f29c0 Clean up BAST pivot: drop POCs, remove recursion, add spec gaps
- Remove the four proposed POCs: they were implementation smoke tests,
  not de-risking probes. The pivot is a backend swap on a proven layout
  engine; byte-identity is already proven and the layout code is
  unchanged, so there is nothing empirical left to de-risk.
- Remove the recursive TreeNode example and the recursion mention in
  design principle 2: recursion is not a binary-layout concern and the
  engine has no cycle detection.
- Add a Spec Gaps section: validate_bytes semantics after keyword
  validator removal (UnionValidator variant dispatch regression),
  arrays of variable-length elements (engine rejects them), and
  field-name discriminator unions (meta-schema cannot express them).
- Reframe the engine change as an accessor-layer refactor: the
  walkers' (kind, field list, annotations) reads change; everything
  beneath them carries over unchanged.
2026-08-15 07:46:03 +00:00
deepseek-v4-pro 37bd5d7b7d Fix BAST pivot: jsonschema remains a direct dependency
Correct the validator split section to clarify that jsonschema is
not removed — it remains the JSON Schema validator for both paths
(BAST meta-schema validation and standard JSON payload validation).
Only the custom keyword registration path is removed. Also fix the
build_validator and validate_json entries in the engine changes
table to reflect that they are repurposed, not removed.
2026-08-14 16:06:26 +00:00
deepseek-v4-pro 004d52d505 Add BAST pivot research document
Proposes replacing custom JSON Schema keywords with a standalone
kind-based JSON format (BAST) using / for composition.
Covers format design, meta-schema, engine changes, validator split,
codegen future, ABI adapter potential, migration path, 7 open
questions, and 4 proposed POCs.
2026-08-14 15:55:17 +00:00
glm-5.2 5a4ed9e8e3 Exclude internal artifacts from crates.io package
Trim the published package from 64 → 52 files (865.7KiB → 715.3KiB,
202.8KiB → 153.0KiB compressed) by excluding internal-only files:
.opencode/ agent configs, docs/reviews/, docs/research/, and
docs/sdd_process.md. Keeps docs/architecture/ ADRs and OQs, which
document the public design.

Verification:
- cargo test --release: 396 tests pass
- cargo clippy --all-targets -- -D warnings: clean
- cargo doc --no-deps: clean
- cargo build --target wasm32-unknown-unknown --release: clean
- cargo publish --dry-run --allow-dirty: clean (52 files, 153.0KiB)
2026-08-11 09:45:15 +00:00
337 changed files with 31659 additions and 7761 deletions

No files matched your search

+10 -1
View File
@@ -1,3 +1,12 @@
target/
node_modules/
.worktrees/
.worktrees/
fuzz/target/
fuzz/artifacts/
fuzz/coverage/
# Grown corpora: the hash-named files the campaigns drop into
# fuzz/corpus/<target>/ are gitignored (the quinn/h2 policy); the
# committed seeds are the seed-* files, kept via the per-dir
# .gitignore un-ignores.
fuzz/corpus/*/[0-9a-f][0-9a-f]*
+33
View File
@@ -5,6 +5,29 @@ auto-loads this file as instructions, overriding the built-in defaults for
this project. Custom agents in `.opencode/agents/` inherit these rules
unless their own prompts say otherwise.
## Session Continuity (keep the agent loop alive)
opencode ends the turn whenever an assistant message contains no tool
call — including messages that are pure analysis. Long reasoning bursts
are welcome in this repo (they pre-catch errors and self-correct), but a
burst that ends as analysis-only text silently stops the session
mid-task. Past sessions documented this repeatedly ("the session
stalled"; the working fix discovered there: "call tools frequently, keep
thinking bursts short"). Keep the depth; change where the burst ends:
1. **Never end a turn with analysis-only text.** Every visible message
must either issue a tool call or be a final report for a genuinely
completed phase/task. When a thinking burst converges on a decision,
act on it (read, edit, bash) in the same turn.
2. **Land work incrementally.** Once a design decision is settled, write
the code before analyzing the next one. Do not emit full-design
essays in a single chat message; reasoning belongs in thinking tokens
or committed docs, not in the transcript.
3. **On a silent turn end, resume without re-deriving.** If the turn
ended after an analysis-only message and the task is incomplete,
pick up from the last settled decision — do not redo the analysis
and do not ask the user whether to continue.
## Git Workflow
**Commit and push when reasonable.** When a change is complete and
@@ -142,8 +165,18 @@ cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # if docs changed
cargo build --target wasm32-unknown-unknown --release # if layout/wasm-relevant code changed
cargo publish --dry-run --allow-dirty # before a release
cargo test --manifest-path fuzz/shared/Cargo.toml # fuzz corpus replay (the fuzz gate)
```
The corpus replay is the standing fuzz gate (the alkcall
`docs/research/fuzzing.md` §7.9 posture): it replays every committed
seed through the same invariant functions the fuzz targets run, on
stable, without nightly. Campaigns (nightly, cargo-fuzz) run manually
via `fuzz/run-detached.sh` — never as a foreground child of an agent
session — before releases, after touching `src/data_access.rs`,
`src/schema.rs`, the compile walks, or the sequential reader. See
`docs/plans/fuzzing.md` and `fuzz/README.md`.
## Architecture Context
- `docs/architecture/` — the authoritative spec. Read it before
+270
View File
@@ -0,0 +1,270 @@
# Changelog
All notable changes to this crate are documented here. The format is
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this crate adheres to [Semantic Versioning](https://semver.org/).
## [0.4.0] - 2026-09-30
The fuzzing release. The fuzz/ workspace (docs/plans/fuzzing.md) put
five libFuzzer targets on the engine — bast_compile, data_access,
read_opseq, layout_build, validate_pair — with release-budget campaigns
across all five. Two genuine engine bugs found and fixed same-day,
plus one upstream pin (docs/plans/fuzzing.md §6, §5 campaign numbers).
### Breaking changes
- **`VariableEncoding` gains `MaxLengthReserved`** (finding W3-3):
the aligned-mode `maxLength` reservation (ADR-003 strategy 2, the
`VARCHAR(N)` pattern) is now recorded as its own encoding variant
instead of masquerading as `LengthPrefixed`. Code matching on
`VariableEncoding` exhaustively must add an arm; in the document
form the strategy is still expressed via `maxLength`, never as an
`encoding` value.
- **`read_field`/`write_field` fix for aligned `maxLength`
reservations** (finding W3-3, commit `a0dd3d2`): an aligned
String/Bytes leaf with a declared `maxLength` is now read as a raw
zero-padded, NUL-trimmed window and written zero-padded — previously
the first four raw bytes of the reservation window were misparsed as
a u32 length prefix, breaking the validate_bytes ⇒ read_field
agreement for every aligned schema declaring `maxLength`.
- **`plan_read_array` bounds check** (finding W2-1, wave 2): a
fixed-stride array whose declared window (count × stride) extends
past the buffer now returns an `Access` error naming the array
field instead of reporting success and deferring the failure to the
next field read (or masking it entirely when the array was last).
### Additions
- **`data_access::read_reservation` / `read_reservation_string` /
`write_reservation`** — the single source of truth for the
`maxLength` reservation leaf semantics, shared with the aligned
materializer.
- **`fuzz/` subtree** — five libFuzzer targets, 257 committed seeds,
the corpus-replay gate (`cargo test --manifest-path
fuzz/shared/Cargo.toml`, AGENTS.md verification checklist), the
detached campaign runner; nightly confined to `fuzz/`, `fuzz/`
excluded from the package. All targets, seeds, and the replay gate
are stable-toolchain safe.
## [0.3.0] - 2026-09-07
The compiled-forms release. The packed read path — the hot path for
stream parsing — is driven by a compile-once `ReadPlan` instead of a
per-call walk of the BAST typed tree; byte validation runs a compiled
`ValidationPlan`; plans and offset maps fingerprint to a stable hash
(ADR-011/ADR-012). Reads of SFTP-shaped packet streams went from
~189× hand-rolled Rust to ~74× (~2.4× faster), fixed-stride chunk
reads from ~18× to ~11×, and `SequentialReader::read_next_borrowed`
makes the per-field hot loop allocation-free.
### Breaking changes
- **`Bast*` types are owned.** `BastDoc`/`BastStruct`/`BastField`/… no
longer borrow from the source `serde_json::Value`; all v0.2.0
lifetimes are gone. `BastDoc::new` parses the root eagerly; `$ref`s
resolve lazily.
- **`OffsetMap::get` / `PackedLayout::get`** return `&OffsetEntry`
(was `Option<OffsetEntry>` by value), backed by an O(log n)
`BTreeMap` path→index (first-occurrence-wins for duplicate names).
- **`SequentialReader::new`** takes the compiled plan; construct via
`AlkTypeEngine::sequential_reader()` (packed mode only).
- **`materialize_packed` / `materialize_aligned`** take the compiled
plan / `(&BastDoc, &OffsetMap)` pair respectively.
- **Field-name-discriminator union wire convention** (ADR-011
addendum): the builder lays out the union's declared `fields`
(shared) first, then the variant's own fields. Variants must not
re-declare the discriminator or any shared field, and the
discriminator field must be the first entry in `fields` — all
enforced at parse with clean `Schema` errors. Schemas relying on
0.2.0's variant-only layout are rejected (they produced
reader↔builder-disagreeing bytes).
- **`maxLength` is string/bytes-only** — rejected at parse on every
other kind (it was silently unenforced there).
- **Aligned-mode `Record` fields reject `offset-indirect`** (the
materializer always walks the inline count-prefixed form — the
annotated shape was never readable).
- **Schema input bounds** (untrusted-schema hardening, AGENTS.md §3):
array `count` ≤ 2^16 and `count × stride` ≤ 2^26 bytes; `align` ≤
4096; `maxLength` ≤ 2^26; cyclic `$ref` graphs and >128-deep nesting
are rejected by every public walker (`OffsetMap::compute`,
`LayoutBuilder::new`, `materialize_aligned` included), not just the
engine.
### Additions
- **`ReadPlan`** (ADR-011) — the compiled packed-read plan, re-exported
with `CompositePlan`/`FieldPlan`/`ReadKind`/`DiscriminatorPlan`.
`ReadPlan::compile` is untrusted-input-safe standalone (depth cap +
cycle set). `fixed_size()` exposes the compile-time-known byte size
for fixed structs.
- **`ValidationPlan`** (ADR-012 §3) — the compiled `validate_bytes`
walker, with `ValidNode`/`ValidVariant` sub-types.
- **`fingerprint()`** on `ReadPlan`/`OffsetMap`/`ValidationPlan` +
`Hash`/`Eq` derives on the plan types (ADR-012 §1/§4) — plan
identity for cache-keying across processes.
- **`OffsetMap` `LeafMeta`** — each entry records whether it is
fixed/length-prefixed/offset-indirect so `read_field`/`write_field`
dispatch without re-walking the schema; `OffsetEntry` type re-exported.
- **`SequentialReader::read_next_borrowed`** — zero-allocation variant
of `read_next` (field name borrowed from the plan).
- **`AlkTypeEngine::validate_bytes`** now runs the compiled
`ValidationPlan` (was an interpretive BAST walk in 0.2.0).
### Fixes (post-release-commit hardening — reviews #006, #007, #008)
All found and fixed before the first crates.io publish of 0.3.0, so
no published version ever exhibited them.
- **Untrusted-input crashes removed.** A huge declared array count
OOM-aborted the process (`Vec::with_capacity(count)` before reading
a byte) — now compile-capped and walked with push-only growth.
Cyclic `$ref` graphs stack-overflowed the three standalone layout
walkers — now guarded by a shared reference-graph check. Deeply
nested stride-0 arrays briefly allowed ~477 MB of simultaneous
allocation from a ~1 KB schema — restored to incremental growth.
- **Cross-consumer divergences closed.** Builder, reader,
materializer, tunion, and the validation plan now agree on
field-disc union layout (shared-then-variant), on the discriminator
field's position (must be first), and on union mapping-key matching
(numeric fast-path dispatch only for canonical keys like `"2"`;
`"01"`/`"+1"` fall back to the string comparison all consumers
share). The legacy BAST walker's field-disc union arm walks shared
fields before the variant (it previously materialized variant fields
from shared fields' bytes).
- **Silently-corrupt layouts rejected.** Aligned record fields with
`maxLength`/`offset-indirect`; non-final inline length-prefixed
fields (records included — the ADR-006 check now sees them);
aligned-mode `maxLength`/`offset-indirect` on records; unions in
aligned mode (pre-existing, now tested).
- **Coverage**: 90.67% lines / 86.32% functions at review #007's
audit, 91.66% after its fixes; every uncovered region outside test
modules read and classified in-tree (docs/reviews/007).
### Non-breaking improvements
- Engine compile is one-shot and allocation-tidy; plans are
`Send + Sync` (statically asserted) and fingerprintable.
- Zero-progress array-element guard on all three array walkers (a
zero-size element makes the declared count unbounded on the wire).
- WASM-clean unchanged: two dependencies (`jsonschema`
default-features off, `serde_json` with `preserve_order`), no
`async`, no `unsafe`, no feature flags.
- Benches (`benches/wire_vs_bast.rs`): read/write chunk streams, an
SFTP-shaped union packet stream, and `validate_bytes` per buffer —
the numbers quoted above and in ADR-007/ADR-011.
## [0.2.0] - 2026-08-17
A breaking release that replaces the v0.1.0 `AlkType:*` custom-keyword
JSON Schema format with BAST (Binary Abstract Syntax Tree) — a JSON
document that describes binary layouts using a `kind`-based vocabulary
with `$defs`/`$ref` for composition. BAST is itself a valid JSON Schema
instance (it has a meta-schema), making it self-validating,
editor-friendly, and trivially consumable from any language with a JSON
parser. The engine works the same way as before: compile a document
once into an `AlkTypeEngine`, then read/write fields at computed offsets
and validate bytes/JSON. The pivot was made now because v0.1.0 has no
real consumers (≈15 crates.io downloads, mostly bots/scanners), so the
custom-keyword wart could be removed cleanly.
### Breaking changes
- **Schema format.** The v0.1.0 `AlkType:*` custom-keyword JSON Schema
format (`{ "AlkType:Struct": true, "fields": [...] }`) is removed.
Schemas are now BAST documents:
`{ "$defs": { "<TypeName>": { "kind": "struct", "fields": [...] } } }`.
The `kind`-based vocabulary covers 18 binary kinds (integers, floats,
bytes, string, struct, union, enum, array, etc.).
- **`AlkTypeEngine::compile` signature.** Now takes
`(bast_doc: &Value, root_name: &str, mode: LayoutMode, json_schema: Option<&Value>)`.
The root type name is a required parameter — it selects which `$defs`
entry is the top-level type (previously the root was implicit from the
single top-level schema object).
- **Builder API output.** `Definitions`/`Schema`/`Discriminator` now
produce BAST JSON via `Definitions::build_doc(name, schema)`. The
builder method names are unchanged; only the emitted JSON shape
changed. `Schema::struct_()` produces a BAST struct;
`Schema::object()` produces a standard JSON Schema (for the
`validate_json` path).
- **Validation split.** v0.1.0 used a single `jsonschema` validator
with 19 custom `AlkType:*` keywords for both bytes and JSON
validation. 0.2.0 splits this into two independent paths:
- `validate_bytes` uses a new BAST-native validator
(`bast_validation`) — a recursive walker over the BAST type tree.
- `validate_json` / `is_valid_json` use a standard
`jsonschema::Validator` built from a consumer-provided JSON Schema
(passed to `compile` as the `json_schema` parameter). No custom
keywords; BAST is not involved — BAST describes bytes, not JSON
shape.
Both paths return `AlkTypeError::Validation` with a uniform
`jsonschema::ValidationError<'static>` payload.
- **Removed.** The v0.1.0 custom-keyword accessor layer
(`AlkTypeKind::FromStr`, `parse_*`, `resolve_ref*`,
`DiscriminatorKind`) is removed. The BAST parser (`bast` module)
exposes a cleaner typed surface (`BastDoc`/`BastDef`/`BastStruct`/
`BastField`/`BastType`/etc.) that borrows from the source
`serde_json::Value` without cloning field data.
- **Public module surface.** New public modules: `bast`, `bast_meta`,
`bast_validation`, `builder`, `materialize`. The `schema` module is
retained but now holds only `Endian`/`AlkTypeKind`/`VariableEncoding`
(the binary-kind vocabulary); the v0.1.0 custom-keyword machinery is
gone.
### Additions
- **BAST meta-schema.** Embedded in the crate as
`BAST_META_SCHEMA` (re-exported from the crate root) and published at
`https://alk.dev/bast/v1/schema`. BAST documents are validated against
it at compile time (`AlkTypeEngine::compile` calls
`validate_bast_doc` before parsing).
- **`materialize` module.** Materializes a `serde_json::Value` tree from
a binary buffer by walking the BAST typed tree. Used by
`AlkTypeEngine::validate_bytes` (ADR-010).
- **Builder for JSON Schemas.** `Schema::object()` produces a standard
JSON Schema object (for the `validate_json` path), complementing
`Schema::struct_()` which produces a BAST struct (for the bytes path).
One builder, two output shapes — the method name selects which.
### Bug fixes vs v0.1.0
- **Enum index bounds are now checked.** The v0.1.0 validator had a
dead constraint: enum variant indices were never bounds-checked
against `values.len()`. The BAST-native validator enforces it
(`validate_enum` checks `idx < values.len()`).
- **Offset-indirect, field-level endian, and aligned materialization**
bugs found during review #003 are fixed.
### Non-breaking improvements
- `$ref` is restricted to `#/$defs/<name>` — one hash lookup, no
`normalize_refs` pass (the v0.1.0 engine needed one).
- Schemas remain untrusted input: every engine path that walks a BAST
document returns `Err` on a malformed document, never `panic!`/
`unreachable!`. Overflow-safe arithmetic (`checked_add`,
`usize::try_from`) on all offset/count casts.
- Still two dependencies (`jsonschema` with `default-features = false`,
`serde_json` with `preserve_order`), no `async`, no `unsafe`, no
platform deps, no feature flags. Compiles to
`wasm32-unknown-unknown`.
### Upgrade notes
There is no migration path from v0.1.0 `AlkType:*` schemas — the format
is incompatible. Rewrite schemas as BAST documents (the `builder` API
produces them; see the README usage example) and update `compile` calls
to pass the root type name and the optional JSON Schema. The
read/write/validate API surface (`read_field`, `write_field`,
`sequential_reader`, `validate_bytes`, `validate_json`,
`is_valid_json`) is unchanged.
## [0.1.0] - 2025-11-10
Initial crates.io release. Custom-keyword JSON Schema format
(`AlkType:*`), single `jsonschema` validator for both bytes and JSON,
`AlkTypeEngine` with packed/aligned layout modes, builder API producing
`serde_json::Value`.
[0.3.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.3.0
[0.2.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.2.0
[0.1.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.1.0
Generated
+188 -1
View File
@@ -27,8 +27,9 @@ dependencies = [
[[package]]
name = "alktype"
version = "0.1.0"
version = "0.4.0"
dependencies = [
"criterion",
"jsonschema",
"serde_json",
]
@@ -39,6 +40,18 @@ version = "0.2.21"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "683d7910e743518b0e34f1186f92494becacb047c7b6bf616c96772180fef923"
[[package]]
name = "anes"
version = "0.1.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4b46cbb362ab8752921c97e041f5e366ee6297bd428a31275b9fcf1e380f7299"
[[package]]
name = "anstyle"
version = "1.0.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "940b3a0ca603d1eade50a4846a2afffd5ef57a9feac2c0e2ec2e14f9ead76000"
[[package]]
name = "autocfg"
version = "1.5.1"
@@ -84,12 +97,107 @@ version = "0.6.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "175812e0be2bccb6abe50bb8d566126198344f707e304f45c648fd8f2cc0365e"
[[package]]
name = "cast"
version = "0.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "37b2a672a2cb129a2e41c10b1224bb368f9f37a2b16b612598138befd7b37eb5"
[[package]]
name = "cfg-if"
version = "1.0.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
[[package]]
name = "ciborium"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "42e69ffd6f0917f5c029256a24d0161db17cea3997d185db0d35926308770f0e"
dependencies = [
"ciborium-io",
"ciborium-ll",
"serde",
]
[[package]]
name = "ciborium-io"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "05afea1e0a06c9be33d539b876f1ce3692f4afea2cb41f740e7743225ed1c757"
[[package]]
name = "ciborium-ll"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "57663b653d948a338bfb3eeba9bb2fd5fcfaecb9e199e87e1eda4d9e8b240fd9"
dependencies = [
"ciborium-io",
"half",
]
[[package]]
name = "clap"
version = "4.6.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "473c7e07f409a8d772161724aa8db6a765a2532a70f9667eeb7b49d3d02fbdca"
dependencies = [
"clap_builder",
]
[[package]]
name = "clap_builder"
version = "4.6.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7b48fea5a88e9ae728a2dcbedbfc0e730f7d60da42e1cb049a83c9fb8b789889"
dependencies = [
"anstyle",
"clap_lex",
]
[[package]]
name = "clap_lex"
version = "1.1.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9"
[[package]]
name = "criterion"
version = "0.7.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e1c047a62b0cc3e145fa84415a3191f628e980b194c2755aa12300a4e6cbd928"
dependencies = [
"anes",
"cast",
"ciborium",
"clap",
"criterion-plot",
"itertools",
"num-traits",
"oorandom",
"regex",
"serde",
"serde_json",
"tinytemplate",
"walkdir",
]
[[package]]
name = "criterion-plot"
version = "0.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9b1bcc0dc7dfae599d84ad0b1a55f80cde8af3725da8313b528da95ef783e338"
dependencies = [
"cast",
"itertools",
]
[[package]]
name = "crunchy"
version = "0.2.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "460fbee9c2c2f33933d720630a6a0bac33ba7053db5344fac858d4b8952d77d5"
[[package]]
name = "data-encoding"
version = "2.11.0"
@@ -107,6 +215,12 @@ dependencies = [
"syn 2.0.119",
]
[[package]]
name = "either"
version = "1.18.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "252afb9ae5eaa683babdc6a068b3f5726eb19e05070c731f9b2a23a7c3e8ed34"
[[package]]
name = "email_address"
version = "0.2.9"
@@ -174,6 +288,17 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "half"
version = "2.7.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6ea2d84b969582b4b1864a92dc5d27cd2b77b622a8d79306834f1be5ba20d84b"
dependencies = [
"cfg-if",
"crunchy",
"zerocopy",
]
[[package]]
name = "hashbrown"
version = "0.16.1"
@@ -304,6 +429,15 @@ dependencies = [
"hashbrown 0.17.1",
]
[[package]]
name = "itertools"
version = "0.13.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "413ee7dfc52ee1a4949ceeb7dbc8a33f2d6c088194d9f922fb8318faf1f01186"
dependencies = [
"either",
]
[[package]]
name = "itoa"
version = "1.0.18"
@@ -479,6 +613,12 @@ version = "1.21.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
[[package]]
name = "oorandom"
version = "11.1.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
[[package]]
name = "outref"
version = "0.5.2"
@@ -628,6 +768,15 @@ version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cf54715a573b99ac80df0bc206da022bcd442c974952c7b9720069370852e21f"
[[package]]
name = "same-file"
version = "1.0.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "93fc1dc3aaa9bfed95e02e6eadabb4baf7e3078b0bd1b4d7b6b0b68378900502"
dependencies = [
"winapi-util",
]
[[package]]
name = "scopeguard"
version = "1.2.0"
@@ -733,6 +882,16 @@ dependencies = [
"zerovec",
]
[[package]]
name = "tinytemplate"
version = "1.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be4d6b5f19ff7664e8c98d03e2139cb510db9b0a60b55f8e8709b689d939b6bc"
dependencies = [
"serde",
"serde_json",
]
[[package]]
name = "unicode-general-category"
version = "1.1.0"
@@ -773,6 +932,16 @@ version = "0.8.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5c3082ca00d5a5ef149bb8b555a72ae84c9c59f7250f013ac822ac2e49b19c64"
[[package]]
name = "walkdir"
version = "2.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "29790946404f91d9c5d06f9874efddea1dc06c5efe94541a7d6863108e3a5e4b"
dependencies = [
"same-file",
"winapi-util",
]
[[package]]
name = "wasip2"
version = "1.0.4+wasi-0.2.12"
@@ -827,12 +996,30 @@ dependencies = [
"unicode-ident",
]
[[package]]
name = "winapi-util"
version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c2a7b1c03c876122aa43f3020e6c3c3ee5c05081c9a00739faf7503aeba10d22"
dependencies = [
"windows-sys",
]
[[package]]
name = "windows-link"
version = "0.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5"
[[package]]
name = "windows-sys"
version = "0.61.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ae137229bcbd6cdf0f7b80a31df61766145077ddf49416a728b02cb3921ff3fc"
dependencies = [
"windows-link",
]
[[package]]
name = "wit-bindgen"
version = "0.57.1"
+16 -3
View File
@@ -1,21 +1,34 @@
[package]
name = "alktype"
version = "0.1.0"
version = "0.4.0"
edition = "2021"
rust-version = "1.85"
license = "MIT OR Apache-2.0"
description = "Binary struct engine: takes a JSON Schema with AlkType:* custom keywords and produces an offset map, read/write functions, and validation"
description = "Binary struct engine: takes a BAST (Binary Abstract Syntax Tree) document and produces an offset map, read/write functions, and validation"
repository = "https://git.alk.dev/alkdev/alktype"
readme = "README.md"
keywords = ["binary", "jsonschema", "wire-format", "serialization", "layout"]
categories = ["encoding", "data-structures", "parsing"]
exclude = [".opencode/", "docs/reviews/", "docs/research/", "docs/sdd_process.md", "Cargo.lock", "AGENTS.md", "fuzz/"]
[lib]
name = "alktype"
[workspace]
members = ["."]
exclude = ["fuzz"]
[features]
default = []
[dependencies]
jsonschema = { version = "0.46", default-features = false }
serde_json = { version = "1", features = ["preserve_order"] }
serde_json = { version = "1", features = ["preserve_order"] }
[dev-dependencies]
serde_json = "1"
criterion = { version = "0.7", default-features = false }
[[bench]]
name = "wire_vs_bast"
harness = false
+184 -91
View File
@@ -1,51 +1,62 @@
# alktype
The binary struct engine: a small Rust crate that takes a JSON Schema
with `AlkType:*` custom keywords and produces an offset map, read/write
The binary struct engine: a small Rust crate that takes a BAST (Binary
Abstract Syntax Tree) document and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
`alktype` is a standalone crate with **two dependencies**: `jsonschema`
(for validation) and `serde_json` (for schema parsing). No tokio, no
platform deps, no `unsafe`. Compiles to `wasm32-unknown-unknown`.
(for JSON validation and BAST meta-schema validation) and `serde_json`
(for BAST document parsing). No tokio, no platform deps, no `unsafe`.
Compiles to `wasm32-unknown-unknown`.
## What it is
A JSON Schema annotated with `AlkType:*` custom keywords serves three
roles simultaneously:
BAST is a JSON document that describes binary data layouts using a
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
itself a valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser. See
[`docs/architecture/bast-format.md`](docs/architecture/bast-format.md)
for the normative format spec.
A BAST document serves three roles simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
| **Validation spec (bytes)** | Compiled `ValidationPlan` walk over the materialized `Value` (ADR-012) | Access time (`validate_bytes`) |
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map / packed layout) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
| **Wire access (packed)** | Compiled `ReadPlan` (ADR-011) — compile-once, no per-read schema walk | Access time (`SequentialReader`) |
No separate format definition, no separate parser, no separate
validator. The schema is the single source of truth for the binary
format. Adding a new field to a protocol is adding a property to the
schema JSON — the engine computes the new offsets automatically.
validator. The BAST document is the single source of truth for the
binary format. Adding a new field to a protocol is adding an entry to
the BAST `fields` array — the engine computes the new offsets
automatically.
This is the same principle as `#[repr(C)]` struct field access, but at
runtime from a portable JSON Schema instead of at compile time from
language-specific annotations. The schema is the ABI contract.
runtime from a portable JSON document instead of at compile time from
language-specific annotations. The BAST document is the ABI contract.
## Usage
Build the schema with the fluent Rust builder (ADR-009), compile it
once into an [`AlkTypeEngine`], then read/write fields at computed
Build the BAST document with the fluent Rust builder (ADR-009), compile
it once into an [`AlkTypeEngine`], then read/write fields at computed
offsets:
```rust
use alktype::{AlkTypeEngine, Endian, LayoutMode, Schema, FieldValue};
use alktype::{AlkTypeEngine, Definitions, Endian, LayoutMode, Schema, FieldValue};
// Channels' 8-byte chunk header: big-endian, packed mode.
let mut schema = Schema::struct_()
let doc = Definitions::new().build_doc("ChunkHeader", Schema::struct_()
.endian(Endian::Big)
.field("channel_id", Schema::uint32())
.field("length", Schema::uint32())
.build();
.field("length", Schema::uint32()));
let engine = AlkTypeEngine::compile(&mut schema, LayoutMode::Packed)?;
// `json_schema: None` — no JSON-validation path needed for a binary-only schema.
let engine = AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)?;
// Write a frame into a buffer. For fixed-size structs, the byte
// positions are a direct read off the layout — channel_id at 0,
@@ -55,8 +66,8 @@ let mut buf = vec![0u8; 8];
alktype::data_access::write_u32(&mut buf, 0, 42, "channel_id", Endian::Big)?;
alktype::data_access::write_u32(&mut buf, 4, 7, "length", Endian::Big)?;
// Validate the bytes against the schema in one call.
engine.validate_bytes(&buf)?; // materializes a Value, then validates
// Validate the bytes against the BAST document in one call.
engine.validate_bytes(&buf)?; // materializes a Value, then runs the BAST-native validator
// Read the frame back sequentially (packed mode is sequential by
// construction — variable-length fields shift subsequent fields).
@@ -67,43 +78,64 @@ assert_eq!(value, FieldValue::U32(42));
# Ok::<(), alktype::AlkTypeError>(())
```
Schemas may also be authored as plain `serde_json::json!{...}` literals
and passed directly to `AlkTypeEngine::compile` — the builder is a
construction convenience, not a requirement.
BAST documents may also be authored as plain `serde_json::json!{...}`
literals and passed directly to `AlkTypeEngine::compile` — the builder
is a construction convenience, not a requirement.
## The 19 `AlkType:*` kinds
```rust
use alktype::{AlkTypeEngine, LayoutMode};
use serde_json::json;
| Kind | Rust type | Size | Notes |
|------|-----------|-----:|-------|
| `AlkType:Int8` | `i8` | 1 | |
| `AlkType:Int16` | `i16` | 2 | endian-sensitive |
| `AlkType:Int32` | `i32` | 4 | endian-sensitive |
| `AlkType:Int64` | `i64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `AlkType:Uint8` | `u8` | 1 | |
| `AlkType:Uint16` | `u16` | 2 | endian-sensitive |
| `AlkType:Uint32` | `u32` | 4 | endian-sensitive; also the enum/string/bytes length-prefix width |
| `AlkType:Uint64` | `u64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `AlkType:Float32` | `f32` | 4 | endian-sensitive; NaN/inf rejected by validator |
| `AlkType:Float64` | `f64` | 8 | endian-sensitive; NaN/inf rejected by validator |
| `AlkType:Boolean` | `bool` | 1 | |
| `AlkType:Enum` | `u32` index | 4 | index into the schema's `"enum"` array |
| `AlkType:String` | length-prefixed UTF-8 | 4 + N | `[length: u32][bytes]` by default |
| `AlkType:Bytes` | length-prefixed raw bytes | 4 + N | `[length: u32][bytes]` by default |
| `AlkType:Timestamp` | length-prefixed RFC 3339 | 4 + N | non-strict string check (see inline docs) |
| `AlkType:Struct` | record of fields | composite | nested; field paths are dotted (`"header.version"`) |
| `AlkType:Union` | tagged union | composite | byte-offset or field-name discriminator |
| `AlkType:Array` | repeated element | composite | fixed-size elements with stride, or variable count |
| `AlkType:Record` | string-keyed map | composite | `[count: u32][key, value]...` |
let doc = json!({
"$defs": {
"ChunkHeader": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "channel_id", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
}
}
});
let engine = AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)?;
# Ok::<(), alktype::AlkTypeError>(())
```
The engine recognizes a kind when the schema object has a key starting
with `AlkType:` whose value is `true` (the boolean shorthand) or an
annotation object (e.g. `{ "AlkType:String": { "encoding": "offset-indirect" } }`).
## The 18 BAST kinds
| `kind` | Rust type | Size | Notes |
|--------|-----------|-----:|-------|
| `int8` | `i8` | 1 | |
| `int16` | `i16` | 2 | endian-sensitive |
| `int32` | `i32` | 4 | endian-sensitive |
| `int64` | `i64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `uint8` | `u8` | 1 | |
| `uint16` | `u16` | 2 | endian-sensitive |
| `uint32` | `u32` | 4 | endian-sensitive; also the enum/string/bytes length-prefix width |
| `uint64` | `u64` | 8 | endian-sensitive; JSON precision caveat (ADR-005) |
| `float32` | `f32` | 4 | endian-sensitive; NaN/inf rejected by validator |
| `float64` | `f64` | 8 | endian-sensitive; NaN/inf rejected by validator |
| `bool` | `bool` | 1 | `0x00`=false, `0x01`=true |
| `enum` | `u32` index | 4 | index into the `values` array; bounds-checked by the BAST-native validator |
| `string` | length-prefixed UTF-8 | 4 + N | `[length: u32][bytes]` by default |
| `bytes` | length-prefixed raw bytes | 4 + N | `[length: u32][bytes]` by default |
| `struct` | record of fields | composite | nested; field paths are dotted (`"header.version"`) |
| `union` | tagged union | composite | byte-offset or field-name discriminator |
| `array` | repeated element | composite | fixed-size elements with stride; `count` required in v1 (D-BAST-004) |
| `record` | string-keyed map | composite | `[count: u32][key, value]...` |
The 18 kinds map to the `AlkTypeKind` Rust enum. `AlkTypeKind::to_bast_str`/
`from_bast_str` convert between the enum and the lowercase BAST strings
(D-BAST-002). Only `struct`, `union`, and `enum` can appear as named
`$defs` entries; primitives, arrays, and records appear as field/element/
value types via [TypeRef](docs/architecture/bast-format.md#typeref).
## Two layout modes
The consumer selects the layout mode at engine construction time via
`AlkTypeEngine::compile(schema, mode)`. The same schema can be compiled
in either mode. Decided in ADR-002.
`AlkTypeEngine::compile(bast_doc, root_name, mode, json_schema)`. The
same BAST document can be compiled in either mode. Decided in ADR-002.
| Mode | Use case | Read API | Write API |
|------|----------|----------|-----------|
@@ -118,60 +150,114 @@ in either mode. Decided in ADR-002.
- **Aligned mode**: a 4-byte length prefix sits at a known offset; the
variable data is not part of the static layout. Offset indirection
(the metatensor blob pattern: `{offset, length}` pointing into a
separate data region) is opt-in via the `encoding` annotation.
separate data region) is opt-in via the field-level `encoding`
annotation. Fixed-size reservation via `maxLength` is also supported.
### TUnion discriminators
### Union discriminators
`AlkType:Union` supports two discriminator kinds (ADR-003):
`kind: "union"` supports two discriminator kinds (ADR-003):
- **Byte-offset** — a fixed-size integer at a known byte offset. The
SFTP `Packet` pattern: byte 0 is the type byte, bytes 1..N are the
variant struct. Mapping keys are stringified integers.
- **Field-name** — a named field within the struct. The TypeBox
- **Byte-offset** — a fixed-size integer (`uint8`/`uint16`/`uint32`) at
a known byte offset. The SFTP `Packet` pattern: byte 0 is the type
byte, bytes 1..N are the variant struct. Mapping keys are stringified
integers. With all-canonical numeric keys the compiled reader
dispatches on the raw integer (no per-read stringification).
- **Field-name** — a named field within the union. The TypeBox
`typedef.ts` pattern. Mapping keys are string values matching the
discriminator field's value.
discriminator field's value. The `fields` array declares the
discriminator field (D-BAST-005), which must be its first entry; the
variant must not re-declare it or any shared field. The builder lays
out the declared `fields` first, then the variant's own fields
(ADR-011 addendum) — builder, reader, materializer, and validator all
agree on that convention.
Variant `$ref`s are resolved lazily — no compile-time inlining step.
## Endianness
Per-schema, default little-endian. Set `"endian": "big"` on the
top-level schema (or via `Schema::endian(Endian::Big)`) and the engine
byte-swaps every multi-byte read/write accordingly. SFTP consumers
specify big-endian; channels' chunk header is big-endian.
Per-schema, default little-endian. Set `"endian": "big"` on the root
struct (or via `Schema::endian(Endian::Big)`) and the engine byte-swaps
every multi-byte read/write accordingly. Field-level `endian` overrides
the struct default. SFTP consumers specify big-endian; channels' chunk
header is big-endian.
## Validation
Two entry points on [`AlkTypeEngine`], one underlying `jsonschema`
validator (ADR-010):
Two entry points on [`AlkTypeEngine`], two validators for two input
types (ADR-VAL-SPLIT):
- `validate_bytes(&[u8])` — for raw byte buffers (channels' chunk
header, SFTP packets). Materializes a `Value` tree from the bytes via
the layout engine, then runs the compiled **`ValidationPlan`** (0.2.0
used an interpretive BAST walker; 0.3.0 compiles the value-domain
constraints — integer ranges, `maxLength`, enum index bounds, union
variant dispatch — once at compile time). No
`jsonschema` involvement; the BAST document is the complete
validation spec for bytes (D-BAST-006).
- `validate_json(&Value)` / `is_valid_json(&Value)` — for already-parsed
JSON (call's payload schemas).
- `validate_bytes(&[u8])` — materializes a `Value` tree from the bytes
via the layout engine, then validates that `Value`. Single-call binary
buffer validation.
JSON (call's `OperationSpec.input_schema` payloads). Validates
against a **standard `jsonschema::Validator`** compiled at
`AlkTypeEngine::compile` time from a consumer-provided JSON Schema
(the `json_schema: Option<&Value>` parameter). BAST is not involved —
BAST describes bytes, not JSON shape (D-BAST-007).
The validator is compiled once at load time; access-time validation is
a fast `is_valid()` check. High-throughput paths can skip validation;
security-sensitive paths can validate every frame.
Both paths return `AlkTypeError::Validation(jsonschema::ValidationError<'static>)`
— one uniform payload, one match arm (D-BAST-009).
Validation is opt-in per operation. High-throughput paths can skip it;
security-sensitive paths can validate every frame. The BAST-native
validator also fixes a v0.1.0 dead constraint: enum index bounds are
now checked (the materializer emits a numeric index; the validator
checks it against `values.len()`).
## BAST document shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003).
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001)
— it selects which `$defs` entry is the top-level type.
- `$ref` is restricted to `#/$defs/<name>` — one hash lookup, no
normalization pass.
The BAST meta-schema is embedded in the crate as `BAST_META_SCHEMA`
(re-exported from the crate root) and published at
`https://alk.dev/bast/v1/schema`. Consumers can validate a BAST
document's structure with any JSON Schema validator. See
[`docs/architecture/bast-format.md`](docs/architecture/bast-format.md)
for the full spec.
## Crate independence
`alktype` does **not** depend on any application or networking crate.
It defines its own types (`AlkTypeError`, `AlkTypeEngine`, `FieldValue`,
etc.) and is usable in contexts where networking doesn't exist — CLI
tools, test harnesses, schema-building utilities, and WASM targets. The
upcoming `alkcall` crate (the `alknet-call` + `alknet-channels`
unification) depends on `alktype` for both binary layout and JSON
payload schemas; `alktype` knows nothing about `alkcall`.
tools, test harnesses, schema-building utilities, and WASM targets.
## Schemas as untrusted input
The crate treats schemas as untrusted input. A malformed schema
returns `AlkTypeError::Schema` / `AlkTypeError::Offset` from any
engine path — never a panic. This matters for hub/spoke topologies
where the remote peer provides the schema (e.g. `alkcall` accepting an
`OperationSpec` from an arbitrary internet peer). All `unreachable!()`
sites in production code were converted to `Err` ahead of v0.1.0
(review #002, L2).
The crate treats BAST documents as untrusted input. A malformed
document returns `AlkTypeError::Schema` from any engine path — never a
panic. This matters for hub/spoke topologies where the remote peer
provides the schema (e.g. `alkcall` accepting an `OperationSpec` from
an arbitrary internet peer). Every `unreachable!()` site in production
code was converted to `Err` ahead of v0.1.0 (review #002, L2); the
BAST parser preserves this invariant — overflow-safe arithmetic
(`checked_add`, `usize::try_from`) on all offset/count casts.
0.3.0 adds compile-time bounds for adversarial schemas: array counts
≤ 2^16 elements, computed array sizes ≤ 2^26 bytes, `align` ≤ 4096,
`maxLength` ≤ 2^26, and a shared reference-graph guard that rejects
cyclic `$ref`s and >128-deep nesting in every public schema walker.
Adversarial buffers fail with `Access` errors at read time — the
materializers never preallocate from declared counts. Reviews #006,
#007, and #008 document the audit trail
([docs/reviews/](docs/reviews/)).
## Documentation
@@ -179,20 +265,27 @@ Architecture documentation lives under [`docs/architecture/`](docs/architecture/
- [Overview](docs/architecture/overview.md) — purpose, "schema is the
format" principle, dependencies, consumers, scope boundaries
- [Schema layer](docs/architecture/schema-layer.md) — the 19 kinds,
jsonschema custom keyword integration, schema annotations
- [BAST format](docs/architecture/bast-format.md) — **normative format
spec**: meta-schema, TypeRef, TypeDef shapes, validation model
- [Schema layer](docs/architecture/schema-layer.md) — the BAST parser
(`BastDoc`/`BastDef`/`BastType` typed tree), the 18 kinds, the
`AlkTypeKind` enum
- [Layout engine](docs/architecture/layout-engine.md) — offset
computation, the two layout modes, alignment, endianness
- [Data access](docs/architecture/data-access.md) — read/write
functions, TUnion dispatch, field paths, zero-copy access
- [Validation](docs/architecture/validation.md) — custom keyword
validators, `AlkTypeError`, load-time vs access-time validation
- [Validation](docs/architecture/validation.md) — the two-validator
model, `AlkTypeError`, load-time vs access-time validation
- [Builder](docs/architecture/builder.md) — fluent Rust API for
constructing alktype JSON Schemas at runtime
constructing BAST documents and standard JSON Schemas at runtime
- [Architecture decisions (ADRs)](docs/architecture/decisions/) —
purpose/scope, two layout modes, schema annotations, error handling,
int64/uint64 kinds, packed-mode read factory, TUnion in aligned mode,
builder API, `validate_bytes`
purpose/scope (ADR-001), BAST format (ADR-BAST), two-validator model
(ADR-VAL-SPLIT), two layout modes (ADR-002), schema annotations
(ADR-003), error handling (ADR-004), int64/uint64 kinds (ADR-005),
non-final inline variable fields (ADR-006), packed-mode read factory
(ADR-007), TUnion in aligned mode (ADR-008), builder API (ADR-009),
`validate_bytes` (ADR-010), compiled read plan (ADR-011), plan
fingerprinting + `ValidationPlan` (ADR-012)
## License
+645
View File
@@ -0,0 +1,645 @@
//! Informal speed comparison: hand-rolled codec logic vs alktype-driven
//! codec over the same wire shapes.
//!
//! History: this bench originated (uncommitted) in `alktty` as the
//! curiosity probe that surfaced review #004's 400x read gap — the
//! finding that drove the 0.3.0 compiled-forms release (ADR-011/012).
//! It now lives here so alktype owns its perf story. The alktty-only
//! async roundtrip group (tokio `ChunkReader`/`ChunkWriter` over a
//! duplex pipe) was dropped — that measures alktty's I/O stack, not
//! this engine.
//!
//! The alktype engine / layout / plans are built **once outside** the
//! measured routine, per the "build cost is paid once" framing.
//!
//! Shapes:
//!
//! - **ChunkHeader** — a 5-byte header (`stream_type: uint8`,
//! `length: uint32` big-endian). The original shape, kept so numbers
//! stay comparable with the historical series (review #004: 400x →
//! 0.3.0: ~18x on read p64).
//! - **Read** — the hand-rolled path mirrors
//! `ChunkReader::read_chunk_after_peek` minus the tokio I/O
//! (identical overhead on both sides): validate `stream_type <= 4`,
//! parse `u32::from_be_bytes`, slice the payload. The alktype path
//! drives `SequentialReader::read_next` over the `ChunkHeader`
//! struct, then slices the payload at the parsed length. Both
//! return a `&[u8]` payload view — no allocation in either measured
//! path.
//! - **Write** — serialize the 5-byte header. Hand-rolled mirrors
//! `ChunkWriter::write_chunk`'s header writes; alktype uses
//! `PackedLayout` offsets (built once) and
//! `data_access::write_u8`/`write_u32` at those offsets.
//! - **Packet** — a byte-offset-discriminator union
//! (`Read {handle, length}` / `Write {handle, length, data: bytes}`),
//! the SFTP-shaped case ADR-011's framing argument was about:
//! exercises `CompositePlan::Union` dispatch, variant walks, and
//! length-prefixed variable reads. The alktype consumer pattern is
//! the documented one: `read_next` on the root yields
//! `FieldValue::Union { discriminator, variant_start }`, the consumer
//! selects the pre-built reader for that variant and walks it over
//! `&buf[variant_start..]`.
//! - **validate_bytes** — `engine.validate_bytes` per buffer
//! (materialize + `ValidationPlan` walk, ADR-010/ADR-012 §3): the
//! read+validate-on-untrusted-stream shape `alkcall` cares about. No
//! hand comparator: a hand-rolled codec validates inline during the
//! (already measured) parse, while `validate_bytes` additionally
//! materializes a `Value` tree per buffer — the honest reading is the
//! absolute per-chunk cost.
//!
//! One-shots (paid once at startup, not per chunk):
//! `alktype_engine_compile` (dominated by BAST meta-schema
//! validation), `alktype_sequential_reader_new` (an `Arc::clone`),
//! `alktype_layout_build`.
//!
//! Two payload sizes (64 B, 4 KiB) so per-chunk fixed overhead is
//! visible separately from payload-copy cost.
//!
//! Run: `cargo bench --bench wire_vs_bast`
use criterion::{criterion_group, criterion_main, BenchmarkId, Criterion};
use std::hint::black_box;
use alktype::{
data_access, AlkTypeEngine, Endian, FieldValue, LayoutBuilder, LayoutMode, PackedLayout,
ReadPlan, SequentialReader,
};
/// Mirrors `alktty::wire::MAX_CHUNK_LEN` — the hand-rolled comparator
/// validates against the same cap the real codec enforces.
const MAX_CHUNK_LEN: u32 = 16 * 1024 * 1024;
/// The `ChunkHeader` BAST definition. The `StreamType` enum is
/// intentionally NOT used — BAST enums encode as `u32`, but the
/// on-wire `stream_type` is a `uint8`; both sides read it as `uint8`.
const CHUNK_HEADER_BAST: &str = r#"{
"$schema": "https://alk.dev/bast/v1/schema",
"$defs": {
"ChunkHeader": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "stream_type", "kind": "uint8" },
{ "name": "length", "kind": "uint32" }
]
}
}
}"#;
/// SFTP-shaped byte-discriminator union: one byte selects the variant,
/// `Write` carries a trailing length-prefixed `bytes` field. The root
/// struct wraps the union (`AlkTypeEngine::compile` requires a struct
/// root); mapping keys are the stringified `uint8` discriminator
/// values.
const PACKET_BAST: &str = r##"{
"$schema": "https://alk.dev/bast/v1/schema",
"$defs": {
"Packet": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "event", "kind": { "$ref": "#/$defs/Event" } }
]
},
"Event": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" }
}
},
"Read": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
},
"Write": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" },
{ "name": "data", "kind": "bytes" }
]
}
}
}"##;
// ---------------------------------------------------------------------------
// ChunkHeader fixtures
// ---------------------------------------------------------------------------
/// One chunk's worth of bytes on the wire: 5-byte header + payload.
fn make_chunk_bytes(stream_type: u8, payload: &[u8]) -> Vec<u8> {
let mut buf = Vec::with_capacity(5 + payload.len());
buf.push(stream_type);
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
buf.extend_from_slice(payload);
buf
}
/// Concatenate `n` chunks into one buffer, each with `payload_len` bytes.
fn make_chunk_stream(n: usize, payload_len: usize) -> Vec<u8> {
let payload = vec![0xA5u8; payload_len];
let mut buf = Vec::with_capacity(n * (5 + payload_len));
for i in 0..n {
let st = (i % 5) as u8;
buf.extend_from_slice(&make_chunk_bytes(st, &payload));
}
buf
}
// ---------------------------------------------------------------------------
// Hand-rolled chunk read: mirrors ChunkReader::read_chunk_after_peek minus
// the tokio I/O. Returns (stream_type, payload) so the compiler can't
// elide the work. Validates stream_type <= 4 and length <= MAX_CHUNK_LEN.
// ---------------------------------------------------------------------------
#[inline]
fn hand_read_header(buf: &[u8]) -> Option<(u8, u32)> {
if buf.len() < 5 {
return None;
}
let stream_type = buf[0];
if stream_type > 4 {
return None;
}
let length = u32::from_be_bytes([buf[1], buf[2], buf[3], buf[4]]);
if length > MAX_CHUNK_LEN {
return None;
}
Some((stream_type, length))
}
#[inline]
fn hand_read_chunk(buf: &[u8]) -> Option<(u8, &[u8])> {
let (st, len) = hand_read_header(buf)?;
let end = 5usize.checked_add(len as usize)?;
if buf.len() < end {
return None;
}
Some((st, &buf[5..end]))
}
/// Drive `hand_read_chunk` across `n` contiguous chunks in `buf`.
/// Returns the total payload bytes consumed (so the loop body is
/// meaningfully used and not optimized away).
fn hand_read_stream(buf: &[u8], n: usize) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
let (st, payload) = match hand_read_chunk(&buf[pos..]) {
Some(v) => v,
None => break,
};
total += payload.len();
pos += 5 + payload.len();
black_box(st);
}
black_box(total)
}
// ---------------------------------------------------------------------------
// alktype chunk read: SequentialReader over ChunkHeader. The reader is
// constructed once per benchmark group and reset() between chunks. After
// the header read, the payload is sliced at the parsed length — same as
// the hand-rolled path. We do NOT re-read a length prefix for the payload
// (that would be the double-prefix problem).
// ---------------------------------------------------------------------------
fn alktype_read_stream(buf: &[u8], n: usize, reader: &mut SequentialReader) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
reader.reset();
let st = match reader.read_next_borrowed(&buf[pos..]) {
Ok(Some((_, FieldValue::U8(v)))) => v,
_ => break,
};
let len = match reader.read_next_borrowed(&buf[pos..]) {
Ok(Some((_, FieldValue::U32(v)))) => v,
_ => break,
};
if len > MAX_CHUNK_LEN {
break;
}
let end = match 5usize.checked_add(len as usize) {
Some(e) if e <= buf.len() - pos => e,
_ => break,
};
let payload = &buf[pos + 5..pos + end];
total += payload.len();
pos += end;
black_box(st);
black_box(payload.as_ptr());
}
black_box(total)
}
// ---------------------------------------------------------------------------
// Hand-rolled chunk write: mirrors ChunkWriter::write_chunk's header
// writes into a caller-provided buffer. Writes `n` contiguous chunks.
// ---------------------------------------------------------------------------
fn hand_write_stream(out: &mut Vec<u8>, n: usize, payload_len: usize) {
let payload = vec![0xA5u8; payload_len];
for i in 0..n {
let st = (i % 5) as u8;
let start = out.len();
out.resize(start + 5 + payload_len, 0);
out[start] = st;
out[start + 1..start + 5].copy_from_slice(&(payload_len as u32).to_be_bytes());
out[start + 5..start + 5 + payload_len].copy_from_slice(&payload);
}
black_box(out.len());
}
// ---------------------------------------------------------------------------
// alktype chunk write: data_access::write_u8 / write_u32 at the
// PackedLayout offsets. The layout is built once per group and reused.
// Payload bytes are copied with the same slice copy as the hand-rolled
// path so the comparison isolates the header-encoding overhead.
// ---------------------------------------------------------------------------
fn alktype_write_stream(out: &mut Vec<u8>, n: usize, payload_len: usize, layout: &PackedLayout) {
let payload = vec![0xA5u8; payload_len];
let st_pos = layout.get("stream_type").expect("stream_type field").offset;
let len_pos = layout.get("length").expect("length field").offset;
for i in 0..n {
let start = out.len();
out.resize(start + 5 + payload_len, 0);
let _ = data_access::write_u8(out, start + st_pos, (i % 5) as u8, "stream_type");
let _ = data_access::write_u32(
out,
start + len_pos,
payload_len as u32,
"length",
Endian::Big,
);
out[start + 5..start + 5 + payload_len].copy_from_slice(&payload);
}
black_box(out.len());
}
// ---------------------------------------------------------------------------
// Packet fixtures: byte-disc union stream, alternating Read/Write
// variants. Wire layout per packet (packed, big-endian):
// Read: disc(1) + handle(4) + length(4) = 9 bytes
// Write: disc(1) + handle(4) + length(4) + len(4)+data = 13 + payload
// ---------------------------------------------------------------------------
fn make_packet_bytes(disc: u8, payload: &[u8]) -> Vec<u8> {
let mut buf = Vec::with_capacity(13 + 4 + payload.len());
buf.push(disc);
buf.extend_from_slice(&0x0102_0304u32.to_be_bytes());
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
if disc == 6 {
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
buf.extend_from_slice(payload);
}
buf
}
fn make_packet_stream(n: usize, payload_len: usize) -> Vec<u8> {
let payload = vec![0xA5u8; payload_len];
let mut buf = Vec::new();
for i in 0..n {
let disc = if i % 2 == 0 { 5u8 } else { 6u8 };
buf.extend_from_slice(&make_packet_bytes(disc, &payload));
}
buf
}
// ---------------------------------------------------------------------------
// Hand-rolled packet read: read the discriminator byte, match the
// variant, parse its fields directly. Returns bytes consumed.
// ---------------------------------------------------------------------------
fn hand_read_packet(buf: &[u8]) -> Option<usize> {
let disc = *buf.first()?;
match disc {
5 => {
if buf.len() < 9 {
return None;
}
let handle = u32::from_be_bytes(buf[1..5].try_into().ok()?);
let length = u32::from_be_bytes(buf[5..9].try_into().ok()?);
black_box((handle, length));
Some(9)
}
6 => {
if buf.len() < 13 {
return None;
}
let handle = u32::from_be_bytes(buf[1..5].try_into().ok()?);
let length = u32::from_be_bytes(buf[5..9].try_into().ok()?);
let data_len = u32::from_be_bytes(buf[9..13].try_into().ok()?);
let end = 13usize.checked_add(data_len as usize)?;
if buf.len() < end {
return None;
}
black_box((handle, length));
black_box(&buf[13..end].as_ptr());
Some(end)
}
_ => None,
}
}
fn hand_read_packet_stream(buf: &[u8], n: usize) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
let Some(consumed) = hand_read_packet(&buf[pos..]) else {
break;
};
total += consumed;
pos += consumed;
}
black_box(total)
}
// ---------------------------------------------------------------------------
// alktype packet read: the documented union consumer contract. The root
// reader walks the wrapping struct; `read_next` returns
// `FieldValue::Union { discriminator, variant_start }`; the consumer
// selects the pre-built reader for that variant and walks it over
// `&buf[pos + variant_start..]` until exhausted.
// ---------------------------------------------------------------------------
/// Walk one variant's fields to exhaustion; returns bytes consumed.
/// Uses `read_next_borrowed` — the zero-alloc hot-loop pattern for
/// consumers that match or discard the field name.
fn alktype_walk_variant(reader: &mut SequentialReader, buf: &[u8]) -> Option<usize> {
reader.reset();
loop {
match reader.read_next_borrowed(buf) {
Ok(Some((name, value))) => {
black_box(name);
black_box(&value);
}
Ok(None) => return Some(reader.position()),
Err(_) => return None,
}
}
}
fn alktype_read_packet_stream(
buf: &[u8],
n: usize,
packet: &mut SequentialReader,
read: &mut SequentialReader,
write: &mut SequentialReader,
) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
packet.reset();
let disc = match packet.read_next_borrowed(&buf[pos..]) {
Ok(Some((_, FieldValue::Union {
discriminator,
variant_start,
}))) => {
pos += variant_start;
discriminator
}
_ => break,
};
let vbuf = &buf[pos..];
let consumed = match disc.as_str() {
"5" => alktype_walk_variant(read, vbuf),
"6" => alktype_walk_variant(write, vbuf),
_ => break,
};
let Some(consumed) = consumed else {
break;
};
total += consumed;
pos += consumed;
}
black_box(total)
}
/// One-time sanity check (outside the measured loops): the union
/// consumer pattern the stream loop relies on — root reader reports the
/// mapping key and the variant start; the variant reader's walk to
/// exhaustion reports exactly the variant's byte size, so
/// `variant_start + consumed` lands on the next packet.
fn assert_packet_reader_parity(
payload_len: usize,
packet: &mut SequentialReader,
read: &mut SequentialReader,
write: &mut SequentialReader,
) {
let payload = vec![0u8; payload_len];
let read_pkt = make_packet_bytes(5, &payload);
packet.reset();
match packet.read_next_borrowed(&read_pkt) {
Ok(Some((_, FieldValue::Union {
discriminator,
variant_start,
}))) => {
assert_eq!(discriminator, "5");
assert_eq!(variant_start, 1, "variant starts after the 1-byte disc");
}
_ => panic!("expected union value for Read packet"),
}
let consumed = alktype_walk_variant(read, &read_pkt[1..]).expect("read variant walk");
assert_eq!(consumed, 8, "Read = handle(4) + length(4)");
assert_eq!(1 + consumed, read_pkt.len(), "Read packet fully consumed");
let write_pkt = make_packet_bytes(6, &payload);
packet.reset();
match packet.read_next_borrowed(&write_pkt) {
Ok(Some((_, FieldValue::Union {
discriminator,
variant_start,
}))) => {
assert_eq!(discriminator, "6");
assert_eq!(variant_start, 1);
}
_ => panic!("expected union value for Write packet"),
}
let consumed = alktype_walk_variant(write, &write_pkt[1..]).expect("write variant walk");
assert_eq!(
consumed,
12 + payload_len,
"Write = handle(4) + length(4) + len-prefix(4) + data"
);
assert_eq!(1 + consumed, write_pkt.len(), "Write packet fully consumed");
}
// ---------------------------------------------------------------------------
// Benchmarks
// ---------------------------------------------------------------------------
fn bench_read(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
let engine =
AlkTypeEngine::compile(&bast, "ChunkHeader", LayoutMode::Packed, None).expect("compile");
let mut reader = engine.sequential_reader().expect("packed reader");
let mut group = c.benchmark_group("read_chunk_stream");
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let n = 1024;
let buf = make_chunk_stream(n, payload_len);
group.bench_with_input(BenchmarkId::new("hand_rolled", label), &n, |b, &n| {
b.iter(|| hand_read_stream(black_box(&buf), n));
});
group.bench_with_input(BenchmarkId::new("alktype", label), &n, |b, &n| {
b.iter(|| alktype_read_stream(black_box(&buf), n, &mut reader));
});
}
group.finish();
}
fn bench_write(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
let builder = LayoutBuilder::new(&bast, "ChunkHeader").expect("builder");
let layout = builder
.build(&std::collections::HashMap::new())
.expect("layout");
let mut group = c.benchmark_group("write_chunk_stream");
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let n = 1024;
group.bench_with_input(
BenchmarkId::new("hand_rolled", label),
&(n, payload_len),
|b, &(n, pl)| {
b.iter(|| {
let mut out = Vec::with_capacity(n * (5 + pl));
hand_write_stream(&mut out, n, pl);
});
},
);
group.bench_with_input(
BenchmarkId::new("alktype", label),
&(n, payload_len),
|b, &(n, pl)| {
b.iter(|| {
let mut out = Vec::with_capacity(n * (5 + pl));
alktype_write_stream(&mut out, n, pl, &layout);
});
},
);
}
group.finish();
}
fn bench_packet_read(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(PACKET_BAST).expect("bast json");
let engine =
AlkTypeEngine::compile(&bast, "Packet", LayoutMode::Packed, None).expect("compile");
let mut packet_reader = engine.sequential_reader().expect("packed reader");
let read_plan = std::sync::Arc::new(ReadPlan::compile(&bast, "Read").expect("read plan"));
let write_plan = std::sync::Arc::new(ReadPlan::compile(&bast, "Write").expect("write plan"));
let mut read_reader = SequentialReader::new(read_plan);
let mut write_reader = SequentialReader::new(write_plan);
// One-time parity check of the union consumer pattern (not measured).
assert_packet_reader_parity(64, &mut packet_reader, &mut read_reader, &mut write_reader);
let mut group = c.benchmark_group("read_packet_stream");
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let n = 1024;
let buf = make_packet_stream(n, payload_len);
group.bench_with_input(BenchmarkId::new("hand_rolled", label), &n, |b, &n| {
b.iter(|| hand_read_packet_stream(black_box(&buf), n));
});
group.bench_with_input(BenchmarkId::new("alktype", label), &n, |b, &n| {
b.iter(|| {
alktype_read_packet_stream(
black_box(&buf),
n,
&mut packet_reader,
&mut read_reader,
&mut write_reader,
)
});
});
}
group.finish();
}
fn bench_validate(c: &mut Criterion) {
let header_bast: serde_json::Value =
serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
let header_engine = AlkTypeEngine::compile(&header_bast, "ChunkHeader", LayoutMode::Packed, None)
.expect("compile");
let packet_bast: serde_json::Value = serde_json::from_str(PACKET_BAST).expect("bast json");
let packet_engine =
AlkTypeEngine::compile(&packet_bast, "Packet", LayoutMode::Packed, None).expect("compile");
let mut group = c.benchmark_group("validate_stream");
let n = 1024;
let headers: Vec<Vec<u8>> = (0..n)
.map(|i| make_chunk_bytes((i % 5) as u8, &[0xA5u8; 64]))
.collect();
group.bench_function("alktype_chunk_header", |b| {
b.iter(|| {
for h in &headers {
header_engine.validate_bytes(black_box(h)).expect("validate");
}
})
});
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let payload = vec![0xA5u8; payload_len];
let packets: Vec<Vec<u8>> = (0..n)
.map(|i| make_packet_bytes(if i % 2 == 0 { 5 } else { 6 }, &payload))
.collect();
group.bench_with_input(
BenchmarkId::new("alktype_packet", label),
&packets,
|b, packets| {
b.iter(|| {
for p in packets {
packet_engine.validate_bytes(black_box(p)).expect("validate");
}
})
},
);
}
group.finish();
}
/// One-shot costs paid once at startup, not per chunk.
fn bench_oneshot(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
c.bench_function("alktype_engine_compile", |b| {
b.iter(|| {
let _ =
AlkTypeEngine::compile(black_box(&bast), "ChunkHeader", LayoutMode::Packed, None)
.expect("compile");
});
});
c.bench_function("alktype_sequential_reader_new", |b| {
let engine = AlkTypeEngine::compile(&bast, "ChunkHeader", LayoutMode::Packed, None)
.expect("compile");
b.iter(|| engine.sequential_reader());
});
c.bench_function("alktype_layout_build", |b| {
let builder = LayoutBuilder::new(&bast, "ChunkHeader").expect("builder");
b.iter(|| builder.build(&std::collections::HashMap::new()));
});
}
criterion_group!(
benches,
bench_read,
bench_write,
bench_packet_read,
bench_validate,
bench_oneshot
);
criterion_main!(benches);
+75 -57
View File
@@ -1,12 +1,12 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype
The binary struct engine: a small Rust crate that takes a JSON Schema
with `AlkType:*` custom keywords and produces an offset map, read/write
The binary struct engine: a small Rust crate that takes a BAST (Binary
Abstract Syntax Tree) document and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
@@ -14,60 +14,75 @@ format definition; the engine is generic.
| Document | Status | Description |
|----------|--------|-------------|
| [overview.md](overview.md) | draft | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [schema-layer.md](schema-layer.md) | draft | The 19 `AlkType:*` kinds, jsonschema custom keyword integration, TypeBox interop, schema annotations |
| [overview.md](overview.md) | accepted | Crate purpose, "schema is the format" principle, dependencies, consumers, scope boundaries |
| [`bast-format.md`](bast-format.md) | accepted | **Normative BAST format specification.** Meta-schema, TypeRef, TypeDef shapes (Struct/Union/Enum/FieldDef), examples, validation model. The format the engine consumes. |
| [schema-layer.md](schema-layer.md) | accepted | The BAST parser (`src/bast.rs`) — the typed tree (`BastDoc`/`BastDef`/`BastType`/…) every engine module walks, the 18 BAST kinds, the `AlkTypeKind` enum, and the foundational annotation types. |
| [layout-engine.md](layout-engine.md) | draft | Offset computation, the two layout modes (packed sequential vs aligned static), alignment, endianness, variable-length handling |
| [data-access.md](data-access.md) | draft | Read/write functions, TUnion dispatch, field paths, zero-copy access, length-prefix reading |
| [validation.md](validation.md) | draft | Custom keyword validators for all 19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine`; `validate_bytes` for binary buffers (ADR-010) |
| [builder.md](builder.md) | draft | Fluent Rust API for constructing alktype JSON Schemas at runtime, producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema (ADR-009) |
| [validation.md](validation.md) | accepted | The two-validator model (BAST-native for `validate_bytes`, standard `jsonschema` for `validate_json`), `AlkTypeError`, load-time vs access-time validation, `AlkTypeEngine` as the compiled form of a BAST document (ADR-010, ADR-VAL-SPLIT). |
| [builder.md](builder.md) | accepted | Fluent Rust API for constructing BAST documents (`struct_()`) and standard JSON Schemas (`object()`) at runtime, producing `serde_json::Value` (ADR-009, D-BAST-008). |
### In-progress work
| Document | Status | Description |
|----------|--------|-------------|
| [BAST pivot — research record](../research/bast-pivot.md) | accepted | Motivation, POC scope and result, decisions D-BAST-001..009, risks for the BAST format pivot. Implemented in steps 1–10. |
| [BAST pivot — implementation plan](../plans/bast-implementation.md) | accepted | Ordered implementation steps, the public-API semver contract, and the ADR-sync checklist for the BAST pivot. Steps 1–10 complete. |
## Applicable ADRs
| ADR | Title | Relevance |
|-----|-------|-----------|
| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` |
| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Concrete JSON shapes for all schema-level annotations |
| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors |
| [001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Purpose, Scope, and the jsonschema Engine | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries. *Format-specific content superseded by ADR-BAST; purpose/scope retained.* |
| [BAST](decisions/bast-bast-format.md) | BAST (Binary Abstract Syntax Tree) as the Schema Format | The BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary. Supersedes ADR-001's format-specific content; records D-BAST-001..009. |
| [VAL-SPLIT](decisions/val-split-two-validator-model.md) | Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema. Records D-BAST-006/007/009. |
| [002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Two Layout Modes — Packed Sequential vs Aligned Static | The most important architectural finding; when to use each mode; `LayoutBuilder`/`SequentialReader` vs `OffsetMap` (format-agnostic — input format changed, modes didn't) |
| [003](decisions/003-schema-annotations.md) | Schema Annotations — Endianness, Alignment, Encoding, TUnion Discriminators | Annotation *semantics* (carry forward unchanged); annotation *location* moved to BAST type-level properties under the pivot |
| [004](decisions/004-error-handling-validation-strategy.md) | Error Handling and Validation Strategy | `AlkTypeError` enum (shape unchanged, D-BAST-009); load-time build, access-time check; field-path-carrying errors. *Validation-strategy section refined by ADR-VAL-SPLIT.* |
| [005](decisions/005-int64-uint64-first-class-kinds.md) | Int64/Uint64 as First-Class Kinds | 64-bit integers (SFTP offsets, metatensor data_offsets); JSON precision caveat |
| [006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Reject Non-Final Inline Length-Prefixed Variable Fields in Aligned Mode | Prevents silent data corruption (inline variable data clobbering subsequent fields) |
| [007](decisions/007-packed-mode-read-factory.md) | Packed-Mode Read API — Engine as SequentialReader Factory | `engine.sequential_reader()` returns an owned reader, not a reference |
| [008](decisions/008-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken |
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate; two methods on one struct, not a trait |
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003. *Output format amended to BAST / standard JSON Schema by ADR-BAST.* |
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate. *Validation step amended to the BAST-native validator by ADR-VAL-SPLIT.* |
| [011](decisions/011-compiled-read-plan-for-packed-mode.md) | Compiled Read Plan for Packed Mode | `ReadPlan` — the packed read-side compiled form, symmetric to `OffsetMap` (aligned) and `PackedLayout` (packed write). Closes review #004's 400x read-path gap; retires ADR-007's "re-parse on demand" framing. *Accepted — implemented in 0.3.0 (phases 1–2).* |
| [012](decisions/012-plan-fingerprinting-and-m1-closure.md) | Plan Fingerprinting, ValidationPlan, and Closing the Deferred M1 Sites in 0.3.0 | `ReadPlan`/`OffsetMap`/`ValidationPlan` `Hash + Eq` + `fingerprint()`; owned `BastDoc` (lifetime removal); `OffsetMap` carries `LeafMeta` to close the aligned-side M1 sites; `ValidationPlan` retires the interpretive `bast_validation` walk (review #005 M3 reversed the original deferral). Bundles with ADR-011 into one 0.3.0 breaking release. *Accepted — fully implemented in 0.3.0 (fingerprinting, owned `BastDoc`, `LeafMeta`, `ValidationPlan`).* |
## Relevant Open Questions
| OQ | Title | Status | Relevance |
|----|-------|--------|-----------|
| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it |
| OQ-001 | Arrays of variable-length-element structs | deferred(scope) | Requires lazy walking logic; blocked on a concrete consumer that needs it. BAST arrays require `count` in v1 (D-BAST-004), aligning with this deferral. |
| OQ-002 | `no_std` + `alloc` support | deferred(scope) | Target `std` for v1; blocked on an embedded use case |
| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Resolved in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) |
| OQ-004 | `Discriminator::Field` name — `&str` or `String` | open | Builder API ownership question; resolve before the SFTP Packet POC's field-name discriminator path |
| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | open | Blocks the SFTP Packet `validate_bytes` POC (next round) |
| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | open | Documentation fix in builder.md; the engine requires `AlkType:Struct` at the top level |
| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | open | Blocks the SFTP use case for `validate_bytes` (binary `handle`/`data` fields) |
| OQ-003 | Builder API for schema construction | resolved (ADR-009) | Shipped in v0.1.0; alkcall is the concrete consumer; see [builder.md](builder.md) |
| OQ-004 | `Discriminator::Field` name — `&str` or `String` | resolved | `String`, for ownership simplicity |
| OQ-005 | `Union` materialization shape — byte-offset vs field-name consistency | resolved | Both kinds return `{ "__discriminator": <value>, ...variant-fields }` |
| OQ-006 | Builder spec Example 3 — wrap `Union` in a `Struct` | resolved | [builder.md](builder.md) Example 3 wraps the union in a `Schema::struct_().field("payload", ...)` |
| OQ-007 | `Bytes` materialization — lossy UTF-8 conversion | resolved | Array of u8: materializer produces `Value::Array` of `Value::Number`; BAST-native validator accepts both `Value::String` and `Value::Array` |
| OQ-008 | `UnionValidator` variant dispatch | resolved | BAST-native validator recurses into the selected variant's BAST definition on `__discriminator` lookup — no custom keywords, no `inline_union_variant_refs` |
## Key Design Principles
1. **The schema is the format.** A JSON Schema with `AlkType:*` custom
keywords is both the validation spec and the layout spec. No separate
format definition, no separate parser, no separate validator. One
schema, three uses: validate, compute offsets, access data. See
[overview.md](overview.md) and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
1. **The schema is the format.** A BAST document is both the layout
spec and the validation spec for bytes. No separate format
definition, no separate parser, no separate validator. One schema,
three uses: validate, compute offsets, access data. See
[overview.md](overview.md), [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md),
and [ADR-BAST](decisions/bast-bast-format.md).
2. **jsonschema is the validation engine, not a custom engine.** The
`jsonschema` crate (v0.46.5, Draft 2020-12) handles validation with
custom keyword support. The novel code is the offset computation, not
the validation. This eliminates ~14,000 lines of hand-rolled schema
engines (typebox-rs, the @alkdev/alktype prototype). See [schema-layer.md](schema-layer.md)
and [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
2. **BAST is a JSON Schema dialect, not a custom format.** A BAST
document is valid JSON conforming to the BAST meta-schema (a
standard Draft 2020-12 JSON Schema). Any JSON Schema validator can
check whether a BAST document is well-formed; editors with JSON
Schema support provide autocomplete for free. See
[`bast-format.md`](bast-format.md) and
[ADR-BAST](decisions/bast-bast-format.md).
3. **Two layout modes for two use cases.** Packed sequential
(`LayoutBuilder`/`SequentialReader`) for protocol wire formats (SFTP,
channels, TTY). Aligned static (`OffsetMap`) for mmap-friendly formats
(metatensor). The consumer selects the mode; the schema is the same.
See [layout-engine.md](layout-engine.md) and
(metatensor). The consumer selects the mode; the BAST document is the
same. See [layout-engine.md](layout-engine.md) and
[ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md).
4. **Variable-length types default to inline length-prefixing.**
@@ -84,38 +99,41 @@ format definition; the engine is generic.
[ADR-003](decisions/003-schema-annotations.md).
6. **Endianness is per-schema, default little-endian.** The engine reads
the `"endian"` annotation and byte-swaps accordingly. SFTP consumers
specify `"endian": "big"`. See [layout-engine.md](layout-engine.md)
and [ADR-003](decisions/003-schema-annotations.md).
the struct-level `"endian"` annotation and byte-swaps accordingly.
SFTP consumers specify `"endian": "big"`. See
[layout-engine.md](layout-engine.md) and
[ADR-003](decisions/003-schema-annotations.md).
7. **Validation is opt-in, built once at load time.** The jsonschema
validator is compiled once at schema load time. Access-time validation
is a fast `is_valid()` check. High-throughput paths can skip
validation; security-sensitive paths can validate every frame. See
7. **Two validators for two input types.** `validate_bytes(&[u8])` uses
the BAST-native validator (a recursive walker over the BAST type
tree — no `jsonschema` involvement). `validate_json(&Value)` uses a
standard `jsonschema::Validator` from a consumer-provided JSON Schema
(BAST is not involved — BAST describes bytes, not JSON shape). One
`AlkTypeError::Validation` variant covers both (D-BAST-009). See
[validation.md](validation.md) and
[ADR-004](decisions/004-error-handling-validation-strategy.md).
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
8. **Not a serialization framework.** The alktype engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use alktype. See [overview.md](overview.md) and
computed offsets — no reflection, no dynamic dispatch per field. For
JSON data, use serde. For binary data with a known BAST document, use
alktype. See [overview.md](overview.md) and
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
9. **Schemas can be built at runtime from Rust (v0.1.0).** A fluent
builder API produces `serde_json::Value` for both AlkType-kind
schemas and standard JSON Schema, covering alkcall's two roles
9. **Schemas can be built at runtime from Rust.** A fluent builder API
produces `serde_json::Value` for both BAST documents (`struct_()`) and
standard JSON Schemas (`object()`), covering alkcall's two roles
(binary layout + JSON payloads) from one module. The builder is
additive — consumers with static schemas continue to load JSON.
See [builder.md](builder.md) and [ADR-009](decisions/009-builder-api.md).
additive — consumers with static BAST documents continue to load
JSON. See [builder.md](builder.md) and
[ADR-009](decisions/009-builder-api.md).
10. **Two validation entry points, one engine (v0.1.0).**
`validate_json(&Value)` for already-parsed JSON (call's payloads);
`validate_bytes(&[u8])` for binary buffers (channels' chunk header).
Same underlying `jsonschema` validator; the bytes path materializes
a `Value` tree via the layout engine, then validates. See
[validation.md](validation.md) and
[ADR-010](decisions/010-generalized-validation-validate-bytes.md).
10. **Two validation entry points, one engine.** `validate_json(&Value)`
for already-parsed JSON (call's payloads); `validate_bytes(&[u8])`
for binary buffers (channels' chunk header). Different validators,
one `AlkTypeError::Validation` variant. See [validation.md](validation.md),
[ADR-010](decisions/010-generalized-validation-validate-bytes.md),
and [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
## References
@@ -135,4 +153,4 @@ format definition; the engine is generic.
> **Note**: The research findings, POC code, and prior-attempt paths above
> refer to the parent `@alkdev/alknet` workspace where this crate originated.
> They are preserved here as historical context for the architectural
> decisions; the artifacts themselves are not part of this standalone repo.
> decisions; the artifacts themselves are not part of this standalone repo.
+701
View File
@@ -0,0 +1,701 @@
---
status: draft
last_updated: 2026-08-15
---
# alktype — BAST Format
**BAST** (Binary Abstract Syntax Tree) is alktype's schema format: a
JSON document that describes binary data layouts using a `kind`-based
vocabulary with `$defs`/`$ref` for composition. BAST replaces the
v0.1.0 `AlkType:*` custom-keyword JSON Schema format.
This document is the **normative format specification**. It is grounded
in the POC on branch `bast-validator-poc` (commit `f371fe4`) and the
decisions D-BAST-001 through D-BAST-009 in
[the pivot research record](../research/bast-pivot.md#decisions). The
implementation plan is
[`docs/plans/bast-implementation.md`](../plans/bast-implementation.md).
Until the BAST pivot lands in code, [`schema-layer.md`](schema-layer.md)
describes the *current* (custom-keyword) schema layer. This document
describes the *target* (BAST) schema layer. They coexist temporarily;
the implementation plan's final step retires `schema-layer.md`'s
custom-keyword content.
## Design Principles
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
Schema). Any JSON Schema validator can check whether a BAST document
is well-formed; editors with JSON Schema support provide autocomplete
and inline validation for free.
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
top-level `$defs` block. `$ref` handles cross-references and union
variant references — the same pattern as JSON Schema's own `$defs`
and TypeBox's `Type.Module`. No custom reference resolution mechanism.
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
This replaces the `AlkType:*` custom-keyword pattern with a flat,
easily-matched string. The 18 `AlkTypeKind` enum variants are
unchanged; `AlkTypeKind::from_str`/`to_str` map between the enum and
the lowercase BAST strings (D-BAST-002).
4. **Order is explicit.** Struct fields are an ordered array, not an
object with `properties`. Field order is unambiguous — no reliance on
`serde_json`'s `preserve_order` for correctness — and matches the
mental model of binary layouts.
5. **Annotations are type-level properties.** Endianness, alignment,
encoding, and discriminators are properties of the type definition
or field, not custom keywords on a separate schema object. Their
*semantics* carry forward unchanged from ADR-003; only their
*location* moves.
## Document Shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry. A bare struct at the top level
would be a different shape with different parsing logic and no home
for additional definitions — rejected.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode)` (D-BAST-001). It
selects which `$defs` entry is the top-level type. Convention (first
entry) is fragile and depends on JSON key order; a `$root` marker is
redundant with an explicit parameter.
## The Meta-Schema
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
validates the *structure* of BAST documents (is it well-formed?). It
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
in the crate for offline use. A different validator — the BAST-native
validator (see [Validation Model](#validation-model) below) — validates
*binary data* against a BAST document (are the bytes a valid instance?).
These are different validators for different inputs.
```json
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://alk.dev/bast/v1/schema",
"title": "Binary Abstract Syntax Tree (BAST) v1",
"description": "Meta-schema for BAST documents. A BAST document describes the binary layout of structured data.",
"type": "object",
"properties": {
"$defs": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeDef" }
}
},
"required": ["$defs"],
"$defs": {
"TypeDef": {
"oneOf": [
{ "$ref": "#/$defs/StructDef" },
{ "$ref": "#/$defs/UnionDef" },
{ "$ref": "#/$defs/EnumDef" }
]
},
"StructDef": {
"type": "object",
"properties": {
"kind": { "const": "struct" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
}
},
"required": ["kind", "fields"],
"additionalProperties": false
},
"FieldDef": {
"type": "object",
"properties": {
"name": { "type": "string", "pattern": "^[a-zA-Z_][a-zA-Z0-9_]*$" },
"kind": { "$ref": "#/$defs/TypeRef" },
"endian": { "enum": ["little", "big"] },
"align": { "type": "integer", "minimum": 1 },
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
"maxLength": { "type": "integer", "minimum": 0 }
},
"if": {
"properties": {
"kind": { "enum": ["string", "bytes"] }
}
},
"else": { "properties": { "maxLength": false } },
"required": ["name", "kind"],
"additionalProperties": false
},
"TypeRef": {
"oneOf": [
{
"description": "Primitive type",
"type": "string",
"enum": [
"int8", "int16", "int32", "int64",
"uint8", "uint16", "uint32", "uint64",
"float32", "float64",
"bool", "string", "bytes"
]
},
{
"description": "Reference to a named $defs entry",
"type": "object",
"properties": {
"$ref": { "type": "string", "pattern": "^#/\\$defs/[a-zA-Z_][a-zA-Z0-9_]*$" }
},
"required": ["$ref"],
"additionalProperties": false
},
{
"description": "Array type (fixed-size only in v1 — count is required)",
"type": "object",
"properties": {
"kind": { "const": "array" },
"element": { "$ref": "#/$defs/TypeRef" },
"count": { "type": "integer", "minimum": 0 }
},
"required": ["kind", "element", "count"],
"additionalProperties": false
},
{
"description": "Record (string-keyed map) type",
"type": "object",
"properties": {
"kind": { "const": "record" },
"values": { "$ref": "#/$defs/TypeRef" }
},
"required": ["kind", "values"],
"additionalProperties": false
}
]
},
"UnionDef": {
"type": "object",
"properties": {
"kind": { "const": "union" },
"endian": { "enum": ["little", "big"] },
"fields": {
"type": "array",
"items": { "$ref": "#/$defs/FieldDef" }
},
"discriminator": {
"oneOf": [
{
"type": "object",
"properties": {
"kind": { "const": "byte" },
"offset": { "type": "integer", "minimum": 0 },
"type": { "enum": ["uint8", "uint16", "uint32"] }
},
"required": ["kind", "offset", "type"],
"additionalProperties": false
},
{
"type": "object",
"properties": {
"kind": { "const": "field" },
"name": { "type": "string" }
},
"required": ["kind", "name"],
"additionalProperties": false
}
]
},
"mapping": {
"type": "object",
"additionalProperties": { "$ref": "#/$defs/TypeRef" }
}
},
"required": ["kind", "discriminator", "mapping"],
"additionalProperties": false
},
"EnumDef": {
"type": "object",
"properties": {
"kind": { "const": "enum" },
"values": {
"type": "array",
"items": { "type": "string" },
"minItems": 1
}
},
"required": ["kind", "values"],
"additionalProperties": false
}
}
}
```
## Type Definitions
### Struct
```json
{
"kind": "struct",
"endian": "big",
"align": 256,
"fields": [
{ "name": "channel_id", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
}
```
- `kind` (required): `"struct"`.
- `endian` (optional): `"little"` (default) or `"big"`. Sets the default
for all fields; field-level `endian` overrides.
- `align` (optional): struct-level alignment, only meaningful in aligned
static mode (ADR-002/003).
- `fields` (required): ordered array of [FieldDef](#fielddef). Array
position is field order — no reliance on JSON key order.
### FieldDef
```json
{ "name": "handle", "kind": "string", "encoding": "offset-indirect", "maxLength": 256 }
```
- `name` (required): identifier, `^[a-zA-Z_][a-zA-Z0-9_]*$`.
- `kind` (required): a [TypeRef](#typeref) — primitive string, `$ref`
object, array object, or record object.
- `endian` (optional): overrides the struct/union default for this field.
- `align` (optional): field-level alignment (aligned mode only).
- `encoding` (optional): `"length-prefixed"` (default) or
`"offset-indirect"`. See [Variable-length encoding](#variable-length-encoding).
- `maxLength` (optional, `string`/`bytes` fields only): byte-length
cap. See [Variable-length encoding](#variable-length-encoding).
Rejected at parse on any other kind (review #006 N3: the annotation
was silently unenforced there — the validation plan bakes `maxLength`
into string/bytes leaves only).
### TypeRef
`TypeRef` is the central mechanism for referencing types. Seven forms:
| Form | Example | Meaning |
|------|---------|---------|
| Primitive string | `"uint32"` | A built-in primitive (see [Primitives](#primitives)) |
| `$ref` object | `{ "$ref": "#/$defs/Read" }` | Reference to a named `$defs` entry |
| Array object | `{ "kind": "array", "element": "uint32", "count": 3 }` | Fixed-size array |
| Record object | `{ "kind": "record", "values": "string" }` | String-keyed map |
| Inline struct | `{ "kind": "struct", "fields": [...] }` | Anonymous struct |
| Inline union | `{ "kind": "union", ... }` | Anonymous union |
| Inline enum | `{ "kind": "enum", "values": [...] }` | Anonymous enum |
The `$ref` form uses standard JSON Pointer syntax **restricted to
`#/$defs/<name>`** — no external references, no fragment-only pointers,
no bare names. The restriction keeps resolution a single hash lookup
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
TypeBox's bare-name refs.
Arrays and records are inline type constructors, not top-level `$defs`
entries. Complex element types use nested `$ref`:
```json
{ "kind": "array", "element": { "$ref": "#/$defs/ComplexElement" }, "count": 4 }
```
Inline `struct`/`union`/`enum` TypeRefs are anonymous composites — a
field, array element, record value, or union variant whose type is
declared inline rather than named in `$defs`. They are structurally
identical to their named counterparts (same `StructDef`/`UnionDef`/
`EnumDef` shape); only the reference mechanism differs. Named composites
are preferred for reuse and for `$ref`-based dispatch; inline composites
are convenient for one-off nested types.
### Primitives
The 13 primitive `kind` strings map to the unchanged `AlkTypeKind`
variants (D-BAST-002 — lowercase strings, PascalCase enum variants):
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|-----------|---------------|-----------|------|----------|
| `int8` | `Int8` | `i8` | 1 | fixed |
| `int16` | `Int16` | `i16` | 2 | fixed |
| `int32` | `Int32` | `i32` | 4 | fixed |
| `int64` | `Int64` | `i64` | 8 | fixed |
| `uint8` | `Uint8` | `u8` | 1 | fixed |
| `uint16` | `Uint16` | `u16` | 2 | fixed |
| `uint32` | `Uint32` | `u32` | 4 | fixed |
| `uint64` | `Uint64` | `u64` | 8 | fixed |
| `float32` | `Float32` | `f32` | 4 | fixed |
| `float64` | `Float64` | `f64` | 8 | fixed |
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
`int64`/`uint64` are alktype additions (not in TypeBox's `typedef.ts`),
required by SFTP `offset: u64` and metatensor `data_offsets`. JSON
precision caveat per ADR-005 applies: integers beyond `2^53` lose
precision in `serde_json::Value::Number`; the binary path is exact.
### Enum
```json
{
"kind": "enum",
"values": ["Ok", "PermissionDenied", "NoSuchFile", "Failure"]
}
```
- `kind` (required): `"enum"`.
- `values` (required): non-empty array of strings, in declaration order.
- Binary representation: a `u32` index into `values` (0-based), encoded
per the struct's endianness. This is a deliberate deviation from
TypeBox's string enum in favor of binary efficiency — a `u32` index is
compact, fixed-size, and sufficient for any realistic enum.
**Bug fix vs v0.1.0:** The v0.1.0 engine has a dead constraint on the
bytes path — the built-in `enum` keyword checks string membership, but
the materializer emits `Value::Number(index)`, which never matches. The
BAST-native validator (see [Validation Model](#validation-model)) checks
the materialized index against `values.len()` bounds, fixing this.
### Union
```json
{
"kind": "union",
"endian": "big",
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "uint8"
},
"mapping": {
"1": { "$ref": "#/$defs/Init" },
"3": { "$ref": "#/$defs/Open" },
"5": { "$ref": "#/$defs/Read" }
}
}
```
- `kind` (required): `"union"`.
- `endian` (optional): default endianness for variant fields.
- `discriminator` (required): one of:
- **Byte-offset**: `{ "kind": "byte", "offset": <N>, "type": "uint8"|"uint16"|"uint32" }`.
The discriminator byte is at `offset`; the variant struct starts at
`offset + discriminator_size`. Mapping keys are stringified integers.
- **Field-name**: `{ "kind": "field", "name": "<field>" }`. The
discriminator is a length-prefixed string field; mapping keys are
string values matching the field's value. The union's `fields` array
(optional, only valid with field-name discriminators per D-BAST-005)
provides the field definitions including the discriminator field.
- `mapping` (required): object mapping discriminator values to
[TypeRef](#typeref) entries (typically `$ref` to `$defs` variants).
Variant `$ref`s are resolved **lazily** by the materializer and
validator — no `inline_union_variant_refs` compile step (removed under
BAST). The validator recurses into the selected variant's BAST
definition on `__discriminator` lookup, recovering the OQ-008
per-variant constraint enforcement (e.g., `maxLength` on a `bytes`
field inside a variant) without custom keywords.
### Array
```json
{ "kind": "array", "element": "float32", "count": 3 }
```
- `kind` (required): `"array"`.
- `element` (required): a [TypeRef](#typeref).
- `count` (required in v1): the fixed element count. **Arrays of
variable-length elements without a `count` are not supported in v1**
(D-BAST-004, aligning with OQ-001). The meta-schema enforces this:
`count` is in `required`. Variable-length collections are available
via `record` instead.
For fixed-size elements with a known count, the array size is
`element_size × count`. For variable-length elements (e.g.,
`"element": "string"`) with a known count, each element carries its own
length prefix — the array is count-prefixed in the sense that the count
is known at schema time, but the total byte size is not.
### Record
```json
{ "kind": "record", "values": "string" }
```
- `kind` (required): `"record"`.
- `values` (required): a [TypeRef](#typeref) for the value type.
Binary layout: a count-prefixed sequence of `(key, value)` pairs —
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
Each key is a length-prefixed UTF-8 string. Each value is encoded per
its `values` type. There is no separate `value_len` prefix — the value's
size is determined by its kind (fixed-size kinds have a known size;
variable-length kinds carry their own length prefix). The count and
key-length prefixes respect the struct's endianness.
## Variable-Length Encoding
The three strategies from ADR-003 carry forward with the same semantics,
expressed as field-level properties instead of keyword-value objects:
| Strategy | BAST syntax | Behavior |
|----------|------------|----------|
| Inline length-prefixed (default) | `{ "name": "handle", "kind": "string" }` | `[u32 length][data]` |
| Fixed-size reservation | `{ "name": "name", "kind": "string", "maxLength": 256 }` | Reserve `maxLength` bytes (aligned mode); validation constraint (packed mode) |
| Offset indirection | `{ "name": "blob", "kind": "bytes", "encoding": "offset-indirect" }` | `{offset: u32, length: u32}` pointing to separate data region |
**Default strategy selection (unchanged from v0.1.0):**
- **Packed sequential mode:** always inline length-prefixing.
`maxLength` is a validation constraint only.
- **Aligned static mode:** fixed-size reservation if `maxLength` is
declared; offset indirection if `"encoding": "offset-indirect"` is
declared; inline length-prefixing otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1
and 3) respects the effective endianness (struct default or field
override). In little-endian mode, `u32::from_le_bytes`; in big-endian
mode, `u32::from_be_bytes`. Ensures SFTP consumers (big-endian) have
consistent byte order for field values and length prefixes.
Applies to variable-length primitive types only: `string` and
`bytes`. The parser rejects `maxLength` (and the meta-schema forbids
it) on every other kind — including `record` (review #006 N3/M5: no
consumer honored it there, so the annotation was either silently
unenforced or, in aligned mode, silently corrupt).
## Endianness
Struct-level or union-level property with per-field override (same
semantics as ADR-003):
- Struct/union-level `"endian"` sets the default for all fields.
- Field-level `"endian"` overrides the struct/union default.
- Default is `"little"` when neither is specified.
- The length prefix for variable-length fields respects the effective
endianness.
```json
{
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "id", "kind": "uint32" },
{ "name": "handle", "kind": "string" },
{ "name": "crc", "kind": "uint32", "endian": "little" }
]
}
```
## Alignment
Struct-level or field-level property, only meaningful in aligned static
mode (same as ADR-003):
```json
{
"kind": "struct",
"align": 256,
"fields": [
{ "name": "header", "kind": { "$ref": "#/$defs/Header" } },
{ "name": "weight", "kind": "float32", "align": 16 }
]
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/
f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length
prefix), 1 for struct/union/array. Unchanged from v0.1.0.
- Ignored in packed sequential mode.
## Validation Model
BAST separates two concerns that the v0.1.0 format conflates, and in
doing so reveals that the engine has **two distinct validation paths**
with different inputs and guarantees. This is the validator split,
decided in D-BAST-006, D-BAST-007, and D-BAST-009. See
[`validation.md`](validation.md) for the current (pre-pivot) validation
layer; this section specifies the target model.
### Two validators, two inputs
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
**`validate_bytes` — bytes in, BAST is the validator.** The materializer
produces a `Value` tree from bytes. By construction, this `Value` is
*structurally correct*: all declared fields are present (the
materializer iterates the field list), types are correct (`read_u32`
produces `Value::Number`), bounds are checked (via `data_access::
check_bounds`), UTF-8 is valid (via `from_utf8`), the discriminator is
in the mapping, and the boolean byte is 0 or 1. What the materializer
does NOT check — and what the validation half checks afterward — are
**value-domain constraints expressed in the BAST document**. Under
ADR-012 §3 these constraints are compiled once into a `ValidationPlan`
at engine-compile time (eager `$ref` resolution, cyclic-graph
rejection); each `validate_bytes` call walks the compiled constraint
tree against the `Value`. The plan's nodes enforce exactly these
constraints (the set is normative; the walker that enforced it
interpretively in 0.2.0 is retired):
| Constraint | `ValidNode` arm |
|------------|-----------------|
| Integer range (Int8..Uint32) | `Int { min, max }` / `Uint { max }` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `I64` / `U64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `Float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `Str { max_len }` — `maxLength` baked in from the owning field at compile time |
| Bytes `maxLength` (array length) | `Bytes { max_len }` — accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `Enum { count }` checks `idx < count` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `Union { variants }` reads `__discriminator`, dispatches on the compiled variant nodes |
| Struct fields | `Struct { fields }` requires each declared field present, recurses |
| Array count | `Array { count, element }` checks `arr.len() == count` and recurses per element |
| Record values | `Record { values }` recurses into each value |
| Boolean | `Bool` (materializer already rejects non-0/1 bytes) |
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
describes both the layout (how to read) and the constraints (what
values are valid). This is the "schema is the format" principle from
ADR-001, now fully realized.
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
**`validate_json` — JSON in, JSON Schema is the validator.** The
consumer provides a JSON `Value` (e.g., an incoming JSON-RPC request).
The BAST document is irrelevant — BAST describes bytes, not JSON shape.
The right validator for a JSON value is a standard
`jsonschema::Validator` built from a standard JSON Schema document the
consumer provides. This is the path alkcall uses for its `OperationSpec`
JSON validation. No custom keywords; BAST is not involved.
### What is removed
Under the BAST pivot, the v0.1.0 validation machinery is removed from
the `validate_bytes` path:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation).
- `inline_union_variant_refs()` — union variant refs are resolved lazily
by the validator and materializer.
- `build_validator()`'s custom-keyword path — repurposed or removed (see
the implementation plan's step 6 for the decision on its fate).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path. The `validate_bytes` path no longer touches `jsonschema` — a
small wasm binary-size win in addition to the architecture
simplification. (Since ADR-012 §3, the interpretation step itself is
also compiled away: see the `ValidationPlan` above.)
### `AlkTypeError::Validation` payload shape
**Decided (D-BAST-009):** Keep
`Validation(jsonschema::ValidationError<'static>)`.
The `validate_bytes` path no longer uses `jsonschema`, so its error
payload is constructed via `jsonschema::ValidationError::custom` purely
to keep the variant's type unchanged. The rationale is consumer
ergonomics on the *combined* path: consumers like alkcall use both
`validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
channels) and handle `AlkTypeError::Validation` in one place. A single
uniform payload type means one match arm covers both sources. The
alternative (`Validation(String)`) would force `validate_json` to
flatten its structured errors (instance path, schema path, keyword) to
a `String` via `Display` — the more information-rich path loses data to
accommodate the less rich one. That is the wrong direction.
The `no_std`/minimal-build angle (OQ-002) that the alternative was
meant to enable is moot: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. The right place to
revisit is when/if OQ-002 is actually pursued.
## Relationship to JSON Schema and TypeBox
### BAST is a JSON Schema dialect
BAST is a specific JSON Schema instance format — like how JSON Schema
itself is a JSON document conforming to the JSON Schema meta-schema.
BAST documents conform to the BAST meta-schema. The entire JSON Schema
tooling ecosystem works with BAST:
- **Validation:** `jsonschema::options().build(&bast_meta_schema)?.validate(&bast_doc)`
- **Editors:** VSCode with `$schema` pointing to the BAST meta-schema URL
- **Documentation:** JSON Schema generators produce human-readable docs
from the meta-schema
### TypeBox interop
TypeBox's `Type.Module({...})` pattern maps naturally to BAST's `$defs`
structure. A TypeBox module defining binary types can serialize to BAST
JSON. The relationship:
- TypeBox → BAST JSON → alktype engine (binary layout)
- TypeBox → standard JSON Schema → jsonschema (JSON validation)
Same TypeBox source, two output formats, two validators.
### Not a replacement for JSON Schema
BAST does not replace JSON Schema for JSON data validation. A BAST
document cannot validate a JSON payload — it describes binary data
layouts and value-domain constraints for bytes. For JSON validation,
consumers use standard JSON Schema documents (which may be derived from
BAST via future codegen, or authored separately). The `validate_json`
path accepts a consumer-provided JSON Schema and uses a standard
`jsonschema::Validator` — BAST is not involved.
This is the split: `validate_bytes` is BAST-native (the BAST document
is both the layout spec and the validation spec for bytes);
`validate_json` is JSON-Schema-native (a standard JSON Schema is the
validation spec for JSON values). One crate, two validators, two input
types.
## Decisions
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
recorded in [the pivot research record](../research/bast-pivot.md#decisions).
The implementation-relevant summary:
| Decision | Summary |
|----------|---------|
| [D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
| [D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
| [D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
| [D-BAST-004](../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
| [D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
| [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
| [D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
| [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
| [D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
## References
- [Pivot research record](../research/bast-pivot.md) — motivation, POC
scope and result, decisions D-BAST-001..009, risks
- [Implementation plan](../plans/bast-implementation.md) — ordered
steps, public-API semver contract, ADR-sync checklist
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
(carry forward unchanged; only location moves)
- [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) — Int64/
Uint64 as first-class kinds; JSON precision caveat
- [`schema-layer.md`](schema-layer.md) — the current (v0.1.0) schema
layer; superseded by this document when the pivot lands
- [`validation.md`](validation.md) — the current (v0.1.0) validation
layer; rewritten for the validator split when the pivot lands
+240 -169
View File
@@ -1,29 +1,41 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype — Builder API
The builder layer: a fluent Rust API for constructing alktype JSON
Schemas (both `AlkType:*`-bearing binary-layout schemas and plain
JSON-Schema-only operation payload schemas) at runtime, producing
`serde_json::Value`. Decided in [ADR-009](decisions/009-builder-api.md);
resolves [OQ-003](questions/003-builder-api-for-schema-construction.md).
The builder layer: a fluent Rust API for constructing BAST documents
(binary-layout schemas) and standard JSON Schemas (JSON-validation
schemas) at runtime, producing `serde_json::Value`. Decided in
[ADR-009](decisions/009-builder-api.md); resolves
[OQ-003](questions/003-builder-api-for-schema-construction.md). The
two-output-format split is D-BAST-008, recorded in
[ADR-BAST](decisions/bast-bast-format.md).
## What
The `builder` module provides a single `Schema` builder type and a
`Definitions` helper for named `$defs`. The builder's `.build()` method
returns a `serde_json::Value` — the same form alktype already consumes
via `AlkTypeEngine::compile` (for `AlkType:*` schemas) and the same form
`OperationSpec.input_schema` / `output_schema` / `error_schemas` hold
(for plain JSON Schema, no `AlkType:*` kinds).
returns a `serde_json::Value` — one of two forms depending on the
constructor used (D-BAST-008):
- **BAST JSON** (binary layout) — `Schema::struct_().field(...).build()`
produces a BAST TypeDef (`{ "kind": "struct", "fields": [...] }`).
Primitive constructors produce bare TypeRef strings (`"uint32"`).
Feed to [`AlkTypeEngine::compile`](validation.md) (the binary-layout
path) → `validate_bytes`.
- **Standard JSON Schema** (JSON validation) — `Schema::object().field(...)`
produces `{ "type": "object", "properties": {...}, "required": [...] }`.
No BAST `kind`, no custom keywords — a plain JSON Schema. Feed to a
standard `jsonschema::Validator` (or `AlkTypeEngine::compile` with a
JSON Schema for the `validate_json` path, D-BAST-007).
The builder covers:
- All 19 `AlkType:*` kinds (binary-layout schemas) — see
[schema-layer.md](schema-layer.md) for the kinds.
- All 18 BAST kinds (binary-layout schemas) — see
[schema-layer.md](schema-layer.md) for the kinds and
[`bast-format.md`](bast-format.md) for the format.
- All standard JSON Schema keywords needed for operation payload
schemas: `type`, `properties`, `required`, `items`, `enum`, `format`,
`additionalProperties`, `minimum`, `maximum`, `minItems`,
@@ -39,16 +51,20 @@ alktype's first consumer and needs to build schemas at runtime from
Rust code, for two roles:
1. **Binary layout schemas** (channels' 8-byte chunk header, future
binary call frames) — `AlkType:*` schemas, fed to
`AlkTypeEngine::compile` (packed mode, big-endian).
binary call frames) — BAST documents, fed to
`AlkTypeEngine::compile` (packed mode, big-endian) →
`validate_bytes`.
2. **JSON payload schemas** (call's `OperationSpec.input_schema` /
`output_schema` / `error_schemas`) — plain JSON Schema, no
`AlkType:*` kinds, validated via the standard `jsonschema` validator.
`output_schema` / `error_schemas`) — plain JSON Schema, no BAST
`kind`, validated via the standard `jsonschema` validator
(`AlkTypeEngine::compile` with a JSON Schema → `validate_json`).
A single builder serving both roles means alkcall imports one module
for schema construction. See [ADR-009](decisions/009-builder-api.md)
for the decision rationale (why `Value` not a typed `Schema` enum, why
both AlkType and standard JSON Schema in one builder).
both BAST and standard JSON Schema in one builder) and
[ADR-BAST](decisions/bast-bast-format.md) for the two-output-format
decision (D-BAST-008).
## Architecture
@@ -62,35 +78,42 @@ duplicating the JSON form that `AlkTypeEngine::compile`,
### Module placement
`src/builder.rs`, re-exported from the crate root. The builder is a
peer of `schema.rs` (which parses schemas) and `engine.rs` (which
peer of `bast.rs` (which parses BAST documents) and `engine.rs` (which
compiles them). The builder constructs; it does not parse or compile.
```rust
// src/lib.rs (additions)
pub mod builder;
pub use builder::{Schema, Definitions};
pub use builder::{Schema, Definitions, Discriminator};
```
### Field order is load-bearing
### Field order is explicit
`serde_json` with `preserve_order` is already a dependency (ADR-001).
The builder's `Value` output uses `serde_json::Map` (which preserves
insertion order under `preserve_order`), so field declaration order in
the builder is the field order in the binary layout. This is critical
for packed mode (ADR-002) where field order determines offsets.
BAST struct fields are an ordered array (BAST design principle #4 —
see [`bast-format.md`](bast-format.md#design-principles)). The builder's
`struct_()`/`union_()` accumulates fields in call order and emits them
as the `fields` array on `.build()`. Field order in the builder is the
field order in the binary layout. This is critical for packed mode
(ADR-002) where field order determines offsets.
(`serde_json`'s `preserve_order` feature remains a dependency, but
layout correctness no longer depends on it — the `fields` array makes
order explicit. `preserve_order` is still load-bearing for the
`mapping` object's iteration order and for `Definitions`' `$defs`
block, which the parser walks in document order.)
## Public API
### `Schema` builder
`Schema` is the single entry point. Constructors for each AlkType kind
and each standard JSON Schema type; setters for annotations and
`Schema` is the single entry point. Constructors for each BAST kind and
each standard JSON Schema type; setters for annotations and
constraints; `.build()` produces `Value`.
#### AlkType kind constructors
#### BAST kind constructors
One constructor per `AlkTypeKind` variant (see [schema-layer.md](schema-layer.md)
§"The 19 AlkType Kinds"):
§"The 18 BAST Kinds"):
```rust
impl Schema {
@@ -112,43 +135,44 @@ impl Schema {
// Variable-length kinds
pub fn string() -> Self;
pub fn bytes() -> Self;
pub fn timestamp() -> Self;
// Composite kinds
pub fn struct_() -> Self; // fields added via .field()
pub fn union_(disc: Discriminator) -> Self; // variants via .mapping()
pub fn array_of(element: Schema) -> Self;
pub fn array_of(element: Schema) -> Self; // .count() required for valid BAST (D-BAST-004)
pub fn record_of(value: Schema) -> Self;
}
```
Each constructor sets the corresponding `"AlkType:<Kind>": true` key.
For example, `Schema::uint32()` produces `{"AlkType:Uint32": true}`.
Primitive constructors produce the bare BAST TypeRef string on
`.build()`. For example, `Schema::uint32().build()` produces `"uint32"`.
Composite constructors produce the BAST object form.
**`enum_of`** sets both `"AlkType:Enum": true` and the standard
`"enum"` keyword with the provided values (declaration order is the
index order — see [schema-layer.md](schema-layer.md) §"TEnum binary
representation"):
**`enum_of`** produces a BAST enum TypeDef (`{ "kind": "enum", "values":
[...] }`); declaration order is the index order — see
[schema-layer.md](schema-layer.md) §"The 18 BAST Kinds"):
```rust
Schema::enum_of(&["read", "write", "execute"])
// -> { "AlkType:Enum": true, "enum": ["read", "write", "execute"] }
Schema::enum_of(&["read", "write", "execute"]).build()
// -> { "kind": "enum", "values": ["read", "write", "execute"] }
```
**`array_of`** and **`record_of`** take the element/value schema as a
nested `Schema`:
nested `Schema`. `array_of` requires `.count(N)` for valid BAST
(D-BAST-004 — arrays of variable-length elements without a count are
deferred, aligning with OQ-001):
```rust
Schema::array_of(Schema::uint32())
// -> { "AlkType:Array": true, "items": { "AlkType:Uint32": true } }
Schema::array_of(Schema::uint32()).count(3).build()
// -> { "kind": "array", "element": "uint32", "count": 3 }
Schema::record_of(Schema::float32())
// -> { "AlkType:Record": true, "values": { "AlkType:Float32": true } }
Schema::record_of(Schema::float32()).build()
// -> { "kind": "record", "values": "float32" }
```
#### Standard JSON Schema type constructors
For plain JSON Schema (no `AlkType:*` kinds) — call's
`input_schema` / `output_schema` / `error_schemas`:
For plain JSON Schema (no BAST `kind`) — call's `input_schema` /
`output_schema` / `error_schemas`:
```rust
impl Schema {
@@ -163,22 +187,24 @@ impl Schema {
}
```
The `_` suffix disambiguates standard JSON Schema types from AlkType
kinds (`string` is the AlkType kind; `string_` is the standard JSON
Schema type — the AlkType kind constructor sets `"AlkType:String":
true`, the standard constructor sets `"type": "string"`). This is
deliberate: the two are distinct schema forms and the builder makes
the distinction visible at the call site.
The `_` suffix disambiguates standard JSON Schema types from BAST kinds
(`string` is the BAST primitive; `string_` is the standard JSON Schema
type — `string()` would produce `"string"` as a BAST TypeRef,
`string_()` produces `{ "type": "string" }` as a standard JSON Schema).
This is deliberate: the two are distinct schema forms and the builder
makes the distinction visible at the call site.
#### Annotation setters
Annotation setters mirror ADR-003. Each setter is named after the
annotation it produces; calling the setter sets the corresponding JSON
key. Setters return `Self` for chaining.
Annotation setters mirror ADR-003 (semantics unchanged; location moved
to BAST type-level properties under the pivot — see
[ADR-BAST](decisions/bast-bast-format.md)). Each setter is named after
the annotation it produces; calling the setter sets the corresponding
JSON key. Setters return `Self` for chaining.
```rust
impl Schema {
/// Schema-level endianness (ADR-003 §1). Default little.
/// Struct/union-level endianness (ADR-003 §1). Default little.
pub fn endian(mut self, endian: Endian) -> Self;
/// Struct or field alignment (ADR-003 §2). Struct-level sets the
@@ -193,12 +219,19 @@ impl Schema {
/// a variable-length type, reserves this many bytes (strategy 2).
/// In packed mode, validation constraint only.
pub fn max_length(mut self, max: usize) -> Self;
/// Array count (D-BAST-004 — required for valid BAST arrays in v1).
pub fn count(mut self, count: usize) -> Self;
}
```
`Endian` and `VariableEncoding` are re-exported from `schema.rs` (no
new types — the builder uses the existing enums). The setters produce
the exact JSON shapes from ADR-003:
new types — the builder uses the existing enums). When applied to a
struct, `endian`/`align` are struct-level; when the `Schema` is used as
a `.field()` argument, the builder extracts `endian`/`align`/`encoding`/
`maxLength` and places them on the *field* object (BAST field-level
properties). The setters produce the exact BAST JSON shapes from
[`bast-format.md`](bast-format.md):
```rust
Schema::struct_()
@@ -207,31 +240,34 @@ Schema::struct_()
.field("length", Schema::uint32())
.build()
// -> {
// "AlkType:Struct": true,
// "kind": "struct",
// "endian": "big",
// "properties": {
// "channel_id": { "AlkType:Uint32": true },
// "length": { "AlkType:Uint32": true }
// }
// "fields": [
// { "name": "channel_id", "kind": "uint32" },
// { "name": "length", "kind": "uint32" }
// ]
// }
```
#### Composite builders
`struct_()`, `union_()`, `array_of()`, `record_of()` are the
composite constructors. `struct_()` and `union_()` need additional
setters to populate their children:
`struct_()`, `union_()`, `array_of()`, `record_of()` are the composite
constructors. `struct_()` and `union_()` need additional setters to
populate their children:
```rust
impl Schema {
/// Add a field to a struct (or object). Field order is load-bearing
/// for binary layouts (packed mode field order = byte order).
/// Add a field to a struct (or a field-name-discriminator union).
/// Field order is load-bearing for binary layouts (packed mode
/// field order = byte order — the `fields` array is ordered).
/// The field's schema is built from the passed `Schema`.
pub fn field(mut self, name: &str, field: Schema) -> Self;
/// Mark fields as required (standard JSON Schema `required` keyword).
/// Can be called multiple times; required names accumulate.
/// Field names must have been added via `.field()`.
/// Only meaningful for `object()` (standard JSON Schema) — BAST
/// structs require all declared fields present (the validator
/// enforces this). Can be called multiple times; required names
/// accumulate.
pub fn required(mut self, names: &[&str]) -> Self;
/// Set the items schema for a standard `array` type.
@@ -247,9 +283,11 @@ impl Schema {
}
```
**`field`** sets `properties[name] = field.build()`. Repeated calls
append. Field order in the built `Value` is the call order (because
`serde_json::Map` preserves insertion order under `preserve_order`).
**`field`** appends a `{ "name": ..., "kind": <field.build()>, ... }`
entry to the struct/union's `fields` array, extracting field-level
annotations (`endian`, `align`, `encoding`, `maxLength`) from the
passed `Schema`. Repeated calls append in order. Field order in the
built `Value` is the call order.
**`required`** sets the standard JSON Schema `"required"` array. The
builder does not check that the named fields exist (that's a
@@ -260,11 +298,12 @@ accumulates names:
```rust
Schema::object()
.field("path", Schema::string_())
.field("path", Schema::string_().max_length(4096))
.field("offset", Schema::integer().minimum(0))
.field("length", Schema::integer().minimum(0))
.required(["path"])
.required(["offset", "length"])
.build()
// -> {
// "type": "object",
// "properties": { "path": {...}, "offset": {...}, "length": {...} },
@@ -280,35 +319,31 @@ For operation payload schemas (call's `input_schema` etc.):
impl Schema {
/// `minimum` (inclusive lower bound for numbers/integers).
pub fn minimum(mut self, min: f64) -> Self;
/// `maximum` (inclusive upper bound for numbers/integers).
pub fn maximum(mut self, max: f64) -> Self;
/// `minLength` (minimum string length).
pub fn min_length(mut self, min: usize) -> Self;
/// `minItems` (minimum array length).
pub fn min_items(mut self, min: usize) -> Self;
/// `maxItems` (maximum array length).
pub fn max_items(mut self, max: usize) -> Self;
/// `format` (e.g. "date-time", "uri", "email").
pub fn format(mut self, fmt: &str) -> Self;
/// `title` (human-readable description).
pub fn title(mut self, t: &str) -> Self;
/// `description` (human-readable description).
pub fn description(mut self, d: &str) -> Self;
}
```
These set the corresponding standard JSON Schema keywords. They apply
to both AlkType-kind schemas and standard JSON Schema type schemas
(e.g., `Schema::string().max_length(4096)` sets `maxLength`, which
serves as both a validation constraint and, in aligned mode, a
fixed-size reservation — ADR-003 §3).
to standard JSON Schema type schemas (e.g.,
`Schema::string_().max_length(4096)` sets `maxLength`, which on the
`validate_json` path is a JSON-Schema validation constraint). On a BAST
schema, `max_length` also serves as the aligned-mode fixed-size
reservation (ADR-003 §3) and the packed-mode validation constraint
(enforced by the BAST-native validator — see
[validation.md](validation.md)).
#### `.build()`
@@ -343,21 +378,22 @@ to compose them. `from_value` wraps the `Value` so it can be passed to
### `Discriminator` for `union_()`
`union_()` takes a `Discriminator` describing the union's dispatch
mechanism. This mirrors `schema.rs::DiscriminatorKind` but with a
builder-friendly shape (the kind enum is re-exported from `schema.rs`,
not duplicated):
mechanism. This mirrors `bast::BastDiscriminator` (the parser's typed
view) but with a builder-friendly shape:
```rust
pub enum Discriminator {
/// Byte-offset discriminator (ADR-003 §4 Kind A).
/// `offset` is the byte position; `disc_type` is the AlkType kind
/// `offset` is the byte position; `disc_type` is the BAST kind
/// of the discriminator (Uint8/Uint16/Uint32).
Byte {
offset: usize,
disc_type: AlkTypeKind, // restricted to Uint8/Uint16/Uint32
},
/// Field-name discriminator (ADR-003 §4 Kind B).
/// `name` is the field holding the discriminator value.
/// `name` is the field holding the discriminator value. The
/// discriminator field and any shared fields are declared via
/// `.field()` on the union builder.
Field {
name: String,
},
@@ -376,9 +412,13 @@ let packet = Schema::union_(Discriminator::Byte {
.mapping("101", Schema::ref_def("Status"))
.build();
// -> {
// "AlkType:Union": true,
// "discriminator": { "kind": "byte", "offset": 0, "type": "AlkType:Uint8" },
// "mapping": { "5": {"$ref":"#/$defs/Read"}, "6": {...}, "101": {...} }
// "kind": "union",
// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
// "mapping": {
// "5": { "$ref": "#/$defs/Read" },
// "6": { "$ref": "#/$defs/Write" },
// "101": { "$ref": "#/$defs/Status" }
// }
// }
```
@@ -386,22 +426,29 @@ let packet = Schema::union_(Discriminator::Byte {
```rust
let event = Schema::union_(Discriminator::Field { name: "type" })
.field("type", Schema::string())
.mapping("read", Schema::ref_def("Read"))
.mapping("write", Schema::ref_def("Write"))
.build();
// -> {
// "AlkType:Union": true,
// "kind": "union",
// "discriminator": { "kind": "field", "name": "type" },
// "fields": [ { "name": "type", "kind": "string" } ],
// "mapping": { "read": {...}, "write": {...} }
// }
```
(Field-name-discriminator unions require a `fields` array declaring the
discriminator field — D-BAST-005. The builder emits `fields` only when
the discriminator is `Field` and at least one field was added.)
### `Definitions` — named `$defs` for cross-reference
`Definitions` is a helper for building named `$defs` that schemas can
`$ref` by name. This is the ergonomics win for alkcall's
`OperationSpec`, where input/output/error schemas reference shared
definitions (e.g., `FileNotFound`, `RateLimited`).
`$ref` by name, and for assembling a complete BAST document. This is
the ergonomics win for alkcall's `OperationSpec`, where
input/output/error schemas reference shared definitions (e.g.,
`FileNotFound`, `RateLimited`).
```rust
pub struct Definitions { /* ... */ }
@@ -410,50 +457,67 @@ impl Definitions {
pub fn new() -> Self;
/// Define a named schema. Returns a `Schema` that produces
/// `{"$ref": "#/$defs/<name>"}` — the JSON Pointer form that
/// `jsonschema` and `AlkTypeEngine::compile` expect (after
/// `normalize_refs`, which the engine runs at compile time).
/// `{"$ref": "#/$defs/<name>"}` — the JSON Pointer form BAST
/// requires (no `normalize_refs` step; refs are always full
/// pointers).
pub fn define(&mut self, name: &str, schema: Schema) -> Schema;
/// Like `define`, but the schema is an existing `Value` (adopted
/// via `Schema::from_value`).
pub fn define_value(&mut self, name: &str, value: Value) -> Schema;
/// Produce the `{"$defs": { ... }}` object to merge into a
/// top-level schema. Call once at the end.
/// Produce the `{"$defs": { ... }}` object.
pub fn build(self) -> Value;
/// Build a complete BAST document with `root_name` as the root
/// type. The root schema is inserted into `$defs` alongside any
/// previously defined entries. The resulting `Value` is ready for
/// `AlkTypeEngine::compile(&doc, root_name, mode, ...)`.
pub fn build_doc(self, root_name: &str, root: Schema) -> Value;
/// Merge the `$defs` into a top-level schema `Value`. If `top`
/// already has a `$defs` object, the definitions are merged into
/// it; otherwise a `$defs` key is inserted. For BAST documents,
/// prefer `build_doc` — it places the root type inside `$defs`
/// (where BAST requires it).
pub fn merge_into(self, top: &mut Value);
}
```
**Usage:**
**Usage (complete BAST document):**
```rust
let mut defs = Definitions::new();
let file_not_found = defs.define("FileNotFound",
Schema::object()
.field("path", Schema::string_())
.field("errno", Schema::integer())
.required(["path", "errno"])
);
defs.define("Init", Schema::struct_().field("version", Schema::uint32()));
defs.define("Read", Schema::struct_()
.field("handle", Schema::bytes())
.field("offset", Schema::uint64())
.field("len", Schema::uint32()));
let rate_limited = defs.define("RateLimited",
Schema::object()
.field("retry_after_ms", Schema::integer().minimum(0))
.required(["retry_after_ms"])
);
let read_file_error = Schema::object()
.field("code", Schema::string_())
.field("details", Schema::any()) // one of the defined errors
.required(["code"])
.build();
// Merge $defs into the top-level schema that references them
let mut top = Schema::object()
.field("error", read_file_error)
.build();
top.as_object_mut().unwrap().insert("$defs".to_string(), defs.build());
let doc = defs.build_doc("Packet", Schema::struct_()
.field("payload", Schema::union_(Discriminator::Byte {
offset: 0,
disc_type: AlkTypeKind::Uint8,
})
.mapping("1", Schema::ref_def("Init"))
.mapping("5", Schema::ref_def("Read"))));
// -> {
// "$defs": {
// "Init": { "kind": "struct", "fields": [ { "name": "version", "kind": "uint32" } ] },
// "Read": { "kind": "struct", "fields": [ ... ] },
// "Packet": { "kind": "struct", "fields": [
// { "name": "payload", "kind": {
// "kind": "union",
// "discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
// "mapping": { "1": { "$ref": "#/$defs/Init" }, "5": { "$ref": "#/$defs/Read" } }
// } }
// ] }
// }
// }
//
// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None)
// then validate incoming frames via engine.validate_bytes(&frame).
```
`define` returns a `Schema` (the `$ref` to the definition), so it can
@@ -483,31 +547,36 @@ For cases where the `Definitions::define` return value isn't handy
## Usage Examples
### Example 1: channels' 8-byte chunk header (binary layout)
### Example 1: channels' 8-byte chunk header (binary layout, BAST)
```rust
use alktype::{Schema, Endian};
use alktype::{Schema, Endian, Definitions};
let chunk_header = Schema::struct_()
.endian(Endian::Big)
.field("channel_id", Schema::uint32())
.field("length", Schema::uint32())
.build();
.field("length", Schema::uint32());
// Build a complete BAST document (single-type — one $defs entry).
let doc = Definitions::new().build_doc("ChunkHeader", chunk_header);
// -> {
// "AlkType:Struct": true,
// "endian": "big",
// "properties": {
// "channel_id": { "AlkType:Uint32": true },
// "length": { "AlkType:Uint32": true }
// "$defs": {
// "ChunkHeader": {
// "kind": "struct",
// "endian": "big",
// "fields": [
// { "name": "channel_id", "kind": "uint32" },
// { "name": "length", "kind": "uint32" }
// ]
// }
// }
// }
//
// Feed to AlkTypeEngine::compile(&mut chunk_header, LayoutMode::Packed)
// Feed to AlkTypeEngine::compile(&doc, "ChunkHeader", LayoutMode::Packed, None)
// then validate incoming frames via engine.validate_bytes(&frame).
```
### Example 2: call's `OperationSpec` input schema (JSON payload)
### Example 2: call's `OperationSpec` input schema (JSON payload, standard JSON Schema)
```rust
use alktype::Schema;
@@ -530,18 +599,19 @@ let read_file_input = Schema::object()
// }
//
// Stored in OperationSpec.input_schema; validated via the standard
// jsonschema validator (validate_json for parsed payloads, or via
// serde_json::from_slice then validate_json for wire frames).
// jsonschema validator (AlkTypeEngine::compile with Some(&read_file_input)
// for the validate_json path, or serde_json::from_slice then
// validate_json for wire frames).
```
### Example 3: SFTP `Packet` union (binary layout, byte discriminator)
The SFTP wire shape is `[type:u8][payload-struct]` — a struct with a
union payload field. The engine requires `AlkType:Struct` at the top
level (`OffsetMap::compute` / `SequentialReader::new` both enforce
this; a `Union` is a field type within a struct, not a top-level
schema). The builder constructs the union wrapped in a struct, and
`$defs` are merged into the top-level schema so `$ref`s resolve:
union payload field. The engine requires a struct at the root
(`OffsetMap::compute` / `SequentialReader::new` both enforce this; a
`Union` is a field type within a struct, not a top-level schema). The
builder constructs the union wrapped in a struct, and `$defs` are
placed inside the document via `build_doc` so `$ref`s resolve:
```rust
use alktype::{Definitions, Discriminator, AlkTypeKind, Schema};
@@ -555,27 +625,21 @@ defs.define("Status", Schema::struct_().field("code", Schema::uint32()).field("m
// A "Packet" is a struct with one field — the union. This mirrors
// SFTP's wire shape: [type:u8][payload-struct].
let mut packet = Schema::struct_()
.field(
"payload",
Schema::union_(Discriminator::Byte {
offset: 0,
disc_type: AlkTypeKind::Uint8,
})
.mapping("1", Schema::ref_def("Init"))
.mapping("3", Schema::ref_def("Open"))
.mapping("5", Schema::ref_def("Read"))
.mapping("6", Schema::ref_def("Write"))
.mapping("101", Schema::ref_def("Status")),
)
.build();
// Merge $defs into the top-level schema so $refs resolve at compile time.
defs.merge_into(&mut packet);
// Feed to AlkTypeEngine::compile(&mut packet, LayoutMode::Packed)
let doc = defs.build_doc("Packet", Schema::struct_()
.field("payload", Schema::union_(Discriminator::Byte {
offset: 0,
disc_type: AlkTypeKind::Uint8,
})
.mapping("1", Schema::ref_def("Init"))
.mapping("3", Schema::ref_def("Open"))
.mapping("5", Schema::ref_def("Read"))
.mapping("6", Schema::ref_def("Write"))
.mapping("101", Schema::ref_def("Status"))));
// Feed to AlkTypeEngine::compile(&doc, "Packet", LayoutMode::Packed, None)
// then validate incoming frames via engine.validate_bytes(&frame).
```
### Example 4: OperationSpec error schemas (named `$defs`)
### Example 4: OperationSpec error schemas (named `$defs`, standard JSON Schema)
```rust
use alktype::{Definitions, Schema};
@@ -610,15 +674,19 @@ let op_errors = vec![
http_status: Some(429),
},
];
// $defs is built once and stored alongside the OperationSpec
// `$defs` is built once and stored alongside the OperationSpec.
// (For the validate_json path, compile with Some(&defs.build()) as the
// json_schema argument — but typically OperationSpec schemas are
// validated directly via jsonschema, not via AlkTypeEngine.)
```
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation shapes the builder's setters produce |
| Builder API for schema construction | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 |
| BAST format + two output formats | [ADR-BAST](decisions/bast-bast-format.md) | `struct_()` → BAST, `object()` → standard JSON Schema (D-BAST-008) |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | The annotation semantics the builder's setters produce (location moved to BAST type-level properties) |
| Load-time validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | The builder does not pre-validate; compile-time is the validation point |
## Open Questions
@@ -634,12 +702,15 @@ and transitively on `Schema::union_`). See
## References
- [ADR-009](decisions/009-builder-api.md) — the decision this spec implements
- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format and the
two-output-format decision (D-BAST-008)
- [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) —
scope boundaries this module extends; "schemas are JSON" principle
- [ADR-003](decisions/003-schema-annotations.md) — the annotation
shapes the builder's setters produce
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds the
builder's constructors produce
semantics the builder's setters produce
- [`bast-format.md`](bast-format.md) — the normative BAST format
specification (the output format for `struct_()`)
- [schema-layer.md](schema-layer.md) — the BAST kinds and parser
- [validation.md](validation.md) — the validation layer that consumes
builder output (via `AlkTypeEngine::compile`)
- `@alkdev/alknet: docs/architecture/crates/call/operation-registry.md`
+24 -19
View File
@@ -100,7 +100,7 @@ impl SequentialReader {
```
`read_field`/`write_field` on `AlkTypeEngine` work for the fixed-size
primitive kinds and the length-prefixed `String`/`Bytes`/`Timestamp`
primitive kinds and the length-prefixed `String`/`Bytes`
fields. Composite kinds (`Struct`, `Union`, `Array`, `Record`) return a
`FieldValue` carrying a layout descriptor (byte range, variant start,
or array stride) for the consumer to recurse on — see §"FieldValue" above.
@@ -238,21 +238,23 @@ pub struct UnionDispatch {
}
```
After dispatch, the consumer calls `tunion::resolve_variant(union_schema, &dispatch.key)`
to get the variant schema, then reads the variant's fields at
After dispatch, the consumer calls `tunion::resolve_variant(union_node, &dispatch.key)`
to get the variant `BastType`, then reads the variant's fields at
`dispatch.variant_offset` using the normal `data_access` functions (or a
fresh `SequentialReader` scoped to the variant).
fresh `SequentialReader` scoped to the variant). `$ref` variant types are
returned as `BastType::Ref`; the caller resolves them via
`BastDoc::resolve_typeref` when a concrete definition is needed.
### Byte-offset discriminator
```rust
/// Read the discriminator value from a byte-offset TUnion. The discriminator
/// is a fixed-size integer (AlkType:Uint8/Uint16/Uint32) at a known byte
/// offset. Returns the mapping key (stringified integer) and the variant
/// struct offset.
/// is a fixed-size integer (uint8/uint16/uint32) at a known byte offset.
/// Returns the mapping key (stringified integer) and the variant struct
/// offset.
pub fn read_byte_discriminator(
buffer: &[u8],
union_schema: &Value,
union_node: &BastUnion<'_>,
endian: Endian,
) -> Result<UnionDispatch, AlkTypeError>;
```
@@ -268,11 +270,11 @@ starts at `offset + discriminator_size`.
/// Read the discriminator value from a field-name TUnion. The
/// discriminator is a named field within the struct — the consumer
/// provides the field's computed offset (from the OffsetMap or
/// LayoutBuilder). Supports AlkType:String, Uint8, and Enum discriminator
/// LayoutBuilder). Supports string, uint8, and enum discriminator
/// fields.
pub fn read_field_discriminator(
buffer: &[u8],
union_schema: &Value,
union_node: &BastUnion<'_>,
disc_field_offset: usize,
endian: Endian,
) -> Result<UnionDispatch, AlkTypeError>;
@@ -287,16 +289,19 @@ the variant's fields starting at the end of the discriminator field.
### Variant resolution
```rust
/// Look up a variant schema from the union's mapping. Inline schemas
/// are returned directly. $ref pointers of the form "#/$defs/<name>"
/// are resolved against the union schema's own $defs block.
pub fn resolve_variant<'a>(union_schema: &'a Value, key: &str)
-> Result<&'a Value, AlkTypeError>;
/// Look up a variant type from the union's mapping. Inline struct/
/// union/enum types are returned directly; `$ref` pointers are
/// returned as `BastType::Ref` — the caller resolves them via
/// `BastDoc::resolve_typeref` when a concrete definition is needed.
pub fn resolve_variant<'a>(
union_node: &'a BastUnion<'a>,
key: &str,
) -> Result<&'a BastType<'a>, AlkTypeError>;
/// Get the discriminator's byte size (1/2/4 for Uint8/16/32) for a
/// Get the discriminator's byte size (1/2/4 for uint8/16/32) for a
/// byte-offset TUnion. Field-name discriminators have no fixed size
/// and produce a AlkTypeError::Schema.
pub fn discriminator_size(union_schema: &Value) -> Result<usize, AlkTypeError>;
/// and produce an `AlkTypeError::Schema`.
pub fn discriminator_size(union_node: &BastUnion<'_>) -> Result<usize, AlkTypeError>;
```
### TUnion in the layout engines
@@ -324,7 +329,7 @@ dispatch to the primitive `data_access` function for the field's kind.
For aligned-mode access, `AlkTypeEngine::read_field(&buffer, "header.version")`
returns `FieldValue` — it looks up the `ByteRange` in the `OffsetMap`, finds
the field's `AlkType:*` kind in the schema, and calls the matching
the field's `AlkTypeKind` in the BAST typed tree, and calls the matching
`data_access::read_*` function. `write_field` is the mirror. Composite
kinds (`Struct`, `Union`, `Array`, `Record`) return a `FieldValue`
carrying a layout descriptor; the consumer recurses with a fresh reader
@@ -1,7 +1,19 @@
# ADR-001: alktype — Purpose, Scope, and the jsonschema Engine
## Status
Accepted
**Superseded (format-specific content) by
[ADR-BAST](bast-bast-format.md).** The crate's purpose, scope
boundaries, and the "schema is the format" principle are **retained and
strengthened** — BAST *is* the format. Only the *concrete format*
(custom-keyword JSON Schema → BAST) and the *validation strategy*
(single `jsonschema` custom-keyword validator → two-validator model)
are superseded: the format-specific content by ADR-BAST, the
validation-strategy content by
[ADR-VAL-SPLIT](val-split-two-validator-model.md). This ADR is kept as
the historical record of the v0.1.0 design and the purpose/scope
decision; read it alongside ADR-BAST and ADR-VAL-SPLIT for the current
state.
## Context
@@ -1,7 +1,14 @@
# ADR-002: Two Layout Modes — Packed Sequential vs Aligned Static
## Status
Accepted
Accepted — unchanged under the BAST pivot
([ADR-BAST](bast-bast-format.md)). Layout modes are format-agnostic:
the input format changed from custom-keyword JSON Schema to BAST, but
the two modes, their alignment/packing rules, and the
`LayoutBuilder`/`SequentialReader`/`OffsetMap` API did not. The layout
engines now walk the BAST typed tree ([`BastDoc`](../schema-layer.md))
instead of raw JSON with `get_alktype_kind*`, but the offset
computation algorithm is identical.
## Context
@@ -1,7 +1,18 @@
# ADR-003: Schema Annotations — Endianness, Alignment, Encoding, and TUnion Discriminators
## Status
Accepted
**Accepted (semantics); amended (location) by
[ADR-BAST](bast-bast-format.md).** The annotation *semantics* decided
here — endianness default, struct/field-level alignment, the three
variable-length encoding strategies, and the two TUnion discriminator
kinds — **carry forward unchanged** under the BAST pivot. Only the
annotation *location* moves: from v0.1.0's custom-keyword objects
(`{"AlkType:String": { "encoding": "..." }}`) to BAST type-level
properties (`{ "name": "handle", "kind": "string", "encoding": "..." }`).
The BAST shapes are normative in
[`bast-format.md`](../bast-format.md#variable-length-encoding); this
ADR is kept as the semantic reference. Read it alongside ADR-BAST.
## Context
@@ -125,9 +136,9 @@ reserving worst-case space.
- `true` is a shorthand for the default (length-prefixed). This keeps
the common case concise and the override explicit.
- The `encoding` annotation and `maxLength` apply to all variable-length
types: `AlkType:String`, `AlkType:Bytes`, `AlkType:Array`,
`AlkType:Record`, `AlkType:Timestamp`.
- The `encoding` annotation and `maxLength` apply to the variable-length
primitive types `AlkType:String` and `AlkType:Bytes`. (`maxLength` on
records was amended out by review #006 N3/M5 — see §3a.)
### 3a. TRecord value type
@@ -153,8 +164,13 @@ the `"values"` property in the schema:
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix).
- The count and key-length prefixes respect the schema's endianness.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).
- ~~In aligned static mode with `maxLength`, the entire record is
reserved at `maxLength` bytes (zero-padded).~~ **Amended (review #006
N3/M5, 2026-09-02):** `maxLength` is rejected at parse on record
fields. The aligned materializer walks the record's inline
count-prefixed form and never honors the reservation (M5: silent
cross-field corruption), and no packed consumer enforced it either
(N3: silently unenforced). `maxLength` is `string`/`bytes`-only.
### 4. TUnion discriminators
@@ -1,7 +1,21 @@
# ADR-004: Error Handling and Validation Strategy
## Status
Accepted
**Accepted (error type); amended (validation strategy) by
[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The `AlkTypeError`
enum, its four variants, the load-time-build / access-time-check split,
and the field-path-carrying errors decided here are **retained
unchanged** under the BAST pivot (D-BAST-009 keeps
`Validation(jsonschema::ValidationError<'static>)`). The "validation
strategy" section — which described v0.1.0's single
`jsonschema`-custom-keyword validator for both paths — is **refined**:
the bytes path now uses the BAST-native validator
(`bast_validation`), the JSON path now uses a standard
`jsonschema::Validator` from a consumer-provided JSON Schema. See
[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the
two-validator model. This ADR is kept as the error-handling reference;
read it alongside ADR-VAL-SPLIT for the current validation strategy.
## Context
@@ -63,11 +63,26 @@ write-side.
### Cost
`SequentialReader::new` clones the top-level struct's field schemas (a
`Vec<(String, Value)>` of the `properties` entries) and clones the
schema itself. This is cheap — a struct has a small number of fields
(SFTP's largest packet has 5). The construction cost is negligible
compared to the cost of reading a buffer.
`SequentialReader::new(Arc<ReadPlan>)` is a refcount bump — 15.7 ns
(measured, alktty `wire_vs_bast` bench, 0.3.0). The reader shares the
engine's compiled [`ReadPlan`](011-compiled-read-plan-for-packed-mode.md)
(the packed read-side compiled form) via `Arc` instead of cloning
schema data; construction cost is negligible compared to reading a
buffer. The engine holds the owned `BastDoc` (ADR-012 §2a) for the
aligned materialize path and the one-shot `*::compile` paths.
> **Historical note**: the original 0.2.0 framing here ("re-parse on
> demand" — the read loop re-parsing `BastDoc::new` per field) was the
> root cause of the 400x read-path gap measured in
> [review #004](../../reviews/004-performance-review.md).
> [ADR-011](011-compiled-read-plan-for-packed-mode.md) (implemented,
> 0.3.0) retired it: the packed read loop walks `Arc<ReadPlan>` (2.27
> µs/chunk → 98 ns/chunk), `sequential_reader()` is an `Arc::clone`,
> and the owned `BastDoc` (ADR-012 §2a) removed the remaining
> per-access re-parse sites in `read_field`/`write_field`/`validate_bytes`.
> The factory decision itself (`sequential_reader() ->
> Option<SequentialReader>`, owned fresh reader, consumer-driven
> cursor) was retained unchanged.
## Consequences
+13 -1
View File
@@ -2,7 +2,19 @@
## Status
Accepted
**Accepted (API surface); amended (output format) by
[ADR-BAST](bast-bast-format.md).** The fluent builder API, the
`Schema`/`Definitions`/`Discriminator` types, the constructor and
setter catalog, and the "produces `serde_json::Value`, not a typed
`Schema` enum" decision decided here are **retained unchanged** under
the BAST pivot. Only the `build()` *output format* changes:
`struct_()` now produces BAST JSON (`{ "kind": "struct", "fields": [...] }`)
instead of v0.1.0's custom-keyword JSON (`{ "AlkType:Struct": true,
"properties": {...} }`); `object()` continues to produce standard JSON
Schema. This is D-BAST-008, recorded in ADR-BAST. The builder examples
in [`builder.md`](../builder.md) reflect the current BAST output. This
ADR is kept as the API-surface decision; read it alongside ADR-BAST
for the output format.
## Context
@@ -2,7 +2,23 @@
## Status
Accepted
**Accepted (two-step concept); amended (validation step) by
[ADR-VAL-SPLIT](val-split-two-validator-model.md).** The
`validate_bytes(&[u8])` entry point, the "materialize `Value` from
bytes, then validate" two-step concept, the mode dispatch, the
field-path-carrying errors, and the "not a `Validator` trait / not
framing-aware / not a binary-payload validator for JSON-only schemas"
scope boundaries decided here are **retained unchanged** under the
BAST pivot. Only the validation *step's implementation* changes: the
materialized `Value` is validated by the **BAST-native validator**
(`bast_validation`) instead of v0.1.0's `jsonschema` custom-keyword
validator. The `jsonschema` crate is no longer touched on the bytes
path (it remains for the `validate_json` path and for BAST meta-schema
validation). The error payload type stays
`Validation(jsonschema::ValidationError<'static>)` (D-BAST-009). See
[ADR-VAL-SPLIT](val-split-two-validator-model.md) for the
two-validator model. This ADR is kept as the `validate_bytes` decision;
read it alongside ADR-VAL-SPLIT for the current validation step.
## Context
@@ -0,0 +1,584 @@
# ADR-011: Compiled Read Plan for Packed Mode
## Status
Accepted. Implemented in 0.3.0 (phases 1–2, 2026-09-02). Closes review
#004 H1 + M1 (packed side) + L1 + L2;
retires the "re-parse on demand" framing from ADR-007. A derisking
POC on branch `readplan-poc` confirmed the `ReadPlan` shape covers
every `BastType` arm in the current read loop before implementation
began (see "POC coverage" at the end).
**Refinements on ADR-012 acceptance (2026-08-20, review #005):** the
`CompositePlan::Union` shape was refined to carry `shared:
Option<Box<ReadPlan>>` (field-disc union shared fields, resolving POC
Finding 1 / review #005 H1) and to drop `VariantPlan`/`VariantKind`
in favor of `variants: Vec<(String, CompositePlan)>` (resolving review
#005 M1 — nested unions now work by `CompositePlan` recursion,
restoring the 0.2.0 capability the POC rejected). The "BastDoc
unchanged" scope statement stands as the ADR-011-only view; ADR-012
§2a subsequently makes `BastDoc` owned. The `ValidationPlan` this
ADR's "Out of scope" originally deferred indefinitely is now in
0.3.0 via ADR-012 §3 (review #005 M3 reversed the deferral). These
are pre-implementation refinements to types that do not yet exist on
`main`; the ADR-011 decision (a compiled `ReadPlan` for packed reads)
is unchanged.
**Addendum — field-disc union wire convention (2026-09-02, review #006
H3):** the packed-mode wire layout for a field-name-discriminator
TUnion is **shared-then-variant**: the union's declared `fields` (the
discriminator field + any shared fields) occupy the union's start
offset in declaration order, and the selected variant's fields follow
immediately after all shared fields. All three packed-mode consumers
now implement this one convention: the reader and materializer already
walked `shared` then the variant (the `shared` sub-plan shape above);
`LayoutBuilder` was corrected in the same pass — it previously laid out
only the selected variant, disagreeing with the read side on span and
field positions (review #006 H3 item 1). The convention requires that
a variant **must not re-declare** the discriminator field or any
shared field — `BastUnion::parse` enforces this at parse time (also:
the discriminator field must be declared in `fields`, and `fields`
must not contain duplicate names), so the shared walk and the variant
walk cover disjoint fields and the wire has exactly one copy of each
shared byte. Schemas whose variants redeclared shared fields were
ambiguous under the old split-convention behavior and are rejected
rather than given a silent meaning; this is a **breaking wire-format
constraint** for any 0.2.0-era schema that relied on re-declaration,
announced with the 0.3.x series. `DiscriminatorPlan::Field`'s disc
read is at the disc field's position within the shared walk (the
materializer's position-correct behavior, review #006 H3 item 2); the
reader's plan-walk reads it there too.
## Context
Review #004 (`docs/reviews/004-performance-review.md`) measured the
packed read path at **~400x slower per chunk** than a hand-rolled codec
(2.27 µs/chunk vs 5.6 ns/chunk), with the cost fixed across payload
sizes — the signature of per-field interpretive overhead, not
payload-copy overhead. The write path is competitive (~1.1x at 4 KiB)
because it uses a compiled form; the read path is not because it
doesn't.
### The three layout-side compiled forms and the one gap
ADR-002 defines two layout modes. Each mode has a write-side and a
read-side. Three of the four slots already have a **compiled form** —
a data structure built once from the schema, held by the engine, and
walked at access time without re-touching the schema:
| mode | write-side | read-side |
|------|-----------|-----------|
| aligned static | `OffsetMap` (used for both) | `OffsetMap` |
| packed sequential | `PackedLayout` (`LayoutBuilder::build`) | *(none)* |
- **`OffsetMap`** (aligned, both sides) — a flat table of
`(field_path, ByteRange)` pairs computed once via
`OffsetMap::compute(&BastDoc)`. Read and write both look up a field's
byte range and call `data_access::read_*`/`write_*` at the known
offset. No schema walk at access time.
- **`PackedLayout`** (packed, write-side) — a flat table of
`(field_path, FieldPosition)` pairs computed once via
`LayoutBuilder::build(&var_sizes)`. The write loop calls
`data_access::write_*` at the precomputed offsets. No schema walk at
write time.
- **packed read-side** — `SequentialReader` walks `BastDoc`
interpretively on every field read. There is no compiled form.
This is the structural reason the read path is 400x slow: it is the
only access path in the engine with no compiled form. Every other
mode/side pair compiles the schema once and reuses the result.
### Root cause: the `BastDoc<'a>` borrow constraint
`BastDoc<'a>` borrows `&'a Value` and `&'a str` throughout
(`src/bast.rs:51-55`). The owning structs that need a parsed tree at
read time — `SequentialReader` (owns a cloned `Value`),
`AlkTypeEngine` (owns `bast_doc: Value`), `LayoutBuilder` (owns
`doc_value: Value`) — cannot store a `BastDoc` that borrows from their
own `Value` field. That would be a self-referential struct, which safe
Rust cannot express.
The workaround chosen in ADR-007 was "re-parse on demand": the engine
and reader retain a clone of the raw `Value` and reconstruct the
`BastDoc` from it whenever the typed tree is needed. ADR-007's "Cost"
section argued this was cheap because construction is a small `Vec` of
field schemas. That is true for *construction* (once), but the decision
did not account for `read_field_at` re-parsing `BastDoc::new` **per
field read** — the cost that actually dominates. For an N-field struct,
reading all fields is O(N²) in parse work (each of N reads re-parses
all N fields).
### Why `OffsetMap` is a flat table but the packed read plan cannot be
`OffsetMap` works as a flat `(path, byte_range)` lookup table because
aligned positions are **data-independent** — field N's offset depends
only on the schema, not on the bytes of fields 0..N-1. Random access
by path is free.
Packed positions are **data-dependent** — a variable-length field's
extent is read from its length prefix at access time, and every
subsequent field's position shifts accordingly. You cannot look up
field N's offset without reading fields 0..N-1 first. So the compiled
form for packed reads cannot be a flat lookup table; it must be a
**read program** — a pre-resolved tree of read instructions that the
read loop walks in order, advancing a cursor. The schema is compiled
into the program once; the bytes are walked against it at read time.
This asymmetry is inherent to packed sequential layout (ADR-002) and is
not a flaw in `OffsetMap`. The two modes need different compiled-form
shapes because they have different position-computation semantics.
## Decision
**Introduce `ReadPlan` — the compiled read-side form for packed mode,
symmetric to `OffsetMap` (aligned read-side) and `PackedLayout` (packed
write-side).**
`AlkTypeEngine::compile` builds the `ReadPlan` once from the `BastDoc`
(in packed mode) and holds it for the life of the engine.
`sequential_reader()` hands out fresh `SequentialReader`s that share
the engine's `Arc<ReadPlan>` — the plan is immutable; only the cursor
state (`field_index`, `position`) is per-reader. The read loop walks
the plan, never touching `BastDoc` or the raw `Value`.
The same `ReadPlan` is consumed by `materialize_packed` (the other
byte-walking path), unifying the two packed read-side consumers on one
compiled form — mirroring how `OffsetMap` unifies the aligned read and
write sides.
### The `ReadPlan` shape
A pre-resolved tree of read instructions. Every `$ref` is resolved, every
endianness is computed (field override or container default), every
union variant is inlined. The read loop indexes into a `Vec`, matches
on a `ReadKind`, and calls `data_access::read_*` with a precomputed
`Endian` — no `resolve_typeref`, no `BastDef::parse`, no JSON node
access on the happy path. (`format!` for error-path attribution may
still occur on the error path; it does not run on the happy path and
is not the cost being removed here.)
```rust
pub struct ReadPlan {
endian: Endian,
fields: Vec<FieldPlan>,
by_name: HashMap<String, usize>,
}
pub struct FieldPlan {
name: String,
kind: ReadKind,
endian: Endian,
encoding: VariableEncoding,
max_length: Option<usize>,
body: Option<CompositePlan>,
}
pub enum ReadKind {
Primitive(AlkTypeKind),
Enum,
Struct,
Union,
Array,
Record,
}
pub enum CompositePlan {
Struct(ReadPlan),
Union {
disc: DiscriminatorPlan,
shared: Option<Box<ReadPlan>>,
variants: Vec<(String, CompositePlan)>,
},
Array {
element: Box<CompositePlan>,
count: usize,
element_stride: usize,
},
Record {
value: Box<CompositePlan>,
},
}
pub enum DiscriminatorPlan {
Byte { offset: usize, disc_type: AlkTypeKind },
Field { name: String, field_index: usize },
}
```
Nested structs share the `ReadPlan` shape (a struct field's `body` is
`CompositePlan::Struct(ReadPlan)`). Union variants are pre-resolved:
each `(key, CompositePlan)` entry carries the variant's compiled body,
so dispatch is a flat lookup + recurse — no `resolve_typeref_as_def` at
read time. A variant may itself be `CompositePlan::Union { ... }`, so
**nested unions** (a union variant that is itself a union, which the
0.2.0 reader supports via `resolve_and_walk_variant`'s `Union` arm) are
covered by ordinary recursion; no separate `VariantKind` enum is
needed. Array element strides are precomputed (`element_stride = 0`
signals variable-length elements, same convention as today's
`FieldValue::Array`).
`Union.shared` carries the union's declared `fields` (the discriminator
field + any shared fields) for the field-name-discriminator case —
`DiscriminatorPlan::Field.field_index` indexes into `shared`, and the
read loop walks `shared` first, then looks up and walks the selected
variant's `CompositePlan` starting after the shared fields. The
byte-offset-discriminator case has no shared fields (`shared: None`):
the discriminator byte is read at `disc.offset` and the variant starts
immediately after the discriminator size. The POC's `plan_read_union`
`Field` arm stub (Finding 1) is replaced by this `shared` sub-plan;
there is no separate `VariantPlan`/`VariantKind` type in the production
shape — the POC's `VariantPlan { kind, plan }` wrapper is dropped in
favor of recursing on `CompositePlan` directly, which is what makes
nested-union support fall out for free.
### Construction
```rust
impl ReadPlan {
pub fn compile(bast_doc: &Value, root_name: &str) -> Result<Self, AlkTypeError>;
}
```
`compile` walks `BastDoc` once, resolves all `$ref`s eagerly, computes
effective endianness at every node, inlines union variants, and
builds the `FieldPlan`/`CompositePlan` tree. Malformed schemas surface
as `AlkTypeError::Schema` — the same untrusted-input discipline
(AGENTS.md §3) and overflow-safe arithmetic (AGENTS.md §4) as
`BastDoc::new`.
### Engine integration
`AlkTypeEngine::compile` builds the `ReadPlan` in packed mode and
stores `Arc<ReadPlan>`. `sequential_reader()` returns
`SequentialReader { plan: Arc::clone(&self.plan), .. }` — an owned
reader (ADR-007's factory decision is retained; the reader owns its
cursor, shares the immutable plan).
`bast_doc: Value` (`src/engine.rs:85`) is retained on the engine
unconditionally. In packed mode it becomes unused by the read path
(both `sequential_reader` and `validate_bytes` consume the plan);
in aligned mode it is still needed for `read_field`/`write_field`/
`validate_bytes`. Keeping it always avoids a mode-conditional field
and costs only a `Value` clone paid once at `compile`. `ReadPlan`
must be `Send + Sync` so `Arc<ReadPlan>` can be shared from the
`Send + Sync` engine (ADR-007); this falls out naturally from the
plan being immutable owned data, but the implementation should add a
`static` bound assertion test to lock it in.
### Public API change (breaking — version bump to 0.3.0)
- `SequentialReader::new(&Value, &str)` → `SequentialReader::new(Arc<ReadPlan>)`.
The old constructor is replaced by `ReadPlan::compile(&Value, &str)`
followed by `SequentialReader::new(Arc::from(plan))`.
- `materialize_packed(&BastDoc<'_>, &[u8])` →
`materialize_packed(&ReadPlan, &[u8])`.
- `materialize_aligned` is unchanged (already takes `&OffsetMap`, a
compiled form).
- `ReadPlan` is a new public type, re-exported from `lib.rs`.
- `BastDoc` and the `Bast*` types are **unchanged by this ADR** — they
remain the validation-side typed tree, borrowed, as today. This is a
smaller breakage than review #004's Option A (which changed
`BastDoc<'a>` → `BastDoc` and every `Bast*` signature).
**Note (added on ADR-012 acceptance):** ADR-012 §2a subsequently
makes `BastDoc` owned, riding the same 0.3.0 bump. That is an
ADR-012 change, not an ADR-011 change; ADR-011's scope statement
stands as the ADR-011-only view. With ADR-012 §3 (ValidationPlan,
now in 0.3.0), `bast_validation` will also stop being the permanent
home of the `BastDoc` walk — see ADR-012.
The crate is pre-1.0 with two in-house downstream consumers
(`alktty`, `alkcall`), both of which will be updated with the bump.
## Scope
### In scope (consumes the `ReadPlan`)
- **`SequentialReader`** — the read loop walks `FieldPlan`/`CompositePlan`
instead of `&BastField`/`&BastType` + `&BastDoc`. The functions
`read_field_value`, `read_typeref_value`, `walk_struct_size`,
`read_union_value`, `read_array_value`, `read_record_value` are
rewritten to take plan nodes. One walker, not two — the review's
Option B concern ("duplicates the `BastType` matching logic") does
not apply because the plan *replaces* the `BastType` matching, not
parallels it.
- **`materialize` (packed mode)** — `materialize_packed` takes
`&ReadPlan` and walks it to produce `serde_json::Value`. Same read
logic, same `data_access` calls, different input type. Unifies the
two packed read-side consumers on one compiled form.
- **`AlkTypeEngine::validate_bytes` (packed mode)** — calls
`materialize_packed(&self.plan, buffer)` instead of reconstructing a
`BastDoc`. Closes M1's `engine.rs:284` re-parse.
### Out of scope (stays on `BastDoc`)
- **`bast_validation`** — the BAST-native value-domain validator walks
`BastDoc` to check constraints (`maxLength`, enum string values,
union variant keys, integer ranges). These are value-domain checks,
not byte-position walks; they don't benefit from a *read* plan and
would require a separate "validation plan" with a different shape.
**Not in scope for ADR-011** — but no longer deferred indefinitely:
ADR-012 §3 brings a `ValidationPlan` into 0.3.0. The "not a hot
loop" framing this paragraph originally relied on was re-evaluated
and rejected (see ADR-012 §3): read+validate on untrusted streams
makes validation hot in the same sense review #004 measured for
the read path. For 0.3.0 as accepted by ADR-011 alone,
`bast_validation` keeps walking `BastDoc`; ADR-012 §3 closes that.
- **`LayoutBuilder` / `PackedLayout`** — the packed write-side already
has a compiled form (`PackedLayout`). `LayoutBuilder::build`
(`src/layout_builder.rs:190`) re-parses `BastDoc::new` per `build()`
call (M1), but the typical pattern is build-once-reuse, so this is
not a hot loop. A future `WritePlan` that lets `LayoutBuilder` cache
the typed tree (review #004 Option A's territory) is additive and can
follow; it is not blocking the read-path fix.
- **`OffsetMap` / aligned mode** — already a compiled form; unchanged.
`materialize_aligned` already takes `&OffsetMap`.
- **Aligned-mode `read_field` / `write_field`** (`engine.rs:334,467`)
re-parse `BastDoc::new` per call (M1). These are one-shot paths, not
hot loops; they can adopt a compiled form later without affecting
the packed read-path decision. Left as-is for now.
## Consequences
### Positive
- **Closes the 400x read-path gap (review #004 H1).** The read loop no
longer touches `BastDoc` or the raw `Value`. Per-field work drops
from "re-parse the typed tree + resolve_typeref + match" to "index
into a `Vec` + match `ReadKind` + `data_access::read_*` with a
precomputed `Endian`." The expected per-chunk cost is in the
hand-rolled codec's ballpark (the `data_access` calls are the same
ones the hand-rolled codec makes).
- **Closes M1's packed-side re-parse (`engine.rs:284`, `validate_bytes`).**
The aligned-side M1 sites (`engine.rs:334,467`,
`layout_builder.rs:190`) are deliberately left as-is — they are not
hot for the packed-codec use case, and Option A remains available as
an additive later fix if an aligned-mode hot loop ever emerges. This
is a reversible bet, not a claim that the aligned-side M1 is a
non-issue.
- **Unifies the two packed read-side consumers on one compiled form.**
`SequentialReader` and `materialize_packed` walk the same `ReadPlan`,
mirroring how `OffsetMap` unifies the aligned read and write sides.
The "two parallel walkers" concern from review #004 Option B does
not apply — the plan replaces the `BastType` matching, not
duplicates it.
- **Symmetric with the other compiled forms.** The engine now has a
compiled form for every mode/side pair: `OffsetMap` (aligned R/W),
`PackedLayout` (packed W), `ReadPlan` (packed R). The "compiled form
of a BAST document" framing in ADR-004/validation.md becomes true for
the read path, not just the write path.
- **ADR-007's factory gets cheaper.** Today
`sequential_reader()` clones `doc_value: Value` (the whole BAST
document) per reader. With the plan, it clones an `Arc<ReadPlan>`
(refcount bump). The plan is immutable and shared across all readers
from one engine. ADR-007's "owned fresh reader" decision is retained;
the reader owns its cursor, shares the plan.
- **Deterministic compile.** `ReadPlan::compile` is a pure function of
the BAST document + root name — same input, same plan. This makes the
"compiled form" visibly deterministic, which is a prerequisite for
future capabilities (fingerprinting the plan for cross-run caching,
disk-cached compiled plans, or schema-version handshakes for
`alkcall`'s hub/spoke topology). Not implemented in this ADR and not
needed to justify the decision; listed here only so a future ADR
doesn't re-derive the prerequisite. See "Future capabilities" below.
- **Retires the "re-parse on demand" framing (L2).** ADR-007's "Cost"
section and `src/engine.rs:112-115`'s doc comment framed re-parse as
the intended design. With the plan, the read path never re-parses;
the framing is retired. ADR-007's "Cost" section and the doc comment
are updated in the same commit.
- **L1 falls out.** The dead `_field_schema: &Value` parameter and the
`Vec<(String, Value)>` field storage (where the `Value` half is
unused) are replaced by `Vec<FieldPlan>`. No dead `Value` clones.
### Negative
- **Breaking public-API change (0.2.0 → 0.3.0).** `SequentialReader::new`
and `materialize_packed` change signatures (take `ReadPlan` instead
of `&Value`/`&BastDoc`). `ReadPlan` is a new public type. Per
AGENTS.md, this is semver-relevant. The crate is pre-1.0 with two
in-house downstream consumers, both updated with the bump. The
breakage is smaller than review #004's Option A (no `Bast*` type
changes — `BastDoc` stays borrowed, stays the validation-side tree).
- **A parallel typed tree, not a flat lookup table.** This is the
honest cost. `OffsetMap` and `PackedLayout` are flat `(path, range)`/
`(path, position)` projections of `BastType`; `ReadPlan` is a full
parallel hierarchy (`CompositePlan` mirrors `BastType`'s
Struct/Union/Array/Record). The maintenance tax is real and higher
than those: when schema semantics change, `BastDoc`/`BastType` and
`ReadPlan`/`CompositePlan` move together. It is worth it because the
perf win on composite-heavy schemas (the SFTP-shaped union-with-`$ref`
-variants packet) justifies it — see the next bullet. This is a
permanent tax accepted in exchange for a ~20–50x composite-dispatch
win on top of the 400x re-parse fix, not a structural symmetry with
the flat compiled forms.
- **Eager `$ref` resolution at compile time.** `ReadPlan::compile`
resolves all `$ref`s eagerly, including union variant refs. This is
correct (the schema is fixed at compile time) and matches the
review's Option B design, but it means a schema with a `$ref` cycle
(which BAST forbids — refs are always `#/$defs/<name>`, no
recursion) would loop forever. The meta-schema (`bast_meta`)
already rejects recursive schemas; `ReadPlan::compile` inherits
that guard. No new failure mode.
- **Nested `Box<CompositePlan>` vs flat `Vec<Op>`.** The plan as
specified uses nested `Box`es — idiomatic, debuggable, easy to
build. A flat `Vec<Op>` with jump indices (true "bytecode") would be
more cache-friendly but harder to build and read. Protocol headers
are small N (SFTP's largest packet has 5 fields); the perf win is
eliminating the re-parse and `resolve_typeref`, not SoA cache
effects. Start nested; go flat only if a bench says otherwise (a
two-way door — the plan is a private internal type; its shape can
change without a semver bump as long as the public `ReadPlan` name
and `compile`/`SequentialReader::new` signatures are stable).
## Scope Boundaries (What This Is Not)
- **Not a `BastDoc` replacement.** `BastDoc` stays as the
validation-side typed tree within ADR-011's scope (borrowed from
`&Value`, unchanged). The `Bast*` types and their signatures are not
touched by ADR-011. Validation (`bast_validation`), aligned one-shot
reads/writes (`engine.rs:334,467`), and `LayoutBuilder::build`
continue to walk `BastDoc` within ADR-011's scope. **ADR-012
subsequently revises two of these:** `BastDoc` becomes owned (§2a)
and `bast_validation` adopts a `ValidationPlan` (§3), both riding
the same 0.3.0 bump. ADR-011's scope statement is the ADR-011-only
view and is not re-litigated here.
- **Not a flat lookup table.** Packed positions are data-dependent;
the plan is a read program (instructions to walk), not a `(path,
offset)` table. This is inherent to packed sequential layout
(ADR-002), not a limitation of this design.
- **Not a validation plan.** `bast_validation`'s value-domain checks
(maxLength, enum values, union variant keys, integer ranges) are a
different concern and a different shape from `ReadPlan`. They are
out of ADR-011's scope; ADR-012 §3 adds a `ValidationPlan` in 0.3.0
rather than leaving validation on an interpretive `BastDoc` walk
indefinitely.
- **Not the review's Option A or Option B.** It is the "compiled form"
path the review pointed at but did not name: Option A (make `BastDoc`
own its data) kills the re-parse but leaves the read loop as a
`BastDoc` tree walk; Option B (precompute an owned read plan in
`SequentialReader::new`) is the surgical subset that closes H1 only.
This ADR is the principled version of B — a public `ReadPlan` built
at `compile` time, shared across readers and `materialize`, symmetric
with `OffsetMap`/`PackedLayout` — and it closes H1 + M1 (packed side)
+ L1 + L2.
## Recommended Order
1. **`ReadPlan` type + `compile`** — the `ReadPlan`/`FieldPlan`/
`CompositePlan`/`ReadKind`/`DiscriminatorPlan` types and the
`ReadPlan::compile(&Value, &str)` constructor. Pure addition; no
existing code touched. Unit-tested against the same BAST fixtures
the `BastDoc` tests use.
2. **`SequentialReader` rewrite** — the read loop walks `&ReadPlan`
instead of reconstructing `BastDoc`. `read_field_value`,
`read_typeref_value`, `walk_struct_size`, `read_union_value`,
`read_array_value`, `read_record_value` take plan nodes. The
existing `sequential_reader.rs` tests (which drive `read_next`/
`read_field`/`reset` over real buffers) pass unchanged — they
exercise the read path through the public API, so they validate
the rewrite without modification.
3. **`materialize_packed` rewrite** — takes `&ReadPlan`, walks the
plan to produce `Value`. The existing `validate_bytes` (packed)
tests cover it end-to-end.
4. **`AlkTypeEngine::compile` integration** — builds `Arc<ReadPlan>`
in packed mode, stores it, `sequential_reader()` hands out
`Arc::clone(&self.plan)`. `validate_bytes` (packed) calls
`materialize_packed(&self.plan, buffer)`.
5. **L2 — retire the "re-parse on demand" framing.** Update
ADR-007's "Cost" section (replace the "re-parse on demand"
paragraph with the `Arc<ReadPlan>` cost) and the
`src/engine.rs:112-115` doc comment. ADR-007's status block stays
"Accepted" for the factory decision; only the cost framing changes.
6. **Public API bump (0.2.0 → 0.3.0).** `lib.rs` re-exports `ReadPlan`;
`SequentialReader::new` and `materialize_packed` signatures change.
Update `alktty`/`alkcall` in the same commit.
7. **Verification block.** `cargo test --release`,
`cargo clippy --all-targets -- -D warnings`, `cargo doc --no-deps`
(new public type), `cargo build --target wasm32-unknown-unknown
--release` (the plan touches `sequential_reader.rs` and
`materialize.rs`, both wasm-relevant). Re-run the `alktty`
`wire_vs_bast` bench to confirm the 400x gap closes.
Steps 1–2 close H1. Step 3 closes the `materialize` half of M1.
Step 4 closes the `validate_bytes` half of M1 (packed side). Step 5
closes L2. L1 falls out at step 2. The aligned-side M1 paths
(`engine.rs:334,467`, `layout_builder.rs:190`) are left as-is per
"Out of scope."
## References
- [Review #004](../../reviews/004-performance-review.md) — the
performance finding (H1, M1, L1, L2) and the three fix options this
ADR supersedes
- [ADR-002](002-two-layout-modes-packed-vs-aligned.md) — the two
layout modes; `ReadPlan` is the packed read-side compiled form that
this ADR adds to the table
- [ADR-007](007-packed-mode-read-factory.md) — the engine as
`SequentialReader` factory; retained (owned fresh reader), with the
"re-parse on demand" framing retired (L2)
- [ADR-004](004-error-handling-validation-strategy.md) — `AlkTypeError`,
load-time build / access-time check, field-path-carrying errors;
`ReadPlan::compile` is a load-time build, the read loop is an
access-time check
- [ADR-010](010-generalized-validation-validate-bytes.md) —
`validate_bytes` (packed) consumes the `ReadPlan` via
`materialize_packed`
## Future capabilities (in 0.3.0 via ADR-012)
The deterministic-compile property of `ReadPlan` is a prerequisite for
several capabilities. ADR-012 ("Plan Fingerprinting, ValidationPlan,
and Closing the Deferred M1 Sites in 0.3.0") picks up all three items
below into the 0.3.0 release so they ship with this ADR's breaking
changes in one round of downstream churn, not two or three:
- **Fingerprinting the plan** (`#[derive(Hash)]` + a `fingerprint()`
method) for cross-run caching of compiled plans, disk-cached plans,
and `alkcall` schema-version handshakes. → **In 0.3.0 (ADR-012 §1).**
- **Closing the deferred M1 sites** via an owned `BastDoc` (lifetime
removal) for `LayoutBuilder` + extending `OffsetMap` with leaf
metadata for the aligned `read_field`/`write_field` paths. → **In
0.3.0 (ADR-012 §2).** Note: ADR-012 reframes the earlier "WritePlan"
candidate listed here as "not a new type — extend the existing
compiled forms (`PackedLayout`/`OffsetMap`) and cache the parse."
- A `ValidationPlan` that follows the same compile-once-walk-many
pattern for `bast_validation`. → **In 0.3.0 (ADR-012 §3).** Different
shape (value-domain, not byte-position) but the same class of
per-buffer re-walk cost on the read+validate-on-untrusted-input
common case. ADR-012 owns the shape decision and the implementation
plan scopes it. (Originally deferred by ADR-012 as "not a hot loop";
review #005 M3 reversed the deferral — see ADR-012 §3.)
None of the in-0.3.0 items justify this ADR; the 400x read-path gap
does. They are listed here as forward references and to record that
the "WritePlan" candidate has been reframed out by ADR-012.
## POC coverage
Before this ADR was accepted, a derisking POC on branch `readplan-poc`
walked the read loop against every `BastType` arm in
`src/sequential_reader.rs:303-381` (and the parallel arms in
`materialize.rs`) and confirmed the `ReadPlan`/`CompositePlan`/
`ReadKind`/`DiscriminatorPlan` shape covers all cases, including the
two spots where a plan arm could subtly miss a case:
- **Union discriminator split (`Byte` vs `Field`).** Both are covered:
`DiscriminatorPlan::Byte { offset, disc_type }` and
`DiscriminatorPlan::Field { name, field_index }`. The field-name case
pre-resolves the discriminator field's `ReadKind` so dispatch reads
it from the plan, not from a re-parsed `BastField`. The
field-disc union's declared `fields` (discriminator + any shared
fields) are carried as a sub-`ReadPlan` on `CompositePlan::Union`'s
`shared` field (a refinement of the POC shape, which stubbed the
`Field` arm — Finding 1); the production read loop walks `shared`
first, then the selected variant's `CompositePlan`. Nested-union
variants (a variant that is itself a union) are covered by ordinary
`CompositePlan` recursion; the POC rejected them, the 0.2.0 reader
accepts them, and the production shape restores parity.
- **Array variable-element-stride (`element_stride = 0`).** Covered:
`CompositePlan::Array { element, count, element_stride }` preserves
the `0`-signals-variable convention, and the read loop walks
sequentially when `element_stride == 0` (matching today's
`walk_variable_array_size`).
The POC also confirmed `ReadPlan: Send + Sync` holds for the planned
shape (immutable owned data, no interior mutability, no lifetimes).
@@ -0,0 +1,611 @@
# ADR-012: Plan Fingerprinting, ValidationPlan, and Closing the Deferred M1 Sites in 0.3.0
## Status
Accepted. Implemented in 0.3.0 — §3's `ValidationPlan` (phase 7,
2026-08-31), §1's fingerprinting (phase 6), §2a's owned `BastDoc`
(phases 3–4), §2b's `LeafMeta` (phase 5); all shipped 2026-09-02.
Bundles three pieces of work into the 0.3.0 release so the
crate ships one round of breaking changes, not two (or three). The
three pieces: (a) fingerprinting `ReadPlan`/`OffsetMap`, (b) closing
the deferred M1 sites via an owned `BastDoc` + `OffsetMap` `LeafMeta`,
and (c) a `ValidationPlan` that retires the interpretive
`bast_validation` walk (added by reversing the original "defer
`ValidationPlan`" decision — see "ValidationPlan — in scope for
0.3.0" below). Companion to [ADR-011](011-compiled-read-plan-for-packed-mode.md)
(the `ReadPlan`) and the [0.3.0 implementation plan](../../plans/030-compiled-forms.md).
§3's concrete shape was scoped by the follow-on design session and is
implemented in `src/validation_plan.rs` — see §3a below.
## Context
ADR-011 accepted the `ReadPlan` as the packed read-side compiled form
and deferred two things to "future capabilities":
1. **Fingerprinting the plan** for cross-run caching, disk-cached
compiled plans, and `alkcall` hub/spoke schema-version handshakes.
2. **A `WritePlan` and/or `ValidationPlan`** following the same
compile-once-walk-many pattern for the deferred M1 sites and the
validation walk.
ADR-011 also explicitly deferred the aligned-side M1 sites
(`engine.rs:334,467` `read_field`/`write_field`;
`layout_builder.rs:190` `LayoutBuilder::build`) as "a deliberate
reversible bet that an aligned-mode hot loop won't emerge."
This ADR retires the deferrals in one release. The reasoning is
timing: 0.3.0 is already a breaking bump (ADR-011 changes
`SequentialReader::new` and `materialize_packed` signatures), and the
crate has no real downstream consumers yet (only `alktty`/`alkcall`,
both in-house). Doing all three pieces now costs one round of
downstream churn instead of two or three, and the fingerprinting work
cuts across both `ReadPlan` and `OffsetMap` — splitting would create
a cross-release dependency that's cleaner in one release. The
`ValidationPlan` inclusion follows the same logic applied to the
validation walk: deferring it would create a *second* breaking change
to `validate_bytes`/`bast_validation` after 0.3.0, which is exactly
the round of downstream churn this release exists to retire.
### Reframing "WritePlan"
ADR-011's "Future capabilities" section listed a `WritePlan` as a
candidate. On inspection, a new public `WritePlan` type is the wrong
shape for the deferred M1 sites, for two reasons:
1. **The packed write-side already has a compiled form: `PackedLayout`.**
`LayoutBuilder::build`'s M1 re-parse is the *builder* re-parsing
`BastDoc::new` on each `build()` call to get the typed tree it
walks. The fix is to cache the parsed tree on the builder at `new()`
time — internal, non-breaking, no new public type. The compiled
form (`PackedLayout`) is unchanged; only its construction stops
re-parsing.
2. **The aligned R/W side already has a compiled form: `OffsetMap`.**
`read_field`/`write_field`'s M1 re-parse is `lookup_leaf_field`
walking `BastDoc` to get leaf metadata (`kind`, `encoding`,
`endian`) that `OffsetMap` doesn't carry. The fix is to extend
`OffsetMap`'s entries with that metadata at `compute` time —
additive fields on an existing public type (breaking, but we're
bumping anyway). No new public type.
A new `WritePlan` type would overlap with `PackedLayout` (packed
write) and `OffsetMap` (aligned R/W) without a clean distinguishing
shape. The honest picture: the packed write-side compiled form is
`PackedLayout`; the aligned R/W compiled form is `OffsetMap`; the M1
fixes are "cache the parse" and "extend the compiled form with leaf
metadata," not "add a third compiled form." This serves the
minimal-public-API-changes goal better than a literal `WritePlan`.
### `ValidationPlan` — in scope for 0.3.0 (no longer deferred)
The BAST-native validator (`bast_validation`) walks `BastDoc` to
check value-domain constraints (enum value sets, integer ranges,
`maxLength` caps, union variant keys). This is a different shape
from `ReadPlan`/`OffsetMap` (value-domain, not byte-position), and
the earlier framing deferred it as "not a hot loop — validation is
opt-in per operation per AGENTS.md."
**That deferral is reversed.** The "not a hot loop" dismissal
under-counted the common case: **read + validate together on
untrusted input.** The downstream `alkcall` consumer accepts schemas
from arbitrary internet peers in a hub/spoke topology (AGENTS.md §3);
the common operation on an incoming frame is "read it, then validate
it before acting." `validate_bytes` (ADR-010) is therefore called
once per incoming buffer, and each call re-walks `BastDoc` for
validation even after ADR-011 makes the *read* half plan-fast. That
is the same class of per-buffer interpretive cost review #004 measured
for the read path (400x per chunk), on a different code path, on the
operation the untrusted-input discipline actually requires.
The cost-of-inaction framing that the original deferral relied on was
also wrong: a `ValidationPlan` introduced *after* 0.3.0 would be a
breaking change to `validate_bytes`'s contract and to the
`bast_validation` public surface, forcing rework of `alktty`/`alkcall`
— the exact downstream-churn this release is supposed to retire, not
create a second round of. Shipping it in 0.3.0 pays the cost once,
alongside the other breaking changes, while there are zero real
consumers. The cost of action now is a static, known quantity; the
cost of action later is the same work plus a second round of
downstream churn plus the risk of the interpretive path being the
one that gets used in the meantime on untrusted bytes.
**Decision: a `ValidationPlan` ships in 0.3.0 as §3 below.** The
shape is a compile-once-walk-many compiled form over the BAST
document's value-domain constraints, symmetric to `ReadPlan` (packed
read-side) and `OffsetMap` (aligned R/W). The concrete shape,
construction, and `validate_bytes` integration are scoped in the
0.3.0 implementation plan (a dedicated phase) and detailed in a
follow-on design session before implementation; this ADR commits the
*decision* (in 0.3.0, not deferred) and the *scope* (a compiled
validation form that retires the interpretive `BastDoc` walk in
`bast_validation`), so the deferral black hole is closed.
## Decision
### 1. Fingerprinting — `ReadPlan: Hash + Eq`, `OffsetMap: Hash + Eq`
Add `#[derive(Hash, Eq)]` (alongside the existing `Debug, Clone, PartialEq`)
to `ReadPlan` and `OffsetMap`, plus their public sub-types
(`FieldPlan`, `CompositePlan`, `ReadKind`, `DiscriminatorPlan`,
`ByteRange`, and the new `LeafMeta` — see §2). `VariantPlan`/
`VariantKind` are not in the production `ReadPlan` shape (ADR-011
was refined on acceptance to drop them — see ADR-011 status), so
they are not derived. `ValidationPlan` (§3) gets `Hash + Eq` + its
own `fingerprint()` as part of its public surface.
**`by_name` representation change.** `ReadPlan.by_name` is currently
`HashMap<String, usize>`. `HashMap` iteration order is non-deterministic
and `HashMap` does not implement `Hash`, which blocks `#[derive(Hash)]`
on `ReadPlan`. Switch `by_name` to `BTreeMap<String, usize>`. Lookup
cost at protocol-header N (~5 fields) is negligible (the `BTreeMap` is
only used by `read_field`'s name→index lookup, not by the sequential
`read_next` hot path). This makes the derived `Hash` cover the full
structural state of the plan.
**Fingerprint contract.** Two plans with equal `Hash` (or equal under
`PartialEq`) produce identical reads over identical bytes. Formally:
`plan1 == plan2 ⟹ ∀ buffer. read(plan1, buffer) == read(plan2, buffer)`.
This is the contract the downstream uses rely on:
- **Cross-run disk cache.** A consumer can hash a `ReadPlan`/
`OffsetMap` and cache the compiled plan keyed by the hash, skipping
`compile` on warm starts. Safe because the contract guarantees a
cache hit produces identical read behavior.
- **`alkcall` hub/spoke schema handshake.** Peers exchange plan
fingerprints instead of full BAST documents. A peer that receives a
fingerprint it has already compiled can skip re-transmitting the
schema. The contract guarantees fingerprint equality implies
behavioral equivalence, so the handshake is sound.
- **Schema-version diagnostics.** A consumer can log a plan
fingerprint alongside read results for reproducibility — two runs
over "the same schema" that produce different fingerprints reveal a
silent schema drift.
The contract is a *behavioral* equivalence, not a structural identity:
two plans with different `by_name` insertion order but the same
`fields` Vec produce the same reads, and after the `BTreeMap` change
they also produce the same `Hash`. The contract is documented on the
`Hash` impl and tested by a property-style test (compile the same
schema twice, assert `plan1 == plan2` and `plan1.hash() ==
plan2.hash()`).
**Fingerprint API.** No new public method is strictly needed —
consumers call `std::hash::Hash` directly. For ergonomics and to make
the contract visible, add a convenience method:
```rust
impl ReadPlan {
/// A stable 64-bit fingerprint of this plan's read behavior.
///
/// Two plans with the same fingerprint produce identical reads
/// over identical bytes (the fingerprint contract).
pub fn fingerprint(&self) -> u64;
}
impl OffsetMap {
/// A stable 64-bit fingerprint of this offset map's read/write
/// behavior. Same contract as `ReadPlan::fingerprint`.
pub fn fingerprint(&self) -> u64;
}
```
Implemented via `std::hash::DefaultHasher` (or a stable hasher like
`FxHasher` if we want cross-version stability — decision belongs to
the implementation step, called out in the plan). The fingerprint is
additive API, not breaking.
### 2. Closing the deferred M1 sites
#### 2a. `LayoutBuilder` — cache the parsed `BastDoc` at `new()`
`LayoutBuilder` currently stores `doc_value: Value` + `root_name: String`
and re-parses `BastDoc::new(&self.doc_value, &self.root_name)` on every
`build()` call (`layout_builder.rs:190`). The fix: store the parsed
typed tree at `new()` time and reuse it in `build()`.
This requires `BastDoc` to be owned (no lifetime borrowing from
`doc_value`). Two options:
- **Option α (smaller):** keep `BastDoc<'a>` borrowing, store
`doc_value: Value` + a *pre-resolved, owned* representation of just
what `build` needs (the field tree with `$ref`s resolved). This is
essentially a `WritePlan` by another name — rejected per the
reframing above.
- **Option β (cleaner):** make `BastDoc` own its data. This is
review #004's Option A, scoped to `LayoutBuilder` only. It's a
larger refactor but eliminates the lifetime entanglement for the
builder and is the prerequisite for any future owning consumer that
wants to cache the parsed tree.
**Decision: Option β, scoped to `LayoutBuilder`.** The `BastDoc<'a>` →
`BastDoc` (owned) refactor is the principled fix and is already
breaking (the `Bast*` types are re-exported from `lib.rs`), so it
rides the 0.3.0 bump. This does *not* change `SequentialReader` or
`materialize_packed` (those consume `ReadPlan` per ADR-011, not
`BastDoc`). It changes `LayoutBuilder::new` to parse once and `build`
to reuse. The `doc_value: Value` field is removed; the builder holds
the owned `BastDoc` directly.
**Note on `BastDoc` ownership scope:** ADR-011 left `BastDoc` borrowed
and unchanged ("the validation-side typed tree"). This ADR changes
that: `BastDoc` becomes owned. The validation-side (`bast_validation`)
and aligned-side (`OffsetMap::compute`, `materialize_aligned`)
consumers adapt to the owned `BastDoc` — they no longer need a
borrowed `&Value` kept alive alongside. This is a net simplification:
one typed-tree type, owned, used by all non-`ReadPlan` consumers. The
POC on `readplan-poc` confirmed `ReadPlan` doesn't need `BastDoc` to
be borrowed (it compiles from `&Value` once and discards the
`BastDoc`), so making `BastDoc` owned doesn't regress the read path.
#### 2b. `OffsetMap` — carry leaf metadata
`OffsetMap` currently stores `Vec<(String, ByteRange)>`. The
`read_field`/`write_field` M1 re-parse is `lookup_leaf_field` walking
`BastDoc` to get `LeafFieldInfo { kind, encoding, endian }`
(`engine.rs:556-600`). The fix: extend `OffsetMap`'s entries to carry
that metadata at `compute` time.
```rust
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub struct LeafMeta {
pub kind: AlkTypeKind,
pub encoding: VariableEncoding,
pub endian: Endian,
}
pub struct OffsetMap {
fields: Vec<(String, ByteRange, LeafMeta)>, // was Vec<(String, ByteRange)>
total_size: usize,
}
```
`OffsetMap::compute` resolves each leaf field's `LeafMeta` during the
walk (it already walks the tree; it just doesn't currently record the
metadata). `read_field`/`write_field` drop the `BastDoc::new` +
`lookup_leaf_field` calls and read `LeafMeta` from the map. The
`LeafFieldInfo` struct in `engine.rs` is removed (replaced by
`OffsetMap`'s `LeafMeta`).
**Breaking changes:**
- `OffsetMap::get` return type: `Option<&ByteRange>` →
`Option<(&ByteRange, &LeafMeta)>` (or a small accessor struct).
Call sites in `alktty`/`alkcall` update with the bump.
- `ByteRange` is unchanged (still `Copy + Hash`).
- `LeafMeta` is a new public type, re-exported from `lib.rs`.
This is additive on the *capability* (the map now answers questions it
previously couldn't) but breaking on the *signature* (`get`'s return
type changes). Rides the 0.3.0 bump.
### 3. `ValidationPlan` — compile-once validation form
`bast_validation` currently walks `BastDoc` interpretively on every
`validate_bytes` call to check value-domain constraints (enum value
sets, integer ranges, `maxLength` caps, union variant keys). After
ADR-011, the *read* half of `validate_bytes` (packed) is plan-fast;
the *validation* half is still an interpretive `BastDoc` walk per
buffer. On the `alkcall` hub/spoke topology, `validate_bytes` is the
gate between "bytes arrived from an untrusted peer" and "act on the
decoded frame," so it runs once per incoming buffer and validation is
hot in the same sense review #004 measured for the read path.
**Decision: a `ValidationPlan` is a compiled form over the BAST
document's value-domain constraints, built once at `compile` time
(symmetric to `ReadPlan`/`OffsetMap`) and walked by
`bast_validation`/`validate_bytes` without re-touching `BastDoc`.**
The shape, construction, and `validate_bytes` integration are scoped
in the 0.3.0 implementation plan as a dedicated phase and detailed in
a follow-on design session before implementation begins. The
properties this ADR commits to (so the plan and any implementing agent
have a fixed contract):
- **Compile-once-walk-many.** `ValidationPlan::compile` walks `BastDoc`
once; `validate_bytes` (both modes) walks the `ValidationPlan` per
buffer, never `BastDoc`. This is the same pattern as `ReadPlan` and
`OffsetMap`; it is the structural reason the per-buffer
interpretive cost goes away.
- **Value-domain, not byte-position.** The plan carries constraint
descriptors (enum allowed-sets, integer range bounds, `maxLength`
caps, union variant keys, and any other value-domain checks
`bast_validation` performs today), keyed for dispatch against the
materialized `Value` tree, not byte offsets. The shape is therefore
different from `ReadPlan`/`OffsetMap`; the *pattern* (compiled form,
immutable, shared via `Arc`) is the same.
- **No new `BastDoc` walk in the hot path.** After this ADR, the only
consumers that walk `BastDoc` interpretively are the one-shot
`compile` paths (`ReadPlan::compile`, `OffsetMap::compute`,
`ValidationPlan::compile`, `LayoutBuilder::new`). The per-buffer
paths (`sequential_reader`, `materialize_packed`,
`materialize_aligned`, `validate_bytes`) all walk compiled forms.
This is the end state ADR-011 pointed at; this ADR closes it.
- **Semver.** `ValidationPlan` is a new public type, re-exported from
`lib.rs`. `validate_bytes`'s *signature* is unchanged (still
`(buffer) -> Result<(), AlkTypeError>`); the change is internal
(walks the plan instead of `BastDoc`). If the `ValidationPlan`
design surfaces a need to change `validate_bytes`'s signature, that
rides the 0.3.0 bump and is recorded in the plan's Semver Contract
table when the shape is scoped. `bast_validation`'s public surface
(`build_validator`, `validate_value`) is reviewed at shape-scope
time; additive changes ride the bump, removals/renames are avoided
unless the shape work shows they're necessary.
- **Fingerprinting.** `ValidationPlan` is `Hash + Eq` with a
`fingerprint()` method, same as `ReadPlan`/`OffsetMap` (§1/§4), so
the downstream uses (cross-run cache, `alkcall` handshake,
schema-version diagnostics) extend to the validation form without
new API. The fingerprint contract generalizes: two validation plans
with equal hashes accept/reject identical `(bytes)` identically.
**What this ADR does *not* decide** (left to the follow-on shape
session + plan phase): the concrete `ValidationPlan` struct/enum
shape, how `maxLength`/range/enum/union-key constraints are
represented, whether `bast_validation`'s `validate_value` is retired
or kept as a convenience wrapper over the plan, and whether the
`AlkTypeKind`-driven dispatch in `bast_validation` collapses into the
plan or stays a thin match over plan-carried descriptors. These are
shape questions, not decision questions; the decision (in 0.3.0,
compiled form, no per-buffer `BastDoc` walk) is fixed here.
### 3a. `ValidationPlan` shape — resolved by the design session
The follow-on design session (0.3.0 phase 7 predecessor) resolved the
open shape questions; implemented in `src/validation_plan.rs`:
- **Shape.** `ValidationPlan { root: ValidNode }`, a compiled
constraint tree — one `ValidNode` arm per value-domain check,
mirroring the interpretive walker's arms one-to-one:
`Int { min, max }` / `I64` / `Uint { max }` / `U64` / `Float` / `Bool`
/ `Str { max_len }` / `Bytes { max_len }` / `Enum { count }` /
`Struct { fields: Vec<ValidField> }` / `Union { variants:
Vec<ValidVariant> }` / `Array { count, element }` / `Record
{ values }`. `ValidField` carries `name + node`; `ValidVariant`
carries `key + node`. The nodes are public (diagnostics access via
`ValidationPlan::root()`); construction is only possible through
`compile`. Note the union node carries *only the variant nodes* — the
declared union `fields` (shared fields) are validated as part of the
variant walk, because the walker dispatches on the materialized
`__discriminator` and validates the whole object against the selected
variant (the `ValidNode::Union` doc comment records this; the
interpretive union arm recursed into the variant the same way).
- **Constraint representation.** Inline scalar fields on the node arms
(ranges as `i64`/`u64` pairs, `maxLength` as `Option<usize>`, enum
bound as `count: u64`, union keys as owned `String`s). `maxLength` is
resolved from the *owning field* at compile time and baked into the
`Str`/`Bytes` leaf — the walk never consults field annotations. It
never crosses a `$ref` (a `$ref` always targets a struct/union/enum
`$defs` entry, so the interpretive walk could never consult it
through one either).
- **`compile` signature.** `ValidationPlan::compile(&BastDoc) ->
Result<Self, AlkTypeError>` — the plan-table's `&str` root-name
parameter was vestigial (the doc already holds its root).
- **`validate_value` disposition.** Retained as a one-shot wrapper:
`compile(doc)` + `validate(value)`. The interpretive walker behind it
is *retired* (deleted) — the wrapper delegates to the plan, so there
is one constraint implementation, not two. `bast_validation.rs` keeps
the shared error helper (`validation_err`) and the
`__discriminator` key constant.
- **Error contract preserved.** The plan walk reproduces the
interpretive error messages byte-identically: a segment stack
(`field` / `[index]` / `[key]`) renders paths only on failure — zero
per-node allocation on the happy path. Numeric `__discriminator`
dispatch matches mapping keys without allocation for the u64/i64
forms (mapping keys are stringified integers; non-integer numbers
fall back to `Number::to_string`).
- **Compile-time rejection of adversarial graphs.** Eager `$ref`
resolution with a definition-level cycle set and a depth cap (128):
a cyclic or self-referential schema is `AlkTypeError::Schema` at
compile, not a stack overflow — the interpretive walker resolved
`$ref`s lazily with no guard and could overflow on recursion.
Diamond (shared, non-cyclic) refs compile fine; the cycle set is
path-scoped.
- **Engine integration.** `AlkTypeEngine` holds `Arc<ValidationPlan>`
built at `compile` time in *both* modes; the accessor
`validation_plan() -> &Arc<ValidationPlan>` is new public API.
`validate_bytes`'s signature is unchanged. The plan compile runs
*before* the layout build: it is the engine's reference-graph gate
(see Consequences).
- **Scope note.** phase-7's `Send + Sync` / `Hash + Eq` /
`fingerprint()` requirements are structural on the types above
(`#[derive(...)]` on plain owned data; the same `DefaultHasher`
fingerprint as phase 6).
### 4. Fingerprinting `OffsetMap` (bundled with §2b)
Since `OffsetMap` is getting new fields (`LeafMeta`) in §2b, its
`#[derive(Hash, Eq)]` (from §1) covers the new fields automatically.
The fingerprint contract for `OffsetMap` is the aligned-side analog
of `ReadPlan`'s: two offset maps with equal hashes produce identical
aligned reads/writes over identical bytes.
## Scope
### In scope
- `ReadPlan: Hash + Eq` + `fingerprint()` method (§1).
- `OffsetMap: Hash + Eq` + `fingerprint()` method (§1, §4).
- `BastDoc<'a>` → `BastDoc` (owned) refactor, scoped to the consumers
that currently hold `doc_value: Value` and re-parse: `LayoutBuilder`,
`bast_validation`, `materialize_aligned`, `OffsetMap::compute`
(§2a). `ReadPlan::compile` and the packed read path are unaffected
(they consume `&Value` once and discard `BastDoc`).
- `LayoutBuilder` caches the owned `BastDoc` at `new()`, `build()`
reuses it — no re-parse (§2a).
- `OffsetMap` carries `LeafMeta`; `read_field`/`write_field` drop
`BastDoc::new` + `lookup_leaf_field` (§2b).
- `LeafMeta` new public type (§2b).
- `BTreeMap` for `ReadPlan.by_name` (§1).
- `ValidationPlan` new public type + `compile` + `Hash + Eq` +
`fingerprint()` (§3). `validate_bytes` (both modes) walks the
`ValidationPlan` instead of re-walking `BastDoc`. `bast_validation`
adopts the plan; the public `validate_value`/`build_validator`
surface is reviewed at shape-scope time and rides the bump only if
the shape work shows a signature change is necessary.
- `ValidationPlan: Hash + Eq` + `fingerprint()` method (§3, §1) — the
fingerprint contract extends to the validation form.
### Out of scope
- Disk-cache or handshake *implementations* — the fingerprint
*contract* and method are in scope (§1, §3); the downstream uses
(cache format, wire protocol) are the consumers' problem, not this
ADR's.
- The `ValidationPlan` shape — **resolved** (§3a). Decided by the
design session and implemented in `src/validation_plan.rs`; the
decision (in 0.3.0, compiled form, no per-buffer `BastDoc` walk) was
fixed here.
- Cycle-guard hardening for the *layout* walkers' own recursion
(`LayoutBuilder`/`OffsetMap` struct recursion) beyond the engine-path
gate described in Consequences — if a non-engine entry point walking
those types on untrusted docs becomes a consumer pattern, the
guards get their own change (the `AlkTypeEngine::compile` gate
covers the supported path today).
- Cross-version fingerprint stability — the fingerprint is stable
within a crate version but may change across versions (a new
`AlkTypeKind` variant, for example, changes the hash). Cross-version
stability is a non-goal; consumers cache within a version. The
implementation step chooses a hasher and documents the stability
contract.
## Consequences
### Positive
- **One breaking release, not two (or three).** ADR-011's `ReadPlan`
+ this ADR's `BastDoc`-owned + `OffsetMap` extension +
`ValidationPlan` all ship together. The two in-house downstream
consumers (`alktty`, `alkcall`) update once.
- **Closes all deferred M1 sites.** `LayoutBuilder::build`
(`layout_builder.rs:190`), `read_field` (`engine.rs:334`),
`write_field` (`engine.rs:467`) all stop re-parsing. The packed-side
`validate_bytes` (`engine.rs:284`) was already closed by ADR-011;
this ADR closes the aligned-side equivalent.
- **Retires the interpretive validation walk.** `validate_bytes` on
untrusted streams (the `alkcall` common case) stops re-walking
`BastDoc` per buffer. This is the latent perf cliff review #005 M3
flagged: the read half was plan-fast after ADR-011, the validation
half was not. Closing it here — while there are zero real consumers
and one breaking bump already paying the downstream-churn cost —
avoids a second breaking change to `validate_bytes`/`bast_validation`
after 0.3.0. (Implemented: a spot benchmark of plan-validate on a
4-field mixed frame puts the validation half at ~0.2 µs/validate;
the compile-per-call one-shot it replaces runs ~2.7x slower before
the walk is even counted — and the full 0.2.0 per-buffer cost
included lazy `$ref` deep-clones that the one-shot no longer pays.
The materialize half, not validation, remains the dominant
`validate_bytes` cost.)
- **Compile-time rejection of cyclic `$ref` graphs.** A side effect of
eager plan compilation: a self-referential document is now a clean
`Schema` error instead of a stack overflow. The plan compile runs
*before* the layout build in `AlkTypeEngine::compile`, making it the
engine's reference-graph gate — `LayoutBuilder`/`OffsetMap`
struct-recursion had no cycle guard and previously could recurse
unboundedly on such a document (a pre-existing untrusted-schema
hazard, surfaced by the phase-7 `compile_rejects_cyclic_ref_graph`
test). (Resolved since: review #006 H2 added the shared
`walk_guard::check_ref_graph` guard at every standalone walker entry,
so the trust boundary no longer depends on the engine path.)
- **Fingerprinting enables downstream uses.** Cross-run plan caching,
`alkcall` schema handshake, and schema-version diagnostics all
become possible without further API work — across `ReadPlan`,
`OffsetMap`, and `ValidationPlan`.
- **`BastDoc` owned is a net simplification.** One typed-tree type,
owned, used by all `*::compile` paths. No more
lifetime-entanglement workarounds. The "re-parse on demand" framing
from ADR-007 is fully retired across read, write, and validation
paths.
- **`OffsetMap` extension is additive capability.** The map now
answers `kind`/`encoding`/`endian` questions it previously couldn't,
enabling future aligned-side tools without re-walking `BastDoc`.
### Negative
- **Breaking public-API changes (0.2.0 → 0.3.0).** `BastDoc<'a>` →
`BastDoc` (owned) changes every `Bast*` signature that took `&'a`.
`OffsetMap::get` return type changes. `LeafMeta` is new public.
`ReadPlan` is new public (from ADR-011). `ValidationPlan` (+ the
`ValidNode`/`ValidField`/`ValidVariant` node types) is new public
(§3a). All ride the bump.
- **`BastDoc` ownership refactor is broad.** Touches `bast.rs` (every
typed node: `&'a str` → `String`/`Arc<str>`, `&'a Value` →
`Value`/`Arc<Value>`) and every consumer (`layout_builder`,
`offset_map`, `materialize`, `bast_validation`, `engine`). This is
review #004's Option A, which ADR-011 deferred — this ADR picks it
up because the `LayoutBuilder` M1 fix requires it and we're bumping
anyway. The refactor is mechanical (lifetime removal, not logic
rewrites); the POC on `readplan-poc` confirmed the read path is
unaffected.
- **Interpretive `validate_value` is compile-per-call.** The retained
one-shot wrapper (`bast_validation::validate_value`) compiles a plan
then validates — fine for one-off/diagnostic use, wrong for per-
buffer use. Per-buffer callers must hold the engine (or a plan) —
the doc comments say so. The walker it replaced had the inverse
trade (no compile, but interpretive per call); the engine path
(compile once) is the one that matters.
- **Aligned-mode `maxLength`-reserved strings/bytes.** Materialization
emits the *full reserved* (zero-padded) data for these fields. A
`ValidationPlan` compiled from a document used in packed mode would
apply `maxLength` to trimmed length, matching packed semantics; the
aligned materializer's zero-padding means the value passed to
validation can carry trailing NULs. This is pre-existing
materialize behavior (not a plan artifact); consumers relying on
trimmed values already see it.
- **Fingerprint cross-version stability is not guaranteed.** A future
`AlkTypeKind` variant changes the hash. Documented as a within-
version contract. Consumers that need cross-version stability
serialize the BAST document and re-compile.
- **`BTreeMap` for `by_name` is a tiny lookup cost.** Negligible at
protocol-header N; irrelevant to the 400x fix.
## Scope Boundaries (What This Is Not)
- **Not a `WritePlan` type.** The packed write-side compiled form is
`PackedLayout`; the aligned R/W compiled form is `OffsetMap`. The
M1 fixes are "cache the parse" (§2a) and "extend the compiled form
with leaf metadata" (§2b), not "add a third compiled form."
- **Not a `ValidationPlan` deferral.** `ValidationPlan` is in scope
(§3) and implemented (§3a); the interpretive walker is retired.
- **Not cross-version fingerprint stability.** Within-version only.
- **Not a disk-cache or wire-protocol spec.** The fingerprint contract
and method are in scope; the downstream uses are the consumers'
concern.
## Recommended Order
See [the 0.3.0 implementation plan](../../plans/030-compiled-forms.md)
for the step-by-step execution order. The high-level grouping:
1. **`ReadPlan` (ADR-011 steps 1–5)** — the packed read-path fix. Closes
H1 + packed-side M1 + L1 + L2.
2. **`BastDoc` owned (§2a)** — the typed-tree ownership refactor. Prerequisite
for the `LayoutBuilder` M1 fix and for `ValidationPlan::compile`.
3. **`LayoutBuilder` M1 fix (§2a)** — cache the owned `BastDoc` at `new()`.
4. **`OffsetMap` extension (§2b)** — carry `LeafMeta`; close the
aligned-side `read_field`/`write_field` M1.
5. **`ValidationPlan` (§3)** — compiled validation form; `validate_bytes`
walks the plan instead of `BastDoc`. Requires the owned `BastDoc`
from step 2 for `ValidationPlan::compile`.
6. **Fingerprinting (§1, §4)** — `Hash + Eq` + `fingerprint()` on
`ReadPlan`, `OffsetMap`, and `ValidationPlan`. Rides on top of the
above.
7. **Public API bump (0.2.0 → 0.3.0)** — `lib.rs` re-exports, version
bump, update `alktty`/`alkcall`.
8. **Verification block** — full suite + wasm + bench.
## References
- [ADR-011](011-compiled-read-plan-for-packed-mode.md) — the
`ReadPlan` (packed read-side compiled form). This ADR extends the
0.3.0 release with fingerprinting, the deferred M1 fixes, and the
`ValidationPlan`.
- [Review #004](../../reviews/004-performance-review.md) — the
performance finding (H1, M1, L1, L2). ADR-011 closed H1 + packed
M1 + L1 + L2; this ADR closes the aligned-side M1.
- [Review #005](../../reviews/005-plan-review-030.md) — the 0.3.0 plan
review whose M3 finding reversed the `ValidationPlan` deferral.
- [ADR-007](007-packed-mode-read-factory.md) — the "re-parse on
demand" framing, retired across read, write, and validation paths
by ADR-011 + this ADR.
- [ADR-002](002-two-layout-modes-packed-vs-aligned.md) — the two
layout modes; `OffsetMap` is the aligned R/W compiled form extended
here with `LeafMeta`.
- [0.3.0 implementation plan](../../plans/030-compiled-forms.md) —
the step-by-step execution plan.
@@ -0,0 +1,316 @@
# ADR-BAST: BAST (Binary Abstract Syntax Tree) as the Schema Format
## Status
Accepted — supersedes the format-specific content of
[ADR-001](001-alktype-purpose-scope-jsonschema-engine.md). ADR-001's
purpose, scope, and "schema is the format" principle are retained and
strengthened; only the concrete format (custom-keyword JSON Schema →
BAST) is superseded by this ADR.
## Context
alktype v0.1.0 embedded binary layout information inside standard JSON
Schema documents via custom keywords:
```json
{
"AlkType:Struct": true,
"type": "object",
"properties": {
"channel_id": { "AlkType:Uint32": true, "type": "integer" },
"length": { "AlkType:Uint32": true, "type": "integer" }
},
"endian": "big"
}
```
This worked for the Rust engine — it walked the tree, detected
keywords, computed offsets. But it created friction for everything
outside Rust:
1. **Cross-language consumption.** A Python, Go, or TypeScript consumer
that wanted to parse an alktype schema had to re-implement custom
keyword detection. The format was not self-describing — you needed
to know that `AlkType:Uint32` meant "4-byte little/big-endian
unsigned integer" out-of-band.
2. **No meta-schema.** The custom keywords were not part of any JSON
Schema dialect, so `jsonschema` itself could not validate an alktype
schema's *structure*. Editors had no autocomplete; a typo in
`AlkType:Uint32` (e.g., `AlkType:UINT32`) was a runtime engine
error, not a schema-validation error.
3. **Awkward composition.** The keyword-value shape (`true` vs an
annotation object) and the `normalize_refs` step needed to bridge
TypeBox's bare-name `$ref` output and `jsonschema`'s JSON Pointer
requirement were engine internals leaking into the format.
4. **Validator coupling.** The v0.1.0 bytes path validated by
registering 19 `jsonschema::Keyword` factories (~200 lines). The
custom keyword integration was the only way to enforce value-domain
constraints (integer ranges, `maxLength`, enum index bounds) on the
materialized `Value`. The built-in `enum` keyword checked string
membership, but the materializer emitted `Value::Number(index)` —
so out-of-bounds enum indices *silently passed* (a dead constraint).
The engine's core logic (layout computation, data access, union
dispatch, two layout modes) was format-agnostic beneath the accessor
layer. A POC on branch `bast-validator-poc` (commit `f371fe4`,
`src/bast_poc.rs`) proved that a `kind`-based vocabulary with
`$defs`/`$ref`, a BAST-native validator, and lazy variant ref
resolution could replace the custom-keyword machinery end-to-end with
no loss of capability and a net reduction in code. The research record
is [`docs/research/bast-pivot.md`](../../research/bast-pivot.md);
decisions D-BAST-001 through D-BAST-009 are recorded there.
## Decision
**alktype's schema format is BAST (Binary Abstract Syntax Tree): a JSON
document that describes binary data layouts using a `kind`-based
vocabulary with `$defs`/`$ref` for composition.** BAST is itself a
valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser.
The normative format specification is
[`docs/architecture/bast-format.md`](../bast-format.md) (meta-schema,
TypeRef, examples, validation model). This ADR records the decision and
its consequences; the spec records the shape.
### Design principles
1. **BAST is a JSON Schema instance.** A BAST document is valid JSON
that conforms to the BAST meta-schema (a standard Draft 2020-12 JSON
Schema). Any JSON Schema validator can check whether a BAST document
is well-formed; editors with JSON Schema support provide autocomplete
and inline validation for free.
2. **`$defs`/`$ref` for composition.** Named type definitions live in a
top-level `$defs` block. `$ref` handles cross-references and union
variant references — the same pattern as JSON Schema's own `$defs`
and TypeBox's `Type.Module`. No custom reference resolution
mechanism.
3. **`kind`-based vocabulary.** Every type has a `kind` field whose
value is a known string (`"uint32"`, `"struct"`, `"union"`, etc.).
This replaces the `AlkType:*` custom-keyword pattern with a flat,
easily-matched string. The 18 `AlkTypeKind` enum variants are
the BAST kinds; `AlkTypeKind::from_bast_str`/`to_bast_str` map between
the enum and the lowercase BAST strings (D-BAST-002).
4. **Order is explicit.** Struct fields are an ordered array, not an
object with `properties`. Field order is unambiguous — no reliance
on `serde_json`'s `preserve_order` for correctness — and matches the
mental model of binary layouts.
5. **Annotations are type-level properties.** Endianness, alignment,
encoding, and discriminators are properties of the type definition
or field, not custom keywords on a separate schema object. Their
*semantics* carry forward unchanged from
[ADR-003](003-schema-annotations.md); only their *location* moves.
### Document shape
Every BAST document has the same top-level shape:
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001).
Convention (first entry) is fragile and depends on JSON key order; an
explicit parameter is used instead.
### TypeRef
`TypeRef` is the central mechanism for referencing types. Four forms:
primitive string (`"uint32"`), `$ref` object
(`{ "$ref": "#/$defs/Read" }`), array object
(`{ "kind": "array", "element": "uint32", "count": 3 }`), and record
object (`{ "kind": "record", "values": "string" }`).
The `$ref` form uses standard JSON Pointer syntax **restricted to
`#/$defs/<name>`** — no external references, no fragment-only pointers,
no bare names. The restriction keeps resolution a single hash lookup
and eliminates the `normalize_refs` step the v0.1.0 engine needed for
TypeBox's bare-name refs.
### Meta-schema
The BAST meta-schema is a standard JSON Schema (Draft 2020-12) that
validates the *structure* of BAST documents (is it well-formed?). It
lives at a stable URL (`https://alk.dev/bast/v1/schema`) and is embedded
in the crate as `BAST_META_SCHEMA` (re-exported from the crate root) for
offline use. A *different* validator — the BAST-native validator (see
[ADR-VAL-SPLIT](val-split-two-validator-model.md)) — validates *binary
data* against a BAST document (are the bytes a valid instance?). These
are different validators for different inputs.
### The typed parser
`src/bast.rs` parses a BAST document into a borrowed typed tree
(`BastDoc`/`BastDef`/`BastStruct`/`BastField`/`BastType`/`BastUnion`/
`BastEnum`/`BastArray`/`BastRecord`/`BastRef`). Three consumers (layout
engines, materializer, BAST-native validator) walk the same tree, so a
typed view pays for itself. See
[`schema-layer.md`](../schema-layer.md) for the parser's surface and
[`bast-format.md`](../bast-format.md) for the format.
Variant `$ref`s (union `mapping` entries) are resolved **lazily** by
the materializer and validator via `BastDoc::resolve_typeref` — no
compile-time inlining.
### Untrusted input
Every path that walks a BAST document returns
`Err(AlkTypeError::Schema)` on a malformed document, never
`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream
`alkcall` consumer accepts schemas from arbitrary internet peers in its
hub/spoke topology).
### Bug fix: enum index bounds
The v0.1.0 engine had a dead constraint on the bytes path — the
built-in `enum` keyword checked string membership, but the materializer
emitted `Value::Number(index)`, which never matched. The BAST-native
validator checks the materialized index against `values.len()` bounds,
fixing this. Net improvement, recorded as intended behavior in the
test suite.
## What is removed
Under the BAST pivot, the v0.1.0 custom-keyword machinery is removed:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation). See [ADR-VAL-SPLIT](val-split-two-validator-model.md).
- `normalize_refs()` / `inline_union_variant_refs()` — BAST refs are
always `#/$defs/<name>`; one hash lookup. Variant refs resolve lazily.
- `get_alktype_kind*` family — superseded by the parser's `kind`-string
dispatch.
- `parse_encoding`/`parse_align`/`parse_max_length`/`parse_endian`/
`parse_discriminator` + `DiscriminatorKind` — replaced by the typed
`BastField`/`BastDiscriminator` views and the parser's internal
BAST-property-form copies.
- `resolve_ref`/`resolve_ref_or_inline` — replaced by
`BastDoc::lookup_def`/`resolve_typeref`.
- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`,
`BYTE_DISCRIMINATOR_TYPES` — replaced by `from_bast_str`/`to_bast_str`
and the parser's typed views.
- The custom-keyword `build_validator` path — `build_validator` is
repurposed to build a *standard* `jsonschema::Validator` from a
consumer-provided JSON Schema (no custom keywords). See
[ADR-VAL-SPLIT](val-split-two-validator-model.md).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path.
## Consequences
### Positive
- **Self-describing, cross-language format.** A BAST document carries
its type vocabulary in a meta-schema'd JSON Schema instance. Any
language with a JSON parser and a JSON Schema validator can validate
BAST document structure without knowing alktype's Rust internals.
Editors with `$schema` support provide autocomplete and inline
validation for free.
- **Simpler `$ref` story.** One restricted form (`#/$defs/<name>`), one
hash lookup, no normalization pass. TypeBox interop is a serialization
concern (TypeBox → BAST JSON), not an engine concern.
- **Explicit field order.** The `fields` array makes byte order
unambiguous — no reliance on `serde_json`'s `preserve_order` for
correctness (it remains a dependency for builder output and for
`mapping` iteration order, but layout correctness no longer depends
on it).
- **Architecture simplification.** The BAST-native validator is a flat
recursive match — no factories, no trait objects, no sub-validator
pre-computation, no `with_keyword` registration. ~200 lines of
custom-keyword validators become ~250 lines of straightforward
pattern matching.
- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed.
- **Wasm binary-size win.** The `validate_bytes` path no longer touches
`jsonschema` for validation (it still uses `jsonschema`'s
`ValidationError::custom` type for the error payload, per
D-BAST-009 — but no validator compilation, no keyword registration,
no sub-validators).
### Negative
- **Breaking change to the v0.1.0 public surface.** `compile`'s
signature changes (new `root_name` param, drops `&mut`, takes a BAST
document not a custom-keyword JSON Schema). `validate_json`/
`is_valid_json` change contract (validate against a consumer-provided
JSON Schema, not the alktype schema). `Schema::build`/
`Definitions::build` output format changes. The ~13 `schema::*`
helper re-exports are removed. `build_validator` is repurposed. The
crate is on crates.io at 0.1.0 with zero real consumers, so the bump
is free — but the contract is explicit (see the implementation plan's
Semver Contract table).
- **Two output formats from the builder.** `struct_()` → BAST,
`object()` → standard JSON Schema. The construction API is the same;
only the serialization differs. This is deliberate (D-BAST-008) but
is a thing consumers must learn.
- **`AlkTypeKind::Display` is backed by `to_bast_str`.** The v0.1.0
`as_str`/`Display` rendered the `"AlkType:Uint8"` keyword; the new
`Display` renders the BAST canonical string (`"uint8"`). Error
messages across six modules surface the new name. This is the right
name to surface now, but it is a visible change in error output.
## Scope Boundaries (What This Is Not)
- **Not a replacement for JSON Schema for JSON validation.** A BAST
document cannot validate a JSON payload — it describes binary data
layouts and value-domain constraints for bytes. For JSON validation,
consumers use standard JSON Schema documents (which may be derived
from BAST via future codegen, or authored separately). See
[ADR-VAL-SPLIT](val-split-two-validator-model.md).
- **Not a code generator.** BAST is a data format, not a Rust source
generator. ADR-001's scope boundary stands.
- **Not a schema-evolution / Value system.** TypeBox's `Value.Diff`,
`Value.Migrate`, `Value.Convert` remain out of scope (ADR-001).
- **Not a framing format.** BAST describes one struct/union/enum
instance; it does not strip length prefixes or handle multi-frame
buffers. Framing stays in the consumer (ADR-010).
## Decisions (D-BAST-001..009)
The BAST format is grounded in decisions D-BAST-001 through D-BAST-009,
recorded in
[the pivot research record](../../research/bast-pivot.md#decisions).
Summary:
| Decision | Summary |
|----------|---------|
| [D-BAST-001](../../research/bast-pivot.md#d-bast-001-root-type-selection) | Root type name is a required `compile()` parameter — explicit, not convention |
| [D-BAST-002](../../research/bast-pivot.md#d-bast-002-primitive-type-string-set) | Lowercase kind strings (`"uint32"`); `AlkTypeKind` variants stay PascalCase |
| [D-BAST-003](../../research/bast-pivot.md#d-bast-003-top-level-defs-requirement) | `$defs` is always required; every document has the same top-level shape |
| [D-BAST-004](../../research/bast-pivot.md#d-bast-004-arrays-of-variable-length-elements-deferred) | Arrays require `count` in v1; variable-length-element arrays deferred (OQ-001) |
| [D-BAST-005](../../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions) | Field-name discriminator unions supported; optional `fields` array on `UnionDef` |
| [D-BAST-006](../../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model) | `validate_bytes` uses the BAST-native validator — no external JSON Schema needed |
| [D-BAST-007](../../research/bast-pivot.md#d-bast-007-validate_json-validation-model) | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema |
| [D-BAST-008](../../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats) | One builder, two build methods: `struct_()` → BAST, `object()` → standard JSON Schema |
| [D-BAST-009](../../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape) | Keep `Validation(jsonschema::ValidationError<'static>)` — uniform payload for both paths |
## References
- [`bast-format.md`](../bast-format.md) — the normative BAST format
specification
- [`schema-layer.md`](../schema-layer.md) — the BAST parser
implementation
- [ADR-VAL-SPLIT](val-split-two-validator-model.md) — the two-validator
model (BAST-native for bytes, standard `jsonschema` for JSON)
- [ADR-001](001-alktype-purpose-scope-jsonschema-engine.md) — purpose,
scope, and the "schema is the format" principle (format-specific
content superseded by this ADR; purpose/scope retained)
- [ADR-003](003-schema-annotations.md) — annotation semantics (carry
forward unchanged; only location moves)
- [ADR-009](009-builder-api.md) — builder API (output format amended
to BAST / standard JSON Schema)
- [ADR-010](010-generalized-validation-validate-bytes.md) —
`validate_bytes` (validation step amended to the BAST-native
validator)
- [BAST pivot research record](../../research/bast-pivot.md) —
motivation, POC scope and result, decisions D-BAST-001..009, risks
- [BAST pivot implementation plan](../../plans/bast-implementation.md)
— ordered steps, semver contract, ADR-sync checklist
@@ -0,0 +1,228 @@
# ADR-VAL-SPLIT: Two-Validator Model — BAST-Native for Bytes, Standard jsonschema for JSON
## Status
Accepted — refines the "validation strategy" section of
[ADR-004](004-error-handling-validation-strategy.md) and the "validation
step" of [ADR-010](010-generalized-validation-validate-bytes.md) for
the BAST pivot. Records decisions D-BAST-006, D-BAST-007, and
D-BAST-009.
## Context
alktype v0.1.0 used a single validation mechanism — the `jsonschema`
crate with 19 custom keyword validators — for both the JSON path
(`validate_json(&Value)`) and the bytes path (`validate_bytes(&[u8])`).
The bytes path materialized a `serde_json::Value` tree from the buffer,
then ran the same `jsonschema::Validator` against it.
Under the BAST pivot ([ADR-BAST](bast-bast-format.md)), the format
changed from custom-keyword JSON Schema to BAST, and the custom-keyword
integration was removed. This forced a re-evaluation of both validation
paths:
1. **The bytes path.** BAST is the complete specification of the binary
format — it describes both the layout (how to read) and the
constraints (what values are valid). An external JSON Schema is not
needed for `validate_bytes`; the BAST document *is* the validation
spec for bytes. The natural validator is a recursive walker over the
BAST type tree that checks the value-domain constraints the
materializer does not (integer ranges, `maxLength`,
enum index bounds, union variant constraints). The POC
(`bast-validator-poc` branch, `src/bast_poc.rs`) proved this out
end-to-end with 20 reference tests.
2. **The JSON path.** BAST describes bytes, not JSON shape. A JSON
`Value` (e.g., an incoming JSON-RPC request) is the wrong input for
a BAST document; the right validator is a standard
`jsonschema::Validator` built from a standard JSON Schema document
the consumer provides. BAST is not involved on this path. This is
the path alkcall uses for its `OperationSpec` JSON validation.
The two paths have different inputs (bytes vs JSON `Value`), different
schema sources (the BAST document vs a consumer-provided JSON Schema),
and different validators (a flat recursive match vs a compiled
`jsonschema::Validator`). But they share the same error variant —
`AlkTypeError::Validation` — so consumers handling both (alkcall uses
`validate_json` for channel 0 JSON-RPC and `validate_bytes` for binary
channels) match one arm.
## Decision
**alktype has two validators for two input types:**
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
### `validate_bytes` — BAST-native validator (D-BAST-006)
`src/bast_validation.rs` is a recursive walker
(`validate_value(doc, &value)`) over the BAST typed tree
([`crate::bast::BastDoc`]/[`BastType`]). The materializer
(`src/materialize.rs`) produces a structurally-correct `Value` tree
from bytes (all declared fields present, types correct, bounds checked,
UTF-8 valid, discriminator in mapping, boolean byte 0 or 1). The
validator enforces only the **value-domain constraints expressed in the
BAST document** — the ones the materializer can't see from the bytes
alone:
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` |
| Float finiteness (Float32/64) | `validate_float` |
| String `maxLength` (byte length) | `check_string` |
| Bytes `maxLength` (array length) | `check_bytes` (accepts `Value::String` and `Value::Array`) |
| Enum index bounds | `validate_enum` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, resolves the variant, recurses |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
The validator is a flat `match` — no factories, no trait objects, no
sub-validator pre-computation, no `jsonschema` involvement. ~250 lines
replace ~200 lines of v0.1.0 custom-keyword factories.
No external JSON Schema is required. The BAST document is the complete
specification of the binary format. An optional external JSON Schema
can be layered on top for constraints BAST doesn't express (cross-field
consistency, regex patterns on string content) — additive, not
load-bearing.
### `validate_json` — standard jsonschema (D-BAST-007)
`validate_json(&Value)` / `is_valid_json(&Value)` validate a JSON
`Value` against a standard `jsonschema::Validator` compiled at
`AlkTypeEngine::compile` time from a consumer-provided JSON Schema
(`compile`'s `json_schema: Option<&Value>` parameter). No custom
keywords, no BAST involvement. The JSON Schema is independent of the
BAST document — BAST describes bytes, not JSON shape. It may be
authored separately or derived from BAST via future codegen.
If no JSON Schema was supplied to `compile`, `validate_json` returns
`AlkTypeError::Schema` and `is_valid_json` returns `false`.
`build_validator` (in `src/validation.rs`) is **repurposed**: it builds
a *standard* `jsonschema::Validator` from a plain JSON Schema (no
custom keywords). The engine calls it internally during `compile` when
`json_schema` is `Some`. Consumers that only need a one-off validator
may call `jsonschema::options().build(schema)` directly; `build_validator`
exists so the engine's error mapping (`jsonschema` build error →
`AlkTypeError::Schema`) is reused. The v0.1.0 custom-keyword
`build_validator` is removed.
The `jsonschema` crate remains a direct dependency for this path and
for validating BAST documents against the BAST meta-schema.
### Error payload (D-BAST-009)
`AlkTypeError::Validation(jsonschema::ValidationError<'static>)` is
**retained** as the error variant for both paths. The bytes path no
longer uses `jsonschema` for validation, so its error payload is
constructed via `jsonschema::ValidationError::custom` purely to keep
the variant's type unchanged. The rationale is consumer ergonomics on
the *combined* path: consumers like alkcall use both `validate_json`
and `validate_bytes` and handle `AlkTypeError::Validation` in one
place. A single uniform payload type means one match arm covers both
sources.
The alternative (`Validation(String)`) was rejected — it would force
`validate_json` to flatten its structured errors (instance path, schema
path, keyword) to a `String` via `Display`. The more information-rich
path would lose data to accommodate the less rich one. That is the
wrong direction.
The `no_std`/minimal-build angle (OQ-002) that the alternative was
meant to enable is moot: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. The right place to
revisit is when/if OQ-002 is actually pursued.
## Consequences
### Positive
- **Right validator for each input.** Bytes are validated by the BAST
document that describes them; JSON values are validated by a JSON
Schema that describes them. No forced isomorphism between two
different input types.
- **No external JSON Schema needed for `validate_bytes`.** The BAST
document is both the layout spec and the validation spec for bytes.
This is the "schema is the format" principle from ADR-001, now fully
realized.
- **Enum index bounds enforced.** The v0.1.0 dead constraint is fixed
— the BAST-native validator checks the materialized index against
`values.len()` directly.
- **Per-variant constraint enforcement (OQ-008) without custom
keywords.** The validator recurses into the selected variant's BAST
definition on `__discriminator` lookup, enforcing every field
constraint the variant declares (e.g., `maxLength` on a `bytes` field
inside a variant struct).
- **Wasm binary-size win.** The `validate_bytes` path no longer touches
`jsonschema` for validation (it still uses `ValidationError::custom`
for the error payload type, per D-BAST-009 — but no validator
compilation, no keyword registration, no sub-validators).
- **Architecture simplification.** ~200 lines of custom-keyword
factories become ~250 lines of straightforward pattern matching. No
`with_keyword` registration; no `inline_union_variant_refs` compile
step.
- **Uniform error payload.** Consumers handle one
`AlkTypeError::Validation` match arm for both paths (D-BAST-009).
### Negative
- **Two validators, not one.** The engine struct carries an
`Option<jsonschema::Validator>` (for `validate_json`) and re-parses
the BAST typed tree on each `validate_bytes` call (the BAST-native
validator is not pre-built — it's a recursive walker over the
on-demand `BastDoc`). This is a small cost; the validators serve
different inputs and don't share structure.
- **`validate_json` requires a consumer-provided JSON Schema.** The
engine no longer builds a validator from the alktype schema; the
consumer must supply a JSON Schema at `compile` time (or accept that
`validate_json` returns `AlkTypeError::Schema`). This is a behavioral
break from v0.1.0, intentional under the pivot.
- **`AlkTypeError::Validation` payload is `jsonschema`'s type even on
the bytes path.** The bytes path constructs it via
`ValidationError::custom`, which is slightly awkward but keeps the
variant uniform. The `no_std` revisit (OQ-002) is the place to
reconsider if a bytes-only minimal build ever materializes.
## Scope Boundaries (What This Is Not)
- **Not a `Validator` trait abstraction.** Two methods on one struct,
not a trait with impls for JSON-only and BAST-binary schemas. The two
impls share little internally (`validate_json` is a single
`jsonschema` call; `validate_bytes` is materialize + BAST-native
walk), so a trait would add a layer without unifying behavior. See
[ADR-010](010-generalized-validation-validate-bytes.md) §"Not a
`Validator` trait abstraction".
- **Not a binary-aware validator that skips the `Value` tree.** The
`Value`-materialization path is the validation path. A future
"validate bytes without materializing" path is a two-way door but
explicitly out of scope for v1 (would re-introduce a hand-rolled
validator, ADR-001).
- **Not framing-aware.** `validate_bytes` validates the bytes of *one*
schema instance. Framing stays in the consumer (ADR-010).
## References
- [`bast-format.md` §Validation Model](../bast-format.md#validation-model)
— the normative validation model
- [ADR-BAST](bast-bast-format.md) — the BAST format decision
- [ADR-004](004-error-handling-validation-strategy.md) — error handling
and validation strategy (load-time build, access-time check,
`AlkTypeError` enum — retained; validation-strategy section refined
by this ADR)
- [ADR-010](010-generalized-validation-validate-bytes.md) —
`validate_bytes` (the two-step concept retained; the validation step
amended to the BAST-native validator by this ADR)
- [`validation.md`](../validation.md) — the validation layer
documentation
- `src/bast_validation.rs` — the BAST-native validator implementation
- `src/validation.rs` — the `build_validator` helper
- [BAST pivot research record](../../research/bast-pivot.md) —
D-BAST-006, D-BAST-007, D-BAST-009
+62 -27
View File
@@ -7,8 +7,8 @@ last_updated: 2026-07-22
The layout engine: offset computation, the two layout modes (packed
sequential vs aligned static), alignment, endianness, and variable-length
field handling. This is the novel code — the recursive walk of the schema
JSON that computes byte positions for each field.
field handling. This is the novel code — the recursive walk of the BAST
typed tree that computes byte positions for each field.
## The Two Layout Modes
@@ -24,8 +24,8 @@ protocols.
**Components:**
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(schema)` (requires `AlkType:Struct` at the top level), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `SequentialReader::new(schema)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(bast_doc, root_name)` (requires a `struct` at the root), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `engine.sequential_reader()` (shares the engine's compiled `ReadPlan` via `Arc` — see [ADR-011](decisions/011-compiled-read-plan-for-packed-mode.md)), then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
**How it works:**
@@ -69,7 +69,7 @@ and safetensors.
**Component:**
- **`OffsetMap`** — constructed via `OffsetMap::compute(schema) -> Result<Self, AlkTypeError>` (requires `AlkType:Struct` at the top level). Walks the schema once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result<Self, AlkTypeError>` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment, and resolves each leaf's `LeafMeta` (kind, encoding, effective endianness — ADR-012 §2b). The output is a flat table of `(field_path, OffsetEntry)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
**How it works:**
@@ -113,19 +113,24 @@ with a `AlkTypeError::Offset` — the `OffsetMap` reserves only 4 bytes
(the length prefix), but `data_access::write_string` writes prefix +
data inline, which would clobber subsequent fields. Non-final variable
fields must use `maxLength` (fixed-size reservation) or
`"encoding": "offset-indirect"`. See
`"encoding": "offset-indirect"` — except `record` fields, for which
neither remedy is available (`maxLength` is rejected at parse — review
#006 N3 — and `offset-indirect` is rejected for records in aligned
mode — review #006 M5), so a non-final record field cannot be repaired
and must move to the last position. See
[ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md).
## Offset Computation Algorithm
The offset computation is a recursive walk of the schema JSON. The
The offset computation is a recursive walk of the BAST typed tree
([`BastDoc`](schema-layer.md#the-bast-parser-bast-module)). The
algorithm is the same for both modes; the difference is whether alignment
padding is inserted between fields.
### Fixed-size types
For each fixed-size type, the algorithm:
1. Determines the type's byte size from the `AlkType:*` kind.
1. Determines the type's byte size from the `AlkTypeKind`.
2. In aligned mode: inserts padding to satisfy the type's alignment
(or the field's `align` annotation, or the struct's `align` default).
3. Records the field's `(start, end)` range.
@@ -133,14 +138,14 @@ For each fixed-size type, the algorithm:
### Composite types
**`TStruct`:** Recurse into the struct's `properties`. The inner fields
**`struct`:** Recurse into the struct's `fields` array. The inner fields
are computed relative to the struct's start offset. The struct's total
size is the sum of its fields' sizes (plus alignment padding in aligned
mode). The struct itself may have an `align` annotation that rounds up
its total size.
**`TUnion`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `TUnion` fields with
**`union`:** TUnion is supported in packed sequential mode only. In
aligned static mode, `OffsetMap::compute` rejects `union` fields with
`AlkTypeError::Offset` — see
[ADR-008](decisions/008-reject-tunion-in-aligned-mode.md). Unions
are the protocol dispatch pattern (SFTP type bytes, call protocol event
@@ -161,12 +166,12 @@ total size. The `SequentialReader` reads the discriminator first, looks
up the variant schema, then reads the variant struct sequentially — it
doesn't need to know the union's total size upfront.
**`TArray` of fixed-size elements:** Element stride = element size (plus
**`array` of fixed-size elements:** Element stride = element size (plus
alignment padding in aligned mode). Element `i` starts at
`array_offset + i × stride`. The array's total size is `count × stride`.
**`TArray` of variable-length-element structs:** Deferred for v1
(OQ-001).
**`array` of variable-length-element structs:** Deferred for v1
(OQ-001, D-BAST-004 — BAST arrays require `count` in v1).
### Variable-length types
@@ -193,6 +198,12 @@ annotation shapes).
only. The engine uses strategy 1 (inline length-prefixing) because
protocols don't benefit from fixed-size reservation.
`maxLength` applies to `string` and `bytes` fields only. The parser
rejects it on any other kind (review #006 N3): the validation plan
bakes it into string/bytes leaves only, so on a record (or any other
kind) the annotation did nothing — and in aligned mode a record
reservation was silently corrupt (review #006 M5).
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
1. The field is a struct `{offset: u32, length: u32}`.
2. The `OffsetMap` records the position of this struct.
@@ -289,7 +300,7 @@ pub struct FieldPosition {
A field's computed position in a packed layout, produced by
`LayoutBuilder::build`. For variable-length fields, `size` is `4` (the
length prefix); for fixed-size fields, `size` is the type's byte size.
`kind` records the field's `AlkType:*` kind so the consumer can dispatch
`kind` records the field's `AlkTypeKind` so the consumer can dispatch
to the correct `data_access` read/write function.
### `PackedLayout` (packed mode)
@@ -311,22 +322,46 @@ discriminators, the discriminator is recorded under the synthetic path
(schema `properties` order, with nested struct fields appearing inline
under their parent's path prefix).
### `OffsetMap` (aligned mode)
A flat table of `(field_path, byte_range)` pairs computed from a schema.
### `LeafMeta` / `OffsetEntry` (aligned mode, 0.3.0)
```rust
impl OffsetMap {
pub fn compute(schema: &Value) -> Result<Self, AlkTypeError>;
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
pub struct LeafMeta {
pub kind: AlkTypeKind,
pub encoding: VariableEncoding,
pub endian: Endian,
}
pub struct OffsetEntry {
pub range: ByteRange,
pub meta: LeafMeta,
}
```
`compute` requires a `AlkType:Struct` at the top level. `total_size`
includes trailing alignment padding. `iter` yields fields in insertion
order (schema `properties` order, nested struct fields appearing inline).
`OffsetMap::compute` resolves each leaf's read/write metadata (kind,
variable-length encoding, effective endianness — field override else
container default, propagated the aligned-materializer way) alongside
its byte range, so `read_field`/`write_field` dispatch on the entry
without re-walking the BAST tree per access (ADR-012 §2b).
### `OffsetMap` (aligned mode)
A flat table of `(field_path, OffsetEntry)` pairs computed from a schema.
```rust
impl OffsetMap {
pub fn compute(doc: &BastDoc) -> Result<Self, AlkTypeError>;
pub fn get(&self, field_path: &str) -> Option<&OffsetEntry>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = (&str, &OffsetEntry)>;
pub fn fingerprint(&self) -> u64;
}
```
`compute` requires a `struct` at the root. `total_size`
includes trailing alignment padding. `iter` yields fields in the BAST
`fields` array order (nested struct fields appearing inline). The map
carries `Hash + Eq` (ADR-012 §1); `fingerprint()` is the
stable-within-version hash for caching and schema handshakes.
## Design Decisions
@@ -355,7 +390,7 @@ See [open-questions.md](open-questions.md) for full details.
the two layout modes decision
- [ADR-003](decisions/003-schema-annotations.md) — schema
annotations
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds and their
- [schema-layer.md](schema-layer.md) — the 18 AlkType kinds and their
byte sizes
- [data-access.md](data-access.md) — read/write functions that use the
computed offsets
+19 -15
View File
@@ -106,31 +106,35 @@ architect's desk" is answerable at a glance.
### OQ-006: Builder spec Example 3 — wrap `Union` in a `Struct` — RESOLVED
- **Status**: resolved. [builder.md](builder.md) Example 3 now wraps
the `Union` in a `Schema::struct_().field("payload", ...)` and merges
`$defs` via `Definitions::merge_into`. Matches the realistic SFTP
wire shape and the engine's `AlkType:Struct`-at-root constraint.
the `Union` in a `Schema::struct_().field("payload", ...)` and builds
a complete BAST document via `Definitions::build_doc`. Matches the
realistic SFTP wire shape and the engine's struct-at-root constraint
(the root `$defs` entry must be a `struct`).
- **Full file**: [OQ-006](questions/006-builder-spec-example-3-wrap-union.md)
### OQ-007: `Bytes` materialization — lossy UTF-8 conversion — RESOLVED
- **Status**: resolved. Array of u8: the materializer produces
`Value::Array` of `Value::Number` (one entry per byte, 0..=255) for
`AlkType:Bytes` fields. The `BytesValidator` accepts both
`Value::String` (for `validate_json`) and `Value::Array` (for
`validate_bytes`). `maxLength` = max byte count. Implemented in
`src/materialize.rs` and `src/validation.rs`.
`bytes` fields. The BAST-native validator (`bast_validation::check_bytes`)
accepts both `Value::String` (for `validate_json`-style inputs) and
`Value::Array` (for `validate_bytes`). `maxLength` = max byte count.
Implemented in `src/materialize.rs` and `src/bast_validation.rs`.
- **Full file**: [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)
### OQ-008: `UnionValidator` variant dispatch — RESOLVED
- **Status**: resolved. `UnionValidator` now builds a sub-validator for
each variant at factory time and dispatches on `__discriminator` at
validation time. `AlkTypeEngine::compile` calls
`schema::inline_union_variant_refs` before `build_validator` to inline
`$ref`s in union `mapping` entries (necessary because the
`union_factory` receives the union node, but `$defs` live at the
schema root). Implemented in `src/validation.rs`, `src/schema.rs`,
and `src/engine.rs`.
- **Status**: resolved. Under the BAST pivot, the BAST-native validator
(`bast_validation::validate_union`) reads `__discriminator`, looks up
the variant `BastType` in the union's `mapping`, and recurses into the
variant's BAST definition via `validate_typeref` — enforcing every
field constraint the variant declares (e.g. `maxLength` on a `bytes`
field inside a variant struct). Variant `$ref`s resolve lazily via
`BastDoc::resolve_typeref` — no `inline_union_variant_refs` compile
step (removed under BAST). No custom keywords, no `jsonschema`
involvement on the bytes path. Implemented in `src/bast_validation.rs`
and `src/bast.rs`. See
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
- **Full file**: [OQ-008](questions/008-unionvalidator-variant-dispatch.md)
## Deferred / Blocked
+114 -77
View File
@@ -1,12 +1,12 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-15
---
# alktype — Overview
The binary struct engine: a small Rust crate that takes a JSON Schema
with `AlkType:*` custom keywords and produces an offset map, read/write
The binary struct engine: a small Rust crate that takes a BAST (Binary
Abstract Syntax Tree) document and produces an offset map, read/write
functions, and validation — all driven by the schema. The schema is the
format definition; the engine is generic.
@@ -16,29 +16,43 @@ Component details are in the sibling documents.
## What
`alktype` is a library crate that consumes JSON Schemas annotated
with `AlkType:*` custom keywords (the same kinds defined in TypeBox's
`typedef.ts`, plus `AlkType:Bytes`, `AlkType:Int64`, and `AlkType:Uint64`
as alktype additions) and produces three capabilities:
`alktype` is a library crate that consumes BAST documents and produces
three capabilities:
1. **An offset map** — walks the schema, computes byte offsets for each
field based on type sizes, field order, and alignment.
1. **An offset map** — walks the BAST typed tree, computes byte offsets
for each field based on type sizes, field order, and alignment.
2. **Read/write functions** — given a `&[u8]` buffer and a field path,
read the field's bytes at its offset (zero-copy for fixed-size types).
Given a `&mut [u8]` buffer, write a value at its offset.
3. **Validation** — via `jsonschema` custom keywords, validates that a
buffer's bytes match the schema's type constraints.
3. **Validation** — two validators for two input types:
- `validate_bytes(&[u8])` uses the BAST-native validator (a recursive
walker over the BAST type tree) to check the value-domain
constraints the materializer doesn't (integer ranges, `maxLength`,
enum index bounds, union variant constraints).
- `validate_json(&Value)` uses a standard `jsonschema::Validator`
compiled from a consumer-provided JSON Schema (BAST is not involved
— BAST describes bytes, not JSON shape).
The heavy lifting is done by the `jsonschema` crate (validation) and
`serde_json` (schema parsing). The novel code is the offset computation
— a recursive walk of the schema JSON that computes byte positions for
each field. The custom keyword implementations are small (a few lines
each, generated from shared macros — see [validation.md](validation.md)).
BAST is a JSON document that describes binary data layouts using a
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
itself a valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser. See [ADR-BAST](decisions/bast-bast-format.md)
and [`bast-format.md`](bast-format.md).
The heavy lifting is done by the `jsonschema` crate (the
`validate_json` path and BAST document meta-schema validation) and
`serde_json` (BAST document parsing). The novel code is the offset
computation — a recursive walk of the BAST typed tree that computes
byte positions for each field — and the BAST-native validator — a flat
recursive match over the same tree. See
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md)
(purpose/scope) and [ADR-BAST](decisions/bast-bast-format.md) (format).
The crate replaces two prior attempts that built their own jsonschema
engines — typebox-rs (~8,400 lines) and the @alkdev/alktype prototype
(~5,600 lines) — with `jsonschema` + an offset map + small custom keyword
implementations. See
(~5,600 lines) — with a BAST parser + an offset map + a BAST-native
validator + the `jsonschema` crate for the JSON-validation path. See
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
## Why
@@ -48,21 +62,20 @@ read or write binary data at computed offsets. Instead of per-protocol
serde structs (russh-sftp's 29 packet types), per-handler wire format
code (TTY's 5-byte format parser), or per-format offset computation
(metatensor's tensor access), all of these become instances of the same
engine with different schemas.
engine with different BAST documents.
The guiding insight:
> **The schema is the format.** A JSON Schema with `AlkType:Float32`,
> `AlkType:Struct`, `AlkType:Union` etc. is both the validation spec and
> the layout spec. No separate format definition, no separate parser, no
> separate validator. One schema, three uses: validate, compute offsets,
> access data.
> **The schema is the format.** A BAST document is both the layout spec
> and the validation spec for bytes. No separate format definition, no
> separate parser, no separate validator. One schema, three uses:
> validate, compute offsets, access data.
This is the convergence of three threads identified in the
call-channels-unification research: the `typedef.ts` schema kinds from
TypeBox, the russh-sftp protocol packets, and the metatensor format. The
common pattern: a JSON Schema describes the shape of binary data, and
the binary data is the struct's bytes at computed offsets.
common pattern: a schema describes the shape of binary data, and the
binary data is the struct's bytes at computed offsets.
The crate was bumped up in the timeline when the call-channels-unification
research surfaced that channels, TTY, and the binary call protocol are
@@ -75,41 +88,45 @@ read/write the binary payload."
## The "Schema Is the Format" Principle
A JSON Schema with `AlkType:*` custom keywords serves three roles
simultaneously:
A BAST document serves three roles simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec** | `jsonschema` custom keywords | Load time (build validator), access time (validate buffer) |
| **Validation spec (bytes)** | BAST-native validator (recursive walker over the BAST type tree) | Load time (parse typed tree), access time (`validate_bytes`) |
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
No separate format definition, no separate parser, no separate validator.
The schema is the single source of truth for the binary format. Adding a
new field to a protocol is adding a property to the schema JSON — the
engine computes the new offsets automatically.
No separate format definition, no separate parser, no separate
validator. The BAST document is the single source of truth for the
binary format. Adding a new field to a protocol is adding an entry to
the BAST `fields` array — the engine computes the new offsets
automatically.
This is the same principle as `#[repr(C)]` struct field access, but at
runtime from a portable JSON Schema instead of at compile-time from
language-specific annotations. The schema is the ABI contract.
runtime from a portable JSON document instead of at compile-time from
language-specific annotations. The BAST document is the ABI contract.
## Dependencies
```
alktype
├── jsonschema (v0.46.5, Draft 2020-12) — validation engine, custom keyword support
├── serde_json (with preserve_order) — schema parsing; field order is load-bearing
└── (no tokio, no platform deps) — WASM-clean by construction
├── jsonschema (v0.46, Draft 2020-12, default-features=false) — validate_json path + BAST meta-schema validation
├── serde_json (with preserve_order) — BAST document parsing; mapping iteration order is load-bearing
└── (no tokio, no platform deps) — WASM-clean by construction
```
`alktype` is dependency-light: `jsonschema` + `serde_json` only.
No tokio, no platform deps. Compiles to `wasm32-unknown-unknown` for
browser use. The `jsonschema` crate is already in the workspace at
`@alkdev/alknet: jsonschema/` — alktype is its first consumer.
browser use. The `validate_bytes` path does not touch `jsonschema` for
validation (it uses `jsonschema::ValidationError::custom` only for the
error payload type, D-BAST-009) — a small wasm binary-size win.
`serde_json` requires the `preserve_order` feature because field order
is load-bearing for binary layouts. The order of properties in the
schema JSON determines the order of fields in the binary struct.
`serde_json`'s `preserve_order` feature remains a dependency. Under
BAST, struct field order is explicit (the `fields` array), so layout
correctness no longer depends on it; but `mapping` iteration order and
the `Definitions` `$defs` block order are still load-bearing for the
parser's lazy resolution and the builder's output.
## Consumers
@@ -133,14 +150,17 @@ schema roles (binary layout + JSON payloads) from one library. See
The russh-sftp case is the most instructive and the highest-value POC
target. The `Packet` enum's `TryFrom<&mut Bytes>` impl is a hand-written
dispatch on a type byte followed by serde deserialization. Under alktype,
the dispatch is `TUnion` with a byte-offset discriminator — the schema
says "byte 0 is the discriminator, bytes 1..N are the variant struct."
The engine reads the discriminator, looks up the variant schema, computes
offsets, reads fields. Same result, no per-packet-type code.
the dispatch is a BAST `union` with a byte-offset discriminator — the
schema says "byte 0 is the discriminator, bytes 1..N are the variant
struct." The engine reads the discriminator, looks up the variant
schema, computes offsets, reads fields. Same result, no per-packet-type
code.
## Scope Boundaries (What This Is Not)
These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md).
These boundaries are decided in
[ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md)
and [ADR-BAST](decisions/bast-bast-format.md).
- **Not metatensor.** alktype is the binary struct *engine*. Metatensor
is a *format* (8-byte header + JSON header + binary data) that uses the
@@ -150,60 +170,73 @@ These boundaries are decided in [ADR-001](decisions/001-alktype-purpose-scope-js
should not do anything that explicitly blocks adding a Value system
later.
- **Not a code generator.** typebox-rs's `codegen/` module is a separate
concern. The alktype engine consumes schemas; it does not generate them.
concern. The alktype engine consumes BAST documents; it does not
generate them.
- **Schema builder is in scope as of v0.1.0.** A fluent Rust API for
constructing schemas at runtime, producing `serde_json::Value`, is
shipped in v0.1.0 ([ADR-009](decisions/009-builder-api.md), resolves
OQ-003). The builder covers AlkType kinds and standard JSON Schema;
see [builder.md](builder.md). Schemas may still be authored in
TypeBox, generated by ujsx components, or hand-written — the builder
is an additional construction path, not a replacement.
constructing BAST documents and standard JSON Schemas at runtime,
producing `serde_json::Value`, is shipped in v0.1.0
([ADR-009](decisions/009-builder-api.md), resolves OQ-003). The
builder covers BAST kinds (`struct_()`) and standard JSON Schema
(`object()`); see [builder.md](builder.md). BAST documents may still
be authored in TypeBox, generated by ujsx components, or hand-written
— the builder is an additional construction path, not a replacement.
- **Not a serialization framework.** The alktype engine is not a
general-purpose serde replacement. It operates on raw byte buffers at
computed offsets — no intermediate `Value` tree, no reflection, no
dynamic dispatch per field. For JSON data, use serde. For binary data
with a known schema, use alktype.
computed offsets — no intermediate `Value` tree (except for the
`validate_bytes` materialization step), no reflection, no dynamic
dispatch per field. For JSON data, use serde. For binary data with a
known schema, use alktype.
- **Not a JSON-payload validator.** BAST describes bytes, not JSON
shape. `validate_json` validates a JSON `Value` against a
consumer-provided standard JSON Schema, not against the BAST document.
See [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
## Architecture (component pointers)
- **[schema-layer.md](schema-layer.md)** — the 19 `AlkType:*` kinds,
jsonschema custom keyword integration, TypeBox interop, schema
annotations (endianness, alignment, encoding, TUnion discriminators).
- **[schema-layer.md](schema-layer.md)** — the BAST parser (the typed
surface every engine module walks), the 18 BAST kinds, the
`AlkTypeKind` enum, and the foundational annotation types.
- **[`bast-format.md`](bast-format.md)** — the normative BAST format
specification (meta-schema, TypeRef, examples, validation model).
- **[layout-engine.md](layout-engine.md)** — offset computation, the two
layout modes (packed sequential vs aligned static), alignment,
endianness, variable-length field handling.
- **[data-access.md](data-access.md)** — read/write functions, TUnion
dispatch, field paths, zero-copy access for fixed-size types,
length-prefix reading for variable-length types.
- **[validation.md](validation.md)** — custom keyword validators for all
19 `AlkType:*` kinds, `AlkTypeError`, load-time vs access-time
validation, `AlkTypeEngine` as the compiled form of a schema.
`validate_json` for JSON values; `validate_bytes` for binary buffers
(ADR-010).
- **[builder.md](builder.md)** — fluent Rust API for constructing
alktype JSON Schemas at runtime, producing `serde_json::Value`.
Covers AlkType kinds and standard JSON Schema (ADR-009).
- **[validation.md](validation.md)** — the two-validator model
(BAST-native for `validate_bytes`, standard `jsonschema` for
`validate_json`), `AlkTypeError`, load-time vs access-time validation,
`AlkTypeEngine` as the compiled form of a BAST document.
- **[builder.md](builder.md)** — fluent Rust API for constructing BAST
documents and standard JSON Schemas at runtime, producing
`serde_json::Value`. Covers BAST kinds and standard JSON Schema
(ADR-009, D-BAST-008).
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries |
| Purpose, scope, and the jsonschema engine | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | What the crate is/isn't; why jsonschema not a custom engine; "schema is the format" principle; scope boundaries (format-specific content superseded by ADR-BAST) |
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | BAST as the schema format; meta-schema, `$defs`/`$ref`/`kind` vocabulary; supersedes ADR-001's format-specific content |
| Two layout modes | [ADR-002](decisions/002-two-layout-modes-packed-vs-aligned.md) | Packed sequential (`LayoutBuilder`/`SequentialReader`) for protocols; aligned static (`OffsetMap`) for mmap formats |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) |
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Endianness (schema-level, default LE), alignment (struct + field-level), encoding (length-prefixed vs offset-indirect), TUnion discriminators (byte-offset vs field-name) — semantics carry forward; location moved to BAST type-level properties |
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping (validation-strategy section refined by ADR-VAL-SPLIT) |
| Two-validator model | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | BAST-native validator for `validate_bytes`; standard `jsonschema::Validator` for `validate_json`; D-BAST-006/007/009 |
| Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (SFTP offsets, metatensor data_offsets) |
| Non-final inline variable fields | [ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) | Rejected in aligned mode (would clobber subsequent fields) |
| Packed-mode read factory | [ADR-007](decisions/007-packed-mode-read-factory.md) | `engine.sequential_reader()` returns an owned fresh reader |
| TUnion in aligned mode | [ADR-008](decisions/008-reject-tunion-in-aligned-mode.md) | Rejected for v1 (broken semantics; no current consumer needs it) |
| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers AlkType kinds + standard JSON Schema; resolves OQ-003 |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate |
| Builder API | [ADR-009](decisions/009-builder-api.md) | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003 (output format amended to BAST / standard JSON Schema by ADR-BAST) |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation on `AlkTypeEngine`; materialize `Value` from bytes, then validate (validation step amended to the BAST-native validator by ADR-VAL-SPLIT) |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-001** (deferred(scope)): Arrays of variable-length-element structs.
- **OQ-001** (deferred(scope)): Arrays of variable-length-element
structs — BAST arrays require `count` in v1 (D-BAST-004), aligning
with this deferral.
- **OQ-002** (deferred(scope)): `no_std` + `alloc` support.
- **OQ-003** (resolved by [ADR-009](decisions/009-builder-api.md)):
Builder API for schema construction. Shipped in v0.1.0; see
@@ -223,8 +256,12 @@ See [open-questions.md](open-questions.md) for full details.
- `@alkdev/alknet: alknet-typedef-poc/` — the POC code (disposable)
- `@alkdev/alknet: typebox-rs/` — prior attempt, replaced by alktype
- `@alkdev/alknet: alktype-prototype/` — prior attempt (the @alkdev/alktype prototype; not to be confused with this crate, which reuses the name but is backed by the `jsonschema` crate)
- [BAST pivot research record](../research/bast-pivot.md) — motivation,
POC scope and result, decisions D-BAST-001..009, risks
- [BAST pivot implementation plan](../plans/bast-implementation.md) —
ordered implementation steps, semver contract, ADR-sync checklist
> **Note**: The research findings, POC code, and prior-attempt paths above
> refer to the parent `@alkdev/alknet` workspace where this crate originated.
> They are preserved here as historical context for the architectural
> decisions; the artifacts themselves are not part of this standalone repo.
> decisions; the artifacts themselves are not part of this standalone repo.
@@ -1,5 +1,19 @@
# OQ-008: `UnionValidator` variant dispatch — validate variant fields against variant schema
> **Note (post-BAST-pivot):** The v0.1.0 resolution below —
> `UnionValidator` + `inline_union_variant_refs` + custom-keyword
> `jsonschema` sub-validators — was superseded by the BAST pivot. The
> current implementation is the BAST-native validator
> (`bast_validation::validate_union`), which reads `__discriminator`,
> looks up the variant `BastType` in the union's `mapping`, and
> recurses via `validate_typeref` into the variant's BAST definition.
> Variant `$ref`s resolve lazily via `BastDoc::resolve_typeref` — no
> `inline_union_variant_refs` compile step (removed under BAST). No
> custom keywords, no `jsonschema` involvement on the bytes path. See
> [ADR-VAL-SPLIT](../decisions/val-split-two-validator-model.md). The
> v0.1.0 resolution text is preserved below as the historical record
> of how the question was originally resolved.
- **Origin**: Raised during the v0.1.0 POC round 2 (SFTP Packet
`validate_bytes` POC). Surfaced when the over-`maxLength` `Bytes`
test failed: the materializer read the bytes correctly, but the
+214 -425
View File
@@ -1,58 +1,60 @@
---
status: draft
last_updated: 2026-07-22
status: accepted
last_updated: 2026-08-15
---
# alktype — Schema Layer
The schema layer: the 19 `AlkType:*` custom type kinds, their mapping to
Rust types and byte sizes, the `jsonschema` custom keyword integration,
TypeBox interop, and the concrete JSON shapes for schema-level
annotations.
The schema layer: the BAST (Binary Abstract Syntax Tree) format and the
typed parser that the layout engines, materializer, and BAST-native
validator walk. BAST replaces the v0.1.0 `AlkType:*` custom-keyword JSON
Schema format decided in ADR-001; the pivot is recorded in
[ADR-BAST](decisions/bast-bast-format.md) and grounded in
[D-BAST-001..009](../research/bast-pivot.md#decisions).
## The 19 AlkType Kinds
The **normative format specification** is
[`bast-format.md`](bast-format.md) (meta-schema, TypeRef, examples,
validation model). This document describes the *implementation* — the
typed parser in `src/bast.rs` and the foundational `AlkTypeKind` enum
in `src/schema.rs` — and points at the format spec for shape details.
These are the custom schema kinds defined in TypeBox's `typedef.ts`
(`@alkdev/alknet: typebox/example/typedef/typedef.ts`, 619 lines) and
ported to Rust via `jsonschema` custom keywords. Each kind carries binary
layout semantics — a known byte size (for fixed-size types) or a known
encoding strategy (for variable-length types).
## The 18 BAST Kinds
| Kind | TypeBox key | Rust type | Size | Category |
|------|-------------|-----------|------|----------|
| `TFloat32` | `AlkType:Float32` | `f32` | 4 | fixed |
| `TFloat64` | `AlkType:Float64` | `f64` | 8 | fixed |
| `TInt8` | `AlkType:Int8` | `i8` | 1 | fixed |
| `TInt16` | `AlkType:Int16` | `i16` | 2 | fixed |
| `TInt32` | `AlkType:Int32` | `i32` | 4 | fixed |
| `TInt64` | `AlkType:Int64` | `i64` | 8 | fixed |
| `TUint8` | `AlkType:Uint8` | `u8` | 1 | fixed |
| `TUint16` | `AlkType:Uint16` | `u16` | 2 | fixed |
| `TUint32` | `AlkType:Uint32` | `u32` | 4 | fixed |
| `TUint64` | `AlkType:Uint64` | `u64` | 8 | fixed |
| `TBoolean` | `AlkType:Boolean` | `bool` (0x00=false, 0x01=true) | 1 | fixed |
| `TString` | `AlkType:String` | length-prefixed UTF-8 | variable | variable |
| `TBytes` | `AlkType:Bytes` | length-prefixed raw bytes | variable | variable |
| `TStruct` | `AlkType:Struct` | record of fields | sum of field sizes | composite |
| `TUnion` | `AlkType:Union` | tagged union | discriminator + variant | composite |
| `TArray` | `AlkType:Array` | repeated element | count × element size | composite |
| `TEnum` | `AlkType:Enum` | u32 index into enum values | 4 (fixed) | fixed |
| `TRecord` | `AlkType:Record` | count-prefixed sequence of (key, value) pairs | variable | variable |
| `TTimestamp` | `AlkType:Timestamp` | length-prefixed RFC 3339 string | variable | variable |
BAST uses lowercase `kind` strings (`"uint32"`, `"struct"`, `"union"`,
etc.). The engine represents them as the `AlkTypeKind` Rust enum — one
variant per kind — providing compile-time exhaustiveness checking and
integer-discriminant dispatch (a jump table) instead of string
comparison at every field access.
`AlkType:Int64` and `AlkType:Uint64` are alktype additions —
TypeBox's `typedef.ts` tops out at 32-bit integers. They are required by
the primary POC targets: SFTP `Read`/`Write` packets have `offset: u64`,
and metatensor `data_offsets` are `u64`. See
| BAST kind | `AlkTypeKind` | Rust type | Size | Category |
|-----------|---------------|-----------|------|----------|
| `int8` | `Int8` | `i8` | 1 | fixed |
| `int16` | `Int16` | `i16` | 2 | fixed |
| `int32` | `Int32` | `i32` | 4 | fixed |
| `int64` | `Int64` | `i64` | 8 | fixed |
| `uint8` | `Uint8` | `u8` | 1 | fixed |
| `uint16` | `Uint16` | `u16` | 2 | fixed |
| `uint32` | `Uint32` | `u32` | 4 | fixed |
| `uint64` | `Uint64` | `u64` | 8 | fixed |
| `float32` | `Float32` | `f32` | 4 | fixed |
| `float64` | `Float64` | `f64` | 8 | fixed |
| `bool` | `Boolean` | `bool` (`0x00`=false, `0x01`=true) | 1 | fixed |
| `string` | `String` | length-prefixed UTF-8 | variable | variable |
| `bytes` | `Bytes` | length-prefixed raw bytes | variable | variable |
| `struct` | `Struct` | record of fields | sum of field sizes | composite |
| `union` | `Union` | tagged union | discriminator + variant | composite |
| `array` | `Array` | repeated element | count × element size | composite |
| `enum` | `Enum` | u32 index into enum values | 4 (fixed) | fixed |
| `record` | `Record` | count-prefixed (key, value) pairs | variable | variable |
`int64`/`uint64` are alktype additions — TypeBox's `typedef.ts` tops
out at 32-bit integers. Required by SFTP `Read`/`Write` `offset: u64`
and metatensor `data_offsets`. See
[ADR-005](decisions/005-int64-uint64-first-class-kinds.md).
### The `AlkTypeKind` enum
The engine represents the 19 kinds as a Rust enum — `AlkTypeKind` — with
one variant per kind (`AlkTypeKind::Float32`, `AlkTypeKind::Struct`, etc.).
The enum provides compile-time exhaustiveness checking and integer
discriminant dispatch (a jump table) instead of string comparison at
every field access. It is `pub` and re-exported from the crate root.
`src/schema.rs` defines the enum:
```rust
pub enum AlkTypeKind {
@@ -60,7 +62,7 @@ pub enum AlkTypeKind {
Uint8, Uint16, Uint32, Uint64,
Float32, Float64,
Boolean, Enum,
String, Bytes, Timestamp,
String, Bytes,
Struct, Union, Array, Record,
}
```
@@ -69,439 +71,226 @@ The enum carries the kind's binary-layout metadata as inherent methods:
| Method | Returns | Notes |
|--------|---------|-------|
| `as_str(self)` | `&'static str` | The JSON Schema keyword, e.g. `"AlkType:Uint8"` |
| `to_bast_str(self)` | `&'static str` | The lowercase BAST kind string (`"uint32"`) |
| `from_bast_str(s)` | `Result<AlkTypeKind, AlkTypeError>` | Parses a lowercase BAST kind string; `AlkTypeError::Schema` for unknowns |
| `type_size(self)` | `Option<usize>` | `Some(N)` for fixed-size kinds; `None` for variable/composite |
| `natural_alignment(self)` | `usize` | 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix), 1 for struct/union/array |
| `is_fixed_size(self)` | `bool` | True for the 12 fixed-size primitive kinds |
| `is_composite(self)` | `bool` | True for Struct, Union, Array, Record |
| `is_variable_length(self)` | `bool` | True for String, Bytes, Timestamp, Record |
| `is_variable_length(self)` | `bool` | True for String, Bytes, Record |
| `needs_endian(self)` | `bool` | True for kinds whose read/write takes an `Endian` parameter |
`AlkTypeKind` implements `Display` (renders the keyword string) and
`FromStr` (parses the keyword string back into the variant, returning
`AlkTypeError::Schema` for unknown kinds). The layout engines and the
validator dispatch on the enum, not on strings.
`AlkTypeKind` implements `Display`, backed by `to_bast_str` so the
layout engines, materializer, validator, and parser surface the
canonical BAST name in error messages. `from_bast_str` is the inverse
and is the dispatch point the BAST parser uses to map a `kind` string
to the enum variant (D-BAST-002).
### Fixed-size types
### Foundational annotation types
`TFloat32`, `TFloat64`, `TInt8`, `TInt16`, `TInt32`, `TUint8`, `TUint16`,
`TUint32`, `TBoolean`, and `TEnum` have known byte sizes. The offset
computation uses these sizes directly. Read/write is zero-copy pointer
cast for these types.
**`TBoolean` byte representation:** `0x00` = false, `0x01` = true. Other
values are invalid and produce a `AlkTypeError::Access` on read.
**`TEnum` binary representation:** A `u32` index into the enum's declared
values, in declaration order. The first declared value is index 0, the
second is index 1, etc. The enum's values are declared via the standard
JSON Schema `"enum"` keyword (e.g., `"enum": ["read", "write", "execute"]`).
The `AlkType:Enum` custom keyword signals that the type is an enum for
layout purposes; the built-in `enum` keyword provides the value list.
**Design note:** TypeBox's `TEnum` is a string enum (variable-length). The
alktype engine uses a `u32` index instead — a deliberate deviation from
TypeBox fidelity in favor of binary efficiency. Most enums have a small
number of variants (e.g., the call protocol's 5 event types); a `u32`
index is compact, fixed-size, and sufficient for any realistic enum. The
JSON representation (for validation) remains a string; the binary
representation is the `u32` index.
The `u32` index follows the schema's endianness annotation (ADR-003), like
all other fixed-size types. In little-endian mode the index is
`u32::from_le_bytes`; in big-endian mode it is `u32::from_be_bytes`.
### Variable-length types
`TString`, `TBytes`, `TRecord`, and `TTimestamp` have variable byte sizes.
The alktype engine supports three strategies for handling variable-length
types in binary layouts, selected by the `encoding` annotation and the
standard JSON Schema `maxLength` keyword:
| Strategy | Encoding annotation | Layout behavior | Use case |
|----------|-------------------|-----------------|----------|
| **Inline length-prefixed** | `"length-prefixed"` (default) | `[length: u32][data]`; shifts subsequent fields in packed mode | Protocol wire formats (SFTP, channels, TTY) |
| **Fixed-size reservation** | (none — uses `maxLength`) | `[data: maxLength bytes]`, zero-padded; fixed offset in aligned mode | mmap-friendly formats where max size is known (database `VARCHAR(N)` pattern) |
| **Offset indirection** | `"offset-indirect"` | `{offset: u32, length: u32}` pointing into a separate data region | Blob tensors, metatensor variable-length data (the blob tensor pattern) |
**Strategy 1: Inline length-prefixing (default).** The field's fixed
portion is a 4-byte length prefix at a computed offset. The variable data
follows immediately after. In packed sequential mode, the length prefix
determines the position of subsequent fields. In aligned static mode, the
length prefix is at a known offset; the variable data is not included in
the static layout. This is the universal pattern used by channels, SFTP,
TTY, and most binary protocols.
**Strategy 2: Fixed-size reservation.** When a variable-length field
declares `maxLength` (a standard JSON Schema keyword), the engine reserves
`maxLength` bytes at a fixed offset in aligned static mode. Data shorter
than `maxLength` is zero-padded; data longer than `maxLength` is a
validation error. This makes the field fixed-size from the layout
perspective — subsequent fields have known, unchanging offsets. This is
the database `VARCHAR(N)` pattern and the metatensor struct-tensor
pattern for fields with known maximum sizes.
In packed sequential mode, `maxLength` is a validation constraint only —
the engine still uses inline length-prefixing (strategy 1) because
protocols don't benefit from fixed-size reservation.
**Strategy 3: Offset indirection.** The field is a struct
`{offset: u32, length: u32}` at a known position. The consumer provides
the data region separately; the engine reads the offset and length, then
slices the data region. This is the metatensor blob tensor pattern — the
index struct lives in one region, the blob data lives in another. Enables
mmap-friendly random access to variable-length data without parsing
length prefixes and without reserving worst-case space.
**Default strategy selection:**
- In packed sequential mode: always strategy 1 (inline length-prefixing).
`maxLength` is a validation constraint only.
- In aligned static mode: strategy 2 (fixed-size reservation) if
`maxLength` is declared; strategy 3 (offset indirection) if
`"encoding": "offset-indirect"` is declared; strategy 1 (inline
length-prefixing) otherwise.
**Length prefix endianness:** The 4-byte length prefix (strategies 1 and 3)
respects the schema's `"endian"` annotation (ADR-003). In little-endian
mode, the length is `u32::from_le_bytes`. In big-endian mode, the length
is `u32::from_be_bytes`. This ensures SFTP consumers (big-endian) have
consistent byte order for both field values and length prefixes.
**`TBytes`:** Raw bytes — no UTF-8 constraint. The payload is `&[u8]`.
Otherwise identical to `TString` in layout (same three strategies).
**Design note:** `AlkType:Bytes` is an alktype addition — it does
not exist in TypeBox's `typedef.ts` (which defines 16 kinds). It is
included because raw byte arrays are a common binary protocol primitive
(SFTP data payloads, channels payloads, tensor data) and are semantically
distinct from UTF-8 strings. In the binary representation, TBytes is raw
bytes with no encoding (not base64, not hex). In the JSON representation
(for validation), TBytes is a string (JSON has no native byte type).
**`TRecord`:** A string-keyed map. The value type is declared via the
schema's `"values"` property (e.g., `"values": { "AlkType:Float32": true }`).
Binary layout is a count-prefixed sequence of `(key, value)` pairs:
`[count: u32][key_len: u32][key_bytes][value]...` repeated `count` times.
The count is the number of entries. Each key is a length-prefixed UTF-8
string. Each value is encoded according to its declared `AlkType:*` kind
— a `Record<Uint32>` value is 4 raw bytes; a `Record<String>` value is
itself a length-prefixed string; a `Record<Struct>` value is the struct's
fields laid out inline. There is **no separate `value_len` prefix** —
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix). The
count and key-length prefixes respect the schema's endianness. In
aligned static mode with `maxLength`, the entire record is reserved at
`maxLength` bytes (zero-padded).
**`TTimestamp`:** An RFC 3339 timestamp string (the internet profile of
ISO 8601). Stored as a length-prefixed UTF-8 string (strategy 1) or
fixed-size reservation (strategy 2 with `maxLength`). The data-access
layer treats timestamps as opaque length-prefixed strings — it does not
parse or validate the timestamp format. The jsonschema custom keyword
validator checks RFC 3339 conformance at the JSON level (see
[validation.md](validation.md)).
`TArray` is variable-length when the element type is variable-length or
when the count is not known at schema time. For fixed-size element arrays
with a known count, the size is `element_size × count`.
**`TArray` count declaration:** The array count is declared via the
standard JSON Schema `"minItems"` and `"maxItems"` keywords. When
`minItems == maxItems`, the array has a fixed count known at schema time.
When they differ or are absent, the count is variable and the array uses
a length-prefixed encoding: `[count: u32][element_0]...[element_N]`.
The count prefix respects the schema's endianness.
### Composite types
`TStruct` and `TUnion` are composite — their size is the sum of their
fields' sizes (plus alignment padding in aligned static mode). The offset
computation recurses into their properties.
## Schema-Layer Public API
The `schema` module exposes the foundational types and functions every
other module depends on. These are re-exported from the crate root.
### `get_alktype_kind` vs `get_alktype_kind_loose`
The engine recognizes a `AlkType:*` kind on a schema node two ways,
because the keyword value may be either a boolean (`true`) or an
annotation object (`{ "encoding": "..." }`):
| Function | Recognizes | Returns |
|----------|------------|---------|
| `get_alktype_kind(node) -> Option<&str>` | Boolean form only (`{ "AlkType:String": true }`) | The keyword string, e.g. `"AlkType:String"` |
| `get_alktype_kind_loose(node) -> Option<&str>` | Boolean form **and** object form | The keyword string |
| `get_alktype_kind_enum(node) -> Option<AlkTypeKind>` | Boolean form only | The parsed enum variant |
| `get_alktype_kind_loose_enum(node) -> Option<AlkTypeKind>` | Boolean form **and** object form | The parsed enum variant |
The boolean-form-only functions are used by the validator factories
(which reject the object form as a schema error) and the top-level
kind-check in `OffsetMap::compute` / `LayoutBuilder::new` / `SequentialReader::new`
(which require `AlkType:Struct` at the root). The "loose" variants are
used by the layout engines during field traversal, so that a variable-
length field with an `encoding` annotation (`{ "AlkType:String":
{ "encoding": "offset-indirect" } }`) is still recognized as a `String`.
### Annotation parsers
Each schema-level annotation has a dedicated parser that reads it from a
`serde_json::Value` node and returns a sensible default when absent:
| Function | Annotation | Default |
|----------|------------|---------|
| `parse_endian(node) -> Endian` | `"endian"` | `Endian::Little` |
| `parse_align(node) -> Option<usize>` | `"align"` | `None` |
| `parse_max_length(node) -> Option<usize>` | `"maxLength"` | `None` |
| `parse_encoding(keyword_value) -> VariableEncoding` | `"encoding"` (within the keyword's value object) | `VariableEncoding::LengthPrefixed` |
| `parse_discriminator(node) -> Result<DiscriminatorKind, AlkTypeError>` | `"discriminator"` | (required — returns `AlkTypeError::Schema` if absent) |
### Public enums
`src/schema.rs` also defines the two annotation enums (semantics
unchanged from ADR-003; only their *location* in the document moved —
see [ADR-BAST](decisions/bast-bast-format.md) and
[bast-format.md §Variable-Length Encoding](bast-format.md#variable-length-encoding)):
```rust
pub enum Endian { Little, Big }
pub enum VariableEncoding { LengthPrefixed, OffsetIndirect }
pub enum DiscriminatorKind {
Byte { offset: usize, disc_type: AlkTypeKind },
Field { name: String },
}
```
`DiscriminatorKind::Byte` carries the byte position (`offset`) and the
discriminator's `AlkType:*` kind (`disc_type`, restricted to `Uint8`/
`Uint16`/`Uint32`). `DiscriminatorKind::Field` carries the discriminator
field's name. See [data-access.md](data-access.md) §"TUnion Dispatch" for
how these drive dispatch.
The `Discriminator` builder enum lives in
[`src/builder.rs`](../../src/builder.rs) (the builder's domain); the BAST
parser's typed discriminator view is
[`BastDiscriminator`](#the-bast-parser-bast-module).
### `$ref` resolution and normalization
## The BAST Parser (`bast` module)
| Function | Purpose |
|----------|---------|
| `normalize_refs(schema: &mut Value)` | Walks the schema; rewrites every `"$ref"` whose value is a bare name (no `#` prefix) to `"#/$defs/<name>"`. Idempotent. Runs once at `AlkTypeEngine::compile` time. |
| `resolve_ref(root, ref_path) -> Option<&Value>` | Resolves a JSON Pointer `$ref` (e.g. `"#/$defs/Read"`) against the root schema. |
| `resolve_ref_or_inline(node, root) -> Option<&Value>` | If `node` has a `"$ref"`, resolves it against `root`; otherwise returns `node` itself (it's an inline schema). |
`src/bast.rs` is the typed surface over a BAST document. Three
consumers walk the same tree — the layout engines
([`offset_map`](layout-engine.md), [`layout_builder`](layout-engine.md),
[`sequential_reader`](layout-engine.md)), the
[`materialize`](data-access.md) layer, and the
[`bast_validation`](validation.md) validator — so a typed view pays for
itself: each walks matched arms over `BastType` instead of re-parsing
raw JSON at every node. Borrowing (not cloning) the source
`serde_json::Value` keeps the parse allocation-free beyond the small
typed nodes themselves.
`normalize_refs` bridges TypeBox's bare-name ref output and `jsonschema`'s
JSON Pointer requirement. The layout engines call `resolve_ref_or_inline`
on every `$ref`-bearing node they encounter during traversal.
### Document shape
## jsonschema Custom Keyword Integration
Every BAST document has the same top-level shape:
The `jsonschema` crate (v0.46.5, Draft 2020-12) supports custom keywords
via the `with_keyword` API. Each `AlkType:*` kind is registered as a
custom keyword:
```rust
let validator = jsonschema::options()
.with_keyword("AlkType:Float32", factory)
.with_keyword("AlkType:Int32", factory)
.with_keyword("AlkType:Struct", factory)
// ... all 19 kinds
.build(&schema)?;
```json
{ "$defs": { "<TypeName>": { ...TypeDef... }, ... } }
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path — enabling cross-keyword awareness. The
`AlkType:Struct` validator, for example, inspects the parent's
`properties` to validate each field against its declared `AlkType:*` kind.
- The `$defs` block is **required** (D-BAST-003). Single-type documents
are a special case with one entry.
- The **root type name** is a required parameter to
`AlkTypeEngine::compile(bast_doc, root_name, mode, ...)` (D-BAST-001).
It selects which `$defs` entry is the top-level type; convention
(first entry) is fragile and depends on JSON key order, so an
explicit parameter is used instead.
Each custom keyword implementation is ~10 lines. The `jsonschema` crate
handles all structural validation (object properties, required fields,
array items, enum values) — the custom keywords only need to validate
the leaf type constraints. See [validation.md](validation.md) for the
validator implementations.
See [`bast-format.md`](bast-format.md) for the normative TypeDef shapes
(Struct, Union, Enum, FieldDef, TypeRef) and the meta-schema.
This is the same pattern as TypeBox's `TypeRegistry.Set` on the JS side.
Same semantics, different language, same JSON Schema wire format. A
TypeBox schema serialized to JSON feeds into the alktype engine after a
single pre-processing step: normalizing `$ref` values (see below).
### Typed tree
## TypeBox Interop
The parser produces a borrowed typed tree:
TypeBox modules render to standard JSON Schema under `$defs`. A TypeBox
schema like:
| Type | Role |
|------|------|
| `BastDoc<'a>` | The parsed document: the root `Value`, the chosen root name, and the parsed root `BastDef`. Entry point via `BastDoc::new(root, root_name)`. |
| `BastDef<'a>` | A named `$defs` entry — `{ name, kind: BastDefKind, source }`. Only `struct`/`union`/`enum` can live at the top level. |
| `BastDefKind<'a>` | `Struct(BastStruct)` / `Union(BastUnion)` / `Enum(BastEnum)`. |
| `BastStruct<'a>` | `{ endian, align, fields: Vec<BastField>, source }`. Field order is the `fields` array order (BAST design principle #4 — no reliance on `serde_json`'s `preserve_order`). |
| `BastField<'a>` | `{ name, ty: BastType, endian, align, encoding, max_length, source }`. Annotations are field-level properties (ADR-003 semantics, BAST location). |
| `BastUnion<'a>` | `{ endian, discriminator, fields, mapping: Vec<(key, BastType)>, source }`. Variant `$ref`s resolve **lazily** — no compile-time inlining. |
| `BastDiscriminator<'a>` | `Byte { offset, disc_type }` / `Field { name }`. The typed view of the `discriminator` object. |
| `BastEnum<'a>` | `{ values: Vec<&'a str>, source }`. Non-empty (enforced). |
| `BastType<'a>` | A TypeRef — `Primitive(AlkTypeKind)` / `Ref(BastRef)` / `Array(BastArray)` / `Record(BastRecord)` / `Struct(...)` / `Union(...)` / `Enum(...)`. The central mechanism for typing fields, array elements, record values, and union variants. |
| `BastRef<'a>` | A `$ref` restricted to `#/$defs/<name>`. Carries just the name. |
| `BastArray<'a>` | `{ element: Box<BastType>, count, source }`. `count` is required in v1 (D-BAST-004). |
| `BastRecord<'a>` | `{ values: Box<BastType>, source }`. |
```typescript
const TensorRef = Type.Object({
dtype: Type.Union([Type.Literal("F32"), Type.Literal("I16")]),
shape: Type.Array(Type.Number()),
data_offsets: Type.Tuple([Type.Number(), Type.Number()])
});
```
All of these are re-exported from the crate root (`pub use bast::{...}`
in `src/lib.rs`).
serialized to JSON is a standard JSON Schema with `type: "object"`,
`properties`, and `required`. That JSON feeds into the alktype engine
after `$ref` normalization. The `AlkType:*` custom keywords are added by
TypeBox's `TypeRegistry.Set` — they appear in the serialized JSON as
additional properties on the schema object.
### `$ref` resolution
### `$ref` normalization
BAST `$ref`s are always full JSON Pointers restricted to
`#/$defs/<name>` — no external references, no fragment-only pointers,
no bare names (rejected by the parser). The restriction keeps
resolution a single hash lookup and eliminates the v0.1.0
`normalize_refs` pass that rewrote TypeBox's bare-name refs.
TypeBox generates bare-name `$ref` values (e.g., `"$ref": "Read"`),
referencing sibling definitions within the same `$defs` block. The
`jsonschema` crate requires full JSON Pointer paths (e.g.,
`"$ref": "#/$defs/Read"`). The alktype engine normalizes TypeBox-style
refs at schema load time via [`normalize_refs`](#ref-resolution-and-normalization)
— a ~20-line recursive walk that rewrites every bare-name `"$ref"` to
`"#/$defs/<name>"`. The normalization is idempotent — full JSON Pointer
refs pass through unchanged. It runs once at `AlkTypeEngine::compile`
time, before the schema is passed to `jsonschema` or the offset
computation.
`BastDoc` exposes three resolution helpers:
**Verification:** The jsonschema crate (v0.46.5) rejects bare-name refs
with `Resource 'Read' is not present in a registry`. Full JSON Pointer
refs (`#/$defs/Read`) resolve correctly. The normalization step bridges
the gap between TypeBox's output and jsonschema's input.
| Method | Purpose |
|--------|---------|
| `lookup_def(name) -> Result<&'a Value, AlkTypeError>` | The single hash lookup into `$defs`. |
| `resolve_ref(r: &BastRef) -> Result<BastDef, AlkTypeError>` | Resolve a `BastRef` to its `BastDef`. |
| `resolve_typeref(ty: &BastType) -> Result<BastType, AlkTypeError>` | Deref one `$ref` level, or return the inline type unchanged. The composite-walkers call this. |
| `resolve_typeref_as_def(ty, path) -> Result<BastDef, AlkTypeError>` | Resolve a `BastType` to a `BastDef`, wrapping inline composites in a synthetic def. Convenient for the validator/materializer. |
The alktype engine does not depend on TypeBox or any JS toolchain. It
consumes JSON — whether that JSON was authored in TypeBox, generated by
a ujsx component, or hand-written. The schema is the interface.
Variant `$ref`s (union `mapping` entries) are resolved **lazily** by
the materializer and validator via these helpers — no
`inline_union_variant_refs` compile step (removed under BAST). The
parser only records the `BastRef` target name.
### Untrusted input
Every path that walks a BAST document returns
`Err(AlkTypeError::Schema)` on a malformed document, never
`panic!`/`unreachable!`/`unwrap` (AGENTS.md §3 — the downstream
`alkcall` consumer accepts schemas from arbitrary internet peers in its
hub/spoke topology). Overflow-safe arithmetic (`checked_add`,
`usize::try_from`) is used for any offset/count cast (AGENTS.md §4).
A malformed document (missing `$defs`, missing `kind`, unknown kind
string, dangling `$ref`, empty `mapping`, non-struct/union/enum at the
top level, a field-name union without a `fields` array, etc.) surfaces
as `AlkTypeError::Schema` with a path-annotated message.
### What the parser does *not* do
- **No meta-schema validation.** `BastDoc::new` parses structurally
(every reachable def parses to a typed `BastDef`) but does not run
the BAST meta-schema. Consumers that want full structural validation
can run the meta-schema via `jsonschema` directly
([`BAST_META_SCHEMA`](bast-format.md#the-meta-schema) is re-exported
from the crate root). The parser's structural checks catch the cases
that matter for layout/materialize/validate; the meta-schema is the
authoritative well-formedness check.
- **No eager full-document parse.** Only the root definition and the
definitions it (transitively) references are parsed eagerly; orphan
`$defs` entries are not checked. Lazy `$ref` resolution reaches the
rest at access time.
- **No annotation interpretation.** The parser *records* `endian`/
`align`/`encoding`/`maxLength` on `BastField`/`BastStruct`; the
layout engines and validator *interpret* them (ADR-003 semantics).
## Schema Annotations
Schema-level annotations control binary layout behavior. These are
decided in [ADR-003](decisions/003-schema-annotations.md).
Annotation *semantics* carry forward unchanged from ADR-003; only their
*location* moved from v0.1.0's custom-keyword objects to BAST
type-level properties. The concrete BAST shapes are in
[`bast-format.md`](bast-format.md):
### Endianness
- [Endianness](bast-format.md#endianness) — struct/union-level `endian`
with field-level override.
- [Alignment](bast-format.md#alignment) — struct/field-level `align`
(aligned mode only).
- [Variable-length encoding](bast-format.md#variable-length-encoding) —
field-level `encoding` and `maxLength`.
- [Union discriminators](bast-format.md#union) — `discriminator` object
on the union def (`byte` or `field`).
Schema-level annotation with a default of little-endian:
The `maxLength` keyword is *not* a BAST invention — it is the standard
JSON Schema `maxLength`, repurposed as a byte-length cap. In aligned
mode it reserves a fixed-size slot; in packed mode it is a validation
constraint only. It applies to `string`/`bytes` fields only: the parser
rejects it on any other kind (review #006 N3 — elsewhere it was
silently unenforced), and in aligned mode a record reservation was
silently corrupt (review #006 M5). See [bast-format.md §Variable-Length
Encoding](bast-format.md#variable-length-encoding) and
[ADR-003](decisions/003-schema-annotations.md).
```json
{ "AlkType:Struct": true, "endian": "big", "properties": { ... } }
```
## What Was Removed
- `"endian": "little"` (default) — read/write in little-endian byte order.
- `"endian": "big"` — read/write in big-endian byte order.
- Applies to the entire schema and all nested types.
The v0.1.0 custom-keyword accessor layer was removed in step 8 of the
BAST pivot. The `schema` module retains only the foundational types
(`AlkTypeKind`, `Endian`, `VariableEncoding`, shared constants); the
BAST parser is the typed surface every engine module walks. Removed:
### Alignment
- `get_alktype_kind` / `get_alktype_kind_enum` / `get_alktype_kind_loose`
/ `get_alktype_kind_loose_enum` — superseded by the parser's
`kind`-string dispatch.
- `normalize_refs` / `inline_union_variant_refs` (+ recursive helpers)
— BAST refs are always `#/$defs/<name>`; one hash lookup, variant
refs resolve lazily.
- `parse_encoding` / `parse_align` / `parse_max_length` / `parse_endian`
— `bast.rs` has its own BAST-property-form copies (internal to the
parser).
- `parse_discriminator` + `DiscriminatorKind` — replaced by
`bast::BastDiscriminator`; the builder has its own `Discriminator`
enum.
- `resolve_ref` / `resolve_ref_or_inline` — replaced by
`BastDoc::lookup_def` / `resolve_typeref`.
- `FromStr` impl, `as_str`, `Endian::from_schema`, `ALKTYPE_PREFIX`,
`BYTE_DISCRIMINATOR_TYPES`, and their unit tests.
Both struct-level and field-level, with field-level overriding:
```json
{
"AlkType:Struct": true,
"align": 256,
"properties": {
"weight": { "AlkType:Float32": true, "align": 16 }
}
}
```
- Struct-level `"align"` sets the default for all fields.
- Field-level `"align"` overrides the struct default.
- Default alignment: 1 for u8/i8/bool, 2 for u16/i16, 4 for u32/i32/f32/
enum, 8 for u64/i64/f64, 4 for variable-length (the u32 length prefix),
1 for struct/union/array.
- Only meaningful in aligned static mode (ADR-002). Ignored in packed
sequential mode.
### Variable-length encoding
The alktype engine supports three strategies for variable-length types
(see §Variable-length types above for full details). The strategy is
selected by the `encoding` annotation and the standard JSON Schema
`maxLength` keyword:
```json
// Strategy 1: Inline length-prefixing (default, shorthand)
{ "AlkType:String": true }
// Strategy 1: Explicit inline length-prefixing
{ "AlkType:String": { "encoding": "length-prefixed" } }
// Strategy 2: Fixed-size reservation (uses standard maxLength)
{ "AlkType:String": true, "maxLength": 256 }
// Strategy 3: Offset indirection (opt-in)
{ "AlkType:String": { "encoding": "offset-indirect" } }
```
- `"encoding": "length-prefixed"` (default) — 4-byte length prefix at
computed offset, variable data follows immediately. Used by protocol
wire formats.
- `maxLength` (standard JSON Schema keyword) — in aligned static mode,
reserves `maxLength` bytes at a fixed offset (zero-padded). Makes the
field fixed-size from the layout perspective. In packed sequential
mode, `maxLength` is a validation constraint only.
- `"encoding": "offset-indirect"` — field is a struct
`{offset: u32, length: u32}` pointing into a separate data region.
The consumer provides the data region separately. Used by metatensor
blob tensors.
- Applies to all variable-length types: `AlkType:String`, `AlkType:Bytes`,
`AlkType:Array`, `AlkType:Record`, `AlkType:Timestamp`.
### TUnion discriminators
Two discriminator kinds: byte-offset (protocol dispatch) and field-name
(typedef.ts pattern).
**Byte-offset discriminator** (SFTP type bytes, call protocol event types):
```json
{
"AlkType:Union": true,
"discriminator": {
"kind": "byte",
"offset": 0,
"type": "AlkType:Uint8"
},
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" },
"101": { "$ref": "#/$defs/Status" }
}
}
```
- `"offset"` — byte position of the discriminator.
- `"type"` — the `AlkType:*` kind of the discriminator (typically
`AlkType:Uint8`).
- Mapping keys are stringified integers. The variant struct starts at
`offset + discriminator_size`.
**Field-name discriminator** (typedef.ts pattern):
```json
{
"AlkType:Union": true,
"discriminator": {
"kind": "field",
"name": "type"
},
"mapping": {
"read": { "$ref": "#/$defs/Read" },
"write": { "$ref": "#/$defs/Write" }
}
}
```
- `"name"` — the field name holding the discriminator value.
- Mapping keys are string values matching the discriminator field's value.
- The discriminator field is just another field in the struct.
Mapping values may be either inline schemas or `$ref` pointers. Both work.
See [ADR-BAST](decisions/bast-bast-format.md) §"What is removed" and
[`bast-format.md` §What is removed](bast-format.md#what-is-removed).
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Concrete JSON shapes for endianness, alignment, encoding, and TUnion discriminators |
| BAST format, meta-schema, `$defs`/`$ref`/`kind` vocabulary | [ADR-BAST](decisions/bast-bast-format.md) | Supersedes ADR-001's format-specific content; records D-BAST-001..009 |
| Schema annotations | [ADR-003](decisions/003-schema-annotations.md) | Annotation semantics (carry forward unchanged; only location moves) |
| Int64/Uint64 kinds | [ADR-005](decisions/005-int64-uint64-first-class-kinds.md) | 64-bit integers as first-class kinds (required by SFTP offsets and metatensor data_offsets) |
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle |
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine; "schema is the format" principle (format-specific content superseded by ADR-BAST) |
## Open Questions
See [open-questions.md](open-questions.md) for full details.
- **OQ-003** (deferred(scope)): Builder API for schema construction.
- **OQ-001** (deferred(scope)): Arrays of variable-length-element
structs — BAST arrays require `count` in v1 (D-BAST-004), aligning
with this deferral.
## References
- `@alkdev/alknet: typebox/example/typedef/typedef.ts` — the TypeBox
schema kinds (619 lines)
- `@alkdev/alknet: jsonschema/` — the jsonschema crate (v0.46.5, Draft
2020-12)
- [ADR-003](decisions/003-schema-annotations.md) — schema
annotation shapes
- [validation.md](validation.md) — custom keyword validator implementations
- [`bast-format.md`](bast-format.md) — the normative BAST format
specification (meta-schema, TypeRef, examples, validation model)
- [ADR-BAST](decisions/bast-bast-format.md) — the BAST format decision
- [ADR-003](decisions/003-schema-annotations.md) — annotation semantics
- [BAST pivot research record](../research/bast-pivot.md) — motivation,
POC scope and result, decisions D-BAST-001..009
- [validation.md](validation.md) — the BAST-native validator and the
`validate_json` JSON-Schema path
- [`src/bast.rs`](../../src/bast.rs) — the parser implementation
- [`src/schema.rs`](../../src/schema.rs) — the `AlkTypeKind` enum and
foundational annotation types
+304 -288
View File
@@ -1,63 +1,163 @@
---
status: draft
last_updated: 2026-08-11
status: accepted
last_updated: 2026-08-31
---
# alktype — Validation
The validation layer: custom keyword validators for all 19 `AlkType:*`
kinds, the `AlkTypeError` enum, load-time vs access-time validation
strategy, and the `AlkTypeEngine` as the compiled form of a schema.
The validation layer: two validators for two input types, the
`AlkTypeError` enum, the load-time-build / access-time-check strategy,
and the `AlkTypeEngine` as the compiled form of a BAST document.
## The Validator Split
BAST separates two concerns that the v0.1.0 format conflated, and in
doing so reveals that the engine has **two distinct validation paths**
with different inputs and guarantees. This is the validator split,
decided in [D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model),
and
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape),
and recorded in [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | Compiled `ValidationPlan` walk | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
### `validate_bytes` — bytes in, BAST is the validator
The materializer produces a `serde_json::Value` tree from bytes
(walking the layout engine). By construction, this `Value` is
*structurally correct*: all declared fields are present (the
materializer iterates the field list), types are correct (`read_u32`
produces `Value::Number`), bounds are checked (via
`data_access::check_bounds`), UTF-8 is valid (via `from_utf8`), the
discriminator is in the mapping, and the boolean byte is 0 or 1.
What the materializer does NOT check — and what the validation half
checks afterward — are **value-domain constraints expressed in the BAST
document**. Since ADR-012 §3 (0.3.0), those constraints are not walked
interpretively per buffer: they are **compiled once** into a
`ValidationPlan` ([`src/validation_plan.rs`](../../src/validation_plan.rs))
at `AlkTypeEngine::compile` time, and each `validate_bytes` call walks
the compiled constraint tree against the materialized `Value` — no
`$ref` re-resolution, no schema re-parse, no per-node path formatting
(error paths render only on failure). The plan's constraint nodes
implement exactly the table below (the constraint set is unchanged from
the retired interpretive walker):
| Constraint | Plan node (`ValidNode`) |
|------------|-------------------------|
| Integer range (Int8..Uint32) | `Int { min, max }` / `Uint { max }` |
| Int64/Uint64 (full range) | `I64` / `U64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `Float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `Str { max_len }` — `maxLength` baked in from the owning field at compile time |
| Bytes `maxLength` (array length) | `Bytes { max_len }` — accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `Enum { count }` checks `idx < count` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `Union { variants }` reads `__discriminator`, dispatches on the compiled variant nodes |
| Struct fields | `Struct { fields }` requires each declared field present, recurses |
| Array count | `Array { count, element }` checks `arr.len() == count` and recurses per element |
| Record values | `Record { values }` recurses into each value |
| Boolean | `Bool` (materializer already rejects non-0/1 bytes) |
The plan is a public type (`ValidationPlan`, `Debug + Clone +
PartialEq + Eq + Hash + Send + Sync`): `engine.validation_plan()`
exposes it for consumers that validate their own materialized `Value`
trees or want its `fingerprint()` for caching / schema handshakes
(ADR-012 §1). The one-shot `bast_validation::validate_value(&doc,
&value)` remains as a convenience wrapper (compile + validate) for
callers holding a BAST document without an engine.
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
describes both the layout (how to read) and the constraints (what
values are valid). This is the "schema is the format" principle from
ADR-001, now fully realized.
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
### `validate_json` — JSON in, JSON Schema is the validator
The consumer provides a JSON `Value` (e.g., an incoming JSON-RPC
request). The BAST document is irrelevant — BAST describes bytes, not
JSON shape. The right validator for a JSON value is a standard
`jsonschema::Validator` built from a standard JSON Schema document the
consumer provides at `AlkTypeEngine::compile` time. This is the path
alkcall uses for its `OperationSpec` JSON validation. No custom
keywords; BAST is not involved.
If no JSON Schema was supplied to `compile`, the JSON-validation
methods return `AlkTypeError::Schema` (`validate_json`) or `false`
(`is_valid_json`).
### `AlkTypeError::Validation` payload shape (D-BAST-009)
The `validate_bytes` path no longer uses `jsonschema`, so its error
payload is constructed via `jsonschema::ValidationError::custom` purely
to keep the `Validation` variant's type unchanged. The rationale is
consumer ergonomics on the *combined* path: consumers like alkcall use
both `validate_json` (channel 0, JSON-RPC) and `validate_bytes` (binary
channels) and handle `AlkTypeError::Validation` in one place. A single
uniform payload type means one match arm covers both sources.
The alternative (`Validation(String)`) would force `validate_json` to
flatten its structured errors (instance path, schema path, keyword) to
a `String` via `Display` — the more information-rich path loses data to
accommodate the less rich one. That is the wrong direction.
### What is removed
Under the BAST pivot, the v0.1.0 validation machinery is removed from
the `validate_bytes` path:
- All 19 `jsonschema::Keyword` implementations (~200 lines of validator
factories) — replaced by the BAST-native validator (~250 lines, a
flat match with no factories, no trait objects, no sub-validator
pre-computation).
- `inline_union_variant_refs()` — union variant refs are resolved lazily
by the validator and materializer.
- The custom-keyword `build_validator` path — `build_validator` is
repurposed to build a *standard* `jsonschema::Validator` from a
consumer-provided JSON Schema (no custom keywords). See
[`build_validator`](#build_validator).
The `jsonschema` crate **remains a direct dependency** for
`validate_json` and for validating BAST documents against the BAST
meta-schema. The only thing removed is the custom keyword integration
path. The `validate_bytes` path no longer touches `jsonschema` — a
small wasm binary-size win in addition to the architecture
simplification.
## Validation Strategy
Validation is delegated to the `jsonschema` crate (v0.46.5, Draft
2020-12). The alktype engine does not implement its own validation —
it registers custom keyword validators for each `AlkType:*` kind and
lets `jsonschema` handle the structural validation (object properties,
required fields, array items, enum values).
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md)
and refined by [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md):
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md):
1. **Load time:** Parse the schema JSON, build the layout engine, build the
jsonschema validator. This is the `AlkTypeEngine::compile(schema)` constructor.
1. **Load time:** Parse the BAST document into the typed tree, compile
the `ValidationPlan` (the value-domain constraint tree), compute
the layout engine, and (optionally) build the standard
`jsonschema::Validator` for the JSON-validation path. This is the
`AlkTypeEngine::compile` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation.
### What validation validates
The jsonschema validator operates on `serde_json::Value` instances — it
validates JSON representations of data, not raw byte buffers. This is
the correct separation of concerns:
- **JSON validation** (jsonschema): validates that a JSON document
conforms to the schema. Used for validating hand-written schemas,
TypeBox output, JSON payloads, or the JSON representation of a binary
struct after deserialization.
- **Binary access validation** (data access layer): the read/write
functions perform type-level validation at access time — range checks
for integers, UTF-8 validity for strings, buffer bounds checking.
These return `AlkTypeError::Access` with field paths.
The "schema is the format" principle means the same schema describes
both the JSON shape and the binary layout. The jsonschema validator
checks the JSON shape; the data access layer checks the binary layout.
A consumer that wants to validate a binary buffer end-to-end reads the
buffer into a `Value` tree via the data access layer, then validates
that `Value` against the jsonschema validator. This is a two-step
process, not a single `validate(buffer)` call.
operations. Validation is opt-in per operation: the validation half
walks the compiled `ValidationPlan`, never the BAST document.
### The `AlkTypeEngine` struct
The `AlkTypeEngine` is the compiled form of a schema. It supports both
layout modes (ADR-002) via an internal `Layout` enum:
The `AlkTypeEngine` is the compiled form of a BAST document. It
supports both layout modes (ADR-002) via an internal `Layout` enum:
```rust
pub struct AlkTypeEngine {
layout: Layout, // packed or aligned (private enum)
validator: jsonschema::Validator, // compiled once at load time
endian: Endian, // parsed from the schema's "endian" annotation
schema: Value, // the normalized schema (refs resolved)
layout: Layout, // packed or aligned (private enum)
json_validator: Option<jsonschema::Validator>, // None when no JSON Schema supplied
validation_plan: Arc<ValidationPlan>, // compiled value-domain constraints (ADR-012 §3)
endian: Endian, // parsed from the root struct's "endian"
bast_doc: Value, // retained for sequential_reader/read_field
root_name: String, // the selected $defs entry
}
// Private — the consumer selects via LayoutMode at compile time.
@@ -68,205 +168,106 @@ enum Layout {
```
The consumer selects the mode at construction time via `LayoutMode`
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The `Layout`
enum is private — the engine exposes mode-appropriate accessors instead:
(see [layout-engine.md](layout-engine.md) §"Mode Selection"). The
`Layout` enum is private — the engine exposes mode-appropriate
accessors instead:
```rust
impl AlkTypeEngine {
pub fn compile(schema: &mut Value, mode: LayoutMode) -> Result<Self, AlkTypeError>;
pub fn compile(
bast_doc: &Value,
root_name: &str,
mode: LayoutMode,
json_schema: Option<&Value>,
) -> Result<Self, AlkTypeError>;
pub fn mode(&self) -> LayoutMode;
pub fn endian(&self) -> Endian;
pub fn offset_map(&self) -> Option<&OffsetMap>; // Some in aligned mode
pub fn layout_builder(&self) -> Option<&LayoutBuilder>; // Some in packed mode
pub fn sequential_reader(&self) -> Option<SequentialReader>; // owned fresh reader (ADR-007)
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // ADR-004
pub fn is_valid_json(&self, instance: &Value) -> bool; // ADR-004
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // ADR-010
pub fn read_field<'a>(&self, buffer: &'a [u8], field_path: &str)
-> Result<FieldValue<'a>, AlkTypeError>; // aligned mode
pub fn write_field(&self, buffer: &mut [u8], field_path: &str,
value: &FieldValue<'_>) -> Result<(), AlkTypeError>; // aligned mode
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // D-BAST-007
pub fn is_valid_json(&self, instance: &Value) -> bool; // D-BAST-007
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // D-BAST-006
pub fn validation_plan(&self) -> &Arc<ValidationPlan>; // compiled constraints (ADR-012 §3)
}
```
`compile` takes `&mut Value` because it normalizes `$ref` values in place
(via [`normalize_refs`](schema-layer.md#ref-resolution-and-normalization))
before computing the layout and building the validator. The `schema`
field retains the normalized schema for `read_field`'s kind lookup and
for `sequential_reader()`'s factory construction. The validator is
mode-agnostic (it operates on `Value`, not raw bytes).
`compile` takes `&Value` (not `&mut Value`) — BAST needs no in-place
`normalize_refs`. `root_name` selects which `$defs` entry is the
top-level type (D-BAST-001). `json_schema` is the optional
consumer-provided standard JSON Schema for the `validate_json` path
(D-BAST-007); pass `None` when JSON validation is not needed. The
engine retains a clone of the BAST `Value` so `sequential_reader` and
`read_field` can re-parse the typed tree on demand without lifetime
entanglement with the caller's `Value`.
The `Layout::Packed` variant stores only the `LayoutBuilder` (write-side).
The `SequentialReader` (read-side) is not stored — it has mutable cursor
state that the consumer owns, so `sequential_reader()` constructs a fresh
reader on each call (ADR-007).
The `Layout::Packed` variant stores only the `LayoutBuilder`
(write-side). The `SequentialReader` (read-side) is not stored — it has
mutable cursor state that the consumer owns, so `sequential_reader()`
constructs a fresh reader on each call (ADR-007).
The `read_field`/`write_field` methods on `AlkTypeEngine` are the
aligned-mode data-access API — see [data-access.md](data-access.md)
§"Higher-level read/write".
## Custom Keyword Validators
## `build_validator`
Each `AlkType:*` kind gets a `Keyword` implementation registered via
`jsonschema::options().with_keyword(...)`. The validators check leaf
type constraints; `jsonschema` handles all structural validation.
### Numeric type validators
**`AlkType:Float32` / `AlkType:Float64`:**
- Value must be a finite number.
- For `Float32`: value must be representable as `f32` (no precision loss
beyond `f32`'s mantissa).
**`AlkType:Int8` / `AlkType:Int16` / `AlkType:Int32`:**
- Value must be an integer within the type's range.
- Int8: -128..127, Int16: -32768..32767, Int32: -2147483648..2147483647.
**`AlkType:Uint8` / `AlkType:Uint16` / `AlkType:Uint32`:**
- Value must be a non-negative integer within the type's range.
- Uint8: 0..255, Uint16: 0..65535, Uint32: 0..4294967295.
### String and binary validators
**`AlkType:String`:**
- Value must be a valid UTF-8 string.
- If `maxLength` is specified in the schema, the string's byte length
must not exceed it.
**`AlkType:Bytes`:**
- Value must be a string (the JSON form for `validate_json` consumers)
or an array of integers 0..=255 (the materialized form for
`validate_bytes`). JSON has no native byte type; the string form is
the JSON convention, the array form is the round-trippable form for
non-UTF-8 bytes (see [OQ-007](questions/007-bytes-materialization-lossy-utf8.md)).
- If `maxLength` is specified, the byte length must not exceed it. For
the string form, this is the string's byte length; for the array
form, this is the array length (one entry per byte).
- **Binary representation:** In the binary layout, `TBytes` is raw bytes
with no encoding (not base64, not hex). The JSON representation (for
validation) uses a string or array; the binary representation (for
data access) uses `&[u8]` directly.
**`AlkType:Enum`:**
- The `AlkType:Enum` custom keyword signals that the type is an enum for
*layout* purposes (the engine needs to know it's a fixed-size u32 index,
not a variable-length string). The built-in `enum` keyword provides the
value list and handles value-membership validation. The custom keyword
validator is a no-op beyond the built-in check — it exists solely for
the layout engine to recognize the type.
**`AlkType:Timestamp`:**
- Value must be a valid RFC 3339 timestamp string (the internet profile
of ISO 8601, e.g., `"2026-07-20T15:30:00Z"`).
### Composite type validators
**`AlkType:Struct`:**
- Value must be an object.
- Each property must match its declared `AlkType:*` kind.
- Required fields must be present.
- The `jsonschema` crate's built-in `properties` and `required` keywords
handle the structural checks — the custom keyword only needs to
validate that each field's value matches its `AlkType:*` kind.
**`AlkType:Union`:**
- The instance must be an object with a `__discriminator` field
carrying the mapping key (stringified discriminator value for
byte-offset discriminators, string value for field-name
discriminators). This is the shape the materializer produces for
`validate_bytes`; `validate_json` consumers produce the same shape
when validating a union instance.
- The `UnionValidator` builds a sub-validator for each variant at
factory time (when the parent validator tree is constructed) and
dispatches on `__discriminator` at validation time, validating the
full instance (including the variant fields) against the selected
variant's schema. This closes the OQ-008 gap: variant field
constraints (e.g. `maxLength` on a `Bytes` field inside a variant)
are checked.
- `$ref`s in the union's `mapping` are inlined by
`schema::inline_union_variant_refs` during `AlkTypeEngine::compile`
(before `build_validator`), so the `union_factory` sees full inline
variant schemas. See [OQ-008](questions/008-unionvalidator-variant-dispatch.md).
**`AlkType:Array`:**
- Value must be an array.
- Each element must match the array's declared element type.
- If `minItems`/`maxItems` is specified, the array length must be within
bounds.
### Other validators
**`AlkType:Boolean`:**
- Value must be `true` or `false`.
**`AlkType:Record`:**
- Value must be an object.
- All values must match the record's declared value type (specified via
the `"values"` property in the schema, e.g.,
`"values": { "AlkType:Float32": true }`).
### Validator implementation pattern
Each custom keyword implementation is ~10 lines. Example for
`AlkType:Float32`:
`src/validation.rs` exposes one function:
```rust
struct Float32Validator;
impl Keyword for Float32Validator {
fn validate<'i>(&self, instance: &'i Value) -> Result<(), ValidationError<'i>> {
match instance {
Value::Number(n) if n.as_f64().map_or(false, |f| f.is_finite()) => Ok(()),
_ => Err(ValidationError::custom("expected finite f32-compatible number")),
}
}
fn is_valid(&self, instance: &Value) -> bool {
instance.as_f64().map_or(false, |f| f.is_finite())
}
}
pub fn build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>;
```
Registration:
Under the pivot this is **repurposed** (D-BAST-007): it builds a
*standard* `jsonschema::Validator` from a consumer-provided plain JSON
Schema — no custom keywords, no BAST involvement. The engine calls it
internally during `compile` when `json_schema` is `Some`. Consumers
that only need a one-off validator may call `jsonschema::options().build(schema)`
directly; `build_validator` exists so the engine's error mapping
(`jsonschema` build error → `AlkTypeError::Schema`) is reused.
```rust
let validator = jsonschema::options()
.with_keyword("AlkType:Float32", |parent, value, path| {
Ok(Box::new(Float32Validator))
})
.build(&schema)?;
```
The factory closure receives the parent schema object, the keyword's
value, and the schema path. This enables cross-keyword awareness — for
example, a `AlkType:Struct` validator can inspect the parent's
`properties` to validate each field against its declared `AlkType:*` kind.
The v0.1.0 custom-keyword `build_validator` (registered 19
`with_keyword(...)` factories) is removed.
## AlkTypeError
A single `AlkTypeError` enum covers all error conditions across the
engine's three phases (schema parsing, offset computation, read/write)
plus validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md).
engine's phases (schema parsing, offset computation, read/write) plus
validation. Decided in [ADR-004](decisions/004-error-handling-validation-strategy.md);
the variant shapes are unchanged under the pivot (D-BAST-009).
```rust
pub enum AlkTypeError {
/// Schema parsing errors (invalid JSON, missing keywords, unknown AlkType kinds).
/// Schema parsing errors (malformed BAST, dangling $ref, unknown kind).
Schema(String),
/// Offset computation errors (field not found, unsupported type).
Offset { field_path: String, reason: String },
/// Read/write errors (buffer too short, invalid UTF-8, value out of range).
Access { field_path: String, reason: String },
/// Validation errors (delegated to jsonschema).
Validation(ValidationError<'static>),
/// Validation errors (both paths — D-BAST-009 uniform payload).
Validation(jsonschema::ValidationError<'static>),
}
```
- **`Schema`** — for errors during `AlkTypeEngine::compile()`. Invalid
JSON, missing required keywords, unknown `AlkType:*` kinds.
- **`Schema`** — for errors during `AlkTypeEngine::compile()` or any
BAST-walking path. Malformed BAST, missing `$defs`, unknown `kind`
string, dangling `$ref`, empty `mapping`, etc.
- **`Offset`** — for errors during offset computation. Field not found
in the schema, type not supported for offset computation, recursive
depth exceeded. Carries the field path.
- **`Access`** — for errors during read/write. Buffer too short, invalid
UTF-8 in a string field, value out of range for the target type.
Carries the field path.
- **`Validation`** — wraps `jsonschema`'s `ValidationError`. The
`'static` lifetime is correct — the validator owns its schema reference
and lives for the lifetime of the `AlkTypeEngine`.
in the BAST tree, type not supported for offset computation. Carries
the field path.
- **`Access`** — for errors during read/write. Buffer too short,
invalid UTF-8 in a string field, value out of range for the target
type. Carries the field path.
- **`Validation`** — wraps a `jsonschema::ValidationError<'static>`.
On the `validate_json` path, this is the `jsonschema` crate's own
structured error. On the `validate_bytes` path, it is constructed via
`jsonschema::ValidationError::custom` from the BAST-native
validator's path + reason string. The `'static` lifetime is correct —
the payload owns its data.
### Field-path-carrying errors
@@ -287,20 +288,64 @@ you exactly which field failed and why.
### Load time: `AlkTypeEngine::compile()`
The expensive work happens once at schema load time:
1. Normalize `$ref` values in the schema (`normalize_refs`).
2. Parse the schema's `"endian"` annotation.
3. Compute the layout (`LayoutBuilder`/`SequentialReader` for packed, `OffsetMap` for aligned).
4. Build the jsonschema validator (`jsonschema::options().with_keyword(...).build(&schema)?`).
The result is a `AlkTypeEngine` that can be used for repeated operations.
1. Parse the BAST document into the typed tree (`BastDoc::new`).
2. Parse the root struct's `"endian"` annotation.
3. Compile the `ValidationPlan` — the value-domain constraint tree,
with eager `$ref` resolution. Its compile walk rejects cyclic `$ref`
graphs with a clean `Schema` error *before* the layout computation.
(The layout walkers now also guard themselves — each standalone
entry point runs the shared reference-graph check
(`walk_guard::check_ref_graph`, review #006 H2) — so the plan-first
ordering is belt-and-suspenders at engine compile, and the trust
boundary no longer depends on the call path.)
4. Compute the layout (`LayoutBuilder` for packed, `OffsetMap` for
aligned).
5. If `json_schema` is `Some`, build the standard
`jsonschema::Validator` via `validation::build_validator`.
The result is an `AlkTypeEngine` that can be used for repeated
operations. The validation half is pre-built: the engine holds an
`Arc<ValidationPlan>` and walks it per buffer without re-touching the
BAST document (ADR-012 §3).
### Access time: `engine.validate_bytes(&[u8])`
For binary-layout schemas, `validate_bytes` runs the two phases in
sequence (D-BAST-006):
1. **Materialize `Value` from bytes.** `materialize::materialize_packed`
or `materialize::materialize_aligned` walks the buffer against the
BAST typed tree and the engine's `Endian`, producing a
`serde_json::Value` tree. Composites are recursed into (`Struct` →
object of field values; `Array` → array of element values; `Union`
→ dispatch then recurse; `Record` → object of key/value entries).
The read phase reuses the existing data-access functions and returns
`AlkTypeError::Access` (with field paths) on read failures.
2. **Validate the `Value` against the `ValidationPlan`.** The
materialized `Value` is walked against the compiled constraint tree
(`engine.validation_plan().validate(&value)`), producing
`AlkTypeError::Validation` on the first violated value-domain
constraint. This is the ADR-012 §3 end state: the only per-buffer
schema-touching step is the materialize half (the bytes must be
decoded against the tree); the validation half is plan-fast.
Mode dispatch:
- **Packed mode** — materializes fields in declaration order.
- **Aligned mode** — uses the `OffsetMap` to read fields at their
computed offsets.
Both modes produce the same `Value` form; the validation plan is
mode-agnostic.
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
Validation is opt-in per operation. The consumer calls
`engine.validate_json(instance)` when validation is desired, or
`engine.is_valid_json(instance)` for a boolean check. The jsonschema
validator is already compiled — these are fast checks against the
compiled validator.
`engine.is_valid_json(instance)` for a boolean check. The
`jsonschema::Validator` is already compiled — these are fast checks
against the compiled validator.
```rust
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>;
@@ -308,79 +353,37 @@ pub fn is_valid_json(&self, instance: &Value) -> bool;
```
The argument is a `serde_json::Value` (the JSON representation of the
data), not a raw byte buffer — see §"What validation validates" above.
To validate a binary buffer end-to-end, the consumer reads it into a
`Value` tree via the data access layer, then validates that `Value`.
data), not a raw byte buffer. `validate_json` validates against the
consumer-provided JSON Schema supplied at `compile` time (D-BAST-007);
the BAST document is not involved. If no JSON Schema was supplied,
`validate_json` returns `AlkTypeError::Schema` and `is_valid_json`
returns `false`.
High-throughput paths can skip validation. Security-sensitive paths
(parsing incoming frames from untrusted peers) can validate every frame.
The choice is the consumer's.
(parsing incoming frames from untrusted peers) can validate every
frame. The choice is the consumer's.
### Access time: `engine.validate_bytes(&[u8])` — binary buffer validation
For binary-layout schemas (schemas declaring `AlkType:*` kinds), the
engine offers a single-call form of the two-step dance: walk the bytes
against the layout to materialize a `Value` tree, then validate that
`Value` against the compiled jsonschema validator. Decided in
[ADR-010](decisions/010-generalized-validation-validate-bytes.md).
```rust
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>;
```
`validate_bytes` runs the existing machinery in sequence:
1. **Materialize `Value` from bytes.** A new internal helper
(`materialize_value`, alongside `SequentialReader::read_field_value`
in `src/sequential_reader.rs`) walks the buffer against the schema
and the engine's `Endian`, producing a `serde_json::Value` tree.
Composites are recursed into (`Struct` → object of field values;
`Array` → array of element values; `Union` → dispatch then recurse;
`Record` → object of key/value entries). The read phase reuses the
existing data-access functions and returns `AlkTypeError::Access`
(with field paths) on read failures.
2. **Validate the `Value`.** The materialized `Value` is passed to the
existing `self.validator.validate(&value)`, producing
`AlkTypeError::Validation` on failure.
Mode dispatch:
- **Packed mode** — walks with a fresh `SequentialReader` (the engine
is already a reader factory per ADR-007), materializing fields in
declaration order.
- **Aligned mode** — uses the `OffsetMap` to read fields at their
computed offsets, then materializes composites by recursing into the
offset map's nested entries.
Both modes produce the same `Value` form; the validator is
mode-agnostic (it operates on `Value`, not bytes — ADR-004).
#### When to use which entry point
### When to use which entry point
| Entry point | Schema form | Input form | When |
|-------------|--------------|------------|------|
| `validate_json(&Value)` | Any (AlkType or plain JSON Schema) | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); TypeBox output; anything off `serde_json::from_slice` / `from_str` |
| `validate_bytes(&[u8])` | AlkType binary-layout schema | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs |
| `validate_json(&Value)` | Consumer-provided standard JSON Schema | Already-parsed `serde_json::Value` | Call's JSON payloads (`OperationSpec.input_schema`); anything off `serde_json::from_slice` / `from_str` |
| `validate_bytes(&[u8])` | BAST document (binary layout) | Raw `&[u8]` buffer | Channels' 8-byte chunk header; future binary call frames; SFTP packet buffers; metatensor index structs |
`validate_bytes` requires the engine's schema to declare `AlkType:*`
kinds — it materializes `Value` via the layout engine, which needs
binary-layout semantics. A pure JSON Schema (call's `input_schema`,
no AlkType kinds) compiled via `AlkTypeEngine::compile` would fail at
the materialize step (no `AlkType:Struct` at the root). For pure JSON
payloads, the consumer uses `serde_json::from_slice` then
`validate_json`. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md)
§"Not a binary-payload validator for JSON-only schemas".
`validate_bytes` requires the engine's root type to be a struct (the
layout engine enforces this) — it materializes `Value` via the layout
engine, which needs binary-layout semantics. For pure JSON payloads,
the consumer uses `serde_json::from_slice` then `validate_json`. See
[ADR-010](decisions/010-generalized-validation-validate-bytes.md) and
[ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
#### What `validate_bytes` is not
- **Not a new validation engine.** It runs the existing `jsonschema`
validator against the existing materialized `Value`. No new
validator code, no parallel validation path (ADR-001).
- **Not framing-aware.** It validates the bytes of *one* schema
instance. It does not strip length prefixes, parse
`[length: u32][payload]` framing, or handle multiple frames in a
buffer. That's the consumer's job. alktype validates what one
schema describes; it does not parse the wire envelope around it.
buffer. That's the consumer's job. alktype validates what one schema
describes; it does not parse the wire envelope around it.
- **Not a `Validator` trait.** Two methods on one struct, not a trait
abstraction. See [ADR-010](decisions/010-generalized-validation-validate-bytes.md)
§"Not a `Validator` trait abstraction".
@@ -390,24 +393,25 @@ payloads, the consumer uses `serde_json::from_slice` then
Validation and data access are independent operations on the same data.
The consumer can:
1. Validate the JSON representation of a buffer to ensure it conforms to
the schema.
1. Validate the bytes of a buffer to ensure it conforms to the BAST
document's value constraints.
2. Read fields from the binary buffer at computed offsets.
3. Both — validate the JSON representation first, then read the binary
buffer (defense in depth).
3. Both — validate first, then read (defense in depth).
The engine does not couple validation and access. A consumer that trusts
its data source can skip validation and go straight to read/write. A
consumer that parses untrusted input can validate the JSON
representation first, then access the binary buffer.
The engine does not couple validation and access. A consumer that
trusts its data source can skip validation and go straight to
read/write. A consumer that parses untrusted input can validate first,
then access the binary buffer.
## Design Decisions
| Decision | ADR | Summary |
|----------|-----|---------|
| Error handling and validation | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Two-validator model (BAST-native + standard jsonschema) | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | `validate_bytes` uses the compiled `ValidationPlan`; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema; D-BAST-006/007/009 |
| Error handling and validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation (materialize `Value` from bytes, then validate); two methods on one struct, not a trait |
| Purpose and scope | [ADR-001](decisions/001-alktype-purpose-scope-jsonschema-engine.md) | Why jsonschema not a custom engine |
| Compiled `ValidationPlan` | [ADR-012](decisions/012-plan-fingerprinting-and-m1-closure.md) | The value-domain constraint tree is compiled once at `compile` (eager `$ref` resolution, cycle rejection) and walked per buffer; `Hash + Eq` + `fingerprint()`; retires the interpretive `BastDoc` walk |
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | The BAST document is the complete binary-format spec (layout + value constraints) |
## Open Questions
@@ -418,11 +422,23 @@ see [builder.md](builder.md).
## References
- `@alkdev/alknet: docs/research/alknet-typedef/findings.md`
§"Validation" — the POC's custom keyword validators for all 17 kinds
- [`bast-format.md` §Validation Model](bast-format.md#validation-model)
— the normative validation model
- [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) — the
two-validator decision
- [ADR-012](decisions/012-plan-fingerprinting-and-m1-closure.md) — the
`ValidationPlan` decision (§3)
- [ADR-004](decisions/004-error-handling-validation-strategy.md) —
error handling and validation strategy
- [schema-layer.md](schema-layer.md) — the 19 AlkType kinds that the
validators check
- [data-access.md](data-access.md) — read/write functions that operate
on the same buffers
- [ADR-010](decisions/010-generalized-validation-validate-bytes.md) —
`validate_bytes` (the collapsed two-step dance)
- [schema-layer.md](schema-layer.md) — the BAST parser that the
plan compiler consumes
- [data-access.md](data-access.md) — read/write functions and the
materializer that produce the `Value` the plan checks
- [`src/validation_plan.rs`](../../src/validation_plan.rs) — the
compiled `ValidationPlan` implementation
- [`src/bast_validation.rs`](../../src/bast_validation.rs) — the
one-shot wrapper (`validate_value`) and shared error helpers
- [`src/validation.rs`](../../src/validation.rs) — the `build_validator`
helper
+983
View File
@@ -0,0 +1,983 @@
---
status: done
created: 2026-08-19
last_updated: 2026-09-02
adr: ADR-011, ADR-012
---
# 0.3.0 — Compiled Forms: ReadPlan, Owned BastDoc, OffsetMap LeafMeta, ValidationPlan, Fingerprinting
This is the execution plan for the 0.3.0 release: the compiled-form
rollup that closes review #004's 400x read-path gap (ADR-011) *and*
the deferred M1 sites (ADR-012) *and* retires the interpretive
validation walk via a `ValidationPlan` (ADR-012 §3, reversing the
original deferral per review #005 M3) *and* adds plan fingerprinting
(ADR-012 §1/§4) in one breaking bump. It is the **entry point** an
implementing agent reads first.
Companion documents:
- [ADR-011](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md)
— the `ReadPlan` decision (packed read-side compiled form).
- [ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md)
— fingerprinting + owned `BastDoc` + `OffsetMap` `LeafMeta` +
`ValidationPlan` (this release's other three pieces).
- [Review #004](../reviews/004-performance-review.md) — the
performance finding being closed.
- [POC findings](../../poc/readplan/FINDINGS.md) (branch `readplan-poc`)
— the derisking POC that confirmed the `ReadPlan` shape and surfaced
two findings (field-disc union read shape; struct-array stride).
**Working order:** read this plan top-to-bottom. The Semver Contract
section is the scope-creep guardrail — consult it before each step.
Each step links to its ADR and lists its verification gate. Implement
phases in order; within a phase, steps are ordered by dependency.
## Phases vs sessions
This plan is deliberately larger than one session's work. The eight
phases are the session boundaries — each phase is a coherent unit
that leaves the tree building and tests green, so any one session
can pick up a phase without needing context from the previous one.
Phase boundaries are also commit boundaries (and push boundaries per
AGENTS.md). If a phase is large enough to span sessions, the steps
within it are the sub-session boundaries.
## Semver Contract
The crate is on crates.io at 0.2.0 with zero real consumers (only
`alktty`/`alkcall`, both in-house path dev-deps). A breaking bump to
0.3.0 is free but the contract is explicit so the implementation
doesn't drift. Per AGENTS.md, the public surface is the `lib.rs`
re-exports.
| Public item (from `lib.rs` re-exports) | Class | Change |
|---|---|---|
| `AlkTypeEngine::compile` | **Breaking (internal)** | Signature unchanged `(bast_doc: &Value, root_name: &str, mode, json_schema) -> Result<Self, AlkTypeError>`. Internally builds a `ReadPlan` (packed) or extended `OffsetMap` (aligned) and stores it. The `bast_doc: Value` clone is retained (ADR-011 §Engine integration). |
| `AlkTypeEngine::sequential_reader` | **Breaking (return type)** | Returns `Option<SequentialReader>` (unchanged type), but the reader is now constructed from `Arc<ReadPlan>`, not from `&bast_doc`. The reader's public methods (`read_next`/`read_field`/`reset`/`position`/`endian`/`schema`) keep their signatures. `schema()` returns the `&Value` the plan was compiled from (retained on the engine). |
| `AlkTypeEngine::read_field` / `write_field` | **Unchanged (signature)** | Still `(buffer, field_path) -> Result<FieldValue, AlkTypeError>`. Internally reads `LeafMeta` from the extended `OffsetMap` instead of re-parsing `BastDoc`. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (signature)** | Packed mode calls `materialize_packed(&self.plan, buffer)` (ADR-011); aligned mode calls `materialize_aligned(&doc, buffer, &self.offset_map)` with the owned `BastDoc`. |
| `LayoutMode`, `AlkTypeEngine` | **Unchanged** | — |
| `BastDoc`, `BastDef`, `BastDefKind`, `BastStruct`, `BastField`, `BastType`, `BastUnion`, `BastDiscriminator`, `BastEnum`, `BastArray`, `BastRecord`, `BastRef` | **Breaking (lifetime removal)** | `BastDoc<'a>` → `BastDoc` (owned). Every `&'a str` → `String` (or `Arc<str>` — decision in phase 3). Every `&'a Value` → `Value` (or `Arc<Value>`). Every method signature that took/returned `&'a` changes. The `Bast*` types are re-exported from `lib.rs` so this is a public break. |
| `OffsetMap` | **Breaking (`get` return type)** | `get(field_path) -> Option<&ByteRange>` → `get(field_path) -> Option<&OffsetEntry>` where `OffsetEntry { range: ByteRange, meta: LeafMeta }` (or two accessors). Additive capability. |
| `ByteRange` | **Unchanged** | Still `Copy + PartialEq + Eq + Hash`. |
| `LeafMeta` | **New public type** | `{ kind: AlkTypeKind, encoding: VariableEncoding, endian: Endian }`, re-exported from `lib.rs`. `Copy + PartialEq + Eq + Hash`. |
| `ReadPlan` | **New public type** | From ADR-011. Re-exported from `lib.rs`. `Debug + Clone + PartialEq + Eq + Hash`. |
| `SequentialReader` | **Breaking (constructor + return type)** | `SequentialReader::new(&Value, &str) -> Result<Self, AlkTypeError>` → `SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible — just stores the `Arc`; the `BastDoc` parse moved to `ReadPlan::compile`). Public methods (`read_next`/`read_field`/`reset`/`position`/`endian`/`schema`) unchanged. `schema()` returns `&Value` retained on the plan (see phase 2 — the plan stores `Arc<Value>`, not `&Value`, to avoid the self-referential struct ADR-011 rejects). `engine.rs`'s `.ok()` on the old `Result` correspondingly goes away. |
| `FieldValue` | **Unchanged** | — |
| `materialize_packed` | **Breaking (signature)** | `materialize_packed(&BastDoc<'_>, &[u8])` → `materialize_packed(&ReadPlan, &[u8])`. |
| `materialize_aligned` | **Breaking (signature)** | `materialize_aligned(&BastDoc<'_>, &[u8], &OffsetMap)` → `materialize_aligned(&BastDoc, &[u8], &OffsetMap)` (owned `BastDoc`, no lifetime). |
| `ValidationPlan` | **New public type** | From ADR-012 §3. Re-exported from `lib.rs`. `Debug + Clone + PartialEq + Eq + Hash`. `compile(&BastDoc, &str) -> Result<Self, AlkTypeError>`, `fingerprint() -> u64`. Shape scoped in phase 7 (and a preceding design session); the contract is fixed in ADR-012 §3. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (signature)** | Still `(buffer) -> Result<(), AlkTypeError>`. Internally walks `&ValidationPlan` (both modes) instead of re-walking `BastDoc` for value-domain checks. The `materialize` half is unchanged from the ADR-011/§2b work (packed: `materialize_packed(&self.plan, ...)`; aligned: `materialize_aligned(&self.doc, ..., &self.offset_map)`). |
| `bast_validation` (`build_validator`, `validate_value`) | **Additive (reviewed at phase 7)** | `validate_value` is expected to become a thin wrapper over `ValidationPlan` (or be retired if the shape work shows it's redundant). Additive changes ride the bump; renames/removals are avoided unless phase 7's shape work shows they're necessary. The BAST meta-schema validator (`build_validator`/`BAST_META_SCHEMA`) used at `compile` time is unaffected. |
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged (signature)** | `LayoutBuilder::new`/`build` signatures unchanged. Internally caches the owned `BastDoc` instead of re-parsing. |
| `AlkTypeKind`, `Endian`, `VariableEncoding` | **Unchanged** | — |
| `AlkTypeError` | **Unchanged** | — |
| `Schema`, `Definitions`, `Discriminator` builders | **Unchanged** | — |
| `UnionDispatch`, `build_validator`, `BAST_META_SCHEMA` | **Unchanged** | — |
| `data_access::*` functions | **Unchanged** | — |
**Net breaking surface:** `BastDoc` and all `Bast*` types (lifetime
removal), `OffsetMap::get` (return type), `SequentialReader::new`
(constructor + `Result` drop), `materialize_packed`/`materialize_aligned`
(signatures). **Net additive:** `ReadPlan`, `LeafMeta`, `ValidationPlan`,
`fingerprint()` methods, `Hash + Eq` derives on `ReadPlan`/`OffsetMap`/
`ValidationPlan`. **Net unchanged:** the builder, `AlkTypeEngine`
accessors (signatures), `FieldValue`, `AlkTypeKind`, `AlkTypeError`,
`data_access`, the `Schema`/`Definitions` builders, `validate_bytes`
(signature).
### Decisions deferred to their implementation phases
1. **`Arc<str>` vs `String` for owned `BastDoc` names** (phase 3): `Arc<str>`
shares allocation for repeated names (e.g. union variant keys appearing
in multiple places); `String` is simpler. The POC used `String`. Lean:
`String` unless a bench shows `Arc<str>` matters — the typed tree is
built once, not hot. Decided in phase 3.
2. **`OffsetMap::get` return shape** (phase 5): `Option<&OffsetEntry>` (a
new accessor struct) vs two methods `range(path) -> Option<&ByteRange>`
+ `meta(path) -> Option<&LeafMeta>`. The struct is fewer calls; the two
methods preserve back-compat shape for callers that only want the range.
Lean: struct — it's a breaking bump anyway and the struct is cleaner.
Decided in phase 5.
3. **Fingerprint hasher** (phase 6): `DefaultHasher` (std, stable within a
version) vs `FxHasher` (faster, also stable). Cross-version stability
is a non-goal (ADR-012). Lean: `DefaultHasher` — no new dep, the
fingerprint isn't hot. Decided in phase 6.
4. **Struct-array stride** (phase 2): the POC found the existing reader
returns `element_stride: 0` for fixed-size struct arrays
(`sequential_reader.rs:567`). The `ReadPlan` correctly computes the
stride. Decision: preserve the existing `0` behavior in `ReadPlan` for
back-compat with `SequentialReader`'s consumer contract, *or* fix it
and document the behavioral change. Lean: fix it — the `0` is a latent
bug, the stride is behaviorally observable, and we're bumping. The
plan step calls this out explicitly. Decided in phase 2.
## Phase 1 — `ReadPlan` type + `compile` (ADR-011 step 1) — **DONE (2026-09-02)**
> **Status: implemented.** `src/read_plan.rs` builds the refined
> `CompositePlan::Union` shape (`shared: Option<Box<ReadPlan>>` +
> `variants: Vec<(String, CompositePlan)>`, no `VariantPlan`/`VariantKind`)
> with eager `$ref` resolution, `BTreeMap` `by_name`, field-disc `shared`
> sub-plans, nested-union variant support, and true array strides
> (deferred decision 4 resolved: fixed struct arrays compute their real
> stride, not the 0.2.0 reader's `0`). Two parity notes recorded as
> compile-behavior locks in tests: (a) the plan propagates the
> *referring field's* effective endianness into nested structs/unions —
> exactly what the 0.2.0 packed reader/materializer do — rather than
> consulting nested containers' own `endian` annotations (the POC baked
> `s.endian()` there; its equivalence tests never covered a nested
> annotation, so the divergence was latent); (b) `compile` carries its
> own depth cap (128) + definition-level cycle set, so standalone
> `ReadPlan::compile` is untrusted-input-safe independent of the
> meta-schema and the engine's `ValidationPlan` gate.
**Goal:** Add the `ReadPlan`/`FieldPlan`/`CompositePlan`/`ReadKind`/
**ADR reference:** [ADR-011 §The `ReadPlan` shape](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#the-readplan-shape),
[ADR-011 §Construction](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#construction).
The `CompositePlan::Union` shape in ADR-011 was refined (vs the
accepted-at-POC shape) to carry the field-disc union's shared fields
and to drop `VariantPlan`/`VariantKind`; this phase implements the
refined shape.
**POC reference:** `poc/readplan/src/lib.rs` (branch `readplan-poc`)
is the reference scaffold. The production version lives in `src/` and
adds doc comments, clippy cleanliness, the field-name-discriminator
union read shape the POC stubbed (POC Finding 1), and nested-union
variant support the POC rejected but 0.2.0 accepts (POC "What this
POC does not cover" → nested unions). Both are resolved by the
refined `CompositePlan::Union` shape — see below.
**Files:** New `src/read_plan.rs`. Update `src/lib.rs` to add
`pub mod read_plan;` and re-export `ReadPlan` (and the plan sub-types
that are part of the public surface — `ReadKind`, `CompositePlan`,
etc. if the ADR's public-API section calls for them; the ADR lists
`ReadPlan` as the public type, sub-types can stay `pub` in-module if
consumers don't need to name them).
**Implementation notes:**
- `compile` walks `BastDoc` once (via the existing borrowed `BastDoc`,
which still exists at this phase — the owned-`BastDoc` refactor is
phase 3). Resolves all `$ref`s eagerly, computes effective endianness
at every node, inlines union variants. Malformed schemas surface as
`AlkTypeError::Schema` (AGENTS.md §3); overflow-safe arithmetic
(AGENTS.md §4).
- `by_name: BTreeMap<String, usize>` (not `HashMap` — ADR-012 §1
requires `Hash` on `ReadPlan`, and `HashMap` blocks derive). The
POC used `HashMap`; swap to `BTreeMap`.
- **Union shape (ADR-011 refined — resolves POC Finding 1 and the
nested-union gap):** `CompositePlan::Union { disc, shared,
variants: Vec<(String, CompositePlan)> }`.
- `shared: Option<Box<ReadPlan>>` — the union's declared `fields`
(the discriminator field + any shared fields) for the
field-name-discriminator case. `DiscriminatorPlan::Field.field_index`
indexes into `shared`. The read loop walks `shared` first, then
looks up the selected variant and walks its `CompositePlan`
starting after the shared fields. The byte-offset-discriminator
case sets `shared: None` (no shared fields; the variant starts
immediately after the discriminator size). This replaces the
POC's stubbed `plan_read_union` `Field` arm.
- `variants: Vec<(String, CompositePlan)>` — **not** the POC's
`Vec<(String, VariantPlan)>`. Dropping `VariantPlan`/`VariantKind`
means a variant's body is just a `CompositePlan`, so **nested
unions** (a variant that is itself a union, which the 0.2.0
reader supports via `resolve_and_walk_variant`'s `Union` arm at
`sequential_reader.rs:800`) work by ordinary `CompositePlan`
recursion — a variant can be `CompositePlan::Union { ... }`. No
separate `VariantKind::Union` arm, no behavioral drop vs 0.2.0,
no Semver Contract entry for a capability regression. The POC's
`VariantPlan { kind, plan }` wrapper is not carried forward.
- `compile_union` must reject a variant that is neither a struct
nor a union with `AlkTypeError::Schema` (mirroring
`resolve_and_walk_variant`'s `other => Err(...)` arm), so the
eager-resolution path keeps the untrusted-input discipline.
- Do *not* wire `ReadPlan` into `SequentialReader` or `materialize` yet
— that's phase 2. Phase 1 is the type + `compile` only, unit-tested
against the same BAST fixtures `bast.rs` uses (the existing `bast.rs`
tests are a ready source of fixtures).
**Verification:** `cargo test --release` (new unit tests for `compile`
covering every `BastType` arm — port the `cov_*` tests from
`poc/readplan/tests/coverage.rs`, **plus** a nested-union-variant
test asserting `compile_union` produces `CompositePlan::Union` whose
variant body is itself `CompositePlan::Union`, restoring the 0.2.0
capability the POC rejected). `cargo clippy --all-targets -- -D
warnings`. `cargo doc --no-deps` (new public type). `cargo build
--target wasm32-unknown-unknown --release` (`read_plan.rs` is
wasm-relevant). **Add a `fn read_plan_is_send_sync()` assertion
test** (a `const _: fn() = || { fn assert_send_sync<T: Send + Sync>()
{}; assert_send_sync::<ReadPlan>(); };` static-bound assertion, as
ADR-011 §"Engine integration" requires) to lock in `ReadPlan: Send +
Sync` so a future change can't break it silently — mirror the POC's
`readplan_is_send_sync` test.
---
## Phase 2 — `SequentialReader` + `materialize_packed` consume `ReadPlan` (ADR-011 steps 2–4) — **DONE (2026-09-02)**
> **Status: implemented.** The packed read loop walks `Arc<ReadPlan>`:
> `SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible; the old
> fallible constructor's work moved to `ReadPlan::compile`), the reader
> holds `plan: Arc<ReadPlan>` + cursor only, `schema()` returns the
> `Arc<Value>` retained on the plan (review #005 H2 closed — no
> self-referential struct), and a new `plan()` accessor exposes the
> shared plan. `materialize_packed(&ReadPlan, &[u8])` walks the same
> plan; the aligned materialize path keeps walking `BastDoc` with the
> retained `dummy_field_for`/`ty_source`/`materialize_typeref_packed`
> helpers (phase 5 Scope Boundary). Engine: `Layout::Packed` carries
> `Arc<ReadPlan>`; `sequential_reader()` is an `Arc::clone` (15.7 ns,
> was a whole-document `Value` clone); packed `validate_bytes` calls
> `materialize_packed(&self.plan, ...)`. The temporary validation
> bridge (reconstruct `BastDoc` for the validator) is still in place —
> phase 7 already retired it on `main`'s ValidationPlan; this phase's
> `validate_bytes` edit merged cleanly onto that state.
>
> **Stride (deferred decision 4):** fixed struct/nested-array elements
> now report their true stride through `FieldValue::Array`
> (0.2.0 returned `0`); doc comment updated; no existing test asserted
> the `0`, so no test needed changing — the plan-compile tests lock the
> new values.
>
> **Two parity subtleties found and preserved** (both invisible to the
> existing test suite, both now locked by tests or by construction):
> (a) the materializer unwraps the plan's anonymous single-field
> wrapper for primitive array elements/record values — without this,
> materialized records/arrays would nest each leaf under a synthetic
> object and `validate_bytes` would fail its own parity suite (caught
> by `materialize_record_packed_count_prefixed_pairs`); (b) the
> field-disc union's materialized key order keeps `__discriminator`
> first (matching 0.2.0's `Map` insertion order, observable under
> `preserve_order`).
>
> **Bench (alktty `wire_vs_bast`, 1024 chunks/stream):** read p64
> 2.27 µs/chunk (review #004) → **98 ns/chunk** (~23x; gap 400x →
> ~17x vs hand-rolled's 5.6 ns); read p4k → 100 ns/chunk. `engine.
> sequential_reader()` construction 15.7 ns (was a full `Value` clone).
> The residual gap is dominated by the per-field `String` allocation
> mandated by the unchanged `(String, FieldValue)` `read_next` return
> signature (2 allocs/chunk) plus `data_access` bounds checks — both
> outside this phase's scope (the signature is pinned by the Semver
> Contract).
**Goal:** Rewrite the packed read loop to walk `&ReadPlan` instead of
reconstructing `BastDoc`. `SequentialReader` stores `Arc<ReadPlan>` +
cursor state; `materialize_packed` takes `&ReadPlan`. Closes H1
(the 400x gap) + the packed-side M1 + L1.
**ADR reference:** [ADR-011 §Scope](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#scope),
[ADR-011 §Recommended Order](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#recommended-order)
steps 2–4.
**Files:** `src/sequential_reader.rs` (rewrite the read loop, change
`new`'s signature), `src/materialize.rs` (`materialize_packed` takes
`&ReadPlan`), `src/engine.rs` (`compile` builds `Arc<ReadPlan>` in
packed mode, `sequential_reader()` hands out `Arc::clone`, packed
`validate_bytes` calls `materialize_packed(&self.plan, ...)`).
**Implementation notes:**
- `SequentialReader::new(&Value, &str) -> Result<Self, AlkTypeError>` →
`SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible — the
fallible `BastDoc` parse moved to `ReadPlan::compile` in phase 1;
`new` just stores the `Arc`). The reader stores `plan: Arc<ReadPlan>`,
`field_index: usize`, `position: usize`. `endian()` reads
`self.plan.endian()`.
- **`schema()` ownership (resolves review #005 H2):** `schema()`
returns `&Value` retained on the plan, but the plan stores an
**`Arc<Value>`**, not a `&Value`. ADR-011 §"Root cause" rejects the
self-referential struct pattern (a `ReadPlan` storing `&Value`
borrowing from the engine's `bast_doc: Value` would make the engine
self-referential — exactly the construction ADR-007 worked around
and ADR-011's `Arc<ReadPlan>` was meant to retire). The fix:
`ReadPlan` carries `schema: Arc<Value>`; `ReadPlan::compile` clones
the input `&Value` into `Arc<Value>` once; `schema()` returns
`&self.schema`. The engine stores `bast_doc: Arc<Value>` internally
(one allocation, shared via refcount between the engine and all
plans it builds) — this is an internal change, not a public
signature change (`compile` still takes `&Value`). `Arc<Value>`
implements `Hash + Eq` (`serde_json::Value: Hash + Eq` as of the
pinned `serde_json` with `preserve_order`; `Map::hash` sorts keys
for determinism), so phase 6's `#[derive(Hash)]` on `ReadPlan` is
not blocked by carrying the schema. **Note:** if a future
`serde_json` version regresses `Value: Hash`, phase 6 would need
`ReadPlan`'s hash to exclude the `schema` field; that is a phase-6
concern, not a phase-2 blocker.
- `read_field_at`/`read_field_value`/`read_typeref_value`/
`walk_struct_size`/`read_union_value`/`read_array_value`/
`read_record_value` are rewritten to take plan nodes
(`&FieldPlan`/`&CompositePlan`/`&ReadKind`) instead of
`&BastField`/`&BastType`/`&BastDoc`. Port `plan_read_field_at`/
`plan_walk_struct_size`/etc. from `poc/readplan/src/lib.rs` — they're
the reference implementations. The `read_union_value` rewrite
handles both discriminator kinds via the refined `CompositePlan::Union`
shape from phase 1: byte-disc reads the discriminator at
`disc.offset` then walks the variant (no `shared`); field-disc walks
`shared` first, reads the discriminator field at
`disc.field_index` within `shared`, looks up the variant, and walks
it starting after the shared fields. Nested unions (variant body is
itself `CompositePlan::Union`) recurse naturally — no special arm.
- **Struct-array stride (deferred decision 4):** the POC computes the
true fixed-struct stride; the existing reader returns `0`. The
production `ReadPlan::compile_array` should compute the true stride
(the POC's `fixed_struct_size` helper). Document the behavioral
change in the `FieldValue::Array` doc comment: `element_stride` is
now the true stride for fixed-size struct elements, not `0`. This is
a breaking behavioral change; rides the bump. Update the
`eq_array_ref_element`-style test to assert the new stride.
- **`materialize_packed` rewrite + packed/aligned split (resolves
review #005 L1 and L2):** `materialize_packed(&BastDoc<'_>, &[u8])`
→ `materialize_packed(&ReadPlan, &[u8])`. The materializer walks the
plan instead of `BastDoc`. **Scoped removal of helpers:** only the
*packed-side* call sites of `dummy_field_for`/`ty_source`
(`materialize.rs:249, 316, 351, 391` — the packed
`materialize_*_packed` paths) go away when packed-materialize moves
to the plan. The helpers themselves **stay**, because
`materialize_leaf_at` (`materialize.rs:631`, which calls
`dummy_field_for`) is on the **aligned** path — it's called by
`materialize_struct_aligned` (`:475`), `materialize_array_aligned`
(`:544`), `materialize_variable_aligned` (`:613`). Aligned
`materialize` keeps walking `BastDoc` through 0.3.0 (see the Scope
Boundary note in phase 5), so `dummy_field_for`/`ty_source` must
stay. **`materialize_typeref_packed` split:** this function is
currently shared by both packed and aligned paths (aligned's
`materialize_leaf_at` calls it to read leaves, and aligned's record
path at `:498-506` calls it directly). After phase 2,
packed-materialize gets a new plan-walking function;
`materialize_typeref_packed` stays for aligned's
`materialize_leaf_at` and the aligned record path (renamed or not —
implementer's choice; the function is private). This is two
mode-specific paths — the existing design — not a "parallel walker"
in the maintenance-tax sense ADR-011 §"Negative" cautions against
(that caution is about packed read-side `SequentialReader` +
`materialize_packed` sharing one plan, which this preserves).
- `AlkTypeEngine::compile` (packed branch): build `Arc<ReadPlan>` via
`ReadPlan::compile(bast_doc, root_name)`, store it in `Layout::Packed`.
`sequential_reader()` returns
`Some(SequentialReader::new(Arc::clone(&self.plan)))`.
`validate_bytes` (packed) calls
`materialize_packed(&self.plan, buffer)` then runs validation on the
materialized `Value`. **Validation path through phase 2:** until
phase 7 (ValidationPlan), `validate_bytes` reconstructs a `BastDoc`
for the validator only — the *read* path uses the plan, the
*validation* path uses `BastDoc`. This is a temporary bridge: phase 7
replaces it with a `ValidationPlan` walk (ADR-012 §3), retiring the
per-call `BastDoc` reconstruction. The bridge is acceptable for
phases 2–6 because the ValidationPlan work is committed in this
release (not deferred), so the bridge has a known removal point in
phase 7.
**Verification:** `cargo test --release` — the existing
`sequential_reader.rs` and `materialize.rs` tests drive `read_next`/
`read_field`/`reset`/`validate_bytes` through the public API, so they
validate the rewrite without modification. If any test breaks, the
rewrite diverged from the existing behavior — investigate before
patching the test. `cargo clippy --all-targets -- -D warnings`.
`cargo build --target wasm32-unknown-unknown --release`. **Re-run the
alktty `wire_vs_bast` bench** to confirm the 400x gap closes (this is
the headline result; record the before/after numbers in the commit
message).
---
## Phase 3 — Owned `BastDoc` (ADR-012 §2a) — **DONE (2026-09-02)**
> **Status: implemented.** Every `Bast*` type dropped its `<'a>`:
> `&'a str` → `String`, `&'a Value` → `Value` (deferred decision 1
> resolved: plain `String`/`Value` — the tree is built once, name
> sharing via `Arc<str>` needs a bench justification that doesn't
> exist). `BastDoc::new(&Value, &str)` still takes references in and
> clones into owned storage; `BastDoc` gained `Clone`. `resolve_*`
> return owned `BastDef`/`BastType`. Consumers adapted:
> `OffsetMap::compute(&BastDoc)`, `materialize_aligned(&BastDoc, ...)`
> (no lifetime), `BuildCtx`/`ComputeCtx` hold `&'d BastDoc`.
> **Engine ownership flip:** `AlkTypeEngine` now holds the owned
> `BastDoc` (replacing `bast_doc: Value` + `root_name: String` —
> `root_name()` delegates to the doc), killing its three per-call
> `BastDoc::new` re-parses (`validate_bytes` aligned path,
> `read_field`, `write_field` — review #004 M1 sites by construction;
> phase 5 removes the `lookup_leaf_field` walk itself). New public
> accessor `AlkTypeEngine::root_name()` (additive). **Bonus cleanup:**
> `materialize_typeref_packed`'s dead `_field` param dropped (phase 2
> left it dangling); under ownership it would have forced a deep
> `Value` clone per array element/record value/union variant via
> `dummy_field_for` — the param, `dummy_field_for`, and `ty_source`
> are gone (no behavior change; the `_field` arg was already
> ignored). `BastField::synthetic` retains an owned-signature
> `#[allow(dead_code)]` definition (no remaining callers; kept for the
> phase-4/5-aligned materializer helpers if they need it). Existing
> `BastDoc` consumers (`LayoutBuilder`'s `doc_value` re-parse cache)
> are unchanged pending phase 4. Engine stays `Send + Sync` with the
> owned doc — the engine's thread-share test now asserts it directly.
**Goal:** Make `BastDoc` own its data (drop the `<'a>` lifetime).
`&'a str` → `String`, `&'a Value` → `Value` (or `Arc<str>`/`Arc<Value>`
— deferred decision 1). This is the prerequisite for the `LayoutBuilder`
M1 fix (phase 4) and simplifies all owning consumers. Broad but
mechanical refactor.
**ADR reference:** [ADR-012 §2a](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2a-layoutbuilder--cache-the-parsed-bastdoc-at-new).
**Files:** `src/bast.rs` (every `Bast*` type), every consumer:
`src/layout_builder.rs`, `src/offset_map.rs`, `src/materialize.rs`,
`src/bast_validation.rs`, `src/engine.rs`, `src/tunion.rs`,
`src/bast_meta.rs` (if it walks `BastDoc`), `src/builder.rs` (if it
consumes `Bast*`). `src/lib.rs` re-exports (signatures change but
names stay).
**Implementation notes:**
- `BastDoc<'a>` → `BastDoc`. Fields: `root: Value` (was `&'a Value`),
`root_name: String` (was `&'a str`), `root_def: BastDef` (was
`BastDef<'a>`). `new(root: &Value, root_name: &str)` takes references
*in* (the caller still owns the input `Value`) but clones into owned
storage. The `&Value` → `Value` clone is the cost of ownership; it
happens once at `compile`/`new`, not per-field.
- `BastDef<'a>` → `BastDef`: `name: String`, `kind: BastDefKind`,
`source: Value`.
- `BastStruct<'a>` → `BastStruct`: `endian: Endian`, `align: Option<usize>`,
`fields: Vec<BastField>`, `source: Value`.
- `BastField<'a>` → `BastField`: `name: String`, `ty: BastType`,
`endian: Option<Endian>`, `align: Option<usize>`,
`encoding: VariableEncoding`, `max_length: Option<usize>`,
`source: Value`. `synthetic` constructor takes owned `BastType` +
`Value`.
- `BastType<'a>` → `BastType`: `Primitive(AlkTypeKind)`,
`Ref(BastRef)`, `Array(BastArray)`, `Record(BastRecord)`,
`Struct(BastStruct)`, `Union(BastUnion)`, `Enum(BastEnum)`.
- `BastUnion<'a>` → `BastUnion`: `endian: Endian`,
`discriminator: BastDiscriminator`, `fields: Vec<BastField>`,
`mapping: Vec<(String, BastType)>` (was `Vec<(&'a str, BastType)>`),
`source: Value`.
- `BastDiscriminator::Field { name: String }` (was `name: &'a str`).
- `BastEnum<'a>` → `BastEnum`: `values: Vec<String>` (was
`Vec<&'a str>`), `source: Value`.
- `BastArray<'a>` → `BastArray`: `element: Box<BastType>`, `count: usize`,
`source: Value`.
- `BastRecord<'a>` → `BastRecord`: `values: Box<BastType>`,
`source: Value`.
- `BastRef<'a>` → `BastRef`: `name: String`.
- **`resolve_typeref` / `resolve_ref` / `lookup_def`** now return owned
`BastType`/`BastDef`/`Value` instead of borrowed. The `clone()` in
the current `resolve_typeref` passthrough (`other => Ok(other.clone())`)
is no longer needed for the borrow case (everything is owned) but
the logic is unchanged — `BastType` is `Clone` either way.
- **Consumers adapt:** any code that held `&'a Value` alongside a
`BastDoc<'a>` (e.g. `LayoutBuilder.doc_value`, `AlkTypeEngine.bast_doc`,
`SequentialReader.doc_value` — though the reader is already on
`ReadPlan` after phase 2) drops the separate `Value` and holds the
owned `BastDoc` directly. `materialize_packed` is already on
`ReadPlan` (phase 2) and doesn't need `BastDoc` — unaffected.
`materialize_aligned` takes `&BastDoc` (owned, no lifetime).
- **`Arc<str>` vs `String` (deferred decision 1):** default to
`String`. The typed tree is built once; name sharing via `Arc<str>`
is a micro-optimization not justified without a bench. If phase 4's
`LayoutBuilder` work shows name allocation is measurable, revisit.
**Verification:** `cargo test --release` — the existing `bast.rs` tests
are the primary validation (they exercise every parser path). All
`Bast*`-consuming tests must pass unchanged (they go through public
APIs that still take `&Value`/`&str` in, just return owned types out).
`cargo clippy --all-targets -- -D warnings`. `cargo doc --no-deps`.
`cargo build --target wasm32-unknown-unknown --release` (`bast.rs` is
wasm-relevant).
---
## Phase 4 — `LayoutBuilder` caches the owned `BastDoc` (ADR-012 §2a) — **DONE (2026-09-02)**
> **Status: implemented.** `LayoutBuilder` stores `doc: BastDoc` +
> `endian` (the `doc_value: Value` + `root_name: String` cache is
> gone); `new` parses once, `build` walks `&self.doc` — the
> `layout_builder.rs` re-parse (M1) is retired. The `build`-time
> root-is-struct re-check replaced its `unreachable!()` with a clean
> `Schema` error (AGENTS.md §3 never-panic; the invariant is
> unchanged — `new` already rejects non-struct roots). Boxing fallout:
> the builder now holds the full owned tree, so `Layout::Packed`
> boxes it (`builder: Box<LayoutBuilder>`) to keep the engine's
> `Layout` enum variant sizes balanced (clippy
> `large_enum_variant`); `layout_builder()` still returns
> `Option<&LayoutBuilder>` via auto-deref, public API unchanged.
**Goal:** `LayoutBuilder::new` parses the owned `BastDoc` once and
stores it; `build` reuses it. Removes the `layout_builder.rs:190`
re-parse (M1). Non-breaking from the public API perspective
(`new`/`build` signatures unchanged); the change is internal.
**ADR reference:** [ADR-012 §2a](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2a-layoutbuilder--cache-the-parsed-bastdoc-at-new).
**Files:** `src/layout_builder.rs`.
**Implementation notes:**
- `LayoutBuilder` currently stores `doc_value: Value` + `root_name:
String` + `endian: Endian`. After phase 3, it stores `doc: BastDoc`
(owned) + `endian: Endian`. `new` calls `BastDoc::new` once;
`build(&self, var_sizes)` uses `&self.doc` directly — no
`BastDoc::new` call inside `build`.
- The `BuildCtx<'d>` struct (currently `doc: &'d BastDoc<'d>`) becomes
`doc: &BastDoc` (no lifetime, or a single lifetime for the borrow
from `&self`). The walk logic is unchanged.
- The `doc_value: Value` clone is removed; the builder holds the owned
`BastDoc` directly. `endian` is read from the doc at `new` time
(already is).
**Verification:** `cargo test --release` — the existing
`layout_builder.rs` tests pass unchanged (they go through
`LayoutBuilder::new` + `build`). `cargo clippy --all-targets -- -D
warnings`. `cargo build --target wasm32-unknown-unknown --release`.
---
## Phase 5 — `OffsetMap` carries `LeafMeta` (ADR-012 §2b) — **DONE (2026-09-02)**
> **Status: implemented.** Prerequisite first (review #005 M2):
> `Hash` added to `Endian`/`VariableEncoding` derives in `schema.rs`
> (additive; both are fieldless `Eq` enums). New public types
> `LeafMeta { kind, encoding, endian }` (`Copy + PartialEq + Eq +
> Hash`) and `OffsetEntry { range, meta }` (with `start()`/`end()`
> convenience accessors), both re-exported from `lib.rs`; `ByteRange`
> gained `Hash` (additive). Storage is `Vec<(String, OffsetEntry)>`
> (deferred decision 2 resolved: struct — `get` returns
> `Option<&OffsetEntry>`, `iter` yields `(&str, &OffsetEntry)`).
> `LeafMeta` is computed at `compute` time with **effective** endian
> threaded through the walk: container default → field override per
> field, propagated into nested-struct probes and array elements via
> the referring field (the same propagation the aligned materializer
> uses). **Parity note:** this replaces `engine.rs`'s
> `lookup_leaf_field` walk, which computed nested-struct defaults from
> the *nested struct's own* `endian` annotation — the two paths
> diverged whenever a nested struct declared `endian` and its
> referring field also declared one (the map now agrees with the
> aligned materializer and the packed `ReadPlan`; the old divergence
> was unreachable through `read_field` only when a nested annotation
> existed, and no test pinned it). `read_field`/`write_field` dispatch
> on the entry's `LeafMeta` — the `BastDoc` re-parse +
> `lookup_leaf_field`/`LeafFieldInfo` per access are gone (the last
> two M1 sites, engine.rs `read_field`/`write_field`). Behavior
> change: `read_field` on a path absent from the map (e.g. a
> whole-struct field path) now errors with `Offset` ("field not found
> in offset map") instead of `Access` ("does not support composite
> types") — the composite-path test already accepted either variant.
> `materialize_aligned`'s four `offset_map.get` call sites updated to
> `.range.start`. `alktty`/`alkcall` untouched (the bench never uses
> `OffsetMap::get`; alkcall has no dependency yet).
**Goal:** Extend `OffsetMap`'s entries with `LeafMeta { kind, encoding,
endian }` computed at `compute` time. `read_field`/`write_field` drop
the `BastDoc::new` + `lookup_leaf_field` calls (M1 aligned-side).
Closes the last two M1 sites (`engine.rs:334,467`).
**ADR reference:** [ADR-012 §2b](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2b-offsetmap--carry-leaf-metadata).
**Files:** `src/offset_map.rs` (extend entries, compute `LeafMeta`),
`src/engine.rs` (rewrite `read_field`/`write_field` to use the map's
`LeafMeta`, remove `lookup_leaf_field` + `LeafFieldInfo`), `src/lib.rs`
(re-export `LeafMeta`).
**Implementation notes:**
- **`Hash` on `Endian`/`VariableEncoding` (resolves review #005 M2 —
do this first, it's a prerequisite):** `src/schema.rs:205` (`Endian`)
and `:212` (`VariableEncoding`) currently derive only `Debug, Clone,
Copy, PartialEq, Eq` — no `Hash`. `LeafMeta` (below) requires all
its fields to be `Hash` for `#[derive(Hash)]`, and phase 6's
`#[derive(Hash)]` on `ReadPlan`/`OffsetMap` requires `FieldPlan`'s
`endian: Endian` + `encoding: VariableEncoding` to be `Hash`. Add
`Hash` to both derives in `src/schema.rs`. Both are fieldless enums
already at `Eq + PartialEq`, so this is additive and semver-safe —
no behavioral change. Trivial, but it's an unstated prerequisite
the original plan omitted.
- New public type `LeafMeta { kind: AlkTypeKind, encoding:
VariableEncoding, endian: Endian }`. `Copy + PartialEq + Eq + Hash`
(all fields are `Copy + Hash` once the sub-step above is done —
`AlkTypeKind` already derives `Hash`; `Endian`/`VariableEncoding`
get it from the sub-step above).
- `OffsetMap` storage: `fields: Vec<(String, ByteRange, LeafMeta)>`
(was `Vec<(String, ByteRange)>`). The `compute` walk already resolves
each leaf's type; add the `LeafMeta` extraction at the point where
the leaf `ByteRange` is recorded.
- **`OffsetMap::get` return type (deferred decision 2):** change to
`get(field_path) -> Option<&OffsetEntry>` where `pub struct
OffsetEntry { range: ByteRange, meta: LeafMeta }`. Add
`OffsetEntry` to `lib.rs` re-exports. Callers that used
`map.get(path).unwrap().start` become
`map.get(path).unwrap().range.start`. Update `alktty`/`alkcall` call
sites (in-house).
- `engine.rs::read_field`/`write_field`: drop the
`BastDoc::new(&self.bast_doc, &self.root_name)?` +
`lookup_leaf_field(&doc, field_path)?` calls. Read `LeafMeta`
from `offset_map.get(field_path)?.meta`. The `kind`/`encoding`/
`endian` match arms in `read_field`/`write_field` are unchanged
(they already dispatch on `AlkTypeKind`/`VariableEncoding`/`Endian`).
- Remove `LeafFieldInfo` and `lookup_leaf_field` from `engine.rs`
(subsumed by `LeafMeta` on the map).
- `materialize_aligned` also uses `OffsetMap` — it currently calls
`offset_map.get(&path)?.start` for leaf reads. Update those call
sites to `.range.start`. The materializer's `resolve_typeref` calls
for composite walks stay (composites aren't in the offset map as
leaves; they're walked recursively). The `BastDoc` argument to
`materialize_aligned` is now owned (phase 3) — no signature change
beyond the lifetime drop.
- **Scope Boundary — aligned `materialize`'s `BastDoc` structure walk
(resolves review #005 L3):** `materialize_struct_aligned`
(`materialize.rs:451-521`) walks `BastDoc` to traverse
struct/array/record *structure*, using `OffsetMap` only for leaf
byte positions. This is the **permanent design for 0.3.0**, not a
deferral: after phase 3 the walk is over owned data (no re-parse,
not O(N²)), and aligned `validate_bytes` is one structure walk per
call (not per-field), so there is no perf driver analogous to review
#004's packed per-chunk gap. ADR-011 §"Out of scope" is half-true
here (aligned materialize takes `&OffsetMap` *and* `&BastDoc`) —
this note owns the decision: aligned materialize keeps walking owned
`BastDoc` for structure through 0.3.0. An `AlignedPlan` that
compiles the structure walk is **not** in scope; if a future bench
shows an aligned-mode hot loop, it gets its own ADR (tracked as an
open question, not a silent gap). Phase 7's `ValidationPlan` does
not change this — validation is value-domain, orthogonal to the
aligned structure walk.
**Verification:** `cargo test --release` — existing `offset_map.rs`
and `engine.rs` `read_field`/`write_field` tests pass (they go through
public APIs). `cargo clippy --all-targets -- -D warnings`. `cargo doc
--no-deps` (new public `LeafMeta`/`OffsetEntry`). `cargo build --target
wasm32-unknown-unknown --release`.
---
## Phase 6 — Fingerprinting `ReadPlan`/`OffsetMap` (ADR-012 §1, §4) — **DONE (2026-09-02)**
> **Status: implemented.** `#[derive(Hash, Eq)]` added to `ReadPlan`,
> `FieldPlan`, `CompositePlan`, `ReadKind`, `DiscriminatorPlan`
> (`ReadPlan`'s `schema: Arc<Value>` hashes fine — `serde_json::Value:
> Hash + Eq` under the pinned `preserve_order` serde_json) and to
> `OffsetMap` (`Clone` added alongside; its `LeafMeta`/`OffsetEntry`/
> `ByteRange` payload gained `Hash` in phase 5 / this phase). The
> POC's `VariantPlan`/`VariantKind` don't exist in the production
> shape (phase 1 dropped them). `fingerprint() -> u64` on both via
> `DefaultHasher` (deferred decision 3 resolved: std `DefaultHasher`,
> no new dep; the fingerprint isn't hot; cross-version stability is a
> non-goal per ADR-012). Fingerprint contract tests on both: same
> schema twice → equal `PartialEq` + equal fingerprint; field-kind
> change, field-order change, and endianness change each → different
> fingerprints; (ReadPlan) different root names over the same document
> → different fingerprints. `ValidationPlan` already carries its own
> `Hash + Eq` + `fingerprint` + contract test (phase 7).
**Goal:** Add `Hash + Eq` derives + `fingerprint() -> u64` to `ReadPlan`
and `OffsetMap`. Enables cross-run caching, `alkcall` schema handshake,
schema-version diagnostics. (`ValidationPlan` gets the same treatment
in phase 7, where it's built — it carries its own `Hash + Eq` +
`fingerprint()` as part of its public surface.)
**ADR reference:** [ADR-012 §1](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#1-fingerprinting--readplan-hash--eq-offsetmap-hash--eq),
[ADR-012 §4](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#4-fingerprinting-offsetmap-bundled-with-2b).
**Files:** `src/read_plan.rs` (derives + `fingerprint`), `src/offset_map.rs`
(derives + `fingerprint`), `src/lib.rs` (no new re-exports — `Hash`/`Eq`
are trait derives, `fingerprint` is an inherent method).
**Implementation notes:**
- `ReadPlan` already uses `BTreeMap` for `by_name` (phase 1), so
`#[derive(Hash, Eq)]` works. Add it alongside the existing
`Debug, Clone, PartialEq`. Same for `FieldPlan`, `CompositePlan`,
`ReadKind`, `DiscriminatorPlan`. (The POC's `VariantPlan`/`VariantKind`
are not in the production shape — phase 1 dropped them — so they are
not derived here.) `ReadPlan` also carries `schema: Arc<Value>` from
phase 2; `Arc<Value>: Hash + Eq` because `serde_json::Value: Hash +
Eq` (with `preserve_order`, `Map::hash` sorts keys deterministically),
so the `schema` field does not block the derive. If a future
`serde_json` version regresses `Value: Hash`, exclude `schema` from
the derived `Hash` via a manual `impl Hash for ReadPlan` that hashes
every field except `schema` — phase-6 concern, not a blocker.
- `OffsetMap` already carries `LeafMeta` (phase 5), and `LeafMeta` is
`Copy + Hash + Eq` (phase 5 added `Hash` to `Endian`/`VariableEncoding`).
Add `#[derive(Hash, Eq)]` to `OffsetMap`,
`OffsetEntry`, `ByteRange` (already `Eq + Hash`), `LeafMeta`.
- **Fingerprint hasher (deferred decision 3):** `DefaultHasher` (std,
no new dep). The fingerprint isn't hot; cross-version stability is a
non-goal. `fingerprint()`:
```rust
pub fn fingerprint(&self) -> u64 {
use std::hash::{Hash, Hasher};
let mut h = std::hash::DefaultHasher::new();
self.hash(&mut h);
h.finish()
}
```
- **Fingerprint contract test:** compile the same schema twice, assert
`plan1 == plan2` and `plan1.fingerprint() == plan2.fingerprint()`.
Compile a schema with one field changed, assert fingerprints differ.
This is the contract test for ADR-012 §1's "two plans with equal
hashes produce identical reads over identical bytes."
**Verification:** `cargo test --release` (new contract tests).
`cargo clippy --all-targets -- -D warnings`. `cargo doc --no-deps`.
---
## Phase 7 — `ValidationPlan` (ADR-012 §3) — **DONE (2026-08-31)**
> **Status: implemented.** The design session ran and the shape landed
> in `src/validation_plan.rs`. Summary of what was decided and built
> (full detail in ADR-012 §3a):
>
> - **Shape:** `ValidationPlan { root: ValidNode }` — a constraint tree
> with one arm per value-domain check (`Int`/`I64`/`Uint`/`U64`/
> `Float`/`Bool`/`Str`/`Bytes`/`Enum`/`Struct`/`Union`/`Array`/
> `Record`), `ValidField { name, node }`, `ValidVariant { key, node }`.
> `maxLength` baked into leaf nodes from the owning field at compile
> time. Unions compile to variant nodes only (the declared union
> fields are validated via the variant walk, matching the interpretive
> arm's dispatch-on-`__discriminator` semantics).
> - **`compile` signature:** `compile(&BastDoc) -> Result<Self,
> AlkTypeError>` (the plan table's `&str` param was vestigial).
> - **`bast_validation`:** the interpretive walker is *retired*
> (deleted, not just bypassed); `validate_value` survives as a
> compile-once-per-call wrapper over the plan (one-shot/diagnostic
> use); shared error helper retained.
> - **Engine:** `Arc<ValidationPlan>` built at `compile` in both modes;
> new accessor `validation_plan()`. `validate_bytes` walks the plan.
> **Bonus:** `ValidationPlan::compile` runs before the layout build and
> serves as the engine's cyclic-`$ref` gate. (As of the review #006 H2
> fix, the layout walkers also carry their own guard —
> `walk_guard::check_ref_graph` runs at each standalone entry — so this
> ordering is now belt-and-suspenders rather than the only defense; the
> plan text below predates that fix.)
> - **Remaining phase-7 bench work** (a `validate_bytes`-stream bench
> in alktty) moves with the bench work into phase 8; a spot check
> during development measured plan-validate at ~0.2 µs/call vs ~0.6
> µs for the compile-per-call one-shot it replaced.
**Goal:** Retire the interpretive `bast_validation` walk. Introduce a
`ValidationPlan` — a compile-once-walk-many compiled form over the
BAST document's value-domain constraints — built once at `compile`
time and walked by `validate_bytes` (both modes) per buffer instead
of re-walking `BastDoc`. Closes the latent perf cliff review #005 M3
flagged: after ADR-011 the *read* half of `validate_bytes` is
plan-fast, but the *validation* half still re-walks `BastDoc` per
buffer, which is hot on the `alkcall` read+validate-on-untrusted-stream
common case.
**ADR reference:** [ADR-012 §3](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#3-validationplan--compile-once-validation-form).
**Predecessor for this phase:** a **design session** to scope the
concrete `ValidationPlan` shape (constraint representation,
`compile`/walk structure, `bast_validation` public-surface review)
**before** implementation begins. ADR-012 §3 fixes the decision (in
0.3.0, compiled form, no per-buffer `BastDoc` walk, `Hash + Eq` +
`fingerprint`) and lists what is *not* decided (the struct/enum
shape, the constraint descriptors, whether `validate_value` is
retired or kept as a wrapper). This phase implements whatever the
design session scopes; the contract below holds regardless of shape.
**Files:** New `src/validation_plan.rs` (the `ValidationPlan` type,
`compile`, walk entry points). `src/bast_validation.rs` (adopt the
plan; `validate_value` either becomes a thin wrapper over the plan
or is retired per the design session's call). `src/engine.rs`
(`compile` builds `Arc<ValidationPlan>` in both modes, stores it;
`validate_bytes` walks `&self.validation_plan` instead of
reconstructing a `BastDoc` for the validator — this removes the
temporary bridge from phase 2). `src/lib.rs` (re-export
`ValidationPlan`).
**Implementation contract (fixed by ADR-012 §3, independent of
shape):**
- **Compile-once-walk-many.** `ValidationPlan::compile` walks the
owned `BastDoc` once (phase 3 made it owned); `validate_bytes`
walks the `ValidationPlan` per buffer, never `BastDoc`. The only
consumers that walk `BastDoc` interpretively after this phase are
the one-shot `*::compile` paths (`ReadPlan::compile`,
`OffsetMap::compute`, `ValidationPlan::compile`,
`LayoutBuilder::new`).
- **Value-domain, not byte-position.** The plan carries constraint
descriptors (enum allowed-sets, integer range bounds, `maxLength`
caps, union variant keys, and any other value-domain checks
`bast_validation` performs today), keyed for dispatch against the
materialized `Value` tree. The shape is different from
`ReadPlan`/`OffsetMap`; the pattern (compiled form, immutable,
shared via `Arc`) is the same.
- **`Send + Sync`.** `ValidationPlan: Send + Sync` (immutable owned
data, no interior mutability) so `Arc<ValidationPlan>` shares from
the `Send + Sync` engine. Add a `static` bound assertion test
mirroring phase 1's `read_plan_is_send_sync`.
- **`Hash + Eq` + `fingerprint()`.** `ValidationPlan` derives
`Debug, Clone, PartialEq, Eq, Hash` and has
`fingerprint() -> u64` (same `DefaultHasher` implementation as
phase 6). The fingerprint contract generalizes: two validation
plans with equal hashes accept/reject identical `(bytes)`
identically. Add a fingerprint contract test (compile the same
schema twice, assert `plan1 == plan2` and `plan1.fingerprint() ==
plan2.fingerprint()`; change one constraint, assert fingerprints
differ).
- **Untrusted-input discipline.** `compile` surfaces malformed
schemas as `AlkTypeError::Schema` (AGENTS.md §3); overflow-safe
arithmetic (AGENTS.md §4). No `unsafe`, no `async`, no new deps,
wasm-clean (AGENTS.md §5–§11).
**What this phase does *not* include (shape-dependent, scoped by the
design session):** the concrete `ValidationPlan` struct/enum, the
constraint-descriptor representation, the `bast_validation`
public-surface decision (`validate_value` retire-vs-wrapper), and any
`AlkTypeError::Validation` variant changes. These are shape questions
the design session resolves; they are *not* a re-opening of the
"ship in 0.3.0" decision, which is fixed in ADR-012 §3.
**Verification:** `cargo test --release` — the existing
`bast_validation.rs` and `engine.rs` `validate_bytes` tests are the
primary validation (they drive validation through the public API and
must pass unchanged, confirming behavioral parity with the
interpretive walk). New unit tests for `ValidationPlan::compile`
covering every constraint kind. New `Send + Sync` assertion test.
New fingerprint contract tests. `cargo clippy --all-targets -- -D
warnings`. `cargo doc --no-deps` (new public type). `cargo build
--target wasm32-unknown-unknown --release` (`validation_plan.rs` is
wasm-relevant). **Re-run the alktty `wire_vs_bast` bench** and, if
the design session scopes one, a `validate_bytes`-on-untrusted-stream
bench alongside `wire_vs_bast` to confirm the validation half of
`validate_bytes` no longer dominates per-buffer.
---
## Phase 8 — Public API bump, docs, verification (ADR-011 step 6, ADR-012) — **DONE (2026-09-02)**
> **Status: implemented.** Version flipped 0.2.0 → 0.3.0; `lib.rs`
> re-exports complete (`ReadPlan` + sub-types, `LeafMeta`,
> `OffsetEntry`, `ValidationPlan` + sub-types from earlier phases).
> Docs: ADR-007 "Cost" section rewritten to the `Arc<ReadPlan>` cost
> (15.7 ns) with the old framing as a historical note (review #004 L2
> closed); ADR-011/012 status blocks flipped to implemented; the
> architecture README ADR table rows updated; `layout-engine.md`
> rewritten for the 0.3.0 surface (`SequentialReader` construction via
> the engine factory, `OffsetMap` `OffsetEntry`/`LeafMeta`/`fingerprint`
> public-types section, `OffsetMap::compute(&BastDoc)` owned signature);
> `SequentialReader` module doc now points at the engine factory.
> Reviews #004 and #005 status flipped to closed. Bench (alktty
> `wire_vs_bast`, re-run on the 0.3.0 tree): read p64 98 ns/chunk
> (hand-rolled 5.7 µs/stream — parity held from phase 2), read p4k
> unchanged, `alktype_layout_build` **180 ns** (was ~1.2 µs — the
> phase-4 owned-doc cache removed the per-build re-parse, ~7x),
> `sequential_reader_new` 15.7 ns (unchanged), write p64 −3%
> (37.9 µs), `engine_compile` 590 µs (unchanged; dominated by
> meta-schema validation). No `validate_bytes`-stream bench was added:
> the phase-7 spot check (~0.2 µs/call plan-validate vs ~0.6 µs
> compile-per-call) stands as the validation-half measurement; a
> dedicated bench remains a follow-up if `alkcall` profiling motivates
> it. Downstream: `alktty` compiles against the path dep unchanged
> (the bench uses `LayoutBuilder::new`/`build` and
> `engine.sequential_reader()` — no touched signatures);
> `alkcall` has no dependency yet.
**Goal:** Flip the version to 0.3.0, update `lib.rs` re-exports, update
the architecture docs (ADR-007 "Cost" rewrite, ADR-011/012 status flip
if not already, README ADR table), update in-house downstream
consumers, run the full verification block.
**ADR reference:** [ADR-011 §Public API change](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#public-api-change-breaking--version-bump-to-030),
[ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md).
**Files:** `Cargo.toml` (version 0.2.0 → 0.3.0), `src/lib.rs`
(re-export `ReadPlan`, `LeafMeta`, `OffsetEntry`, `ValidationPlan`),
`docs/architecture/` (README ADR table, ADR-007 "Cost" section,
ADR-011/012 status), `docs/architecture/validation.md` /
`layout-engine.md` (mention `ReadPlan`/`LeafMeta`/`ValidationPlan`
where relevant), in-house downstream repos (`alktty`, `alkcall` —
update call sites for `OffsetMap::get`, `SequentialReader::new`,
`materialize_packed`, `BastDoc` owned, `validate_bytes` internal
change if any signature change surfaced in phase 7's shape work).
**Implementation notes:**
- **ADR-007 "Cost" section (L2 from review #004):** rewrite the
"re-parse on demand" paragraph to describe the `Arc<ReadPlan>` cost
and the owned-`BastDoc` cache. The factory decision itself stays
"Accepted." This is the last loose end from review #004.
- **`src/engine.rs:112-115` doc comment (L2):** rewrite the "re-parse
the typed tree on demand" comment to describe the compiled-form
architecture (`ReadPlan` for packed reads, `OffsetMap`+`LeafMeta`
for aligned, `ValidationPlan` for validation, owned `BastDoc` for
the builder and the `*::compile` paths).
- **`lib.rs` re-exports:** add `ReadPlan`, `LeafMeta`, `OffsetEntry`,
`ValidationPlan`. `BastDoc` and `Bast*` stay re-exported (signatures
changed in phase 3, names unchanged). `materialize_packed`/
`materialize_aligned` stay re-exported (signatures changed).
`SequentialReader` stays re-exported (`new` signature changed).
- **Downstream updates:** `alktty`'s bench (`benches/wire_vs_bast.rs`)
updates `SequentialReader::new` call + any `OffsetMap::get` usage.
`alkcall` updates similarly. Both are in-house path dev-deps; the
updates ride this release's commits (or follow-on commits in those
repos — they're separate repos, but the path dev-dep means a local
update is immediate).
- **`Cargo.toml` version bump:** `0.2.0` → `0.3.0`. The workspace
section added for the POC (`[workspace] members = ["poc/readplan"]`)
stays on the `readplan-poc` branch and is *not* merged to main — the
POC branch is derisking-only, like `bast-validator-poc`. If the POC
files ever merge to main, drop the workspace section (the POC is
disposable).
**Verification block (run all, all must pass):**
```bash
cargo test --release # full suite
cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # new public types
cargo build --target wasm32-unknown-unknown --release # wasm-clean
cargo publish --dry-run --allow-dirty # before publish
```
Plus: **re-run the alktty `wire_vs_bast` bench** and record the
before/after numbers in the release commit message. The 400x gap
should close to within ~2–5x of hand-rolled (the `data_access` calls
are the same; the remaining gap is the `match` dispatch + `Arc` refcount
vs hand-rolled's direct calls). The SFTP-shaped union case (the one
ADR-011's framing argument cared about) should close further because
eager `$ref` resolution removes the `resolve_typeref_as_def` per-
variant dispatch cost.
---
## Cross-phase invariants
- **The tree builds and tests pass at every phase boundary.** No phase
leaves the crate in a non-compiling state. Phases 1 (add `ReadPlan`),
6 (add `Hash`/`Eq` derives to `ReadPlan`/`OffsetMap`), and 7 (add
`ValidationPlan`) are pure additions; phases 2–5 are rewrites that
must leave tests green; phase 8 is the bump/docs.
- **The POC on `readplan-poc` is the reference scaffold for phases 1–2.**
It is *not* merged to main; it stays on the branch as the derisking
record, like `bast-validator-poc`. If a phase 1–2 implementation
question arises about the plan shape, consult the POC. (The POC's
`VariantPlan`/`VariantKind` and its nested-union rejection are
**not** carried forward — phase 1's refined `CompositePlan::Union`
shape supersedes both; see phase 1.)
- **Review #004 is the closure target.** H1 → phase 2; packed M1 →
phase 2; aligned M1 → phases 4–5; L1 → phase 2 (falls out); L2 →
phase 8 (doc rewrite). The review's status flips to "closed" in the
phase 8 commit.
- **Review #005 is the closure target for the plan-spec issues.** H1
→ phase 1 (refined union shape in ADR-011 + plan); H2 → phase 2
(`Arc<Value>` on the plan); M1 → phase 1 (nested unions via
`CompositePlan` recursion, no behavioral drop); M2 → phase 5 (`Hash`
on `Endian`/`VariableEncoding`); M3 → ADR-012 §3 + phase 7
(`ValidationPlan` shipped in 0.3.0, deferral reversed); L1/L2/L3 →
phase 2 / phase 5 Scope Boundary; N1 → typo; N2 → phase 1
`Send + Sync` assertion test; N3 → Semver Contract table row.
- **No `unsafe`, no `async`, no new deps, no feature flags** (AGENTS.md
§5–§11). The owned-`BastDoc` refactor uses `String`/`Value`, not
`unsafe` self-referential tricks. `DefaultHasher` is std. Wasm-clean
throughout. `ValidationPlan` follows the same constraints.
- **`preserve_order` stays load-bearing** (AGENTS.md §8). The owned-
`BastDoc` refactor must not sort schema object keys anywhere; field
order in the `Value` still determines byte order in packed mode and
iteration order in both modes. `ValidationPlan::compile` inherits
this — value-domain checks that depend on field ordering (e.g. union
discriminator field lookup) respect `preserve_order`.
## What this plan is *not*
- **Not a disk-cache or wire-protocol spec.** The fingerprint contract
and method are in scope (phase 6 for `ReadPlan`/`OffsetMap`, phase 7
for `ValidationPlan`); downstream uses are the consumers' concern.
- **Not cross-version fingerprint stability.** Within-version only
(ADR-012). The fingerprint may change across versions if a new
`AlkTypeKind` variant is added; consumers cache within a version.
- **Not a perf bench.** The bench lives in alktty; this plan re-runs it
at phase 2 and phase 8 to confirm the gap closes. The plan itself
only asserts correctness/coverage.
- **Not an `AlignedPlan`.** Aligned `materialize`'s `BastDoc` structure
walk is the permanent 0.3.0 design (phase 5 Scope Boundary). An
`AlignedPlan` is out of scope; if a future bench motivates one, it
gets its own ADR.
+549
View File
@@ -0,0 +1,549 @@
---
status: complete
created: 2026-08-15
last_updated: 2026-08-15
---
# BAST Pivot — Implementation Plan
**Status: complete.** All 10 steps are implemented and pushed to
`origin/main` (steps 1–8 in commits `66ab9d7` → `54fd112`; step 9 was
a no-op — steps 4–8 converted the tests as they went, leaving only the
intentional `from_bast_str` rejection test referencing the
`"AlkType:Uint32"` string; step 10 synced the architecture docs and
ADRs in this commit). The two new ADRs
([ADR-BAST](../architecture/decisions/bast-bast-format.md),
[ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md))
record the decisions; the amended ADRs (001, 002, 003, 004, 009, 010)
carry supersession/amendment notes. The research record
([`bast-pivot.md`](../research/bast-pivot.md)) is flipped to
`implemented`. What follows is the original plan, preserved as the
historical execution record.
---
This is the execution plan for the BAST pivot: replacing alktype's
v0.1.0 `AlkType:*` custom-keyword JSON Schema format with the BAST
(Binary Abstract Syntax Tree) format. It is the **entry point** an
implementing agent reads first.
Companion documents:
- [`docs/architecture/bast-format.md`](../architecture/bast-format.md) —
the normative BAST format spec (meta-schema, TypeRef, examples,
validation model). Read this for *what* the format is.
- [`docs/research/bast-pivot.md`](../research/bast-pivot.md) — the
research record: motivation, POC scope and result, decisions
D-BAST-001..009, risks. Read this for *why* and *what was proved*.
The POC lives on branch `bast-validator-poc` (commit `f371fe4`) as
`src/bast_poc.rs` — reference scaffolding, deliberately not merged.
**Working order:** read this plan top-to-bottom. The Semver Contract
section is the scope-creep guardrail — consult it before each step.
Each step links to the specific spec section it implements and the
relevant D-BAST-* decision anchor. Implement steps in order; each step
lists its verification gate.
## Semver Contract
The crate is on crates.io at 0.1.0 with zero real consumers, so a
breaking bump is free — but the contract is explicit so the
implementation doesn't drift. Per AGENTS.md, the 0.1.0 public surface
is the items re-exported from `src/lib.rs`. This table is the
authoritative scope-creep guardrail for the pivot.
| Public item (from `lib.rs` re-exports) | Class | Change |
|---|---|---|
| `AlkTypeKind` (enum + variants + methods) | **Additive** | Unchanged. 19 variants, same methods. New `from_str()`/`to_str()` mapping for lowercase BAST kind strings (`"uint32"` ↔ `AlkTypeKind::Uint32`) — additive methods. |
| `Endian`, `VariableEncoding`, `DiscriminatorKind` | **Unchanged** | — |
| `AlkTypeEngine::compile` | **Breaking** | Signature: `compile(schema: &mut Value, mode)` → `compile(bast_doc: &Value, root_name: &str, mode)`. Adds required `root_name` param (D-BAST-001); drops `&mut` (BAST needs no in-place `normalize_refs`); input is a BAST document, not a custom-keyword JSON Schema. |
| `AlkTypeEngine::validate_json` | **Breaking (behavioral)** | Signature unchanged `(instance: &Value) -> Result<...>`, but the validator it runs is now a standard `jsonschema::Validator` from a consumer-provided JSON Schema, not a custom-keyword validator built from the alktype schema. The *contract* of what schema validates the instance changes. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (contract)** | Same signature. Internally the validation step switches from `jsonschema::Validator` to the BAST-native validator. Error type unchanged (D-BAST-009). |
| `AlkTypeEngine::is_valid_json` | **Breaking (behavioral)** | Same caveat as `validate_json` — validates against the consumer JSON Schema, not the alktype schema. |
| `AlkTypeEngine` accessors (`endian`, `mode`, `offset_map`, `layout_builder`, `sequential_reader`, `read_field`, `write_field`, etc.) | **Unchanged** | Layout-layer accessors are format-agnostic. |
| `LayoutMode`, `OffsetMap`, `ByteRange` | **Unchanged** | — |
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged** | — |
| `SequentialReader`, `FieldValue` | **Unchanged** | — |
| `UnionDispatch` | **Unchanged** | — |
| `data_access::*` functions | **Unchanged** | — |
| `AlkTypeError` (all 4 variants) | **Unchanged** | D-BAST-009 keeps `Validation(jsonschema::ValidationError<'static>)`. |
| `Schema` builder (`struct_`, `object`, `field`, `build`, all setters) | **Breaking (output format)** | Public method signatures unchanged. `build()` output changes from custom-keyword JSON to BAST JSON (for `struct_`) / standard JSON Schema (for `object`). Callers that introspect the built `Value` break; callers that pass it straight to `compile` are source-compatible once `compile` takes BAST. |
| `Definitions` builder (`new`, `define`, `define_value`, `build`, `merge_into`) | **Breaking (output format)** | Same as `Schema` — signatures unchanged, `build()`/`merge_into()` output shape changes to BAST `$defs`. |
| `Discriminator` builder enum | **Unchanged** | — |
| `build_validator` (from `validation`) | **Breaking (signature or removal)** | Currently `build_validator(schema: &Value) -> Result<jsonschema::Validator, AlkTypeError>` builds a custom-keyword validator. Under the pivot it either (a) is removed (consumers call `jsonschema` directly for standard JSON Schema) or (b) is repurposed to build a standard `jsonschema::Validator` from a consumer-provided standard JSON Schema (no custom keywords). Decision belongs to step 6. Either way the current signature's contract breaks. |
| `get_alktype_kind`, `get_alktype_kind_enum`, `get_alktype_kind_loose`, `get_alktype_kind_loose_enum`, `normalize_refs`, `inline_union_variant_refs`, `resolve_ref`, `resolve_ref_or_inline`, `parse_align`, `parse_discriminator`, `parse_encoding`, `parse_endian`, `parse_max_length` | **Breaking (removal or rework)** | All currently re-exported from `lib.rs`. `normalize_refs` and `inline_union_variant_refs` are removed (BAST needs neither). The `get_alktype_kind*` family is removed (replaced by direct `kind` parsing). The `parse_*` and `resolve_*` functions are reworked to read BAST properties instead of keyword-value objects, or removed if subsumed by the BAST parser. **Open: which of these stay public vs become internal.** Current leaning — drop all from `lib.rs` re-exports (they're engine-internal accessors, not consumer API); the BAST parser exposes a new typed surface instead. Confirmed during step 3. |
**Net breaking surface:** `compile`, `validate_json`/`is_valid_json`
(contract), `Schema::build`/`Definitions::build` (output format),
`build_validator` (signature/removal), and the ~13 `schema::*` helper
re-exports. **Net additive:** BAST parser, BAST-native validator,
`AlkTypeKind::from_str`/`to_str`. **Net unchanged:** the entire layout
+ data-access + materialize + tunion layer, `AlkTypeError`, the
`Discriminator` builder, `AlkTypeKind` variants.
### Decisions deferred to their implementation steps
These are small enough to decide when the step is reached, but are
flagged here so they don't become drive-by semver changes:
1. **`validate_json` JSON Schema source** (step 6): does the consumer
pass the JSON Schema to `compile` (engine carries a second
validator) or to `validate_json` at call time? The former preserves
the current single-call ergonomics; the latter is more flexible. Not
semver-relevant either way if `validate_json`'s signature can absorb
a new param or stay as-is — needs the call-site analysis.
2. **`build_validator` fate** (step 6): removed vs repurposed. If
repurposed, its signature stays but its contract (no custom
keywords) changes — a behavioral break, not a type break.
3. **`schema::*` helper re-exports** (step 3): drop from `lib.rs`
(engine-internal) vs keep public for consumers that walk schemas.
Leaning: drop — they're accessors for the old format, and the BAST
parser exposes a cleaner typed surface. Confirmed during step 3.
## Steps
### Step 1 — Add `AlkTypeKind::from_str`/`to_str` for BAST kind strings
**Goal:** Add the lowercase-string mapping (`"uint32"` ↔
`AlkTypeKind::Uint32`) that the BAST parser and validator dispatch on.
This is the additive-only, zero-risk foundation — no existing code
changes.
**Spec reference:** [bast-format.md §Primitives](../architecture/bast-format.md#primitives),
[D-BAST-002](../research/bast-pivot.md#d-bast-002-primitive-type-string-set).
**Files:** `src/schema.rs` (the `AlkTypeKind` impl block). No `lib.rs`
change needed — the methods are inherent on the already-re-exported
enum.
**Implementation notes:**
- `to_str(self) -> &'static str` returns the lowercase BAST string.
- `from_str(s: &str) -> Result<AlkTypeKind, AlkTypeError>` returns
`AlkTypeError::Schema` for unknown strings. This is a new inherent
method, distinct from the existing `FromStr` impl that parses the
v0.1.0 `"AlkType:Uint32"` keyword form. Do not remove the existing
`FromStr` yet — step 8 removes the v0.1.0 accessors.
- Cover all 14 primitive kinds plus `struct`, `union`, `array`,
`record`, `enum` (19 total, matching the enum variants). The
lowercase strings are in the [primitives table](../architecture/bast-format.md#primitives);
composite kinds are `"struct"`, `"union"`, `"array"`, `"record"`,
`"enum"`.
**Verification:** `cargo test --release` (new unit tests for the
mapping, both directions; existing tests unaffected). `cargo clippy
--all-targets -- -D warnings`.
---
### Step 2 — Embed the BAST meta-schema
**Goal:** Embed the BAST meta-schema as a `serde_json::Value` constant
in the crate, available for validating BAST documents at compile time
and for publishing at `https://alk.dev/bast/v1/schema`.
**Spec reference:** [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
**Files:** New `src/bast_meta.rs` (or a `const` in `src/schema.rs` —
match existing module conventions). Re-export the meta-schema `Value`
from `lib.rs` if consumers should be able to validate BAST documents
themselves (likely yes — additive, not semver-relevant).
**Implementation notes:**
- The meta-schema JSON is in [bast-format.md §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
Copy it verbatim into a `serde_json::json! {...}` macro invocation or
parse it from an embedded string via `serde_json::from_str`.
- No feature flags (AGENTS.md §6). The meta-schema is a compile-time
constant, no I/O.
- WASM-clean: no `include_str!` of an external file is needed if the
`json!` macro is used; either way is wasm-safe.
**Verification:** `cargo test --release`. `cargo build --target
wasm32-unknown-unknown --release` (meta-schema is a `Value` constant —
wasm-relevant). `cargo clippy --all-targets -- -D warnings`.
---
### Step 3 — BAST document parser
**Goal:** Implement the BAST document parser that the layout engines
and materializer use instead of the `get_alktype_kind*` custom-keyword
accessors. This is the natural entry point for the pivot — the largest
step, and the one the rest of the steps build on.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(the whole document — the parser implements the format spec).
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection),
[D-BAST-003](../research/bast-pivot.md#d-bast-003-top-level-defs-requirement),
[D-BAST-005](../research/bast-pivot.md#d-bast-005-field-name-discriminator-unions).
**Files:** New `src/bast.rs` (the parser). The existing `src/schema.rs`
stays for now — steps 4–8 migrate callers off it. Update `src/lib.rs`
to add `pub mod bast;` and re-export the parser's public surface.
**Implementation notes:**
- The parser reads `kind`/`fields`/annotation properties from BAST
nodes. It produces a typed surface (a small `BastNode` enum or
equivalent) that the layout engines, materializer, and validator can
walk without re-parsing the raw JSON at every node. The POC parsed
lazily from raw JSON in both passes to keep the model honest; a typed
tree is a straightforward follow-on optimization (POC observation 5).
Either is acceptable for the production version; the typed tree is
recommended since three consumers (layout, materialize, validate)
walk the same tree.
- `$ref` resolution: `#/$defs/<name>` only — a single hash lookup. No
`normalize_refs` (BAST refs are always full JSON Pointers), no
`inline_union_variant_refs` (union variant refs resolved lazily by
the validator and materializer). See [bast-format.md §TypeRef](../architecture/bast-format.md#typeref).
- Untrusted input: every path that walks a BAST document must return
`Err(AlkTypeError::Schema)` on a malformed document, never `panic!`/
`unreachable!` (AGENTS.md §3). The POC's
`malformed_document_produces_schema_error_not_panic` test is the
template.
- **Decide deferred decision #3 here:** drop the `schema::*` helper
re-exports from `lib.rs`, or keep them public. Leaning: drop. The
BAST parser exposes a cleaner typed surface; the v0.1.0 accessors
are engine-internal and not consumer API.
**Verification:** `cargo test --release` (port the POC's parser tests
— the malformed-document test, the type-ref resolution tests).
`cargo clippy --all-targets -- -D warnings`. The layout engines don't
use the parser yet (step 4 wires it in), so the existing suite still
passes on the old path.
---
### Step 4 — Wire `compile()` to accept a BAST document + root name
**Goal:** Change `AlkTypeEngine::compile` to the new signature and
have it use the BAST parser instead of the custom-keyword accessors.
The layout engines (`offset_map`, `layout_builder`,
`sequential_reader`) consume the BAST parser's typed output instead of
walking raw JSON with `get_alktype_kind*`.
**Spec reference:** [bast-format.md §Document Shape](../architecture/bast-format.md#document-shape),
[D-BAST-001](../research/bast-pivot.md#d-bast-001-root-type-selection).
Semver contract: `compile` is **Breaking**.
**Files:** `src/engine.rs` (the `compile` signature and body). The
layout modules (`src/offset_map.rs`, `src/layout_builder.rs`,
`src/sequential_reader.rs`) — their schema-walking code changes from
`get_alktype_kind*` calls to BAST parser calls. `src/lib.rs` if the
parser's public surface needs re-exporting (step 3 may have done this).
**Implementation notes:**
- New signature: `pub fn compile(bast_doc: &Value, root_name: &str,
mode: LayoutMode) -> Result<Self, AlkTypeError>`. Note `&Value` (not
`&mut Value`) — BAST needs no in-place `normalize_refs`.
- The engine stores the BAST document (or the parsed typed tree) for
`sequential_reader()`'s factory construction and `read_field`'s kind
lookup. The `Layout` enum and mode dispatch are unchanged.
- `parse_endian`, `parse_align`, `parse_encoding`, `parse_discriminator`
are reworked to read BAST properties (struct/field-level) instead of
keyword-value objects. Their *semantics* are unchanged (ADR-003);
only their *input location* moves. Whether they stay as free
functions or become methods on the typed `BastNode` is an
implementation choice — the POC read properties inline.
- The layout engines are format-agnostic beneath the accessors
(checked offset arithmetic, the two modes, union dispatch). This
step is an accessor swap, not a layout-engine rewrite.
**Verification:** `cargo test --release` (test inputs must be converted
to BAST format — see step 9 for the full test conversion; this step
converts the layout tests as a sanity check). `cargo clippy
--all-targets -- -D warnings`. `cargo build --target
wasm32-unknown-unknown --release` (layout/wasm-relevant).
---
### Step 5 — BAST-native validator (production version)
**Goal:** Port the POC's BAST-native validator into a production module
and wire it into `validate_bytes` as the validation step, replacing the
`jsonschema::Validator` call on the bytes path.
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
[D-BAST-006](../research/bast-pivot.md#d-bast-006-validate_bytes-validation-model),
[D-BAST-009](../research/bast-pivot.md#d-bast-009-alktypeerrorvalidation-payload-shape).
POC reference: `src/bast_poc.rs` on branch `bast-validator-poc`.
**Files:** New `src/bast_validation.rs`. `src/engine.rs`
(`validate_bytes` body — swap the `self.validator.validate(&value)` call
for the BAST-native validator). `src/lib.rs` — add `pub mod
bast_validation;` (the validator is engine-internal; whether it's
re-exported is an implementation choice, leaning no).
**Implementation notes:**
- The POC is the reference. The validator is a single recursive
function (`validate_typeref`) that dispatches on the BAST `kind`. The
constraint table is in [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model).
- Construct `AlkTypeError::Validation` via
`jsonschema::ValidationError::custom` — the variant's payload type is
unchanged (D-BAST-009). The bytes path no longer touches `jsonschema`
for validation, but the error type retains the `jsonschema` type for
uniformity with the `validate_json` path.
- The validator and materializer share the BAST-walking code structure.
If step 3 produced a typed `BastNode` tree, both consume it. If step
3 parses lazily, the validator parses lazily too (POC approach).
- Enum index bounds: check the materialized index against
`values.len()` — this **fixes the v0.1.0 dead constraint** (the
built-in `enum` keyword checked string membership, but the
materializer emits `Value::Number(index)`, which never matched). Net
improvement.
- Union variant dispatch: read `__discriminator`, look up the variant's
BAST definition, recurse. Recovers OQ-008 per-variant constraint
enforcement without custom keywords.
**Verification:** `cargo test --release` — the existing `validate_bytes`
tests are the regression target (test *inputs* change to BAST format
in step 9; expected validation outcomes must be identical). The POC's
20 tests are the reference. `cargo clippy --all-targets -- -D warnings`.
`cargo build --target wasm32-unknown-unknown --release`.
---
### Step 6 — `validate_json` against a consumer-provided JSON Schema
**Goal:** Update `validate_json`/`is_valid_json` to validate against a
standard `jsonschema::Validator` compiled from a consumer-provided JSON
Schema, not a custom-keyword validator built from the alktype schema.
**Spec reference:** [bast-format.md §Validation Model](../architecture/bast-format.md#validation-model),
[D-BAST-007](../research/bast-pivot.md#d-bast-007-validate_json-validation-model).
Semver contract: `validate_json`/`is_valid_json` are **Breaking
(behavioral)**; `build_validator` is **Breaking (signature or
removal)**.
**Files:** `src/engine.rs` (`validate_json`/`is_valid_json` bodies, and
the engine's stored validator field if the JSON Schema is supplied at
compile time). `src/validation.rs` (`build_validator` — repurposed or
removed). `src/lib.rs` (the `build_validator` re-export if removed).
**Implementation notes:**
- **Decide deferred decision #1 here:** does the consumer pass the JSON
Schema to `compile` (engine carries a second validator) or to
`validate_json` at call time? The former preserves single-call
ergonomics; the latter is more flexible. Needs the alkcall call-site
analysis. Not semver-relevant either way if the signature can absorb
the change.
- **Decide deferred decision #2 here:** `build_validator` removed vs
repurposed. If repurposed, its signature stays but its contract
changes (no custom keywords) — a behavioral break. If removed, drop
the `lib.rs` re-export.
- The `jsonschema` crate remains a direct dependency (for `validate_json`
and for validating BAST documents against the meta-schema). Only the
custom keyword integration is removed.
- The engine may carry two validators: the BAST-native validator (for
`validate_bytes`, from step 5) and the standard `jsonschema::Validator`
(for `validate_json`, from this step). Or `validate_json` takes the
JSON Schema at call time and builds a transient validator. The
decision shapes the engine struct's fields.
**Verification:** `cargo test --release` (new tests for the
consumer-provided JSON Schema path; existing `validate_json` tests
converted — their schemas were custom-keyword, now standard). `cargo
clippy --all-targets -- -D warnings`.
---
### Step 7 — Builder API produces BAST JSON
**Goal:** Update the builder's `build()` methods to produce BAST JSON
(for `struct_()`) and standard JSON Schema (for `object()`). Public
method signatures are unchanged; only the output `Value` shape changes.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(the output format), [D-BAST-008](../research/bast-pivot.md#d-bast-008-builder-api--two-output-formats).
Semver contract: `Schema::build`/`Definitions::build` are **Breaking
(output format)**.
**Files:** `src/builder.rs`. `src/lib.rs` if the builder's public
surface changes (it shouldn't — method signatures are unchanged).
**Implementation notes:**
- `Schema::struct_().field(...).build()` → BAST JSON (a `$defs` entry
with `kind: "struct"`, ordered `fields` array, type-level
annotations).
- `Schema::object().field(...).build()` → standard JSON Schema (no
`AlkType:*` keywords, no BAST `kind` — just `type`/`properties`/
`required`).
- `Definitions::build()`/`merge_into()` → a BAST `$defs` block.
- The builder already distinguishes AlkType kinds from JSON Schema types
via naming conventions (`string()` vs `string_()`). The construction
API is the same; only the serialization differs.
- The `Discriminator` builder is unchanged (semver contract:
**Unchanged**).
**Verification:** `cargo test --release` (builder tests assert on the
output `Value` — update the expected shapes). `cargo clippy
--all-targets -- -D warnings`.
---
### Step 8 — Remove v0.1.0 custom-keyword machinery
**Goal:** Remove the dead code now that all callers use the BAST parser
and BAST-native validator.
**Spec reference:** [bast-format.md §What is removed](../architecture/bast-format.md#what-is-removed).
Semver contract: the ~13 `schema::*` helper re-exports are **Breaking
(removal or rework)** (decision #3, confirmed in step 3).
**Files:** `src/schema.rs` (remove `get_alktype_kind*`,
`normalize_refs`, `inline_union_variant_refs`; rework or remove
`parse_*`/`resolve_*`). `src/validation.rs` (remove the 19
`jsonschema::Keyword` implementations if not already removed in step 5/6).
`src/lib.rs` (drop the removed items from the `pub use` block).
**Implementation notes:**
- Remove: all 19 `jsonschema::Keyword` implementations (~200 lines),
`normalize_refs()`, `inline_union_variant_refs()`, the
`get_alktype_kind*` family.
- Rework or remove: `parse_align`, `parse_discriminator`,
`parse_encoding`, `parse_endian`, `parse_max_length`, `resolve_ref`,
`resolve_ref_or_inline`. If the BAST parser subsumes them (likely),
remove them. If any remain useful as free functions over the typed
`BastNode`, keep them internal (not re-exported from `lib.rs`).
- The `jsonschema` crate's `with_keyword(...)` registration calls are
removed from `compile`/`build_validator`. The crate itself stays.
- Drop the removed items from `lib.rs`'s `pub use schema::{ ... }`
block. The BAST parser's public surface replaces them.
**Verification:** `cargo test --release`. `cargo clippy --all-targets
-- -D warnings`. `cargo doc --no-deps` (the public API surface
changed — doc comments must build). `cargo build --target
wasm32-unknown-unknown --release` (removing code shouldn't add
platform deps).
---
### Step 9 — Convert all tests to BAST format
**Goal:** Update the full test suite to use BAST format for inputs.
Test assertions (expected validation outcomes, expected offsets,
expected materialized values) must be identical — only the input
schema shape changes.
**Spec reference:** [bast-format.md](../architecture/bast-format.md)
(input format).
**Files:** `tests/*.rs` (integration tests), `src/*.rs` inline `#[cfg(test)]`
modules (unit tests).
**Implementation notes:**
- This may be partially done by steps 4–8 (each step converts the tests
it touches as a sanity check). This step is the sweep: every test
using `AlkType:*` keywords converts to BAST `kind`/`fields`.
- The POC's 20 tests are the reference for BAST-shaped test inputs.
- Expected validation outcomes are the regression target. The
enum-index-bounds test is new behavior (the v0.1.0 dead constraint
is now enforced) — that test's expectation *changes* (was: silently
passed; now: `AlkTypeError::Validation`). This is the intended fix,
not a regression.
- Coverage: 310 crate + 86 integration tests (~396 total). All must
pass.
**Verification:** `cargo test --release` (the full suite — this is the
gate). `cargo clippy --all-targets -- -D warnings`.
---
### Step 10 — Sync architecture docs and ADRs
**Goal:** Sync the descriptive docs and ADRs to the shipped code. This
is the final step — per AGENTS.md, ADRs are written post-implementation,
grounded in shipped code.
**Spec reference:** [Semver Contract §ADR impact](#adr-impact-checklist)
below.
**Files:** `docs/architecture/README.md`, `docs/architecture/overview.md`,
`docs/architecture/schema-layer.md` (rewrite for the BAST parser),
`docs/architecture/validation.md` (rewrite for the validator split),
`docs/architecture/builder.md` (update `build()` output examples),
`src/lib.rs` (module doc comment). New ADRs: ADR-BAST, ADR-VAL-SPLIT.
Amended ADRs: 001 (superseded), 003, 004, 009, 010.
**Implementation notes:**
- Rewrite `schema-layer.md` to describe the BAST parser (replaces the
custom-keyword accessor walk-through). The current `schema-layer.md`
content is the v0.1.0 reference; `bast-format.md` already contains
the target spec. Either fold `bast-format.md` into `schema-layer.md`
or keep both with `schema-layer.md` pointing at `bast-format.md` for
the format and describing the parser module.
- Rewrite `validation.md` for the validator split (the [bast-format.md
§Validation Model](../architecture/bast-format.md#validation-model)
content moves here, expanded with the production validator's
details).
- Update `builder.md` output examples to BAST JSON.
- Update `src/lib.rs` module doc comment: "Takes a JSON Schema with
`AlkType:*` custom keywords" → "Takes a BAST document".
- Update `docs/architecture/README.md` index — the document table, the
ADR table (new ADRs, superseded ADR-001), the key design principles
(#1, #2, #7, #10 change wording).
- Remove stale TODOs referencing custom-keyword normalization,
`inline_union_variant_refs`, or the rejected bare-name-ref design
(AGENTS.md §"Architecture Context").
- `docs/research/bast-pivot.md` is the research record — its status
flips from `draft` to `accepted`/`implemented` and it gains a pointer
to the ADRs that superseded its decisions.
**Verification:** `cargo doc --no-deps` (doc comments build).
Cross-reference check: every link in this plan, `bast-format.md`, and
the new/updated ADRs resolves. `cargo test --release` (no code change,
but the doc sweep shouldn't break anything).
## ADR Impact Checklist
Sync these ADRs when step 10 lands. Per AGENTS.md, ADRs are written
post-implementation, grounded in shipped code.
| ADR | Action | Reason |
|---|---|---|
| [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) (purpose, scope, "schema is the format") | **Supersede** | The "schema is the format" principle is retained and strengthened (BAST *is* the format), but the concrete format changes from custom-keyword JSON Schema to BAST. A new ADR (ADR-BAST) records the BAST format as the realization of the principle. ADR-001 Status → Superseded by ADR-BAST. |
| [ADR-002](../architecture/decisions/002-two-layout-modes-packed-vs-aligned.md) (two layout modes) | **Unchanged** | Layout modes are format-agnostic. One-line note that the input format changed but the modes didn't. |
| [ADR-003](../architecture/decisions/003-schema-annotations.md) (annotations) | **Amend** | Annotation *semantics* carry forward unchanged; annotation *location* moves from custom-keyword objects to BAST type-level properties. Amend the "where annotations live" sections, keep the semantics. |
| [ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md) (error handling, validation strategy) | **Amend** | Error enum shape unchanged (D-BAST-009). The "validation strategy" section updates: bytes path uses BAST-native validator, JSON path uses standard `jsonschema`. The load-time/access-time split is retained. |
| [ADR-005](../architecture/decisions/005-int64-uint64-first-class-kinds.md) (Int64/Uint64) | **Unchanged** | Kinds carry forward; JSON precision caveat unchanged. |
| [ADR-006](../architecture/decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md) (reject non-final inline in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
| [ADR-007](../architecture/decisions/007-packed-mode-read-factory.md) (packed-mode read factory) | **Unchanged** | Reader factory semantics are format-agnostic. |
| [ADR-008](../architecture/decisions/008-reject-tunion-in-aligned-mode.md) (reject TUnion in aligned mode) | **Unchanged** | Layout rule, format-agnostic. |
| [ADR-009](../architecture/decisions/009-builder-api.md) (builder API) | **Amend** | Public method surface unchanged; `build()` output format changes (BAST for `struct_`, standard JSON Schema for `object`). Amend the "output format" section; keep the method catalog. |
| [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) (`validate_bytes`) | **Amend** | The two-step concept (materialize → validate) is retained. The validation step's *implementation* changes from `jsonschema` custom keywords to the BAST-native validator. Amend the "validation step" section; add a pointer to D-BAST-006/D-BAST-009 and ADR-VAL-SPLIT. |
**New ADRs to write (post-implementation, grounded in shipped code):**
- **ADR-BAST** — the BAST format, meta-schema, and `$defs`/`$ref`/
`kind` vocabulary. Supersedes ADR-001's format-specific content.
- **ADR-VAL-SPLIT** (or fold into ADR-004's amend) — the two-validator
model: BAST-native for `validate_bytes`, standard `jsonschema` for
`validate_json`. Records D-BAST-006, D-BAST-007, D-BAST-009.
**Descriptive docs to sync (post-implementation):**
- `docs/architecture/schema-layer.md` — rewrite for the BAST parser
(replaces the custom-keyword accessor walk-through).
- `docs/architecture/validation.md` — rewrite for the validator split.
- `docs/architecture/builder.md` — update the `build()` output examples
to BAST JSON.
- `src/lib.rs` module doc comment — update the "Takes a JSON Schema
with `AlkType:*` custom keywords" preamble to BAST.
- `docs/architecture/README.md` — update the document table, ADR table,
and key design principles for the pivot.
- `docs/architecture/overview.md` — update the "what" and "why" for
BAST (the crate now takes a BAST document, not a custom-keyword JSON
Schema).
**Stale TODOs to remove:** any TODO referencing custom-keyword
normalization, `inline_union_variant_refs`, or the rejected
bare-name-ref design — align with the ADRs as AGENTS.md §"Architecture
Context" requires.
## Verification Commands
Run these before committing each step. All must pass. Per AGENTS.md:
```bash
cargo test --release # full suite (~396 tests: 310 crate + 86 integration)
cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # if docs changed (step 8, step 10)
cargo build --target wasm32-unknown-unknown --release # if layout/wasm-relevant code changed (step 2, 4, 5, 8)
cargo publish --dry-run --allow-dirty # before a release (post-step 10)
```
+655
View File
@@ -0,0 +1,655 @@
# Plan: alktype fuzzing
Adopted from alkhttp's `docs/plans/fuzzing.md` (the pattern is operational
in six sibling crates: alkcall, alktty, alktunnels, alksocks, alkhttp —
each with the same `fuzz/` layout, the detached runner, and the
corpus-replay-as-plain-test gate). The rationale research lives in
alkcall's `docs/research/fuzzing.md` (tool landscape, comparable-crate
survey, the no-hosted-CI policy §7.9); this plan stays focused on what
alktype fuzzes and in what order.
Rationale for this crate in one paragraph (the detailed version applies
by reference from the two docs above): alktype is the binary engine the
alk* family consumes — alkcall's hub/spoke accepts BAST schema documents
from arbitrary internet peers, and those documents flow into this
crate's compile paths downstream; both untrusted-input shapes exist here
(attacker-shaped JSON BAST docs → `AlkTypeEngine::compile`, and
attacker-shaped byte buffers read according to a schema →
`validate_bytes` / `SequentialReader` / `tunion` dispatch /
`materialize`), and the byte side is hand-rolled decode
(`data_access`, `tunion` discriminators, indirect `{offset,length}`
pairs) — exactly the shapes where example tests miss off-by-one bugs.
The crate is fully synchronous, so targets are simpler than
alkcall/alkhttp's (no current-thread runtime shims anywhere).
**Status:** waves 1–3 implemented and verified (2026-09-30). The
pre-fuzzing inventory (§6) is verified against the code at 0.3.0.
All five targets and the full infrastructure are in-tree (commits
`ed41d77`, wave 1; `16b9023`, wave 2; `aef8d9f` + findings commits,
wave 3); the smoke campaigns ran clean (§5); the release-budget
campaigns across all five targets are complete (§5) — two targets
surfaced real findings: the wave-2 `read_opseq` packing bug (§6
candidate 6, fixed same-day) and three wave-3 `validate_pair`
findings (W3-1 harness pin, W3-2 upstream serde_json pin, W3-3 a real
engine bug — read_field/write_field misread aligned maxLength
reservations — fixed same-day with regression tests).
Progress log:
- **2026-09-30 — wave 3 landed.** Target 5 (`validate_pair`): the
two-input harness (10-lane schema menu incl. a raw-JSON-bytes lane
fused with the buffer), 48 committed seeds, decode-pin tests, and
the §3 target-5 invariants (mode agreement incl. the documented
aligned-rejection taxonomy, the materialize⇄validate_bytes verdict
lattice with verbatim error propagation, unknown-path echo,
non-finite-float Access pin, enum-Validation pin, record spin
bound). Pre-campaign hand-drives all held. **Release-budget
campaigns (§5):** 45–46 min across all five targets. bast_compile
818k execs / 19,894 edges (still growing at budget end);
data_access 11.3M execs / 684 edges (a value-profile run lifted the
saturated corpus from 379 to 684 edges); read_opseq 63k execs /
17,519 edges; layout_build 63k execs / 17,659 edges; validate_pair
37k execs / 10,850 edges. **Finding W3-1 (harness invariant
corrected — the fuzzer fired an over-assertion):** the harness
claimed validate_bytes Ok ⇒ every offset-map leaf's range.end ≤
buffer.len() — false for offset-indirect entries, whose pair points
absolutely into the buffer while a maxLength window may dwarf the
validated buffer; the wave-1 data_access bounds partition is the
real contract. Fixed the invariant, pinned the artifact bytes as
seed-044 + a named regression test (commit `9ca9922`).
**Finding W3-2 (upstream, pinned with slack — no alktype bug):**
serde_json's non-`float_roundtrip` parser drifts one ulp when
re-parsing its own emitted shortest repr of adversarial f64 values
(probe: 0x5bffffffffffffff emits 1.4536774485912136e+135 and
parses back one ulp low; std's parser and ryu's own float parse are
exact — the concise reparse is the drift). The harness replaced
bare `Value` equality with a one-ulp structural comparison; the
artifact bytes are seed-047 (commits `a8e955c`/`b7ead99`).
**Finding W3-3 (real engine bug — fixed, commit `a0dd3d2`):**
`AlkTypeEngine::read_field`/`write_field` treated an aligned
`maxLength` reservation (ADR-003 strategy 2, `VARCHAR(N)`: raw
zero-padded window, NUL-trimmed on read — exactly what the
materializer and `validate_bytes` implement) as length-prefixed,
parsing the window's first four raw bytes as a u32 length. Every
aligned schema declaring `maxLength` broke the
validate_bytes⇒read_field lattice whenever the reservation's first
bytes looked like a large prefix (validate Ok, read_field Access
with bogus bounds; write_field wrote prefix+data into a raw
window). Third crash artifact (W3-1's shape family: the campaign
re-found the disagreement space after W3-1's invariant was
corrected). Fix: `VariableEncoding` gains `MaxLengthReserved`
(additive variant, ADR-003 strategy 2); `OffsetMap::compute`
records it for maxLength fields with the default encoding
(`maxLength`+`offset-indirect` stays `OffsetIndirect`, preserving
W3-1's combination semantics); `read_field`/`write_field` dispatch
through new `data_access::read_reservation{,_string}/
write_reservation` (single source of truth with the materializer);
three engine regression tests + the W3-1/W3-2 artifacts as
committed corpus seeds. Post-fix restart: 30-min validate_pair
campaign clean to budget end (37k execs, 10,850 edges, exit 0).
- **2026-09-30 — wave 2 landed.** Targets 3 (`read_opseq`) and 4
(`layout_build`) with 73 committed seeds (58 + 15), decode-pin tests
for the arbitrary 1.4.2 derive encoding, and the plan §6 candidate-1
spin bound encoded as the explicit End-op assertion. Smoke campaigns:
see §5. **Finding W2-1 (fixed same session):** `plan_read_array`
returned `Ok` for a fixed-stride array whose declared window
(count × stride) extended past the buffer — the bounds check was
missing entirely from the array arm (struct/union arms had theirs).
A truncated array reported success with the failure deferred to the
*next* field read (wrong field path), or masked entirely when the
array was the last field. Found by hand-running the target-3 drive
before the campaign (the first semantic fixture); fixed in
`plan_read_array` with an end-vs-buffer bounds check + Access error
naming the array field, regression test
`array_truncated_below_fixed_stride_window_is_access_error_not_ok`.
**Finding W2-2 (pinned, not a bug):** packed mode compiles
`"encoding": "offset-indirect"` fields but the sequential reader
always reads them inline length-prefixed — this matches
bast-format.md §"Default strategy selection" ("Packed sequential
mode: always inline length-prefixing"), so the annotation is a
no-op in packed mode. Pinned as a corpus-replay invariant
(`packed_mode_is_always_inline_length_prefixed`) so any future
change to the packed reader's encoding awareness is deliberate. The
aligned materializer honors the encoding correctly. An upstream
question — should packed compile either honor the encoding or reject
the declaration — is recorded in §6 candidate 7.
- **2026-09-30 — wave 1 landed (commit `ed41d77`).** The `fuzz/`
workspace, 136 committed seeds (38 `bast_compile` + 98
`data_access`), the detached runner, the corpus-replay gate
(AGENTS.md checklist gains the line), nightly pinned subtree,
explicit root `[workspace]` exclusion, publish-exclude gain,
`json.dict`. `cargo fuzz build` clean; corpus replay 4/4 green;
main crate untouched (569 tests, clippy `-D warnings` clean). One
dict-format fix on the way: libFuzzer's dictionary parser does not
accept `\u` escapes (`"\u0000"` → the `\xAB` form) and needs fully
quoted lines — caught by the campaign launcher, not the fuzzer.
- **2026-09-30 — smoke campaigns clean (see §5).** `data_access`
saturated (pure decode core, the alkcall `chunk_header` profile);
`bast_compile` still discovering coverage at budget end (longer
campaigns keep paying). No crate findings — the §6 candidates
(record-count loops, indirect pairs) held under the parser-level
drives; both remain encoded as wave-2/3 invariants in the stateful
targets.
- **2026-09-30 — wave-2 smoke campaigns (see §5).** Results recorded
there alongside the wave-1 numbers.
---
## 1. Why alktype fuzzes (the if)
1. **Downstream of the trust boundary.** alktype is compiled against in
alkcall (the integration crate), whose peers are untrusted and whose
wire payloads carry schema-shaped JSON. A panic on a malicious BAST
doc or bytes read under one is the quinn-CVE class
(RUSTSEC-2026-0037) at one further hop: the alk* stack parses
documents it never vetted, and alktype is where they get walked.
2. **Both input shapes, one crate.** Sibling crates each had mostly one
parse shape (wire bytes); alktype has the schema-JSON shape *and*
the raw-buffer shape, plus two-input combined paths
(`read_field`/`write_field`, `materialize_aligned` are doc+bytes).
3. **Infrastructure is proven and cheap; the crate is the simplest
consumer yet.** Six siblings run the layout; alktype is sync, has
zero `unsafe`, zero `unwrap`/`expect` outside tests, and no
allocation-from-wire-count anywhere (grep-verified inventory). The
marginal cost is target logic only.
**Honest caveat (alkcall §1's shape):** the code is already well
hardened — `checked_add`/`check_bounds` everywhere, parse-time caps
(`MAX_ARRAY_ELEMENTS`, `MAX_ARRAY_BYTES`, `MAX_ALIGN` = 4096,
`MAX_LENGTH`), `MAX_GRAPH_DEPTH`/`MAX_COMPILE_DEPTH` = 128, meta-schema
gate before any walker. Expected yield is low-moderate: the residual
candidates in §6 are the first things to probe; a clean first campaign
is the successful negative result — "we think the engine is robust"
converted into a demonstrated property.
## 2. Infrastructure (identical to the siblings)
Layout (copy of alkcall/alkhttp):
```
fuzz/
├── Cargo.toml alktype-fuzz (nightly-only bins; own [workspace])
├── rust-toolchain.toml pins nightly + llvm-tools for this subtree only
├── fuzz_targets/ thin fuzz_target! wrappers (3 lines each)
├── shared/ alktype-fuzz-shared — STABLE-toolchain library:
│ invariant logic + corpus-replay tests
├── corpus/<target>/ committed seeds (generated by gen_fuzz_seeds.py)
├── artifacts/ gitignored crash/oom/timeout artifacts + logs
├── gen_fuzz_seeds.py deterministic seed generator (quiche pattern)
├── json.dict JSON/BAST token dictionary (bast_compile)
├── run-detached.sh detached campaign runner (copied from the siblings)
└── README.md operational cheat-sheet
```
Load-bearing details (all six siblings hit these; alkcall's doc is the
deep reference):
- **Invariant logic lives in `fuzz/shared/`**, not the target binaries.
The stable-toolchain shared crate replays every committed seed through
the identical invariant functions as plain `cargo test` — the standing
fuzz gate (alkcall §7.9 tier-3 deliverable; no hosted CI in this repo
by policy). The `fuzz_target!` binaries are thin wrappers.
- **Root `Cargo.toml` needs an explicit `[workspace]` table**
(`members = ["."]`, `exclude = ["fuzz"]`); without it auto-discovery
pulls `fuzz/shared/` into the main workspace and the stable toolchain
builds nightly-consumed dev-deps. alkcall hit this trap.
- **`fuzz/rust-toolchain.toml` pins nightly** so `cargo fuzz build`
works from any CWD; nightly stays confined to `fuzz/`, MSRV 1.85
untouched here. `fuzz/` joins the publish `exclude` list.
- **`.gitignore` additions**: `fuzz/artifacts/`, grown-corpus dirs
(committed seeds stay).
- **Detached runner (non-negotiable operating rule).** Campaigns never
run as a foreground child of an agent session; the runner pins
`-fork=1 -rss_limit_mb=2048 -malloc_limit_mb=2048 -timeout=25` and
detaches via `setsid` + `nohup` + log redirect; the agent polls the
log and artifact directory, never waits. Copied from the siblings.
- **No feature-gating needed in `fuzz/shared`**: alktype has
`default = []` and no feature flags, so the shared crate rides the
main crate build unconditionally (unlike alkhttp's gated
`openapi`/`mcp`).
- **Exposure needs are minimal.** The inventory found every target
entry point already `pub` (`compile`, `data_access::*`,
`SequentialReader`, `LayoutBuilder`, `tunion::*`, `materialize::*`,
`validate_bast_doc`, `build_validator`). No `#[cfg(fuzzing)]` hub is
expected — the first choice remains a minimal `fuzzing` hub only if a
needed item turns out `pub(crate)`, per the sibling pattern (alkcall
never needed one).
**Verification-gate change:** `cargo test --manifest-path
fuzz/shared/Cargo.toml` (corpus replay) joins AGENTS.md's verification
checklist, as the siblings did.
## 3. Target inventory (5, in waves)
| # | Target | Drives | Input style | Status |
|---|---|---|---|---|
| 1 | `bast_compile` | `AlkTypeEngine::compile` both modes (via `bast_meta` → `BastDoc` → plans → layout → validator) | raw bytes → serde_json → BAST doc | implemented 2026-09-30 |
| 2 | `data_access` | the hand-rolled decode core (`src/data_access.rs`, read + write side) | raw `&[u8]` + chosen (offset, endian) | implemented 2026-09-30 |
| 3 | `read_opseq` | stateful `SequentialReader` op sequences over hostile bytes under a fixed plan | `#[derive(Arbitrary)]` op enum | implemented 2026-09-30 |
| 4 | `layout_build` | `LayoutBuilder::build` with adversarial `var_sizes` | `#[derive(Arbitrary)]` map shapes | implemented 2026-09-30 |
| 5 | `validate_pair` | two-input structured: compile a schema once per exec, hammer hostile bytes through `validate_bytes`/`read_field`/`materialize` | `#[derive(Arbitrary)]` (doc, bytes) pair | implemented 2026-09-30 |
### Target 1 — `bast_compile` (the whole schema side, one choke point)
`AlkTypeEngine::compile` (`src/engine.rs:127-176`) fans out through the
entire untrusted-JSON surface: `bast_meta::validate_bast_doc` →
`BastDoc::new` → `ValidationPlan::compile` → `LayoutBuilder::new` +
`ReadPlan::compile` (packed) or `OffsetMap::compute` (aligned) →
`validation::build_validator` (when `json_schema` is `Some`).
- **Drives:** raw bytes → `serde_json` → `compile(value, root_name,
mode, None)` in both modes; a second lane feeds `Some(schema)` with a
second attacker-shaped JSON value for the jsonschema-build path.
- **Invariants:**
- no-panic on any JSON document, both modes;
- compile is always `Result` — every rejection is a clean
`AlkTypeError` (`Schema`/`Offset`/`Validation` payload classes,
`src/error.rs:11-28`), never a panic or a silent bogus engine;
- meta-schema gate ordering: any doc that fails
`validate_bast_doc` must surface `Schema(...)` and must never reach
layout/plan walks (shape partition);
- parse-time caps hold exactly: `align > 4096`, arrays above
`MAX_ARRAY_ELEMENTS`/`MAX_ARRAY_BYTES`, `maxLength >
MAX_LENGTH`, depth > 128, and `$ref` cycles all reject at compile
with the documented error classes (the walk-guard
`check_ref_graph`, compile-depth, and cycle-`seen` machinery
pinned by adversarial corpus entries);
- if compile fails in packed it must also fail in aligned (mode
independence of the schema-gate layer — the parse layers are
shared; divergence means a mode-specific parse bug);
- a successfully compiled engine's `endian()` equals the root
struct's declared endianness.
- **Seeds:** the full BAST feature menu (each kind, endian ×2, TUnion
byte/field/enum discriminators, records, arrays, string/bytes
encodings, `$ref` diamond), each reject-class corpus entry, plus the
hostile menu in §4.
### Target 2 — `data_access` (the decode core)
Every byte-touching decode funnels through `read_array<const N: usize>`
(`src/data_access.rs:48-75`): `checked_add(N)` → `check_bounds` →
`.get(..)` → `try_into`. The widest attacker-influenced values in the
crate are `read_bytes_indirect`'s absolute `{offset,length}` pair
(`src/data_access.rs:328-350`).
- **Drives:** the `pub` read functions directly with the fuzzer
choosing buffer, offset (including far-past-end and huge values),
and endianness; lanes for `read_bytes`/`read_string` (u32 length
prefix), `read_bytes_indirect`/`read_string_indirect` (the
`{offset,length}` pair), `read_enum`, `read_bool` strictness, and
each fixed-width kind from the macro family.
- **Invariants:**
- no-panic for any (buffer, offset, endian) triple;
- `bool` accepts exactly 0x00/0x01 and rejects everything else
(`:134-144` — the strictness is contract, pin it);
- invalid UTF-8 in `read_string` errors (`Access`), never a lossy
silently-corrupting parse (`:195-208`);
- bounds partition: an error implies `checked_add`-overflow or
`end > buffer_len` with the offending `field_path` named; an Ok
implies the field sits fully inside the buffer;
- nothing before/end-of-buffer is read: the decode consumes
exactly its declared width (offset unchanged on error paths);
- `read_bytes_indirect`'s data region always satisfies
`data_offset + data_length ≤ buffer_len` on Ok, and neither
field can push arithmetic past the buffer without an error
(the two `u32` widening casts at `:334, :341` widening-only,
verified by the partition).
### Target 3 — `read_opseq` (stateful, wave 2)
`SequentialReader` is stateful against attacker bytes (mutable cursor:
`field_index`, `position`; `src/sequential_reader.rs:144-148`) and
fuzzer-reachable operations are `read_next`, `read_next_borrowed`,
`read_field` (out-of-order names), `reset` (`:157-300`).
- **Drives:** `#[derive(Arbitrary)]` op sequences (Next, Field(name
choice), Reset, End) against a compiled plan — the plan built once
per exec from a fixed small schema menu, bytes adversarial.
- **Invariants:**
- no-panic over any op interleaving and any buffer;
- cursor discipline: a failed read leaves the reader usable (a
subsequent `reset` restores the exact initial state; cursor never
exceeds the buffer);
- `read_next` returns fields exactly in plan order and `None`
exactly at plan end; interleaved `read_field` for any field at
any cursor state never panics and never mutates the sequential
cursor (its offset argument comes from the plan, not the reader);
- record-count spin bound: wire-controlled `count` loops
(`src/materialize.rs:439-442`, `:817-820`,
`src/sequential_reader.rs:985-988`) consume ≥ 4 verified bytes per
iteration, so iterations are bounded by
`remaining_bytes / 4` — a hostile count fails fast with `Access`
(encode as an explicit per-exec assertion, not just
no-panic/OOM);
- engine-issued readers are independent: two readers over the same
plan and buffer never observe each other's cursors
(ADR-007's owned-fresh-reader contract).
### Target 4 — `layout_build` (wave 2)
`LayoutBuilder::new` parses once (`src/layout_builder.rs:154-156`);
`build(&HashMap<String, usize>)` (`:189`) is repeatable with
attacker-shaped `var_sizes` driving write-position arithmetic in
`walk_struct`.
- **Drives:** a fixed schema menu containing every variable-width
encoding × `#[derive(Arbitrary)]` `var_sizes` maps and write values
(`FieldValue` shapes).
- **Invariants:**
- no-panic across adversarial size maps (zero, huge, mismatched
with `max_length`/`count` declarations);
- every failed write leaves the buffer untouched (byte-equal to the
pre-call snapshot) or documented-partial exactly where the
contract allows — pin the actual contract the code implements;
- field positions from a successful `build` are disjoint and
in-bounds for the reported total size;
- `data_offset/length` pairs written by
`write_string_indirect`/`write_bytes_indirect` always satisfy the
read-side `read_*_indirect` bounds partition above — the write
side and the read side of the pair are one contract
(round-trip pair; `:373-417` guards verified by
`:733-750`-style assertions).
### Target 5 — `validate_pair` (two-input structured, wave 3)
The integration target: schema and bytes are both adversarial.
- **Drives:** `#[derive(Arbitrary)]` (doc, bytes) — compile once per
exec with whichever mode the fuzzer picks, then drive
`validate_bytes`, `read_field` (arbitrary field paths, including
junk paths), `materialize_packed`/`materialize_aligned`, and
`read_next` under the compiled plan.
- **Invariants:**
- no-panic for any (doc, bytes) pair, either mode;
- validate/read/materialize agreement lattice: `validate_bytes` Ok
⇒ `materialize_*` Ok and every `read_field` over a declared path
Ok; `materialize_*` error ⇒ `validate_bytes` error on the same
buffer (exact agreement direction pinned per the code's actual
contract — determine the strict/loose ordering from the
`validate_bytes` implementation, don't assume);
- non-finite floats (NaN/Inf) surfaced by `materialize` are always
`Access` errors, never silently `Null`/`0.0`
(`src/materialize.rs:876-883`);
- unknown field-path strings always error with `Access` naming the
path, never panic, never index the map by substring drift;
- `Value` output is serde-safe: `materialize_*` output round-trips
through `serde_json::to_vec` and back to a structurally equal
`Value` (structural only — this crate's serde_json builds with
`preserve_order`, so object key order is preserved; byte-identity
round-trips are acceptable only where the docs say lossless,
per the alkcall §7.3 false-positive trap when they don't).
## 4. Corpus policy
Committed hand-made seeds per target, generated by
`fuzz/gen_fuzz_seeds.py` (deterministic, in-tree, quiche pattern);
grown corpora and artifacts gitignored. Seed menus:
- `bast_compile`: a minimal valid packed doc and aligned doc; every
`AlkTypeKind` once; each reject class (`align` 4097/65536/u32-max,
`count` over cap, `maxLength` over cap, depth-129 nesting both
inline-nested and via `$ref` chains, `$ref` cycle, `$ref` to
missing def, root not a struct, missing `type`, unknown kind
string, duplicate field names first-wins probe, non-object doc,
deeply-nested JSON at serde_json's own 128 limit); `json.dict`
carries the BAST token set.
- `data_access`: minimal valid encodings per kind per endianness;
truncation at every prefix length (1..N-1 for each width);
`len = 0` / `MAX_LENGTH` / `u32::MAX` prefixes; the indirect pair at
{0,0}, {len, big}, {big, 0}, {u32::MAX, u32::MAX}; offset one-past-
end, offset u32-magnitude; 0x02 bool byte; invalid UTF-8 in a
string; enum value out of range; NaN/Inf bytes.
- `read_opseq` / `layout_build` (wave 2, **done 2026-09-30**): the
semantic fixtures — full sequential walk, reset-mid-walk then full
walk again, failed read then reset, record loop with a hostile count
under a real buffer, junk field paths; zero/huge/mismatched
`var_sizes`; overwrite-everything write; indirect-pair overflow
write, plus the write-then-read pair fixture. `read_opseq` seeds
also encode truncation at every prefix of the full-walk buffer. The
hand-written seeds are byte-encoded against the pinned arbitrary
1.4.2 derive layout and *pinned by decode tests* (both targets carry
a `decode_lands_on_the_intended_variants` replay test, the alkhttp
target-4 pattern).
- `validate_pair` (wave 3, **done 2026-09-30**): hostile-schema/
valid-bytes, valid-schema/hostile-bytes, valid/valid — plus the §6
candidate shapes as pinned reproducers. 48 committed seeds incl.
the W3-1/W3-2 artifact bytes, the aligned maxLength fixtures, and
the mode-agreement pins (ADR-006/ADR-008 lanes).
Stateful `Arbitrary` seeds are hand-encoded against the `arbitrary`
1.4.x derive layout with per-element keep-going bytes, pinned by
seed-decode tests (alkhttp's target-4 pattern).
## 5. Campaign + gate policy
Identical to the siblings (alkcall §7.9 posture; no hosted CI in this
repo):
- **Corpus replay is the standing fuzz gate:** `cargo test
--manifest-path fuzz/shared/Cargo.toml` — joins AGENTS.md's
verification checklist.
- **Campaigns run detached** via `fuzz/run-detached.sh`; budget 10 min
per target for a smoke campaign, 30–45 min before a release or after
touching `src/data_access.rs`, `src/schema.rs`, the compile walks, or
the sequential reader.
- **Grown corpora stay gitignored** (hash-named files ignored via
pattern; committed `seed-*` files stay); merge worthy entries into
seeds only deliberately.
- libFuzzer flags worth pinning: `-rss_limit_mb=2048`,
`-malloc_limit_mb=2048`, `-timeout=25`, `-max_len=65536`,
`-use_value_profile=1`, and `-dict=json.dict` on the JSON targets
(alkcall §7.2's set, minus the CI-tier concerns).
- Nightly stays confined to `fuzz/`; the main crate's stable build,
MSRV, and wasm target are untouched — `cargo fuzz build` must never
be a prerequisite for `cargo test`/`clippy`/`build`.
Release-budget campaign results (2026-09-30, 45–46 min per target,
§5 policy, detached, `-use_value_profile=1`, `-dict=json.dict` on the
JSON lanes): all five exited 0.
- `bast_compile` — 818k execs at ~300–500/s (2,756 s), coverage
19,894 edges / 86,455 features / 4,439 in-memory corpus entries,
**still growing at budget end** (2× the wave-1 45-min smoke edge
count). Longest campaigns keep paying on the schema side.
- `data_access` — 11.3M execs at ~3.7–4.7k/s (2,761 s), coverage 684
edges / 6,208 features — value-profile lifted the "saturated" 379
edges to 684 (the wave-1 ceiling was the no-profile ceiling).
- `read_opseq` — 63k execs at ~23/s (2,737 s), coverage 17,519 edges /
65,059 features / 1,668 entries; the stateful search kept adding
features across the full budget (assertion-throughput-bound).
- `layout_build` — 63k execs at ~23/s (2,751 s), coverage 17,659
edges / 71,477 features / 1,873 entries; write-side contracts held.
- `validate_pair` — three crashes over two runs before and one clean
run after the W3-1/W3-2/W3-3 fixes (final post-fix campaign 30 min):
37k execs (final 1,820 s run), coverage 10,850 edges / 38,364
features / 1,544 entries, `oom/timeout/crash: 0/0/0` at budget end.
The two-input harness is the highest-yield target in the crate: 1
real engine bug + 2 pinned contracts in its first campaigns.
- Artifact totals: the three `validate_pair` crash artifacts are all
pinned as committed seeds/regression tests (W3-1 `seed-044`, W3-2
`seed-047`, W3-3 reproduced by the aligned maxLength fixtures);
`bast_compile`/`data_access`/`read_opseq`/`layout_build` artifacts
empty. No OOM, timeout, or leak on any fork job of any target.
Wave-2 smoke campaign results (2026-09-30, 10 min per target, §5
policy, detached):
- `read_opseq` — 52.1k+ execs at ~90/s, exit 0, empty artifact dir,
no oom/timeout/crash on any fork job. The typed-input decode plus
the per-op assertion work makes this the slowest target per exec in
the crate so far; coverage ~9,365 edges / 15,797 features / 293
in-memory corpus entries. The heavy semantic invariants
(replay-after-failure, full-walk spin bounds, reader independence,
plan-order assertion) run per op, not per exec — the campaign is
assertion-throughput-bound, not coverage-bound; the stateful search
(op × cursor × buffer shape) was still adding features at budget end.
- `layout_build` — 42.4k+ execs at ~84/s, exit 0, empty artifact dir,
no oom/timeout/crash on any fork job; coverage 9,432 edges / 19,204
features / 181 in-memory corpus entries. The adversarial `var_sizes`
map space over five schema menus exercised the missing-size /
unknown-discriminator / overflow rejection paths; the write-side
contracts (failed write leaves buffer byte-identical, positional
disjointness and bounds) held everywhere.
Wave-1 smoke campaign results (2026-09-30, 10 min per target, both
exited 0, artifact dirs empty — no crash/hang/OOM/leak):
- `bast_compile` — 543k execs at ~1.1k/s (each exec compiles two
engines through the whole fan-out — meta gate, parse, plans, layout
walks — in both modes, so per-exec work is heavy), coverage 10,830
edges / 25,326 features / 1,293 in-memory corpus entries, **still
growing at budget end** — longer campaigns keep paying.
- `data_access` — 3.5M+ execs at ~7–8k/s, coverage saturated at 379
edges / 574 features / 37 corpus entries (the pure-decode-core
ceiling is fully enumerated; the alkcall `chunk_header` profile).
- Both exited 0 with empty artifact directories; `oom/timeout/crash:
0/0/0` on every fork job.
- **Seeds regenerate deterministically:** `python3
fuzz/gen_fuzz_seeds.py`.
- Toolchain notes live in `fuzz/README.md`; nightly stays confined to
`fuzz/`.
## 6. Pre-fuzzing candidate findings (confirm or refute)
These are pre-fuzzing code-review findings from the 0.3.0 inventory,
verified against the code. They define what the targets must encode as
invariants and are the first corpus entries to add; the fuzzer
confirms or refutes them.
1. **Wire-controlled `Record` count loops** (`src/materialize.rs:439-
442`, `:817-820`, `src/sequential_reader.rs:985-988`): the only
buffer-derived loop counts (`read_u32(..)? as usize` then
`for i in 0..count`). Each iteration performs at least one
bounds-checked read, so a hostile count should fail fast with
`Access` — bounded by `remaining_bytes / 4`, no allocation, no
spin. Correct as designed *if and only if* that holds; the
`read_opseq` target encodes it as an explicit assertion (§3
target 3) so the fuzzer can break it the moment any per-entry
cost stops being `≥ 4 verified bytes`.
2. **Attacker-controlled absolute `{offset,length}` pairs**
(`read_bytes_indirect`/`read_string_indirect`,
`src/data_access.rs:307-350`): the widest attacker-influenced
values in the crate (each `u32`, up to 2³²−1). The pattern
(widening cast → `checked_add` pair-sum → full bounds check)
looks correct; campaign confirms the bounds partition on every
(buffer, pair) input, both sides of the write/read contract
(§3 target 2 / target 4).
3. **Duplicate field-path tolerance by design** (`OffsetMap::build_
index`, `src/offset_map.rs:199-205`: first-wins; `BastStruct::
parse` does not reject duplicates): ambiguous lookups are
documented behavior. Encode as an invariant — a successful
engine's `read_field` resolves duplicates deterministically
(first wins) — so a future "reject duplicates" change shows up
as a deliberate contract change, not silent drift.
4. **`align_up`/`round_up` plain `+` arithmetic** (`src/offset_map.rs:
618-639`): un-checked `+ align - rem` in a field of
schema-controlled values. Overflow-infeasible today because
align is capped at parse (`MAX_ALIGN` = 4096) and the running
offset is monotonically checked — a `bast_compile` corpus entry
with align at the cap pinning the boundary keeps it that way if
the cap ever moves.
5. **`bast_validation::validate_value` recompiles a `ValidationPlan`
per call** (`src/bast_validation.rs:64-65`): a repeated-op DoS
surface if any consumer compiles-per-call. Not a bug in this
crate's API; fuzz targets compile once per exec, and the doc
records the compile cost as the consumer's responsibility.
6. **Missing bounds check in `plan_read_array` — CONFIRMED and FIXED
(wave 2, finding W2-1, 2026-09-30).** The struct arm
(`plan_bounds_check`) and the union fixed-size arm both verify the
computed end against the buffer; `plan_read_array`
(`src/sequential_reader.rs:898` pre-fix) computed `end` and
returned `Ok` without the check. A truncated fixed-stride array
reported `Ok(Some(...))` with `element_start + count × stride` past
the buffer; the failure surfaced at the *next* field (naming the
wrong field in the error), or never when the array was the last
field — a completed walk returning `Ok` over a short buffer. Found
by hand-running the `read_opseq` drive before the campaign; fixed
with an end-vs-buffer check returning `Access` naming the array
field; regression test in `src/sequential_reader.rs`. The wave-2
seeds' array-truncation fixtures keep the boundary pinned.
7. **Packed-mode `encoding: offset-indirect` is a silent no-op —
PINNED, open design question (wave 2).** `BastField::parse`
records the annotation; the packed reader
(`plan_read_primitive`) ignores it and reads inline
length-prefixed — correct per bast-format.md's "Default strategy
selection" table, but nothing rejects the declaration in packed
mode (aligned mode consumes it; aligned rejects it on record
fields only). The wave-2 replay test
`packed_mode_is_always_inline_length_prefixed` pins the current
behavior. If the packed reader ever grows encoding awareness, the
pin flips deliberately; if the format instead wants the
declaration rejected in packed mode, that is a schema-gate change
with the mode-agreement invariant to re-verify.
8. **Aligned `maxLength` reservation read paths disagreed —
CONFIRMED as a real engine bug and FIXED (wave 3, finding W3-3,
found by the `validate_pair` campaign).** The materializer and
`validate_bytes` implement ADR-003 strategy 2 correctly (raw
zero-padded window, NUL-trimmed), but `AlkTypeEngine::read_field`
dispatched every aligned String/Bytes leaf through the
length-prefixed `data_access::read_string`/`read_bytes` — parsing
the reservation window's first four raw bytes as a u32 length —
and `write_field` wrote prefix+data into the raw window. Any
aligned schema declaring `maxLength` whose first reservation bytes
looked like a large prefix broke the validate_bytes ⇒ read_field
lattice (validate Ok, read Access with bogus bounds). Fixed by
recording the strategy in `LeafMeta`'s encoding
(`VariableEncoding::MaxLengthReserved`, additive variant) and
dispatching read/write through new
`data_access::read_reservation{,_string}`/`write_reservation`
sharing the materializer's exact semantics. Regression tests in
`src/engine.rs` + `data_access.rs`; the W3-1 artifact bytes are
the committed reproducer (`seed-044`).
None of these rises to the alkcall §6.2 / alkhttp FWD-20 class; they
are boundary-confirmations, which is exactly the expected profile of
this crate (§1 honest caveat). Smoke-campaign evidence: no panics or
OOM/timeouts on any of the 383k+ wave-1/2 combined execs; the
release-budget campaigns added 12.3M+ execs with the W3 findings
above. Candidates 1, 2, and 4 held under the parser-level drives —
candidates 1 and 2 got their explicit stateful assertions in waves
2–3 (targets 3–5), and 4 keeps its boundary corpus entry. Candidate 3
(duplicate first-wins) is pinned by the wave-2 `layout_build` index
assertions and the existing crate tests; candidate 6 was confirmed as
a genuine bug and fixed in wave 2; candidate 7 is pinned with an open
design question; candidate 8 was confirmed as a genuine bug and fixed
in wave 3.
## 7. Sequencing
1. ✅ **Wave 1 (2026-09-30, commit `ed41d77`)** — infra (`workspace`
exclude, toolchain pin, runner, seed generator, README,
`.gitignore`) + targets 1–2 + corpora (136 seeds) + corpus replay +
AGENTS.md gate + smoke campaigns (clean, see §5).
2. ✅ **Wave 2 (2026-09-30)** — targets 3–4 (stateful `read_opseq`,
`layout_build`) + 73 seeds + decode-pin tests + smoke campaigns.
One real bug found and fixed first-session (`plan_read_array`
bounds check, §6 candidate 6); packed-mode offset-indirect pinned
as a documented no-op (§6 candidate 7).
3. ✅ **Wave 3 (2026-09-30)** — target 5 `validate_pair` (the
two-input structured harness) + 48 seeds + decode-pin tests +
45-min release-budget campaigns across all five targets. Three
findings: W3-1 (harness over-assertion corrected + pinned), W3-2
(upstream serde_json one-ulp f64 parse drift pinned with slack),
W3-3 (real engine bug — aligned maxLength reservation misread by
`read_field`/`write_field` — fixed with `VariableEncoding::
MaxLengthReserved` + `data_access::read_reservation*`/
`write_reservation`, commit `a0dd3d2`). Post-fix
validate_pair campaign clean to budget end. Fuzzing complete per
§2 scope: corpus replay 30/30 is the standing gate.
4. Everything else inherited verbatim: no hosted CI, OSS-Fuzz out,
no Actions/workflow files anywhere in the repo (alkcall §7.9).
## 8. References
- alkcall `docs/research/fuzzing.md` — rationale, tool landscape,
campaign containment (§7.6), no-hosted-CI policy (§7.9)
- alkhttp `docs/plans/fuzzing.md` — the live pattern this plan copies
(waves, findings log, `fuzzing` hub convention)
- RUSTSEC-2026-0037 / CVE-2026-31812 (quinn-proto) — the remote-DoS
class this crate's peers are exposed to
- Internal: ADR-002 (layout modes), ADR-004 (error/validation
strategy), ADR-006 (aligned-mode variable-field rejection), ADR-008
(aligned-mode TUnion rejection), ADR-010 (`validate_bytes`,
materialize-then-validate)
+551
View File
@@ -0,0 +1,551 @@
---
status: implemented
created: 2026-08-14
last_updated: 2026-08-15
---
# BAST Pivot — Research Record
**Status: implemented.** The BAST pivot landed in steps 1–10 of the
[implementation plan](../plans/bast-implementation.md) (commits
`66ab9d7` → `54fd112` on `origin/main`). The decisions D-BAST-001..009
are recorded in two new ADRs —
[ADR-BAST](../architecture/decisions/bast-bast-format.md) (the format)
and [ADR-VAL-SPLIT](../architecture/decisions/val-split-two-validator-model.md)
(the two-validator model) — which supersede the format-specific content
of [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md)
and refine the validation strategy of
[ADR-004](../architecture/decisions/004-error-handling-validation-strategy.md)
and [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md).
The normative format spec is
[`docs/architecture/bast-format.md`](../architecture/bast-format.md);
the parser is documented in
[`docs/architecture/schema-layer.md`](../architecture/schema-layer.md).
What follows is the original research record — the *why* and *what was
proved*, preserved as the historical grounding for the decisions.
---
Replace alktype's custom JSON Schema keywords (`AlkType:Uint32`,
`AlkType:Struct`, etc.) with a standalone JSON format — BAST (Binary
Abstract Syntax Tree) — that describes binary data layouts using a
`kind`-based vocabulary with `$defs`/`$ref` for composition. BAST is
itself a valid JSON Schema instance (it has a meta-schema), making it
self-validating, editor-friendly, and trivially consumable from any
language with a JSON parser.
The engine's core logic (layout computation, data access, union
dispatch, two layout modes) is unchanged. Only the schema-walking
accessor layer changes: instead of detecting `AlkType:*` keywords
scattered through a JSON Schema tree, the walkers read `kind`/`fields`/
annotation properties from a purpose-built format.
The builder API's public surface stays the same; only the JSON output
format changes internally.
> **Document role.** This is the research record: motivation, POC
> scope and result, decisions, risks, references. The normative format
> specification lives in [`docs/architecture/bast-format.md`](../architecture/bast-format.md).
> The execution plan — ordered implementation steps, the public-API
> semver contract, and the ADR-sync checklist — lives in
> [`docs/plans/bast-implementation.md`](../plans/bast-implementation.md).
> Those documents supersede the format-spec, what-changes, and
> migration-path sections that previously lived here; this record
> keeps the *why* and *what was proved*, not the *how to implement*.
## Motivation
### Current state
alktype v0.1.0 embeds binary layout information inside standard JSON
Schema documents via custom keywords:
```json
{
"AlkType:Struct": true,
"type": "object",
"properties": {
"channel_id": { "AlkType:Uint32": true, "type": "integer" },
"length": { "AlkType:Uint32": true, "type": "integer" }
},
"endian": "big"
}
```
This works for the Rust engine — it walks the tree, detects keywords,
computes offsets. But it creates friction for everything outside Rust:
1. **Cross-language consumption.** A Python, Go, or TypeScript consumer
that wants to parse an alktype schema must re-implement custom keyword
detection. The format is not self-describing — you need to know that
`AlkType:Uint32` means "4-byte unsigned integer" and that it can
appear as either `true` or `{ "encoding": "..." }`.
2. **Code generation.** Generating Rust/TypeScript/Python readers and
writers from a schema requires walking an arbitrary JSON Schema tree
looking for custom keywords. A `kind`-based format with known keys
makes this a straightforward structural walk.
3. **Tooling.** Editors, linters, and schema validators don't understand
`AlkType:*` keywords. A BAST document with a published meta-schema
gets autocomplete, validation, and documentation in any JSON Schema-
aware editor for free.
4. **Two concerns in one document.** The current format conflates binary
layout (what the engine needs) with JSON validation (what jsonschema
needs). A `type: "object"` with `properties` and `required` is a JSON
validation concern; `AlkType:Uint32` is a binary layout concern. They
live in the same JSON object but serve different masters.
### The downstream pain is real
The alkcall agent's review identified that the channels 8-byte chunk
header is hand-rolled with manual bit shifts — alktype's binary layout
capability is unused because the custom-keyword format is awkward to
integrate for a simple 2-field struct. alktty plans to hand-roll its
5-byte TTY chunk format for the same reason. SFTP's 29 packet types
were proven byte-identical with alktype in the POC, but the production
path requires defining 29 schemas in the custom-keyword format.
All three cases are the same pattern: a small binary struct that needs
a schema-driven reader/writer. BAST makes this trivial — a 10-line JSON
file replaces hand-rolled bit shifts.
### Timing
v0.1.0 was published but has zero real consumers (only bots/scanners
have downloaded it). A breaking change now is free. Waiting until
adoption creates migration cost.
## POC Scope
The layout swap needs no POC — it is a backend swap (custom keywords →
`kind`-based format) on top of a proven layout engine. The layout
engine's byte-identity is already proven (alknet-typedef-poc,
alktype-builder-poc) and the layout code is unchanged, so there is
nothing empirical to de-risk there.
One targeted POC **was** needed to de-risk the validation model. The
risk was specific and falsifiable: can a BAST-native validator — a
recursive walker over the BAST type tree — fully replace the 19 custom
keyword validators on the `validate_bytes` path, including the OQ-008
union variant dispatch, without regression?
**The POC has been run and succeeded.** The outcome is recorded in
[POC Result](#poc-result--bast-native-validator) below. The POC code is
on branch `bast-validator-poc` (commit `f371fe4`), not merged to main —
it is reference scaffolding superseded by the production module in
[implementation step 5](../plans/bast-implementation.md#step-5--bast-native-validator-production-version).
### POC: BAST-native validator for `validate_bytes`
**Hypothesis:** A recursive walker over the BAST type tree can enforce
all value-domain constraints that the 19 custom keyword validators
currently enforce, recovering the OQ-008 union variant dispatch
behavior, and fixing the enum-membership dead constraint on the bytes
path — all without `jsonschema` custom keywords and without requiring
the consumer to provide an external JSON Schema.
**Scope:**
1. Implement the BAST-native validator as a new module
(`src/bast_validation.rs` or similar)
2. The validator walks a materialized `Value` tree against the BAST
type definitions, checking:
- Integer ranges (Int8..Uint64)
- Float finiteness (Float32/64)
- String `maxLength` (UTF-8 byte length)
- Bytes `maxLength` (byte length)
- RFC 3339 timestamp shape (non-strict, matching current behavior)
- Enum index bounds (0..values.len()-1 — **fixes the dead constraint**)
- Union variant dispatch (read `__discriminator`, look up variant
BAST definition, recurse)
- Boolean validity (materializer already checks, but the validator
should confirm)
3. Wire it into `validate_bytes()` as the validation step (replacing
the `jsonschema::Validator` call)
4. Run the **existing test suite** — the tests encode all current
expected validation behavior. If they pass, the POC succeeds.
**Success criteria:**
- All existing `validate_bytes` tests pass without modification to their
assertions (test *inputs* will change to BAST format, but the
expected validation outcomes must be identical)
- The union variant dispatch tests (OQ-008) pass — `maxLength` on a
`Bytes` field inside a union variant is enforced
- The enum index-bounds validation works (new behavior — currently
broken, so this is a fix, not a regression)
**Failure path:** If the POC reveals that the BAST-native validator
cannot cleanly express some constraint (e.g., a constraint that relies
on JSON Schema's structural keywords in a way that's hard to
reimplement), the fallback is the "structural-only + external JSON
Schema" model from the original Gap 1 — but this is unlikely given that
the materializer already guarantees structure, leaving only value-domain
checks.
**Out of scope for this POC:**
- `validate_json` — this path uses a standard `jsonschema::Validator`
from a consumer-provided JSON Schema, not the BAST-native validator.
No POC needed; it's a standard `jsonschema` usage.
- BAST document parsing / meta-schema validation — the parser is
straightforward JSON walking; no empirical risk.
- Layout computation — unchanged, already proven.
## POC Result — BAST-native validator
**Status: succeeded.** The POC is on branch `bast-validator-poc` in
`src/bast_poc.rs` (20 tests, all passing; full crate suite — 416 tests —
green; `cargo clippy --all-targets -- -D warnings` clean;
`cargo build --target wasm32-unknown-unknown --release` clean).
The POC implements the `validate_bytes` validation model from
[D-BAST-006](#d-bast-006-validate_bytes-validation-model) as a
self-contained module that does **not** touch the production schema /
materializer / validator paths. It reuses only `data_access` (read
primitives), `AlkTypeError` (error type), and `Endian`. The BAST
document parser, a packed-mode materializer, and the BAST-native
validator are all implemented from scratch — that is the point: prove
the model works end-to-end before refactoring the production code.
### What the POC proves
The hypothesis from the [POC section](#poc-bast-native-validator-for-validate_bytes)
is confirmed: a recursive walker over the BAST type tree fully
replaces the 19 custom keyword validators on the `validate_bytes` path,
including the OQ-008 union variant dispatch, and fixes the
enum-membership dead constraint — all without `jsonschema` custom
keywords and without an external JSON Schema.
The hard cases that were the actual de-risking targets all pass:
- **Union byte-offset discriminator + `maxLength` inside a variant
(OQ-008).** `union_byte_disc_max_length_inside_variant_enforced`
materializes a union with two `$ref` variants, dispatches on a
byte-offset `uint8` discriminator, and enforces `maxLength` on a
`bytes` field inside the selected variant. The validator reads
`__discriminator`, looks up the variant's BAST definition, and
recurses — same behavior as the current `UnionValidator`'s
per-variant sub-validators, but with no `jsonschema` involvement.
- **Union field-name discriminator + `maxLength` inside a variant.**
`union_field_disc_max_length_inside_variant_enforced` covers the
typedef.ts-style discriminator (a length-prefixed string field
selects the variant). Same recursion model.
- **Enum index-bounds fix.** `enum_index_out_of_bounds_rejected`
exercises the constraint that is **broken in the current engine**
(the built-in `enum` keyword checks string membership; the
materializer emits `Value::Number(index)`, which never matches — a
dead constraint). The BAST-native validator checks the materialized
index against the `values` array bounds (0..len-1), which is the
correct validation for a binary enum encoded as an index. Net
improvement, not a regression.
- **Nested struct wrapping a union wrapping a struct.**
`nested_struct_with_union_variant` confirms the recursion composes
through multiple type layers.
- **Arrays of fixed-size structs with `count`.**
`array_of_structs_with_count` covers the `Vector3`-style array
(D-BAST-004).
- **Records (count-prefixed string-keyed maps).**
`record_of_uint32` covers the `TRecord` shape.
- **Untrusted schema input.** `malformed_document_produces_schema_error_not_panic`
confirms a malformed BAST document surfaces as `AlkTypeError::Schema`,
not a panic (AGENTS.md §3).
- **Basic cases** (chunk header, int8/uint32 ranges, string/bytes
`maxLength`, timestamp, bool, short buffer) all pass — if a couple of
basic examples work, all of them do, since the validator is a flat
per-kind dispatch with no per-kind special-casing beyond the range
bounds.
### How the validator works
The validator is a single recursive function (`validate_typeref`) that
dispatches on the BAST `kind`. Each arm checks the value-domain
constraint for that kind and, for composites, recurses into the child
type definitions. The materializer (also implemented in the POC)
guarantees structural correctness — bounds, UTF-8, bool byte,
discriminator lookup, all fields present — so the validator only
enforces what the materializer cannot. The full constraint table is in
[`bast-format.md` §Validation Model](../architecture/bast-format.md#validation-model).
### Observations for the production implementation
1. **No `jsonschema` dependency for `validate_bytes`.** The validator
only needs `serde_json` (for `Value`) and the BAST document. The
`jsonschema` crate is still a direct dependency for `validate_json`
and for validating BAST documents against the BAST meta-schema, but
the `validate_bytes` path no longer touches it. This is a small wasm
binary-size win in addition to the architecture simplification.
2. **`$ref` resolution is a single hash lookup.** The POC's
`resolve_ref_or_inline` handles only `#/$defs/Name` pointers — the
only form BAST allows. The current engine's `normalize_refs` /
`inline_union_variant_refs` / `resolve_ref_or_inline` machinery for
bare-name refs and inlined union variants is no longer needed: BAST
`$ref`s are always full JSON Pointers, and union variant refs are
resolved lazily by the validator (the materializer already does this
for the read path). The `inline_union_variant_refs` compile step can
be removed entirely.
3. **The validator is ~250 lines.** The 19 custom keyword validators
(`src/validation.rs`) plus the macro definitions are ~500 lines and
require the `jsonschema::Keyword` trait plumbing (factory closures,
`Box<dyn Keyword>`, sub-validator construction at factory time). The
BAST-native validator is a flat match — no factories, no trait
objects, no sub-validator pre-computation. The recursion is direct.
4. **The `AlkTypeError::Validation` variant still wraps
`jsonschema::ValidationError<'static>`.** The POC uses
`jsonschema::ValidationError::custom` to construct these so the
error type is unchanged. This is now the decided shape for the
production refactor — see [D-BAST-009](#d-bast-009-alktypeerrorvalidation-payload-shape).
The rationale is consumer ergonomics: a single uniform payload type
means one match arm covers both `validate_json` and `validate_bytes`
errors downstream, and `validate_json`'s structured errors are worth
preserving rather than flattening to a `String`.
5. **The materializer and validator share the BAST-walking code
structure.** Both walk the same `kind`/`fields`/`mapping` tree. The
production refactor could share a typed BAST tree (a small
`BastNode` enum) between them so the walk is parsed once. The POC
parses lazily from the raw JSON in both passes to keep the model
honest; a typed tree is a straightforward follow-on optimization, not
a risk.
### Verdict
The "how do we reproduce the same behavior?" question is answered:
walk the BAST tree the same way the materializer does, checking the
same value-domain constraints the custom keyword validators check
today. The model is a strict simplification — fewer moving parts, no
`jsonschema` integration on the bytes path, no compile-time
`inline_union_variant_refs` step, no factory closures or trait
objects, and the enum dead-constraint is fixed as a side effect.
The POC does not wire into `AlkTypeEngine::validate_bytes` — that is
the production refactor ([implementation step 5](../plans/bast-implementation.md#step-5--bast-native-validator-production-version)),
which replaces `validation::build_validator` usage on the bytes path
with the BAST-native validator. The POC's job was to de-risk the model
before that refactor; that job is done.
## Decisions
The following were open questions in earlier drafts. Each is now
resolved. They are recorded here as decisions, not re-litigated. The
normative format specification that realizes these decisions is in
[`bast-format.md`](../architecture/bast-format.md); the semver
classification of each is in the
[implementation plan's Semver Contract](../plans/bast-implementation.md#semver-contract).
### D-BAST-001: Root type selection
**Decision:** Explicit. The root type name is a required parameter to
`compile()`: `AlkTypeEngine::compile(bast_doc, "ChunkHeader", Packed)`.
This is unambiguous and matches how consumers think about it ("compile
the ChunkHeader schema"). Convention (first entry in `$defs`) is fragile
and depends on JSON key order; a `$root` marker is redundant with an
explicit parameter.
### D-BAST-002: Primitive type string set
**Decision:** Lowercase (`"uint32"`, `"int8"`, `"float64"`, `"bool"`,
`"string"`, `"bytes"`, `"timestamp"`). Matches JSON Schema's own
convention (`"string"`, `"integer"`, `"boolean"`), is easier to type,
and is the convention in the TypeBox research examples. The
`AlkTypeKind` enum variants remain PascalCase in Rust — the mapping is
a simple `from_str()` impl.
### D-BAST-003: Top-level `$defs` requirement
**Decision:** Always `$defs`. Every BAST document has the same
top-level shape: `{ "$defs": { ... } }`. Single-type documents are a
special case with one entry. The `$defs` block is the namespace; the
root type name (D-BAST-001) selects the entry point. A bare struct at
the top level would be a special case with different parsing logic and
no home for additional definitions.
### D-BAST-004: Arrays of variable-length elements (deferred)
**Decision:** Arrays of variable-length elements are **not supported
in v1**. The meta-schema requires `count` on all array types, making
arrays fixed-size only. This matches the engine's current behavior (it
rejects arrays of variable-length elements) and aligns with OQ-001
(deferred, blocked on a concrete consumer needing interleaved
variable-stride arrays).
Variable-length collections are still available via `record` (a
count-prefixed string-keyed map), which the engine supports. If a
consumer needs a variable-length array of fixed-size elements, they
can use a record with integer-stringified keys as a workaround, or
wait for OQ-001 to be addressed.
### D-BAST-005: Field-name discriminator unions
**Decision:** Supported. The meta-schema includes an optional `fields`
array on `UnionDef`. When `discriminator.kind == "field"`, the
`fields` array provides the union's field definitions (including the
discriminator field). When `discriminator.kind == "byte"`, `fields` is
absent — the union's layout is purely the variant layout. This
preserves a feature the engine already supports. The meta-schema is in
[`bast-format.md` §The Meta-Schema](../architecture/bast-format.md#the-meta-schema).
### D-BAST-006: `validate_bytes` validation model
**Decision:** BAST-native validator. The `validate_bytes` path uses a
recursive walker over the BAST type tree to check value-domain
constraints on the materialized `Value` — no external JSON Schema
needed. This recovers the OQ-008 union variant dispatch behavior (the
validator recurses into the variant's BAST definition) and fixes the
enum-membership dead constraint (the validator checks the materialized
index against the `values` array bounds). See [`bast-format.md` §Validation
Model](../architecture/bast-format.md#validation-model) and the [POC](#poc-bast-native-validator-for-validate_bytes).
An optional external JSON Schema can be layered on top for constraints
BAST doesn't express (cross-field consistency, regex patterns on string
content). This is additive, not load-bearing.
### D-BAST-007: `validate_json` validation model
**Decision:** Standard JSON Schema validator. `validate_json` on
`AlkTypeEngine` validates a consumer-provided JSON `Value` against a
`jsonschema::Validator` compiled from a standard JSON Schema document
the consumer provides at compile time. No custom keywords. The BAST
document is not involved in this path — BAST describes bytes, not JSON
shape. This preserves the `validate_json` / `validate_bytes` symmetry
from ADR-010, but the two paths now use different validators (standard
`jsonschema` for JSON, BAST-native for bytes), reflecting their
different inputs and guarantees.
The JSON Schema may be authored separately or derived from BAST via
future codegen. For alkcall's channel 0 (JSON-RPC), the JSON Schema is
the `OperationSpec` schema, authored independently of any BAST
document.
### D-BAST-008: Builder API — two output formats
**Decision:** One builder, two build methods. The construction API is
the same (field names, types, annotations); only the output format
differs. `Schema::struct_().field(...).build()` → BAST JSON (binary
layout). `Schema::object().field(...).build()` → standard JSON Schema
(JSON validation). The builder already distinguishes AlkType kinds from
JSON Schema types via naming conventions (`string()` vs `string_()`).
Both output formats live in the same crate. This is the point of
alktype: one small wasm-compatible codebase that handles both binary
layout and JSON validation for protocol crates. alkcall uses both —
channel 0 is JSON (standard JSON Schema), binary channels use BAST.
Future crates (alktty, tunnels, sftp, git) will use BAST for their
binary formats. The codegen feature (future) will generate
readers/writers from BAST documents for these crates.
### D-BAST-009: `AlkTypeError::Validation` payload shape
**Status: decided.** Keep `Validation(jsonschema::ValidationError<'static>)`.
`AlkTypeError::Validation` currently wraps
`jsonschema::ValidationError<'static>`. Under the BAST pivot the
`validate_bytes` path no longer uses `jsonschema` at all (confirmed by
the [POC](#poc-result--bast-native-validator) — observation 1), so the
error payload on that path is constructed via
`jsonschema::ValidationError::custom` purely to keep the variant's type
unchanged. The two options were:
1. **Keep `Validation(jsonschema::ValidationError<'static>)`.** Simplest —
`ValidationError::custom` is public and `'static`, so the bytes path
can construct it without a real `jsonschema` validator. Cost: the
error type retains its `jsonschema` dependency even though the bytes
path no longer drives it. `validate_json` still uses `jsonschema`, so
the dependency isn't removable either way — but the error type
carries `jsonschema` only for one of its two callers.
2. **Introduce `Validation(String)` (or a small structured payload).**
Drops the `jsonschema` type from the public error enum. This is a
**semver-relevant public-API change** (the `Validation` variant's
payload type changes), so per AGENTS.md it requires an explicit
decision, not a drive-by. Benefit: the error type is
`jsonschema`-free, which matters if a future `no_std`/minimal build
wants to drop `jsonschema` from the bytes-only path (relates to
OQ-002).
**Rationale for option 1:** The deciding factor is consumer ergonomics
on the *combined* path. Consumers like alkcall use both `validate_json`
(channel 0, JSON-RPC) and `validate_bytes` (binary channels) and handle
`AlkTypeError::Validation` in one place. A single uniform payload type
means one match arm covers both sources — no `Validation(jsonschema_err)
vs Validation(string)` branching downstream. Option 2 would force
`validate_json` to flatten its structured errors (instance path, schema
path, keyword) to a `String` via `Display` just to match a bytes-path
shape — the more information-rich path loses data to accommodate the
less rich one. That is the wrong direction.
The `no_std`/minimal-build angle (OQ-002) that option 2 was meant to
enable is moot in practice: `validate_json` requires `jsonschema`
regardless, so a bytes-only `no_std` build already has to give up
`validate_json` as a separate, larger decision. Dropping the type from
one error variant does not unlock that build — the dependency is load-
bearing on the other validation path. The right place to revisit this is
when/if OQ-002 is actually pursued, not preemptively.
The POC already used option 1 (via `ValidationError::custom`); the
production refactor ([implementation step 5](../plans/bast-implementation.md#step-5--bast-native-validator-production-version))
follows the same construction pattern. No semver-relevant change to the
`Validation` variant.
## Risks and Mitigations
| Risk | Mitigation |
|------|-----------|
| BAST format doesn't cover all 19 type kinds | The format is designed to cover all 19. The meta-schema is the spec — if a kind can't be expressed, the meta-schema is wrong. |
| `$ref` resolution complexity moves from engine to schema authoring | BAST `$ref` values are always full JSON Pointers (`#/$defs/Name`). No normalization, no bare names. Resolution is a single hash lookup. |
| BAST-native validator misses a constraint the custom keywords enforced | The [POC](#poc-bast-native-validator-for-validate_bytes) runs the existing test suite, which encodes all current expected validation behavior. If a constraint is missed, a test fails before the pivot lands. |
| Losing OQ-008 union variant dispatch | The BAST-native validator recurses into the variant's BAST definition on `__discriminator` lookup — same behavior, no custom keywords. Covered by the POC. |
| Enum membership broken on bytes path | Already broken today (dead constraint). The BAST-native validator fixes it by checking the materialized index against the `values` array bounds. Net improvement. |
| `validate_json` loses custom keyword validation | `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema. Consumers that relied on custom keywords for JSON validation need to provide equivalent standard JSON Schema keywords. No real consumers exist yet. |
| Builder API output format change breaks consumers | No real consumers exist yet (v0.1.0 has zero adoption). The builder's public methods are unchanged; only the JSON output format changes. |
| Meta-schema maintenance burden | The meta-schema is small and changes rarely. It's embedded in the crate and published at a stable URL. |
## Future Directions
These are enabled by BAST but out of scope for the pivot itself. They
are mentioned to show that BAST makes them possible, not to commit to a
specific implementation timeline.
### Codegen
BAST enables code generation that the custom-keyword format makes
awkward. A codegen module (feature-gated behind `codegen`) would walk
`$defs` entries, inspect `kind` values, and emit Rust/TypeScript/Python
readers and writers via Handlebars templates. The typebox-rs `codegen/`
module (`/workspace/@alkimiadev/typebox-rs`) is the reference
architecture: `SchemaRegistry` for named types, `RustGenerator`/
`TypeScriptGenerator` wrapping `Handlebars`, `schema_to_rust_type()`/
`schema_to_ts_type()` mapping functions. alktype's codegen would follow
the same pattern but walk BAST `kind` values. The handlebars-rs
dependency is WASM-compatible. The pivot changes the schema format;
codegen builds on top of the new format.
### ABI Adapter
A BAST document describes the binary interface of a protocol — it is
essentially an ABI specification in JSON. This enables version
negotiation (two peers exchange BAST documents to agree on a protocol
version; the engine detects mismatches and either rejects or adapts),
schema migration (a consumer with schema v1 can read data written by
schema v2 if the changes are compatible), and WASM interop (a WASM
component can export its BAST schema as part of its WIT interface).
## References
- [`docs/architecture/bast-format.md`](../architecture/bast-format.md) —
the normative BAST format spec (meta-schema, TypeRef, examples,
validation model)
- [`docs/plans/bast-implementation.md`](../plans/bast-implementation.md) —
the execution plan (ordered steps, semver contract, ADR-sync checklist)
- [ADR-001](../architecture/decisions/001-alktype-purpose-scope-jsonschema-engine.md) — current "schema is the format" principle (to be superseded by ADR-BAST post-implementation)
- [ADR-003](../architecture/decisions/003-schema-annotations.md) — annotation semantics (carry forward to BAST unchanged)
- [ADR-009](../architecture/decisions/009-builder-api.md) — builder API (public surface unchanged, output format changes)
- [ADR-010](../architecture/decisions/010-generalized-validation-validate-bytes.md) — `validate_bytes` (unchanged in concept)
- `/workspace/research/typebox_research/ujsx/jpath.gen.ts` — TypeBox `Type.Module` pattern (the `$defs`/`$ref` model BAST follows)
- `/workspace/research/typebox_research/ujsx/mdast.gen.ts` — TypeBox cross-module references and composite types
- `/workspace/research/typebox_research/codegen/ts-to-module.ts` — TypeScript-to-TypeBox codegen (reference for future BAST codegen)
- `/workspace/@alkimiadev/typebox-rs/src/codegen/` — Rust/TypeScript codegen from schemas (reference architecture)
- `/workspace/alknet-typedef-poc/tests/sftp_roundtrip_test.rs` — SFTP POC proving byte-identical output (to be replicated with BAST)
- `/workspace/@alkdev/alkcall/src/channels/wire.rs` — hand-rolled chunk header (target for BAST replacement)
+419
View File
@@ -0,0 +1,419 @@
---
status: open
last_updated: 2026-08-15
reviewed_artifacts:
- src/lib.rs
- src/bast.rs
- src/bast_meta.rs
- src/bast_validation.rs
- src/builder.rs
- src/data_access.rs
- src/engine.rs
- src/error.rs
- src/layout_builder.rs
- src/materialize.rs
- src/offset_map.rs
- src/schema.rs
- src/sequential_reader.rs
- src/tunion.rs
- src/validation.rs
- src/macros.rs
- tests/{engine_integration,error_paths,poc_roundtrip,tunion_dispatch}.rs
- Cargo.toml
tool: manual source read + cargo test/clippy + cargo-llvm-cov
reviewer: post-BAST-pivot code review
---
# Code Review #003 — Post-BAST-Pivot Review
## Purpose
First logic/correctness review after the BAST pivot (the v0.1.0
`AlkType:*` custom-keyword JSON Schema format was replaced with the BAST
format; see `docs/plans/bast-implementation.md`). The pivot touched
every schema-walking path, so this pass re-reads the whole crate for
correctness, code smell, panic safety, and coverage — the same scope as
review #002, but against the new BAST surface.
Two things motivated this review beyond the routine sweep:
1. The pivot was a large, multi-step change (10 steps, 8 commits). A
couple of pre-existing bugs were fixed *during* the pivot (the enum
index-bounds dead constraint, the `write_bytes` u32 truncation), so
the same class of bug could be lurking in the newly-rewritten paths.
2. The publisher asked specifically for a coverage pass
(`cargo-llvm-cov`) with an eye toward *important* things being
covered rather than raw numbers.
## Methodology
- Full read of all 16 `src/*.rs` files (production + test modules) and
all 4 integration test files.
- `cargo test --release`, `cargo clippy --all-targets -- -D warnings`.
- `cargo llvm-cov --release` (summary + per-file + uncovered-lines) to
attribute coverage gaps to specific code paths.
- Targeted reproduction of the suspicious paths (field-level endian
override, aligned-mode variable-length/array materialization) via
throwaway integration tests.
- Cross-reference every error path against its caller to confirm errors
propagate (not swallowed) and carry useful attribution.
- Read `docs/reviews/002-code-review.md` for prior context and
resolved/unresolved items.
## Verification Baseline
All verification run on the reviewed tree (commit `562284f`):
- `cargo test --release`: **389 tests pass** (312 crate unit tests +
77 integration tests across 4 files). Zero failures.
- `cargo clippy --all-targets -- -D warnings`: **clean**.
- `cargo llvm-cov --release`: **90.14% line coverage** (7903/8682),
**86.68% function coverage** (743/842). Per-file breakdown below.
- No `unsafe` anywhere in the crate.
- No `TODO`/`FIXME`/`HACK`/`XXX` markers in source.
- All `unwrap`/`expect`/`panic!`/`unreachable!` are confined to
`#[cfg(test)]` modules, verified by line-context cross-reference.
### Coverage breakdown
| Module | Lines | Functions |
|---|---:|---:|
| bast.rs | 88.3% | 90.3% |
| bast_validation.rs | 93.0% | 90.5% |
| builder.rs | 91.3% | 87.6% |
| data_access.rs | 83.5% | 76.2% |
| engine.rs | 96.9% | 98.3% |
| layout_builder.rs | 91.6% | 81.8% |
| materialize.rs | **81.8%** | **76.1%** |
| offset_map.rs | 89.5% | 79.2% |
| sequential_reader.rs | **84.6%** | **75.0%** |
| tunion.rs | 92.7% | 91.7% |
| **TOTAL** | **90.1%** | **86.7%** |
The low-function-count modules are not test-helper noise — they are
exactly where the correctness bugs below live. The uncovered lines in
`materialize.rs` and `sequential_reader.rs` are the aligned-mode
variable-length/array paths and the field-level-endian paths, which are
**untested and broken** (see M1, M2). The `data_access.rs` 76% function
coverage is mostly the `read_*_indirect` family, which has no production
caller (see L1).
## Summary Statistics
| Severity | Count |
|----------|------:|
| Critical | 0 |
| Medium | 3 (M1, M2, M3) |
| Low | 2 (L1, L2) |
| Nit | 4 (N1, N2, N3, N4) |
No critical findings. The crate is in good shape, but the three Medium
findings are **silent data-corruption / silent-misinterpretation bugs**
in the newly-rewritten paths — they must be fixed before the next
release. The Low findings are dead code and a robustness gap; the Nits
are hygiene.
---
## Findings
### M1. Field-level `endian` override is ignored by the reader and aligned materializer
**Files**: `src/sequential_reader.rs:290`, `src/engine.rs:338,455`,
`src/materialize.rs:461`
**Problem**: The spec documents per-field endian override
(`docs/architecture/bast-format.md` §Endianness — "Field-level `endian`
overrides the struct/union default"), and the packed materializer honors
it (`materialize.rs:120` uses `field.effective_endian(endian)`). But
three paths use only the struct-level endian:
- `sequential_reader.rs:290` `read_field_value` — uses `self.endian`,
never `field.effective_endian`.
- `engine.rs:338` `read_field` and `engine.rs:455` `write_field` —
`let endian = self.endian;`.
- `materialize.rs:461` `materialize_struct_aligned` — passes the struct
`endian` to `materialize_leaf_at`, never `field.effective_endian`.
**Reproduction** (throwaway integration test, confirmed): a big-endian
struct with a `"crc": { "kind": "uint32", "endian": "little" }` field
reads `0x01020304` as `67305985` (big-endian interpretation) in both
`sequential_reader` and `read_field`. The bytes are correct; the
interpretation is wrong — silent data corruption.
**Fix**: thread `field.effective_endian(endian)` through all three paths.
`read_field_value` already receives the `BastField`; `engine::read_field`
/ `write_field` need to look up the field's effective endian (they
already walk the BAST tree via `lookup_field_kind`); `materialize_struct_aligned`
needs to pass `field.effective_endian(endian)` to `materialize_leaf_at`
instead of the struct default.
**Lift**: closes a silent-corruption path on a documented feature. Small
effort (~10 lines + regression tests).
---
### M2. Aligned-mode `validate_bytes` is broken for arrays, `maxLength` fields, and `offset-indirect` fields
**File**: `src/materialize.rs:461-512` (`materialize_struct_aligned`)
**Problem**: `materialize_struct_aligned` routes fixed-size and
variable-length leaves through `materialize_leaf_at`, which calls
`materialize_typeref_packed` — i.e. it always reads a **length-prefixed**
value. But the aligned `OffsetMap` stores three different shapes:
- `maxLength` fields are a raw reservation (no length prefix) — the
materializer reads the first 4 bytes of the *data* as a length prefix.
Reproduced: `"hello"` in an 8-byte reservation read a length of
`1819043180` and failed with a bounds error.
- `offset-indirect` fields are an 8-byte `{offset, length}` pair — read
as a length prefix, garbage.
- arrays are recorded as `vals[0]`/`vals[1]` entries only, so
`offset_map.get("vals")` returns `None` → `Offset` error. Reproduced.
Only the default inline length-prefixed variable field (and only as the
final field, per ADR-006) works in aligned mode. There are **no tests**
covering aligned `validate_bytes` with arrays or non-default variable
encodings — that is why this slipped through the pivot.
**Fix**: two options, decide with the publisher:
1. **Implement** aligned materialization for the three shapes: read
`maxLength` fields as a fixed-size slice, `offset-indirect` fields via
`read_*_indirect` (with a data region), and arrays by iterating the
`vals[i]` offset-map entries.
2. **Reject** these combinations at compile time (return
`AlkTypeError::Schema`/`Offset` from `OffsetMap::compute` or
`compile`) if they are out of scope for v1, so the failure is loud
and at load time rather than a silent misread at access time.
Option 2 is the smaller, safer fix and matches the existing ADR-006/
ADR-008 pattern of rejecting unsupported aligned-mode combinations. The
`offset-indirect` encoding is already dead code on the read path (see
L1), which argues for rejecting it in aligned mode until it is actually
implemented.
**Lift**: closes a silent-misread path. Medium effort either way.
---
### M3. The BAST meta-schema is never applied at compile time; annotation parsers silently tolerate malformed values
**Files**: `src/engine.rs:123` (`compile`), `src/bast.rs:899-926`
(`parse_endian_opt`, `parse_align`, `parse_encoding`, `parse_max_length`)
**Problem**: `BAST_META_SCHEMA` is exported and self-tested, but
`AlkTypeEngine::compile` never validates the document against it. The
parser (`bast.rs`) is the only gate, and it silently tolerates malformed
annotations:
- `parse_endian_opt` (`bast.rs:899`) — `"endian": "middle"` → `None` →
silently defaults to little.
- `parse_encoding` (`bast.rs:921`) — unknown encoding → silently
`LengthPrefixed`.
- `parse_align` / `parse_max_length` — non-integer / negative → silently
`None`.
These are exactly the cases the meta-schema's `enum` / `minimum`
constraints exist to reject. Per AGENTS.md §3, schemas are untrusted
input (the `alkcall` consumer accepts them from arbitrary internet
peers). A malicious peer can send `"endian": "bogus"` and get a
silently-misinterpreted layout instead of a `Schema` error.
**Fix**: validate the document against `BAST_META_SCHEMA` in `compile`
(one-time, cheap — the meta-schema is a `LazyLock<Value>`), *or* make
the annotation parsers return `Err(AlkTypeError::Schema)` on unknown
values. The meta-schema route is preferred: it is the single source of
truth and catches the whole class of malformed-annotation bugs at once.
**Lift**: closes a silent-misinterpretation path on untrusted input.
Small effort (~5 lines + tests).
---
### L1. `offset-indirect` is dead code on the read path
**Files**: `src/data_access.rs:307-355`, `src/engine.rs:388-399`
**Problem**: `data_access::read_string_indirect` / `read_bytes_indirect`
have no production caller (only their own unit tests). `engine.read_field`
always calls `read_string` / `read_bytes` (length-prefixed) regardless of
the field's `encoding`. So a schema declaring
`"encoding": "offset-indirect"` compiles and lays out correctly in the
offset map, but can never be read back.
This is the same root cause as M2's `offset-indirect` arm. Decide
together with M2: either wire `read_*_indirect` into the read path (and
the aligned materializer), or drop the `offset-indirect` encoding
entirely until a consumer needs it. Leaving it half-wired is the worst
state — it looks supported but silently misreads.
**Lift**: removes dead code or completes a feature. Small effort.
---
### L2. `materialize_packed` / `materialize_aligned` take a dead `endian` parameter
**File**: `src/materialize.rs:44-88`
**Problem**: both functions take `endian: Endian` and immediately
`let _ = endian;`, using `struct_node.endian()` instead. The caller's
`self.endian` (from `engine.rs:285,287`) is ignored. The signature is
misleading — a reader assumes the passed endian is honored.
**Fix**: drop the parameter and read the endian from the root struct
inside the function (it already does). ~4 lines. Purely a clarity fix;
no behavior change.
---
### N1. `number_from_f64` maps NaN/Inf to `Value::Null`
**File**: `src/materialize.rs:451-455`
**Problem**: `serde_json::Number::from_f64` returns `None` for NaN/Inf,
so `number_from_f64` substitutes `Value::Null`. A NaN float in the buffer
then surfaces as "expected a number" from `validate_float`
(`bast_validation.rs:176`) rather than "expected a finite number". The
error is misleading, though the outcome (rejection) is correct.
**Fix** (optional): have the materializer propagate a non-finite float
as an `AlkTypeError::Access` at read time, or leave as-is and accept the
slightly-off error message. Not a correctness bug.
---
### N2. `BastType::alk_kind()` returns `Struct` for any `$ref`
**File**: `src/bast.rs:715-725`
**Problem**: `BastType::Ref(_) => AlkTypeKind::Struct` is documented but
a footgun — a `$ref` to a union/enum misreports its kind unless the
caller resolves first. Most callers do resolve first, but the invariant
is fragile and easy to break in a future edit.
**Fix** (optional): leave as-is (documented) or make `alk_kind` return
`Option<AlkTypeKind>` / require resolution. Defer unless it bites.
---
### N3. `check_bytes` accepts both `String` and `Array` forms
**File**: `src/bast_validation.rs:213-256`
**Problem**: `check_bytes` handles `Value::String` and `Value::Array`,
but the materializer only ever emits `Array` for bytes
(`materialize.rs:212`). The `String` arm is dead/legacy. Harmless, but
it widens the accepted surface for no reason.
**Fix** (optional): drop the `String` arm, or keep it if a future
materializer emits bytes as a string. Defer.
---
### N4. Stale ADR references in doc comments
**Files**: `src/error.rs:3` ("ADR-098"), `src/tunion.rs:1` ("ADR-097"),
`src/engine.rs:51,196` ("ADR-101"), `src/offset_map.rs:1`,
`src/layout_builder.rs:1`, `src/sequential_reader.rs:1` ("ADR-096")
**Problem**: none of these ADR numbers exist in
`docs/architecture/decisions/` (which has 001–010 + `bast-bast-format` +
`val-split-two-validator-model`). The pivot renumbered/renamed ADRs but
the code comments were not synced. A reader following the reference hits
a dead end.
**Fix**: map each stale reference to the correct ADR (e.g. "ADR-096" →
ADR-002 for the two layout modes, "ADR-101" → ADR-007 for the packed
read factory, "ADR-098" → ADR-004 for error handling, "ADR-097" →
ADR-003 for annotations) and update the comments. ~6 lines.
---
## The `Timestamp` kind
`AlkTypeKind::Timestamp` is a first-class kind that is byte-identical to
`String` everywhere (length-prefixed UTF-8), and its only distinguishing
behavior is `is_rfc3339_timestamp` (`bast_validation.rs:406`) — a
hand-rolled, non-strict check that the doc itself admits "Feb 31 passes;
seconds range isn't checked; leap seconds aren't handled." The parsing
is fragile (the `rfind('-')` timezone-offset heuristic, no
fractional-second handling).
It is documented as matching v0.1.0, so it is not a regression, but it
is the weakest part of the validator and adds a 19th kind plus a
`needs_endian` / `is_variable_length` / `natural_alignment` arm, all to
validate a string that a consumer could validate with a standard JSON
Schema `format: "date-time"` on the `validate_json` path.
**Decision (publisher)**: remove it. It is a residual from an early
research reference that included a timestamp; it is largely irrelevant
at the BAST level, and JSON-level timestamp validation is `jsonschema`'s
job, not alktype's. Tracked as a follow-up task, not part of this
review's findings.
---
## What's Good
The crate is in notably good shape after the pivot. Highlights:
- **The BAST parser is clean and defensive.** `bast.rs` returns
`AlkTypeError::Schema` on every malformed-document path, never panics,
and uses `checked_add` / `usize::try_from` for all count/offset casts.
The typed tree (`BastDoc`/`BastDef`/`BastType`) is a real improvement
over the v0.1.0 raw-JSON accessors.
- **The enum index-bounds fix is correct.** `validate_enum`
(`bast_validation.rs:274`) checks the materialized index against
`values.len()`, closing the v0.1.0 dead constraint. Well-tested.
- **Overflow safety is thorough.** `checked_add` everywhere in the hot
paths; the `write_bytes` u32 truncation from review #002 (M2) is
fixed and the guard pattern is now the norm.
- **Error attribution is excellent.** Every `Access`/`Offset` error
carries a `field_path`; the BAST parser errors carry a dotted path
into the document (`"bast: struct at .fields[2] ..."`).
- **The two-validator split is clean.** `bast_validation` (bytes) and
`validation` (JSON) are clearly separated, and the
`AlkTypeError::Validation` payload stays uniform across both
(D-BAST-009).
- **Tests are strong where they exist.** 389 tests, good coverage of
error paths, both endiannesses, short buffers, unknown discriminators,
invalid UTF-8. The gaps are precisely the paths M1/M2 identify.
- **No `unsafe`, no `TODO`/`FIXME`** — clean codebase hygiene.
---
## Recommended Order
1. **M1** (field-level endian override) — ~10 lines + tests, closes a
silent-corruption path on a documented feature. Smallest and
highest-value.
2. **M2** (aligned-mode materialization) — decide implement-vs-reject
with the publisher; the reject option is small and matches the
ADR-006/ADR-008 pattern.
3. **M3** (meta-schema at compile time) — ~5 lines + tests, closes a
silent-misinterpretation path on untrusted input.
4. **L1** (offset-indirect dead code) — decide together with M2.
5. **L2** (dead `endian` parameter) — ~4 lines, clarity only.
6. **N1–N4** — hygiene; N4 (stale ADR refs) is worth doing in the same
pass as the `Timestamp` removal since both touch doc comments.
The `Timestamp` removal is a separate, self-contained task the publisher
has already decided on; it can be done independently of the above.
---
## Notes
- All line numbers refer to the tree at commit `562284f` (the last
commit on `main` at review time).
- The coverage numbers are from `cargo llvm-cov --release` on the same
tree. The `--summary-only` and `--show-missing-lines` outputs were
used to attribute gaps; the full HTML report is at
`target/llvm-cov/html`.
- This review does not cover documentation quality (README, inline docs,
docs.rs rendering) beyond the stale-ADR-reference nit (N4). Per the
publisher's workflow, that is a separate sweep.
- Findings M1 and M2 were confirmed by throwaway integration tests that
were removed after reproduction; the regression tests for the fixes
should be added to the permanent suite.
+412
View File
@@ -0,0 +1,412 @@
---
status: closed
last_updated: 2026-09-02
reviewed_artifacts:
- src/sequential_reader.rs
- src/bast.rs
- src/engine.rs
- src/layout_builder.rs
- src/materialize.rs
- src/offset_map.rs
- src/data_access.rs
- src/lib.rs
- docs/architecture/decisions/007-packed-mode-read-factory.md
- ../@alkdev/alktty/benches/wire_vs_bast.rs
tool: manual source read + downstream criterion bench (`cargo bench --bench wire_vs_bast` in alktty)
reviewer: read-path performance review (triggered by alktty wire_vs_bast bench)
---
# Review #004 — Read-Path Performance: the `BastDoc` Re-Parse Gap
## Purpose
A downstream bench in `alktty` (`benches/wire_vs_bast.rs`) compared a
hand-rolled `ChunkHeader` codec against an alktype-driven codec built
from the same 2-field BAST struct (`stream_type: uint8`,
`length: uint32`, big-endian). The bench builds the engine / layout /
reader **once** outside the measured loop, then measures per-chunk read
and write over 1024 contiguous chunks.
The write path is competitive (~2.8x at 64 B, ~1.1x at 4 KiB — the
per-chunk overhead is just `data_access::write_u8`/`write_u32` at
precomputed `PackedLayout` offsets, and the payload copy dominates at
4 KiB). The read path is **~400x slower per chunk** (2.27 µs/chunk vs
5.6 ns/chunk hand-rolled), and the cost is fixed — it dominates at 64 B
*and* at 4 KiB.
This review traces the gap to its source, confirms it is an
implementation gap (not inherent to the design), and lays out the fix
options. The bench itself is honest — its in-file note
(`wire_vs_bast.rs:36-41`) already points at the root cause; this review
formalizes the finding and the remediation plan.
## Methodology
- Full read of the read-path code (`sequential_reader.rs`, `bast.rs`,
`engine.rs`), the write path (`layout_builder.rs`, `data_access.rs`),
and ADR-007 (the packed read factory decision).
- Cross-reference every `BastDoc::new` call site in `src/` to map the
full re-parse surface.
- Trace the lifetime/ownership constraint that forces the re-parse
(`BastDoc<'a>` borrows `&'a Value`; `SequentialReader` owns its
`Value` — self-referential struct, cannot cache the parsed tree).
- Read the downstream bench to confirm the measurement is honest (the
re-parse is inside the measured routine; the engine/reader are built
once outside it).
- Read `docs/reviews/003-code-review.md` for prior context — review #003
did not flag the re-parse (it was a correctness/coverage pass, not a
performance pass).
## Verification Baseline
The bench numbers below are from the alktty downstream tree
(`benches/wire_vs_bast.rs`, criterion 0.7), run against alktype at
commit `cab4932` (v0.2.0). The alktype tree itself is unchanged — this
is a review of existing code, not a fix.
| path | payload=64 B | payload=4 KiB |
|---|---|---|
| hand-rolled read | 5.7 µs (5.6 ns/chunk) | 12.1 µs (11.8 ns/chunk) |
| alktype read | 2.32 ms (2.27 µs/chunk) | 2.38 ms (2.33 µs/chunk) |
| hand-rolled write | 14.3 µs | 321 µs |
| alktype write | 39.8 µs | 355 µs |
One-shot startup costs (paid once): `AlkTypeEngine::compile` 573 µs,
`engine.sequential_reader()` 2.84 µs, `LayoutBuilder::build` 1.30 µs.
The read gap is ~400x per chunk and is **fixed** (does not shrink as
the payload grows), which is the signature of per-field overhead, not
payload-copy overhead.
## Summary Statistics
| Severity | Count |
|----------|------:|
| High | 1 (H1) |
| Medium | 1 (M1) |
| Low | 2 (L1, L2) |
| Nit | 1 (N1) |
The High finding is the read-path re-parse — a severe performance
regression that blocks the primary intended use case (alktype as a
runtime codec for protocol wire formats). It is not a correctness bug;
the re-parse produces the same result every time, it is just
catastrophically wasteful. The Medium finding is the same root cause
manifesting in four one-shot paths. The Low/Nit findings are dead code
and a stale doc comment exposed while tracing the root cause.
---
## Findings
### H1. `SequentialReader::read_field_at` re-parses the BAST typed tree on every field read
**File**: `src/sequential_reader.rs:262`
**Problem**: `read_field_at` calls `BastDoc::new(&self.doc_value,
&self.root_name)?` on every field read. `BastDoc::new`
(`src/bast.rs:69`) recursively parses the root def: `lookup_def_raw`
(hash lookup into the `serde_json` Map), `BastDef::parse` →
`BastStruct::parse` → allocates a `Vec<BastField>`, iterates fields,
`BastField::parse` each (several `node.get().as_str()`/`as_u64()`
calls + a `format!` allocation for the field path), `BastType::parse`
each. For the 2-field `ChunkHeader`, that is ~1 µs per call.
`read_next` calls `read_field_at` once per field, so a 2-field header
incurs **two** `BastDoc::new` calls per chunk = ~2.27 µs/chunk, matching
the bench. For an N-field struct, reading all fields is O(N²) in parse
work (each of N reads re-parses all N fields) — the gap widens with
struct width.
The re-parse is purely redundant: `SequentialReader::new`
(`src/sequential_reader.rs:130`) already parsed the `BastDoc` once at
construction and extracted the field list. `read_field_at` re-derives
the same `field_node` (the `BastField` at `index`) and the same `doc`
that were already in hand at construction time.
**Root cause**: a lifetime/ownership tension, not a logic error.
`BastDoc<'a>` borrows `&'a Value` and `&'a str` throughout
(`src/bast.rs:51-55`). `SequentialReader` **owns** a cloned `Value`
(`doc_value`, `src/sequential_reader.rs:111`). To cache the parsed
`BastDoc` in the reader, the `BastDoc` would have to borrow from
`self.doc_value` — a self-referential struct, which Rust's borrow
checker forbids and safe Rust cannot express without a self-referential
crate (`self_cell`/`ouroboros`, both rejected: new dep per AGENTS.md
§7, and `self_cell`'s soundness relies on `unsafe` the crate avoids per
AGENTS.md §11). So the only way to get a `BastDoc` at read time is to
re-parse from the owned `Value`. The engine's own doc comment
(`src/engine.rs:112-115`) acknowledges this design explicitly:
> "The engine retains a clone of the BAST `Value` so that
> `sequential_reader` and `read_field` can re-parse the typed tree on
> demand without lifetime entanglement with the caller's `Value`."
The "re-parse on demand" framing was a lifetime-entanglement workaround
that did not anticipate the hot-loop cost. ADR-007's "Cost" section
(`decisions/007-packed-mode-read-factory.md:64-70`) argues construction
is cheap ("a `Vec<(String, Value)>` of the `properties` entries ... a
struct has a small number of fields") — true for *construction* (once),
but the decision did not account for a per-field re-parse inside the
read loop.
**Why the write path is fine**: `LayoutBuilder::build`
(`src/layout_builder.rs:190`) also re-parses `BastDoc::new`, but it is
called **once** per write — the resulting `PackedLayout` offsets are
cached and reused. The write loop then does only `data_access::write_*`
at fixed offsets. There is no per-field re-parse on the write hot path,
which is why write is ~1.1x at 4 KiB.
**Fix**: cache the parsed typed tree so the read path does not re-parse.
See "Fix Options" below for the three approaches and the
recommendation.
**Lift**: closes a ~400x read-path gap and unblocks alktype as a runtime
codec for protocol wire formats (its stated purpose per ADR-001). This
is the difference between "plausible runtime codec" and "not viable."
Large effort depending on the chosen option.
---
### M1. The same re-parse pattern exists in four one-shot paths
**Files**: `src/layout_builder.rs:190`, `src/engine.rs:284,334,467`
**Problem**: `BastDoc::new(&self.doc_value, &self.root_name)?` /
`BastDoc::new(&self.bast_doc, &self.root_name)?` is re-called in:
- `LayoutBuilder::build` (`src/layout_builder.rs:190`) — once per
`build()` call. Re-parsing on each build is wasteful if a builder is
reused across writes, but the typical pattern is build-once-reuse,
so this is mild.
- `AlkTypeEngine::validate_bytes` (`src/engine.rs:284`) — once per
validate call. For a stream of buffers, this is a per-buffer re-parse.
- `AlkTypeEngine::read_field` (`src/engine.rs:334`, aligned mode) — once
per field read. Same class as H1 but for aligned random access, and
one parse per field read (not the O(N²) of the sequential reader).
- `AlkTypeEngine::write_field` (`src/engine.rs:467`, aligned mode) —
once per field write.
These are less acute than H1 (one parse per operation, not per-field-
in-a-loop), but they share the same root cause: the engine/builder own
a `Value` and cannot cache a borrowing `BastDoc`. Any fix that makes
`BastDoc` cacheable on an owning struct (Fix Option A) closes these for
free; a read-path-only fix (Option B) leaves them as-is, which is
acceptable since they are not hot loops.
**Lift**: removes redundant parse work on the validate/aligned paths.
Free with Option A; deferred with Option B.
---
### L1. `read_field_value` carries a dead `_field_schema` parameter; `fields` stores dead `Value` clones
**Files**: `src/sequential_reader.rs:114,294`
**Problem**: `SequentialReader` stores `fields: Vec<(String, Value)>`
where the `Value` is `f.source().clone()` per field
(`src/sequential_reader.rs:145`). `read_field_at` passes this as
`_field_schema` to `read_field_value` (`src/sequential_reader.rs:294`),
where it is unused (prefixed `_`). The raw `Value` clone per field is
dead weight — only the `String` name is used (for `read_next`'s return
and `read_field`'s lookup). This is a minor allocation cost on top of
H1's re-parse, and it will be removed naturally when the read plan is
precomputed (Option B) or the `BastDoc` is cached (Option A), since
both replace `Vec<(String, Value)>` with typed/owned field data.
**Lift**: trivial; falls out of the H1 fix.
---
### L2. ADR-007 "Cost" section and the engine doc comment understate the re-parse
**Files**: `docs/architecture/decisions/007-packed-mode-read-factory.md:64-70`,
`src/engine.rs:112-115`
**Problem**: ADR-007's "Cost" section argues `sequential_reader()` is
cheap because it clones a small `Vec` of field schemas. That is true for
the factory call (once). But the decision did not anticipate that
`read_field_at` would re-parse `BastDoc::new` per field — the cost that
actually dominates. The engine doc comment at `src/engine.rs:112-115`
explicitly frames re-parse-on-demand as the intended design ("re-parse
the typed tree on demand without lifetime entanglement"), which is the
root cause H1 traces.
**Fix**: whichever fix option is chosen, update ADR-007's "Cost" /
"Consequences" section and the engine doc comment to reflect that the
parsed tree is now cached (Option A) or precomputed into a read plan
(Option B), and that the "re-parse on demand" framing is retired.
**Lift**: documentation accuracy; prevents the same framing from
misleading a future edit.
---
### N1. `BastType::alk_kind()` returns `Struct` for any `$ref` (carry-forward from review #003 N2)
**File**: `src/bast.rs:715-725`
**Problem**: flagged in review #003 N2 and left as "defer unless it
bites." It does not bite here — `read_field_value` always calls
`doc.resolve_typeref(ty)` before matching on `BastType`, so the
misreporting `alk_kind` is never consulted on a `Ref`. Noting it only
because this review re-read the same path; no new action beyond review
#003's deferral.
---
## Fix Options
The core constraint: `BastDoc<'a>` borrows `&'a Value` / `&'a str`; an
owning struct (`SequentialReader`, `AlkTypeEngine`, `LayoutBuilder`)
cannot store a `BastDoc` that borrows from its own `Value` field
(self-referential). Three ways to break the constraint:
### Option A — Make the typed tree own its data (principled fix)
Change `BastDoc<'a>` → `BastDoc` (no lifetime), `&'a str` → `Arc<str>`
(or `String`), `&'a Value` → `Arc<Value>` (or `Value`). Then
`SequentialReader`, `AlkTypeEngine`, and `LayoutBuilder` each hold a
`BastDoc` directly (built once at construction), and `read_field_at`
uses `&self.doc` — no re-parse, anywhere.
- **Closes**: H1, M1 (all four one-shot paths), and the engine/reader
lifetime entanglement that ADR-007 worked around. The engine's
`bast_doc: Value` clone (`src/engine.rs:85`) becomes redundant with
the owned `BastDoc`.
- **Tradeoff**: broad refactor. Touches `bast.rs` (every typed node)
and every consumer (`layout_builder`, `offset_map`, `materialize`,
`sequential_reader`, `tunion`, `bast_validation`, `engine`). The
`Bast*` types are re-exported in `lib.rs:57-60`, so this is a
**breaking public-API change** — `BastDoc<'a>` becomes `BastDoc`,
and every method signature that took `&'a` changes. Per AGENTS.md,
this is semver-relevant and would need a version bump (0.2.0 → 0.3.0).
- **Dependency cost**: `Arc<str>`/`Arc<Value>` add `alloc` (already in
use via `Vec`/`String`); no new external deps. Stays wasm-clean. The
`preserve_order` serde_json feature remains load-bearing (AGENTS.md
§8) — owning the `Value` does not change field-order semantics.
- **Effort**: large but mechanical. The borrow-based design was chosen
for "allocation-free beyond the small typed nodes" (`src/bast.rs:11-
18`), but the re-parse-per-field already defeats that goal by
allocating a fresh `Vec<BastField>` per read. Owning the data makes
the "parse once, walk many times" invariant actually hold.
### Option B — Precompute an owned read plan in `SequentialReader::new` (surgical fix)
Keep `BastDoc<'a>` borrowing for the other consumers. In
`SequentialReader::new`, parse the `BastDoc` once, resolve all `$ref`s
eagerly, and build a flat, owned tree of read instructions
(`Vec<FieldPlan>`) that the read loop walks with no `BastDoc`
involvement. Each `FieldPlan` carries the field name, the resolved
`AlkTypeKind`, the effective `Endian`, and for composites a nested
plan (struct → sub-plans; union → discriminator + per-variant plans;
array → element plan + count; record → value plan).
- **Closes**: H1 only. M1 (the one-shot re-parses) remains, which is
acceptable since they are not hot loops.
- **Tradeoff**: non-breaking (internal to `sequential_reader.rs`; the
public `SequentialReader` type and its methods keep their
signatures). Duplicates some of the `BastType` matching logic that
`read_field_value`/`read_union_value`/etc. already encode, so there
are two parallel walkers to maintain.
- **Effort**: medium. Self-contained in one file but non-trivial
(composites require recursively resolving and pre-flattening the
type tree, including `$ref` chains into `$defs`).
### Option C — Cache `BastDoc` on the engine, reader borrows (rejected)
Have the engine own the parsed `BastDoc` and return a
`SequentialReader<'_>` that borrows from `&self`. This requires
`BastDoc` to be owned (Option A prerequisite) *and* changes
`sequential_reader() -> Option<SequentialReader>` to
`-> Option<SequentialReader<'_>>` — a breaking public-API change that
also contradicts ADR-007's "owned fresh reader" decision. Strictly
worse than Option A (same refactor cost, more API churn, contradicts
an ADR). Rejected.
### Recommendation
**Option A**, given the publisher's stated willingness to make breaking
changes ("no one is using this except us yet; ... we can change things
now"). It is the only option that closes H1 *and* M1 and retires the
lifetime-entanglement workaround that caused both. The refactor is
broad but mechanical (lifetime removal, not logic rewrites), and the
crate is pre-1.0 with only two in-house downstream consumers
(`alktty`, `alkcall`), so the breakage cost is bounded and known.
Option B is the fallback if the Option A refactor is deferred — it
closes the acute H1 gap non-breakingly while leaving M1 for later. It
is not the recommended path because it leaves a second parallel type
walker in the crate and does not address the root cause (the borrow-
based `BastDoc` design), which will keep forcing re-parses anywhere a
new owning consumer wants to cache the parsed tree.
Regardless of the chosen option, ADR-007's "Cost"/"Consequences"
section and the `src/engine.rs:112-115` doc comment should be updated to
retire the "re-parse on demand" framing (L2).
---
## What's Good
- **The bench is honest.** The alktty bench builds the engine/reader
once outside the measured loop and correctly isolates the per-chunk
logic. Its in-file note (`wire_vs_bast.rs:36-41`) already points at
the `read_field_at` re-parse and labels it "the honest current cost of
the alktype read path, not a bench bug." This review confirms that
assessment.
- **The write path is already competitive.** Once `PackedLayout` is
built, the write loop is just `data_access::write_*` at fixed offsets
— no schema walk, no re-parse. This validates the "build once, reuse"
pattern that the read path should also adopt.
- **`resolve_typeref` for primitives is cheap.** `BastType::Primitive`
is `Copy` (`AlkTypeKind: Copy`, `src/schema.rs:25`), so
`resolve_typeref` (`src/bast.rs:117-125`) returns `other.clone()`
without allocation for the common case. The re-parse cost is entirely
in `BastDoc::new`, not in the per-field type resolution — so caching
the `BastDoc` alone closes the gap without restructuring
`resolve_typeref`.
- **Overflow safety and error attribution are unaffected.** The
`checked_add` / `usize::try_from` discipline (AGENTS.md §4) and the
`field_path`-carrying errors (review #003 "What's Good") are in the
read functions, not the parser — a caching fix preserves them.
---
## Recommended Order
1. **H1 + M1 (Option A)** — the owned-typed-tree refactor. Decide
first (this is a one-way door: breaking public-API change, version
bump to 0.3.0). If approved, this is one refactor that closes both.
2. **L2** — update ADR-007 and the engine doc comment in the same
commit as the H1 fix, since the "re-parse on demand" framing is
being retired.
3. **L1** — falls out of the H1 fix (the dead `Value` clones are
replaced by the cached/owned field data).
4. **N1** — remains deferred per review #003.
If Option A is deferred, **H1 (Option B)** is the standalone
alternative — non-breaking, closes the acute gap only.
---
## Notes
- All line numbers refer to the tree at commit `cab4932` (v0.2.0, the
BAST pivot release).
- The bench is in the `alktty` downstream repo
(`/workspace/@alkdev/alktty/benches/wire_vs_bast.rs`), not in alktype.
alktty depends on alktype as a path dev-dep for the bench only; it
does not use alktype at runtime. The path dep means
`cargo publish --dry-run` for alktty would complain (the bench is
exploratory and uncommitted in alktty; the alktype crate itself has
no bench dependency).
- This review does not cover the wasm build (`cargo build --target
wasm32-unknown-unknown`) because the fix is not yet implemented; the
verification block for the fix should include it per AGENTS.md, as
the typed-tree ownership change touches `bast.rs` which is
wasm-relevant.
- Review #003 (post-BAST-pivot correctness review) did not flag the
re-parse — its scope was correctness, coverage, and panic safety, not
performance. The re-parse is not a correctness regression; the parsed
tree is identical across calls. This review complements #003 by
adding the performance axis.
+659
View File
@@ -0,0 +1,659 @@
---
status: closed
last_updated: 2026-09-02
resolved_findings: 2026-08-20 (all 11 — see "Resolution" at the end)
reviewed_artifacts:
- docs/plans/030-compiled-forms.md
- docs/architecture/decisions/011-compiled-read-plan-for-packed-mode.md
- docs/architecture/decisions/012-plan-fingerprinting-and-m1-closure.md
- docs/reviews/004-performance-review.md
- src/lib.rs
- src/bast.rs
- src/engine.rs
- src/sequential_reader.rs
- src/materialize.rs
- src/offset_map.rs
- src/layout_builder.rs
- src/schema.rs
- poc/readplan/{src/lib.rs, FINDINGS.md} (branch readplan-poc)
tool: manual source read + plan-vs-codebase cross-check + POC branch inspection
reviewer: 0.3.0 implementation plan review (triggered before phase 1)
---
# Review #005 — 0.3.0 Plan Review: Compiled Forms
## Purpose
The 0.3.0 implementation plan
([`docs/plans/030-compiled-forms.md`](../plans/030-compiled-forms.md)) is
the entry point an implementing agent reads first. It rolls up
[ADR-011](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md)
(the `ReadPlan` packed read-side compiled form),
[ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md)
(fingerprinting + owned `BastDoc` + `OffsetMap` `LeafMeta`), and the
fingerprinting work into one breaking bump. The plan is deliberately
structured as seven phases so each can be picked up by a fresh session
without prior context.
This review's purpose is to find planning-spec mistakes — factual
errors, contradictions, undocumented behavioral changes, hedges into an
unplanned future — *before* a phase-by-phase implementation starts,
because fresh-session implementations are reliable precisely when the
spec is accurate. A spec that contradicts the code or an ADR forces the
agent to either reverse-engineer the actual intent or guess, and the
failure rate goes up.
The review explicitly scans for the "deferral black hole" pattern: a
plan or ADR puts work off into a "future version/phase/downstream" with
no concrete reactivation condition, the next agent inherits the gap,
and the gap festers until something forces an untangle. This is a
known LLM-planning quirk distinct from classic planning mistakes, and
a default scan for it is part of this review's methodology.
## Methodology
- Full read of the plan and its two companion ADRs (011, 012), the
performance review (#004) the plan closes, and the POC findings on
branch `readplan-poc`.
- Cross-check every line-number reference and `src/` claim in the plan
against the actual codebase at `main` (commit `2310f6c`, v0.2.0).
Verified: `engine.rs:112-115,284,334,467`; `sequential_reader.rs:567`;
`layout_builder.rs:190`; `bast.rs:51-55`; `lib.rs` re-export list;
`Cargo.toml` version; presence of `poc/` (absent on main, present on
`readplan-poc` as expected); presence of `alktty`/`alkcall` downstream
path dev-deps.
- Cross-check the POC's `ReadPlan`/`CompositePlan` shape against both
ADR-011's shape section and the plan's phase-1 description.
- Cross-check the plan's phase 2 rewrite claims (`dummy_field_for`/
`ty_source` "are removed") against `materialize.rs`'s actual call
sites across both packed and aligned paths.
- Verify the derives the plan relies on (`Hash` on `LeafMeta`,
`ReadPlan`, `OffsetMap`) are reachable from the derives on their
constituent types in `src/schema.rs`.
- Scan for the deferral pattern by flagging every "future/deferred/
later/downstream/if needed" occurrence and asking: (a) is there a
concrete reactivation trigger? (b) is the decision owned or silent?
(c) does inaction have a cost that the deferral framing hides?
## Verification Baseline
The plan and both ADRs were read at the tree state at commit `2310f6c`
("Propose ADR-012 + 0.3.0 implementation plan"), which is `main` HEAD.
The codebase is v0.2.0 (`Cargo.toml`); the POC lives on branch
`readplan-poc` and is not merged, as the plan states. All line-number
references in the plan were verified correct against this tree.
## Summary Statistics
| Severity | Count |
|----------|------:|
| High | 2 (H1, H2) |
| Medium | 3 (M1, M2, M3) |
| Low | 3 (L1, L2, L3) |
| Nit | 3 (N1, N2, N3) |
The two High findings are correctness/contradiction issues that would
block or mislead an implementing agent. The Mediums are either
undocumented behavioral drops, missing implementation prerequisites, or
a deferral worth re-evaluating. Lows and Nits are wording/typo-level.
---
## Findings
### H1. Phase 1's field-name-discriminator union shape exists in neither ADR-011 nor the POC
**File**: `docs/plans/030-compiled-forms.md:144-152`
**Problem**: The plan describes the field-name-discriminator union read
shape as:
> `CompositePlan::Union` carries the union's declared `fields` as a
> sub-`ReadPlan` (the discriminator field + any shared fields), and the
> variant plans are laid out *after* the shared fields.
But ADR-011 §"The `ReadPlan` shape"
(`011-compiled-read-plan-for-packed-mode.md:144-158`) defines:
```rust
pub enum CompositePlan {
Struct(ReadPlan),
Union {
disc: DiscriminatorPlan,
variants: Vec<(String, VariantPlan)>,
},
...
}
```
There is no field for shared/declared fields on the `Union` variant.
The POC (`readplan-poc:poc/readplan/src/lib.rs`) matches the ADR's
shape — `CompositePlan::Union { disc, variants }` only — and its
`compile_union` does not carry shared fields. The POC's `FINDINGS.md`
Finding 1 (the same one the plan cites at lines 143-152) explicitly
says:
> `plan_read_union`'s `Field` arm is a stub that returns an error.
> ... The plan needs a sub-struct for the union's declared fields,
> separate from the variant plans.
So the plan describes a shape that exists in **neither** the accepted
ADR **nor** the reference POC, and presents it as "the production
version must implement it" within the existing `CompositePlan::Union`
shape. An implementing agent reading ADR-011 + plan + POC gets three
different `CompositePlan::Union` shapes and no guidance on where the
shared-fields sub-`ReadPlan` goes (a new `shared: Option<Box<ReadPlan>>`
field? a wrapper enum? two-variant split?).
This is a shape extension to an accepted ADR's public type. The plan
either needs to flag it as an ADR-011 refinement (with the ADR updated
first) or specify the concrete shape the agent should build.
**Lift**: unblocks phase 1. Without resolution, the agent will either
guess the shape and likely diverge from intent, or stop and ask.
---
### H2. `schema()` returning `&Value` from a `&Value` "stored on the plan" is a self-referential struct
**File**: `docs/plans/030-compiled-forms.md:188-191`
**Problem**: Phase 2 says:
> `schema()` returns a `&Value` retained on the plan (the plan stores
> the `&Value` it was compiled from — see ADR-011 §Engine integration;
> the `&Value` outlives the plan because the engine owns both).
The "the plan stores the `&Value` it was compiled from" is the
self-referential struct pattern ADR-011 §"Root cause"
(`011-...md:50-58`) explicitly identifies as impossible in safe Rust
and rejects. The engine owns `bast_doc: Value` and `Arc<ReadPlan>`. If
`ReadPlan` stores `&Value` borrowing from the engine's `bast_doc`, the
engine is self-referential — exactly the construction ADR-007 worked
around with "re-parse on demand" and ADR-011's `Arc<ReadPlan>` was
meant to retire. ADR-011 line 204 specifies the plan is "immutable
**owned** data"; it does not say the plan stores a `&Value`.
`schema()`'s current contract (`src/sequential_reader.rs:247`) is to
return the raw BAST `Value` the reader was built from. To preserve
that contract on `Arc<ReadPlan>` without a self-referential borrow,
`ReadPlan` must store an `Arc<Value>` (engine builds `Arc<Value>` at
compile time, hands a clone to the plan) or an owned `Value`. Then
`schema()` returns `&self.plan.value`. The plan should specify which.
**Lift**: prevents an agent from getting stuck in phase 2 trying to
make `&Value` in `Arc<ReadPlan>` work, which the borrow checker will
reject.
---
### M1. Nested unions: POC rejects a schema 0.2.0 accepts — undocumented behavioral drop
**Files**: `poc/readplan/FINDINGS.md` ("What this POC does not cover"),
`docs/plans/030-compiled-forms.md` (silent), `src/sequential_reader.rs:789-815`
**Problem**: The POC's `compile_union` rejects a union variant that is
itself a union with `AlkTypeError::Schema`. The existing reader
supports this: `resolve_and_walk_variant` at
`src/sequential_reader.rs:800` has a live `BastDefKind::Union` arm
that recurses via `read_union_value`. So 0.2.0 accepts and reads
nested-union schemas; phase 1's `ReadPlan::compile` (per the POC the
plan cites as the reference scaffold) would reject the same schema.
The plan's phase 1 calls out two POC findings explicitly (field-disc
union shape → H1 above, struct-array stride → deferred decision 4)
and says "the production version must implement/decide these." It does
**not** call out the nested-union rejection. An agent following the
plan would inherit the POC's reject-nested-unions behavior by default,
silently dropping a 0.2.0 capability — a behavioral regression that
rides the 0.3.0 bump without being listed in the Semver Contract
table.
This is also the cleanest example of the deferral-black-hole pattern
in the plan: the POC says "if a real schema needs it, the
implementation step adds a `VariantKind::Union` read path. Not
blocking — no current schema exercises it." The "if needed" framing
has no trigger, no OQ, no tracking — it's a black hole. The next agent
inherits the gap.
**Lift**: either (a) add `VariantKind::Union` read path in phase 1
(small — mirrors the existing `resolve_and_walk_variant` Union arm,
~20 lines), or (b) list it in the Semver Contract table as a
behavioral drop with a one-line OQ tracking the deferral. Given the
plan says there are zero real consumers, (b) is defensible, but it
must be *stated*, not silent. (a) is cheap and avoids the regression.
---
### M2. `Endian` and `VariableEncoding` don't derive `Hash` — phases 5/6 will not compile
**Files**: `src/schema.rs:205,212`, `docs/plans/030-compiled-forms.md:62,357,411-415`
**Problem**: `src/schema.rs:205` (`Endian`) and `:212`
(`VariableEncoding`) both derive only `Debug, Clone, Copy, PartialEq,
Eq` — no `Hash`. The plan requires:
- Phase 5 (line 62, 357): `LeafMeta { kind, encoding, endian }` as
`Copy + PartialEq + Eq + Hash`.
- Phase 6 (lines 411-415): `#[derive(Hash, Eq)]` on `ReadPlan`/
`OffsetMap`, and `FieldPlan` carries `endian: Endian` + `encoding:
VariableEncoding`.
Both derives will fail to compile: `#[derive(Hash)]` on a struct
requires all fields to be `Hash`. The plan never mentions adding
`Hash` to these two enums. The fix is trivial (both are fieldless
enums, already `Eq + PartialEq`, so adding `Hash` is semver-safe —
additive, no behavioral change), but it's a prerequisite the plan
omits. An agent working phase 5 will hit a compile error and have to
diagnose why.
**Lift**: trivial. Add a sub-step to phase 5 (or 6): "Add `Hash` to
`Endian` and `VariableEncoding` derives in `src/schema.rs`." This is
additive and safe to do earlier if convenient.
---
### M3. `ValidationPlan` deferral worth re-evaluating — the read+validate common case
**Files**: `docs/architecture/decisions/012-...md:64-73` ("Deferring
`ValidationPlan`"), `docs/plans/030-compiled-forms.md:526-528`,
`docs/reviews/004-performance-review.md` (the read-path perf review)
**Problem**: ADR-012 defers a `ValidationPlan` as "different shape
(value-domain, not byte-position), not a hot loop, separate ADR if a
bench motivates it." The plan inherits this deferral ("Not a
`ValidationPlan`" at lines 526-528). The deferral framing is "if a
bench motivates it" — a concrete trigger exists, so this is not a
black-hole hedge in the M1 sense.
Flagged for re-evaluation, not because the shape argument is wrong
(it's correct — value-domain checks are structurally different from
byte-position walks), but because the *hot-loop* dismissal may under-
account a common case: **read + validate together on untrusted input.**
Review #004 found the packed read path was 400x slow per chunk due to
per-field `BastDoc` re-parse. ADR-011 closes that. But
`validate_bytes`'s packed path (ADR-010) is `materialize_packed` →
`bast_validation::validate_value` over the materialized `Value`.
After ADR-011, `materialize_packed` walks the `ReadPlan` (fast).
`bast_validation::validate_value` still walks `BastDoc` to check
value-domain constraints — once per `validate_bytes` call, over the
full tree, on every buffer.
For a stream of N untrusted buffers (the `alkcall` hub/spoke topology
accepts schemas from arbitrary internet peers — AGENTS.md §3 — and
the common case is "read incoming frame, validate it before acting"),
`validate_bytes` is called N times. Each call does one `BastDoc`
walk for validation. After ADR-011, the *read* half of `validate_bytes`
is plan-fast; the *validation* half is still a `BastDoc` walk per call.
If validation is the common companion to read on untrusted input,
then skipping validation is risky (accepting untrusted bytes
unchecked) and running it re-walks `BastDoc` per buffer — the same
class of cost review #004 measured for the read path, just on a
different code path.
The argument is not "ValidationPlan has the same shape as ReadPlan"
(it doesn't). The argument is: ADR-012's "not a hot loop" dismissal
may be incomplete, because read+validate on untrusted streams makes
validation hot in the same sense read was hot. The deferral's
trigger ("if a bench motivates it") should be sharpened: either (a)
add a `validate_bytes`-on-untrusted-stream bench to alktty alongside
`wire_vs_bast` and let the bench decide, or (b) reason from the
existing review #004 numbers that the validation walk is
non-trivial and should be planned, not deferred.
This is not a request to implement `ValidationPlan` in 0.3.0. It's a
request to *own the decision*: either the trigger fires (and a
follow-on ADR/phase is scoped, possibly 0.4.0) or it doesn't (and the
deferral stands with a sharper justification than "not a hot loop").
As written, the deferral leaves the cost in the superposition where
it can neither be confirmed nor dismissed.
**Lift**: removes a latent perf cliff for the read+validate-on-
untrusted-input case that 0.3.0 is supposed to make viable.
---
### L1. `dummy_field_for`/`ty_source` are used in aligned `materialize`, not just packed
**Files**: `docs/plans/030-compiled-forms.md:209-210`,
`src/materialize.rs:249,316,351,391,631,650-663`
**Problem**: Phase 2 says:
> The `dummy_field_for`/`ty_source` helpers in `materialize.rs` are
> removed (the plan carries everything).
This is factually wrong. `dummy_field_for` is called at
`src/materialize.rs:631` inside `materialize_leaf_at`, which is called
by the **aligned** path: `materialize_struct_aligned` (line 475),
`materialize_array_aligned` (line 544), `materialize_variable_aligned`
(line 613). Aligned `materialize` keeps walking `BastDoc` through
0.3.0 (plan lines 379-385 confirm), so `dummy_field_for`/`ty_source`
must stay. Only the packed-side call sites (lines 249, 316, 351, 391)
go away when packed-materialize moves to the plan.
**Lift**: doc accuracy. An agent following the plan literally would
remove the helpers and break aligned `materialize`.
---
### L2. `materialize_packed` rewrite scope underspecified — packed-vs-aligned split of `materialize_typeref_packed`
**Files**: `docs/plans/030-compiled-forms.md:207-210`,
`src/materialize.rs:122-200, 498-506, 619-637`
**Problem**: `materialize_typeref_packed` is shared by both packed and
aligned paths — aligned's `materialize_leaf_at` (line 619-637) calls
`materialize_typeref_packed` to read leaves, and aligned's record path
(line 498-506) calls it directly. Phase 2 says
`materialize_packed(&ReadPlan, &[u8])` walks the plan instead of
`BastDoc` but does not state what happens to
`materialize_typeref_packed`.
The honest resolution: packed-materialize gets a new plan-walking
function; aligned keeps `materialize_typeref_packed` via
`materialize_leaf_at`; the function stays (renamed or not) for aligned.
This is two mode-specific paths — the existing design — not a
"parallel walker" in the maintenance-tax sense ADR-011 §"Negative"
(cautioning against) discusses. ADR-011's "one walker" claim (lines
234-237) is specifically about packed read-side (`SequentialReader` +
`materialize_packed` sharing the plan), not packed-vs-aligned, so
there's no ADR contradiction — just an underspecification in the plan.
**Lift**: prevents the agent from having to discover the split
mid-rewrite. Add one line to phase 2: "packed-materialize gets a new
plan-walking function; `materialize_typeref_packed` stays for
aligned's `materialize_leaf_at` and the aligned record path."
---
### L3. `materialize_aligned`'s `BastDoc` structure walk is silent in the plan
**Files**: `docs/plans/030-compiled-forms.md` (silent on this),
`src/materialize.rs:451-521`, `docs/architecture/decisions/011-...md:264`
**Problem**: `materialize_struct_aligned` walks `BastDoc` to traverse
struct/array/record structure, using `OffsetMap` only for leaf byte
positions. ADR-011 §"Out of scope" says "aligned mode is unchanged;
`materialize_aligned` already takes `&OffsetMap`" — which is half
true: it takes `&OffsetMap` for positions but also `&BastDoc` for
structure. The plan inherits the half-truth silently: there's no
statement anywhere that aligned materialize keeps walking `BastDoc`
for structure.
After phase 3 (owned `BastDoc`) + phase 5 (`LeafMeta`), the walk is
over owned data, no re-parse, not O(N²), and aligned `validate_bytes`
is one walk per call (not per-field). There's no perf driver
analogous to review #004's packed per-chunk gap. But the absence of
a driver is not the same as a decision: leaving it silent is a
deferral-by-omission. An implementing agent or future reader can't
tell whether the silence is "this is the permanent design" or "we'll
fix this later."
The decision should be owned. Either (a) add a "Scope Boundary" note
that aligned materialize keeps walking owned `BastDoc` for structure
as the permanent design (with an OQ if a future bench motivates an
`AlignedPlan`), or (b) if a bench motivation is plausible, scope an
OQ to track it. (a) is recommended — no perf driver, and after phase
3 the walk is over owned data, so it's not the re-parse pattern.
**Lift**: removes a silent gap that future agents would otherwise
have to reverse-engineer.
---
### N1. Typo: "back-comat" → "back-compat"
**File**: `docs/plans/030-compiled-forms.md:104-105`
**Problem**: "back-comat" in deferred decision 4.
**Lift**: trivial.
---
### N2. Phase 1 verification omits the `Send + Sync` assertion test ADR-011 requires
**Files**: `docs/plans/030-compiled-forms.md:158-163`,
`docs/architecture/decisions/011-...md:204-206`
**Problem**: ADR-011 §"Engine integration" says "the implementation
should add a `static` bound assertion test to lock it in" for
`ReadPlan: Send + Sync`. Phase 1's verification block lists `cargo
test`, `clippy`, `doc`, `wasm` but no mention of adding the assertion
test. An agent following the plan literally won't add it; the
property is currently true by construction but not asserted, so a
future change could break it silently.
**Lift**: add "add a `fn read_plan_is_send_sync()` assertion test" to
phase 1's verification, mirroring the POC's
`readplan_is_send_sync` test.
---
### N3. `SequentialReader::new` return-type change (`Result` drop) undocumented
**Files**: `docs/plans/030-compiled-forms.md:64`,
`src/sequential_reader.rs:129`, `src/engine.rs:205`
**Problem**: Currently `new(&Value, &str) -> Result<Self,
AlkTypeError>` — fallible (BastDoc parse). After phase 2,
`new(Arc<ReadPlan>)` is infallible (just stores the Arc) → returns
`Self`, not `Result<Self>`. The Semver Contract table (line 64) lists
only the argument-type change, not the `Result` drop.
`engine.rs:205`'s `.ok()` call correspondingly goes away. Minor, but
it's a signature change beyond what's listed.
**Lift**: add a row to the Semver Contract table noting the `Result`
drop.
---
## Deferral-pattern scan (LLM-planning quirk)
As part of the methodology, every "future/deferred/later/downstream/if
needed" occurrence in the plan and its ADRs was flagged and tested
for: (a) concrete reactivation trigger, (b) decision owned or silent,
(c) hidden cost of inaction.
| Item | Trigger? | Owned? | Cost of inaction | Finding |
|---|---|---|---|---|
| `ValidationPlan` (ADR-012) | "if a bench motivates it" | Yes (ADR + plan "What this is not") | Possible perf cliff on read+validate untrusted streams | M3 above — sharpen the trigger |
| Nested-union `ReadPlan` support | "if a real schema needs it" (POC) | No (POC only, plan silent) | Silent 0.2.0 capability drop | M1 above — state it |
| `materialize_aligned` structure walk | None — silent | No (silent) | Future agent ambiguity | L3 above — own the decision |
| `Arc<str>` vs `String` (decision 1) | "if phase 4 shows it's measurable" | Yes (deferred decision 1) | None | OK — has trigger, decided in phase 3 |
| `OffsetMap::get` shape (decision 2) | "decided in phase 5" | Yes (deferred decision 2) | None | OK |
| Fingerprint hasher (decision 3) | "decided in phase 6" | Yes (deferred decision 3) | None | OK |
| Struct-array stride (decision 4) | "decided in phase 2" | Yes (deferred decision 4) | None | OK |
| `BastDoc` `Arc<Value>` vs `Value` | None — silent | No (plan doesn't address) | Agent gets stuck (H2) | H2 above |
| Field-disc union shape (POC Finding 1) | "production version must implement" | Yes (plan phase 1) | None, but shape is undefined | H1 above — shape not in ADR |
The four explicit "deferred decisions" in the plan (items 4-7) all
have concrete triggers and decision points — these are the *good*
pattern. The black-hole pattern appears where deferrals lack triggers
(items 1-3, 8-9): three of those became findings (M1, L3, H2), and M3
is a deferral worth sharpening even though it has a trigger.
The general signal: a deferral is healthy when it has a concrete
reactivation condition and is tracked (OQ, ADR, or in-plan deferred
decision). A deferral is a black hole when it has no trigger, no
tracking, and the next agent inherits the gap by default.
---
## What's Good
- **Line-number accuracy is perfect.** Every `src/` reference in the
plan (`engine.rs:112-115,284,334,467`;
`sequential_reader.rs:567`; `layout_builder.rs:190`;
`bast.rs:51-55`; `lib.rs` re-exports) checks out against the v0.2.0
tree. This is unusual for a plan of this length and worth noting.
- **The Semver Contract table is a strong scope-creep guardrail.**
Walking every public `lib.rs` re-export against the table, the
classifications (Breaking / Unchanged / New) are correct for every
item, with the exceptions noted in N3 (the `Result` drop on `new`)
and M1 (the nested-union behavioral drop not listed).
- **The four explicit "deferred decisions" are the right pattern.**
Each has a trigger and a decision point in a named phase. This is
what deferrals should look like.
- **Phases are coherent session boundaries.** Phases 1 (pure
addition), 6 (pure addition), 7 (docs/bump) are small and clean.
Phases 3 (broad but mechanical), 4 (single file), 5 (single file +
engine) are well-scoped. Phase 2 is the largest and the plan
sanctions sub-session splits at the step level (lines 40-42), which
is the right escape valve.
- **Cross-phase invariants are stated and checkable.** "Tree builds
and tests pass at every phase boundary" is the right invariant; the
POC-on-`readplan-poc`-only convention is clearly separated from
production code; AGENTS.md §5-§11 constraints (no `unsafe`, no
`async`, no new deps, `preserve_order` load-bearing) are
reaffirmed.
- **The plan honestly scopes what it is not.** "Not a `ValidationPlan`",
"Not cross-version fingerprint stability", "Not a perf bench" —
these boundaries are stated rather than left implicit, which helps
an implementing agent resist scope creep. (M3 above is about
sharpening one of these, not removing the boundary.)
- **The POC reference is disciplined.** The plan is explicit that the
POC is "not production code," lives only on the branch, and is the
reference scaffold for phases 1-2 only. This matches how
`bast-validator-poc` was handled and avoids the POC leaking into
`main`.
---
## Recommended Order
1. **H1 (field-disc union shape)** — update ADR-011's
`CompositePlan::Union` to include the shared-fields sub-`ReadPlan`
(or document the wrapper shape), then update the plan's phase 1 to
reference the corrected ADR shape. Do this before phase 1 starts;
otherwise the implementing agent has to guess.
2. **H2 (`schema()` `&Value` on `Arc<ReadPlan>`)** — edit the plan's
phase 2 to specify `ReadPlan` stores `Arc<Value>` (or owned
`Value`), and `schema()` borrows from that. One-line edit to the
plan; avoid a phase-2 stuck point.
3. **M1 (nested unions)** — decide (a) implement `VariantKind::Union`
in phase 1, or (b) list as behavioral drop + OQ. Edit the plan and
(if b) the Semver Contract table accordingly. Decide before phase
1.
4. **M2 (`Hash` on `Endian`/`VariableEncoding`)** — add a sub-step
to phase 5 or 6. Trivial.
5. **M3 (`ValidationPlan` re-evaluation)** — either add a
`validate_bytes`-on-untrusted-stream bench to alktty (alongside
`wire_vs_bast`) and let the bench decide, or sharpen ADR-012's
"not a hot loop" justification. Does not block 0.3.0; can be
resolved in parallel with phase 1-7 work. **Flagged for
re-evaluation, not for implementation in 0.3.0.**
6. **L1, L2, L3** — edit the plan's phase 2 to fix the
`dummy_field_for` wording (L1), state the packed-vs-aligned
materialize split (L2), and add a Scope Boundary note for
aligned-materialize's `BastDoc` structure walk (L3). All three are
phase-2 doc edits.
7. **N1, N2, N3** — typo, `Send + Sync` assertion test, `Result`-drop
Semver row. Minor plan edits.
Items 1-3 must be resolved before phase 1 starts (they affect the
`ReadPlan` shape or 0.2.0 behavioral surface). Items 4-7 can be
resolved any time before their phase begins. Item 5 (M3) is
non-blocking and can run in parallel.
---
## Notes
- All line numbers refer to the tree at commit `2310f6c` (the plan's
commit) for `src/` files, and to the plan/ADR markdown as committed
at the same tree.
- The POC on `readplan-poc` was inspected via
`git show readplan-poc:poc/readplan/{src/lib.rs,FINDINGS.md}`; it is
not merged to `main` and the plan correctly states this.
- `alktty` and `alkcall` downstream repos exist as path dev-deps
(`/workspace/@alkdev/alktty`, `/workspace/@alkdev/alkcall`); the
plan's claim that they're in-house and updated with the bump is
verifiable, though this review did not inspect their call sites
in detail.
- This review does not re-litigate ADR-011 or ADR-012's accepted
decisions. H1 and H2 are about the plan *contradicting* the ADRs or
being unsound, not about the ADR decisions themselves; M3 is about
sharpening a deferral, not about re-deciding it.
- The deferral-pattern scan is a methodology experiment: a
pre-declared scan for LLM-specific planning quirks (deferral black
holes) alongside classic planning mistakes. It surfaced M1 and L3
that a conventional severity-only review would have missed or
under-weighted. Worth retaining as a default scan for future plan
reviews.
---
## Resolution (2026-08-20)
All 11 findings resolved in one docs-only edit pass to ADR-011,
ADR-012, and the 0.3.0 plan. No source changed; the crate still
builds/tests at v0.2.0. The M3 deferral reversal is the one
substantive decision change (per user direction: ship ValidationPlan
in 0.3.0, no more hedging); the rest are spec corrections or
pre-implementation refinements to types that do not yet exist on
`main`.
- **H1 (union shape):** ADR-011 §"The `ReadPlan` shape" refined —
`CompositePlan::Union` now carries `shared: Option<Box<ReadPlan>>`
(field-disc shared fields) and `variants: Vec<(String,
CompositePlan)>` (dropping `VariantPlan`/`VariantKind`). Plan
phase 1 rewritten to implement the refined shape. The shape
refinement is pre-implementation (the types don't exist on `main`).
- **H2 (`schema()` `&Value`):** plan phase 2 rewritten — `ReadPlan`
stores `schema: Arc<Value>` (not `&Value`); `schema()` returns
`&self.schema`. Verified `serde_json::Value: Hash + Eq` holds with
`preserve_order` (`Map::hash` sorts keys deterministically), so
phase 6's `#[derive(Hash)]` on `ReadPlan` is not blocked.
- **M1 (nested unions):** resolved as the review's option (a) —
nested-union support falls out of the H1 shape refinement (a
variant can be `CompositePlan::Union`), so no behavioral drop vs
0.2.0 and no Semver Contract entry for a capability regression.
Plan phase 1 adds a nested-union-variant test.
- **M2 (`Hash` on `Endian`/`VariableEncoding`):** plan phase 5
rewritten with an explicit first sub-step to add `Hash` to both
derives in `src/schema.rs` (additive, semver-safe). The inaccurate
"all fields are `Copy + Hash`" parenthetical on `LeafMeta` is
corrected.
- **M3 (`ValidationPlan`):** deferral **reversed** per user
direction. ADR-012 §"Deferring `ValidationPlan`" rewritten as
"ValidationPlan — in scope for 0.3.0"; new ADR-012 §3 commits the
decision (compiled form, no per-buffer `BastDoc` walk, `Hash + Eq`
+ `fingerprint()`) and lists the shape questions deferred to a
follow-on design session + the plan's new phase 7. Plan gains a
new phase 7 (ValidationPlan); old phase 7 (bump) renumbered to
phase 8. ADR-011's "Out of scope" `bast_validation` bullet and
"Scope Boundaries" `Not a validation plan` bullet updated to point
at ADR-012 §3. Plan's "What this plan is *not*" first bullet
removed. The deferral-black-hole pattern this review's methodology
flagged is closed: the work is committed in the plan with a
concrete reactivation trigger (the shape session before phase 7),
not hedged into an unplanned future.
- **L1 (`dummy_field_for`/`ty_source`):** plan phase 2 rewritten —
only the packed-side call sites go away; the helpers stay for the
aligned `materialize_leaf_at` path.
- **L2 (`materialize_typeref_packed` split):** plan phase 2
rewritten — packed-materialize gets a new plan-walking function;
`materialize_typeref_packed` stays for aligned's
`materialize_leaf_at` and the aligned record path.
- **L3 (aligned-materialize `BastDoc` structure walk):** plan phase 5
gains a Scope Boundary note — the walk is the permanent 0.3.0
design; an `AlignedPlan` is out of scope, tracked as an open
question if a future bench motivates it.
- **N1 (typo):** "back-comat" → "back-compat" in deferred decision 4.
- **N2 (`Send + Sync` assertion test):** plan phase 1 verification
rewritten to add the `read_plan_is_send_sync` static-bound
assertion test ADR-011 §"Engine integration" requires.
- **N3 (`Result` drop on `SequentialReader::new`):** Semver Contract
table row updated to note the constructor return-type change
(`Result<Self, AlkTypeError>` → `Self`) alongside the argument-type
change.
The deferral-pattern scan's general signal (healthy deferrals have a
concrete reactivation condition + tracking; black holes have neither)
is reaffirmed by the M3 reversal: the original "if a bench motivates
it" trigger was a black hole because no bench was ever going to be
run against a path that didn't exist yet, and the cost of inaction
(a second breaking change to `validate_bytes`/`bast_validation` after
0.3.0) was hidden by the "not a hot loop" framing.
File diff suppressed because it is too large. Load diff
+454
View File
@@ -0,0 +1,454 @@
---
status: resolved (F1, F2, C1, C2, C3, L1, L2 resolved 2026-09-02; N1/N2a/N3a/N4a are classified-no-action / deferred-by-design)
last_updated: 2026-09-02
reviewed_artifacts:
- src/materialize.rs
- src/sequential_reader.rs
- src/read_plan.rs
- src/offset_map.rs
- src/layout_builder.rs
- src/engine.rs
- src/data_access.rs
- src/bast.rs
- src/builder.rs
- src/tunion.rs
- src/validation_plan.rs
- src/walk_guard.rs
- src/bast_meta.rs
- tests/poc_roundtrip.rs
- tests/tunion_dispatch.rs
- tests/error_paths.rs
- tests/engine_integration.rs
- docs/reviews/006-implementation-review-030.md (post-fix coverage re-check)
tool: cargo-llvm-cov 0.8.4 (--release, per-line text) + manual classification of every uncovered production line + disposable probe tests (run in-session, then deleted)
reviewer: post-review-#006 coverage audit (session request — check test coverage for weak spots, meaningful tests, non-happy-path posture)
---
# Review #007 — Post-#006 Coverage Audit
## Purpose
Review #006 closed every finding and its M4 coverage map, but the
session-level posture (M4's item: "fold a coverage check into each fix
session") had never been run as a *whole-tree* pass after all those
fixes landed. This audit re-measures coverage after the eleven #006
commits, reads every uncovered production line, and classifies it —
the same "untested-but-fine / load-bearing / unreachable" discipline
M4's map used. Two probes were run in disposable tests (deleted after
the session, per #006's no-reproducer rule; neither was a crash
hazard — both reproduce cleanly inside the default harness).
## Methodology
- `cargo llvm-cov --release` (0.8.4, same tool as #006): summary +
per-line text. TOTAL **90.67% lines / 86.32% functions** — stable
with #006's post-M4 numbers (90.60%), the expected drift after the
N3 fix sessions added parse gates + tests.
- Per-file (worst first): `materialize.rs` 85.72, `data_access.rs`
80.32, `sequential_reader.rs` 86.19, `bast.rs` 87.41,
`builder.rs` 91.28, `layout_builder.rs` 91.42, `tunion.rs` 92.02,
`offset_map.rs` 92.93, `read_plan.rs` 90.69, `engine.rs` 96.44,
`validation_plan.rs` 93.97, `walk_guard.rs` 98.04,
`bast_meta.rs` 98.92, `bast_validation.rs`/`error.rs`/`schema.rs`/
`macros.rs`/`validation.rs` 100.
- Every uncovered line *outside* `#[cfg(test)]` modules (806 raw
lines) was read and classified. Lines inside test modules (the
`panic!("expected X, got {other:?}")` helpers) were excluded — they
distort per-file numbers (e.g. `bast.rs`'s 87.41% is really ~96%
production once its 60 helper lines are excluded).
- Two suspicions were probe-verified with disposable tests:
the F1 cross-consumer divergence and the F2 unbounded-`maxLength`
compile. Probe transcripts quoted verbatim in the findings.
- Happy-path posture audit: cross-checked which *error arms* adjacent
to covered code are 0-execution, and which public surfaces have only
success-path tests.
## Baseline
Audited at `main` HEAD `bb28ba3` ("Resolve N3"), 0.3.0, working tree
clean. Full suite green (548 tests static + 2 ignored doctests, per
#006's bookkeeping).
## Summary Statistics
| Severity | Count | Status |
|----------|------:|--------|
| High | 2 (F1, F2) | both resolved 2026-09-02 |
| Medium | 3 (C1, C2, C3) | all resolved 2026-09-02 |
| Low | 2 (L1, L2) | all resolved 2026-09-02 |
| Info | 4 (N1, N2a, N3a, N4a) | classified: N1 artifact, N2a/N4a no-action, N3a deferred to pre-release review |
**Resolution log:**
- **F1 + F2 (2026-09-02):** resolved in one commit — see the
resolution blocks on each finding. 477 lib tests green (511
static + 2 ignored doctests across all targets), clippy
`-D warnings` clean, wasm build green.
- **C1 + C2 + C3 (2026-09-02):** resolved in one commit — see the
resolution blocks. 481 lib tests green, clippy `-D warnings` clean,
wasm build green; `compile_variant`'s cycle arm confirmed executed
in the post-fix coverage run.
- **L1 + L2 (2026-09-02):** resolved in one commit — see the
resolution blocks. 488 lib tests green, clippy `-D warnings` clean,
wasm build green.
---
## Findings
### F1. Zero-progress guard missing in the plan materializer — `validate_bytes` accepts what `SequentialReader` rejects (cross-consumer divergence)
**Files**: `src/materialize.rs:227-253` (`materialize_plan_array` — no
guard), contrast `src/sequential_reader.rs:772-795`
(`plan_walk_variable_array_size` — has the guard) and
`src/materialize.rs:636-663` (`materialize_array_packed` — has the
guard)
**Problem**: The H1 fix session added the zero-progress runtime guard
("array element consumed 0 bytes") to two of the three array walkers:
the compiled reader's variable-array size walk and the legacy BAST
walker's packed array arm. The *plan-based* packed materializer —
`materialize_plan_array`, the walker `validate_bytes` actually uses in
packed mode (engine.rs:310-312) — got no guard.
A stride-0 array whose elements consume 0 bytes (empty-struct elements
are legal: the meta-schema's `StructDef` has no `minItems` on
`fields`) compiles with `element_stride: 0` and loops `count` times
materializing empty objects without reading a single buffer byte:
```
PROBE validate_bytes([]): OK — zero-progress guard MISSING in plan materializer
PROBE reader.read_next([]): Err(access error at items[0]: array element 0 consumed 0 bytes;
a zero-size element makes the declared count unbounded on the wire)
```
Schema: `{ "items": { "kind": "array", "element": { "kind":
"struct", "fields": [] }, "count": 8 } }`, packed mode, empty buffer.
Same schema, same buffer, opposite verdicts — the exact
cross-consumer-disagreement shape review #006 existed for (H3, M6).
Severity High by #006's own keying (AGENTS.md §3): `validate_bytes`
is the flagship untrusted-input path, and it silently accepts a
buffer the same engine's reader rejects. The H1 resolution text
("plan_walk_variable_array_size (reader) and materialize_array_packed
(materializer) now error") lists only two of the three walkers — the
plan materializer was missed because it is *not* the legacy walker
that finding named.
**Not a #006 regression**: the H1 fix text itself specified only the
reader and legacy-walker sites; the plan materializer predates the
guard and was outside that fix's blast radius. But the divergence is
new information — the guard's *invariant* ("a zero-progress element
makes the declared count unbounded") belongs to the array-walk
concept, not to two specific functions.
**Fix**: hoist the same guard into `materialize_plan_array`'s loop
(compare `*offset` before/after `materialize_plan_composite`; error
with the same wording the other two walkers use so downstream
matching sees one shape). Add a locking test driving the same schema
through BOTH paths asserting the verdicts agree (both reject an empty
buffer; both accept a buffer where the elements make progress —
empty-struct elements never do, so the acceptance half needs a
non-empty variant struct alongside).
**Resolution (2026-09-02):** the guard, hoisted verbatim from the two
existing sites (`*offset == before` after the element walk, same
"array element {i} consumed 0 bytes…" wording so downstream matching
sees one shape). Tests (3, in `materialize.rs`):
`f1_zero_progress_array_rejected_by_all_three_walkers` (plan
materializer + the record-value fallback path, both asserting the
`Access` error with the guard's wording),
`f1_validate_bytes_and_reader_agree_on_zero_progress_array` (the
cross-consumer agreement the probe showed was missing —
`validate_bytes` and `SequentialReader::read_next` both reject the
same schema+buffer with the same error class),
`f1_nonempty_variant_struct_array_still_materializes` (the
false-positive check: elements that consume bytes still walk).
Verified: 477 lib tests green, clippy `-D warnings` clean, wasm build
green.
### F2. `maxLength` is unbounded — the N2 analog
**Files**: `src/bast_meta.rs:81` (`"maxLength": { "type": "integer",
"minimum": 0 }` — no maximum), `src/bast.rs:1038-1043`
(`parse_max_length` — no cap, and silently drops non-`usize` values),
contrast `src/schema.rs` `MAX_ALIGN`/`parse_align` (the N2 pattern)
**Problem**: N2 bounded `align` at 4096 with a clean parse error plus
a meta-schema `"maximum"`. `maxLength` has the identical shape and
was not covered by that fix:
```
PROBE aligned maxLength 1e12 compiles; total_size = 1099511627776
```
A one-field schema declares a 1 TiB reservation; `total_size` in that
range is meaningless output the consumer may act on (N2's argument
(a)). Unlike align, no `Access` error follows at read time (an empty
buffer still fails buffer bounds first), so this is layout-meaningless
output, not a crash — the exact severity N2 recorded. Additionally,
`parse_max_length` returns `Option` and silently *drops* values that
overflow `usize` (`.and_then(|n| usize::try_from(n).ok())`) — on a
32-bit target a 5 GiB `maxLength` becomes "no maxLength", changing
layout semantics without telling the consumer.
**Fix**: the N2 playbook verbatim. A `MAX_LENGTH` cap in
`schema.rs` (value TBD — `align`'s 4096 is page granularity; a
reservation cap in the tens-of-megabytes range fits honest layouts;
suggest `2^26 = 67_108_864`, matching `MAX_ARRAY_BYTES`'s rationale),
enforced in `parse_max_length` (converted to `Result<Option<usize>>`,
clean `Schema` error naming the path/value/maximum — no silent drop),
plus `"maximum": 67108864` in the meta-schema's `maxLength` property
so the published contract matches the parser (the N2 dual-layer
pattern).
**Resolution (2026-09-02):** the N2 playbook, cap = `MAX_LENGTH`
(2^26 = 67_108_864, matching `MAX_ARRAY_BYTES`'s rationale: a single
fixed reservation no larger than the largest legal array):
1. `MAX_LENGTH` added to `schema.rs`, documented with the F2 probe
arithmetic.
2. `parse_max_length` converted to `Result<Option<usize>>`: non-integer
→ clean `Schema` error; `usize` overflow → clean `Schema` error (the
silent `.and_then(try_from().ok())` drop is gone); over-cap → clean
`Schema` error naming the path, value, and maximum.
3. Meta-schema `maxLength` property gains `"maximum": 67108864` — the
published contract matches the parser.
4. Tests (5, in `offset_map.rs`, mirroring the `n2_` family):
above-cap rejection naming value+maximum (bytes and string),
at-cap acceptance (`total_size == 67108864`), u64::MAX-scale value
rejected-not-silently-dropped (cap arm on 64-bit, overflow arm on
32-bit — one test covers whichever fires), and the meta-schema
dual-layer check (above-cap rejected, at-cap accepted).
Verified with F1's commit: 477 lib tests green, clippy clean, wasm
green.
### C1. Packed `validate_bytes` has never decoded a wide primitive
**Files**: `src/materialize.rs:121-169` (`materialize_plan_primitive`'s
Int16/Int32/Int64/Uint64/Float64/Boolean arms — all 0-execution),
`src/sequential_reader.rs:345-383` (the reader's same arms are covered
via `read_next` tests, but the materializer's are not)
**Problem**: every packed `validate_bytes` test feeds u8/uint32/
string-shaped data. The i16/i32/i64/u64/f64/bool arms of the plan
materializer — the code every untrusted packed wire buffer flows
through — have never executed through any test. Probe (in-session)
confirmed the BE i16/bool path works; the arms are correct, just
unexercised. This is the flagship decode path for `alkcall`'s packed
frames; one battery test closes it (mirror the aligned
`read_field` battery, tests/engine_integration.rs:130-190, which
already covers all twelve primitive kinds on the aligned side).
**Resolution (2026-09-02):** two tests in `engine.rs`:
`c1_validate_bytes_packed_decodes_all_twelve_primitives_le` (the full
eleven-field battery — i8..bool — plus a corrupted-bool rejection arm)
and `c1_validate_bytes_packed_decodes_big_endian_subset` (BE i16/u64/
f64 through the same public path). Both green.
### C2. Aligned `validate_bytes` never exercises the default inline encoding for string/bytes
**Files**: `src/materialize.rs:1037-1039`
(`materialize_variable_aligned`'s `LengthPrefixed`-else branch —
0-exec through the public path)
**Problem**: the aligned `validate_bytes` tests use records, unions,
maxLength reservations, and offset-indirect encodings. The *default*
encoding — an inline length-prefixed string or bytes field, the most
common real shape — reaches `read_field` (engine_integration.rs:192)
but never `validate_bytes`. The aligned `validate_bytes` surface has
thus never decoded the single most likely field kind through its
public path.
**Fix**: one aligned `validate_bytes` test with a trailing inline
string (and a bytes variant or arm), asserting acceptance plus a
short-buffer rejection.
**Resolution (2026-09-02):**
`c2_validate_bytes_aligned_inline_string_and_bytes_default_encoding`
in `engine.rs`. One wrinkle the test wrote itself into: ADR-006
allows an inline length-prefixed variable field only in the final
position, so the string and bytes shapes get separate one-field
schemas (string after a fixed `id`; bytes as a lone field). Asserts
acceptance for both plus a short-buffer rejection for the string.
### C3. `ReadPlan::compile`'s union-variant cycle arm is untested standalone
**Files**: `src/read_plan.rs:509-514` (`compile_variant`'s
`cycle_err` arm — 0-exec)
**Problem**: the H2 test family exercises `check_ref_graph` (walk
guard) via `OffsetMap::compute`/`LayoutBuilder::new`/
`materialize_aligned`, and `ValidationPlan::compile`'s cycle arm is
covered (`validation_plan.rs:297` shows executions, via the
`shared_refs_compile_without_false_cycle`/cycle tests). But
`ReadPlan::compile`'s own cycle rejection — the defense the *packed
read plan* relies on when driven standalone (its doc explicitly
promises untrusted-input safety) — has no test driving a two-def
cycle through it. The depth cap is tested
(`deep_nesting_beyond_depth_cap_is_schema_error`); the cycle arm is
shadowed in every engine-path test by the ValidationPlan gate running
first (engine.rs:153).
**Fix**: a `read_plan_compile_two_def_cycle_rejected` test calling
`ReadPlan::compile` directly on a two-def cycle, mirroring
`validation_plan.rs`'s existing standalone cycle test.
**Resolution (2026-09-02):**
`c3_cycle_through_union_mapping_variant_is_schema_error` in
`read_plan.rs`. Writing the test sharpened the finding: the
*field-level* cycle arm (`compile_typeref`, :362) was already covered
(2 execs) by `cyclic_ref_through_two_defs_is_schema_error`; the
0-exec arm was `compile_variant`'s own check (:513), reachable only
when the cycle closes through a **union mapping entry**. The new
test's shape (`A → B → U(mapping: "1" → $ref B)`) trips exactly that
arm — verified post-fix at the line level (1 execution).
### L1. `builder.rs`'s JSON-Schema conveniences are entirely untested
**Files**: `src/builder.rs:268-290` (`array()`, `number()`,
`boolean_()`, `null()`), `:465-479` (`items()`,
`additional_properties()`), `:508-565` (`maximum()`, `min_length()`,
`min_items()`, `max_items()`, `format()`, `title()`,
`description()`), `:440` (the `field()`-on-standard-repr path)
**Problem**: only the BAST-side builders have tests. The standard
JSON-Schema side feeds `jsonschema::build_validator` (the
`json_schema` parameter of `AlkTypeEngine::compile`), so a typo'd or
misplaced keyword would ship silently — the builder emits the JSON,
`jsonschema` interprets it, and nothing checks the translation. One
table-style test asserting each convenience produces the expected
JSON key/value closes the surface cheaply.
**Resolution (2026-09-02):** five tests in `builder.rs`:
`l1_standard_type_constructors_produce_type_keyword` (all eight
standard constructors, exact-JSON assertions),
`l1_field_on_standard_object_builds_properties` (`field()` on the
standard repr + `required()`),
`l1_items_and_additional_properties_on_standard_types`,
`l1_constraint_keywords_emit_expected_json_keys` (minimum/maximum/
minLength/minItems/maxItems/format/title/description, each asserted
on its exact keyword), and
`l1_standard_built_schema_compiles_as_json_validator` (the end of the
translation chain: the emitted JSON builds a `jsonschema` validator
and the constraints actually bite — valid passes, over-maximum/
missing-required/below-minimum fail).
### L2. `tunion::read_field_discriminator`'s enum arm is 0-exec
**Files**: `src/tunion.rs:166-169`
**Problem**: N1's resolution extended tunion to match the reader's
kind set and added uint16/uint32 tests both endians — but skipped the
enum arm, which is in the documented kind set
(tunion.rs:106-112 names "string / uint8 / uint16 / uint32 / enum").
The reader's enum arm is tested (`m4_field_disc_enum_dispatches_on_index`);
tunion's is not. One test locks parity on the last arm.
**Resolution (2026-09-02):** `l2_read_field_discriminator_enum_
dispatches_on_index` and `l2_read_field_discriminator_enum_big_endian`
in `tunion.rs` — enum index 0 (LE) and 1 (BE) dispatch with
`variant_offset == 4` / `discriminator_size == 4`.
### N1. `OffsetEntry::start()`/`end()` 0-execution in the combined run is a merge artifact, not a hole
**Files**: `src/offset_map.rs:79-86`
The combined `cargo llvm-cov --release` run reports these 0-exec;
`tests/poc_roundtrip.rs:187-189` calls `start()` (and the
`big_endian_round_trip_via_offset_map` test calls `end()`). Per-test
coverage confirms both execute (32/2 calls respectively in a
poc_roundtrip-only run). llvm-cov's profile merge does not attribute
integration-test-binary executions to the library in every run
configuration. Recorded so nobody "fixes" this by deleting the
accessors or writing a redundant in-module test. (Caveat for future
audits: when a combined run shows 0-exec on something an integration
test visibly calls, re-run per-test-target before classifying.)
### N2a. `data_access.rs`'s remaining uncovered lines are the documented >4 GiB guards — fine to leave
**Files**: `src/data_access.rs:54-95, 223-291, 336-414`
All are `checked_add` overflow arms and u32-truncation guards needing
multi-GiB slices or near-`usize::MAX` offsets — already documented as
defensively-unreachable on 64-bit test hardware in #006 M4 item 3's
resolution. (The `read_array`/`write_array` arms at :54-95 are
additionally unreachable-after-`check_bounds` belt-and-suspenders.)
No action.
### N3a. `bast.rs` dead-or-orphaned surface — flag for the pre-release review
**Files**: `src/bast.rs:212-214, 305-307, 404-406, 506-508, 741-743,
924-926, 973-975` (`source()` accessors — zero callers anywhere in
src or tests), `:353-363` (`BastField::synthetic`,
`#[allow(dead_code)]`, zero callers), `:149-167`
(`resolve_typeref_as_def`'s inline struct/union/enum arms — both call
sites pass `$ref`-only variants since H3's parse rules forbid
re-declaration; plausibly dead now)
Three small deletions-or-justifications. Not fixed this session (the
`source()` accessors are public API — removal is a semver decision
for the pre-release review, and AGENTS.md's semver exception list
says renames/removals need an explicit ask). Recorded so the
pre-release review session has the list.
### N4a. Internal-shape error arms are structurally unreachable — fine to leave
**Files**: `src/materialize.rs:241,309,422,858`,
`src/sequential_reader.rs:295,473,533,832`, `src/offset_map.rs:352`,
`src/layout_builder.rs:194,282`, `src/read_plan.rs:402`
The `"internal: …"` arms that dispatch on a `match` the caller
already narrowed (e.g. "union body at X is not CompositePlan::Union"
inside a function only reachable from a `Union` match arm). They are
honest defensive code — deleting them would force `unwrap()` — and
forcing them in tests would require constructing mid-walk corruption.
Leave uncovered; the pattern is consistent across the codebase.
---
## What's Good
- The #006 fix sessions left the tree in genuinely good shape: 90.67%
lines with every high-traffic wire path (reader dispatch, plan
compiler, offset map, walk guard) in the mid-90s or better.
- The untrusted-input discipline is visible in the coverage: every
parse-level gate added in #006 (H1 caps, N2 align cap, N3
string/bytes-only maxLength, H2 cycle rejections at all three
standalone walkers) has both rejection and boundary tests.
- The `#[cfg(test)]` helper noise is the only thing making
`bast.rs`/`data_access.rs` look worse than they are — the
production coverage of both is materially higher than the raw
per-file number.
## Recommended Order
1. ~~**F1** — guard hoist + cross-consumer agreement test~~
**resolved 2026-09-02** (with F2).
2. ~~**F2** — `MAX_LENGTH` cap, N2's dual-layer playbook verbatim~~
**resolved 2026-09-02** (with F1).
3. ~~**C1 + C2 + C3** — one locking test each~~ **resolved 2026-09-02**.
4. ~~**L1 + L2** — posture tests~~ **resolved 2026-09-02**.
5. **N3a** — defer to the pre-release review (semver decision).
## Notes
- Probe tests were run as `tests/zzz_probe.rs` in-tree during the
session and deleted before any commit (the #006 pattern). Neither
probe was a crash hazard; both reproduce safely in the default
harness.
- Per-file numbers are from a single `cargo llvm-cov --release`
run; the N1 merge artifact means integration-test-only calls
(e.g. `OffsetEntry::start()`) can show 0-exec in the combined
report — the classification above already accounts for that.
- The coverage holes fixed this session (C1-C3, L1, L2) were chosen
because each is load-bearing *and* one-test-cheap; the remaining
uncovered mass is dominated by N2a/N3a/N4a, which are documented
rather than forced.
- Static test counts at the review-#007 commits: 477 (F1/F2),
481 (C1-C3), 488 (L1/L2) — +14 net from the pre-audit 474.
- Post-fix coverage (same tool, full run): TOTAL **91.66% lines /
87.64% functions** (from 90.67/86.32). Per-file movement:
`builder.rs` 91.28→99.33, `engine.rs` 96.44→96.76,
`materialize.rs` 85.72→87.74, `read_plan.rs` 90.69→90.94,
`tunion.rs` 92.02→92.86. The remaining mass is the documented
N2a/N3a/N4a classes.
+296
View File
@@ -0,0 +1,296 @@
---
status: resolved (F1, F2 fixed 2026-09-07; N1, N2, N3 classified; N4 fixed 2026-09-07)
last_updated: 2026-09-07
reviewed_artifacts:
- src/read_plan.rs
- src/sequential_reader.rs
- src/materialize.rs
- src/bast.rs
- benches/wire_vs_bast.rs
- docs/reviews/007-coverage-audit.md (N3a disposition)
tool: manual diff review of post-#007 commits (dea96f0, d4635d2) + disposable probe tests (run in-session, then deleted) + cargo bench + counting-allocator peak-RSS probe
reviewer: pre-publish review #008 (session request — audit the two post-#007 perf/bench commits, then gate the 0.3.0 publish)
---
# Review #008 — Pre-Publish Review: Post-#007 Perf Commits
## Purpose
0.3.0's release commit (`9949f91`) predates review #006 entirely; the
fix sessions for #006 and #007 landed eleven more commits, and *after*
#007 closed, two more commits landed unreviewed: `dea96f0` (bench port
from alktty) and `d4635d2` (the perf commit — fixed-size struct fast
path, integer union dispatch, `read_next_borrowed`). The perf commit
touches the flagship packed read path, which every untrusted wire
buffer flows through. This review audits those two commits before the
first crates.io publish of the 0.3.x line (0.1.0 and 0.2.0 are
published; 0.3.0 never was — every post-release fix can legally ride
inside the first published 0.3.0, no semver conflict).
It also disposes of review #007's N3a — the one finding explicitly
deferred to "the pre-release review", which this session is.
## Methodology
- Full diff read of `d4635d2` (perf) and `dea96f0` (bench port),
cross-checked against the invariants the earlier reviews established:
cross-consumer dispatch agreement (#006 H3, #007 F1), the H1
no-count-sized-prealloc rule, and the N2/F2 dual-layer cap pattern.
- Disposable probe tests (`tests/zzz_probe*.rs`, deleted after the
session; none was a crash hazard) to confirm/deny the three
behaviors code reading flagged: the `int_keys` non-canonical-key
divergence, the engine's nested-array acceptance envelope, and the
`with_capacity` amplification.
- A counting-`GlobalAlloc` probe (peak-bytes metric) to measure the
worst-case simultaneous allocation of the amplification shape
precisely — RSS timing proved too noisy to separate the two test
cases.
- `cargo bench --quick` before/after the fixes to confirm the perf
commit's wins survive.
- N3a dispositions probed where cheap (inline-struct union variants).
## Baseline
Audited at `main` HEAD `d4635d2`, 0.3.0, working tree clean. 566
tests green (488 lib + 17 + 34 + 15 + 12, + 2 ignored doctests),
clippy `-D warnings` clean, wasm build green — per the perf commit's
verification block.
## Summary Statistics
| Severity | Count | Status |
|----------|------:|--------|
| High | 0 | — |
| Medium | 2 (F1, F2) | both fixed 2026-09-07 |
| Info | 3 (N1, N2, N3) | classified |
| Fix | 1 (N4) | fixed 2026-09-07 |
No Highs: both Mediums are probe-verified cross-consumer divergences
and resource-bound violations, but neither aborts the process
(`with_capacity` is now bounded per array by `MAX_ARRAY_ELEMENTS`, so
H1's 1 TB SIGABRT class does not return). Both were fixed in-session
because they violate AGENTS.md §3 (divergent verdicts on untrusted
input; unbounded-count-shaped allocation) — the publish gate.
**Resolution log:**
- **F1 + F2 (2026-09-07):** fixed in one commit — see the resolution
blocks. 569 tests green (491 lib + 17 + 34 + 15 + 12, + 2 ignored),
clippy `-D warnings` clean, doc 0 warnings, wasm green.
- **N4 (2026-09-07):** fixed with F1/F2 — see the block.
---
## Findings
### F1. `int_keys` integer dispatch breaks cross-consumer agreement on non-canonical mapping keys
**Files**: `src/read_plan.rs` (`compile_int_keys`, introduced by
`d4635d2`), contrast `src/materialize.rs` (`materialize_plan_union`'s
byte-disc arm — stringifies), `src/validation_plan.rs`
(`validate_union_numeric` — stringifies), `src/tunion.rs`
(`read_byte_discriminator` — stringifies)
**Problem**: `d4635d2` added a pre-parsed `(u64, variant_index)`
dispatch table for byte-discriminator unions: when every mapping key
parses as `u64`, the reader matches the raw discriminator integer
instead of stringifying per read. But the meta-schema does not
constrain mapping-key shape beyond "object property name", and
`key.parse::<u64>()` accepts **non-canonical** decimal strings:
```
PROBE1 reader: field=msg disc="01" (DISPATCHED)
PROBE1 validate_bytes: Err(access error at msg: union discriminator value 1 not in mapping)
```
With mapping key `"01"` (and discriminator `1` on the wire): the
reader's numeric dispatch **matches** (`"01".parse::<u64>() == 1`) and
dispatches — returning `discriminator == "01"` — while the
materializer (`1.to_string() == "1" ≠ "01"`), the validation plan, and
tunion all **reject** the identical buffer. Pre-`d4635d2`, all four
consumers stringified and all four rejected — agreement held (both
verdicts "reject", same error class). The perf commit flipped the
reader to accept-while-everyone-else-rejects: the exact
cross-consumer-divergence shape #006 H3 and #007 F1 exist for, on the
flagship path. `"+1"` parses as `u64` too (Rust's `from_str_radix`
accepts a leading `+`) — same class. The returned key string also
became schema-quirk-dependent: the reader reports `"01"` where the
materializer's `__discriminator` for a *matched* key would report the
stringified form.
**Not a #007 regression**: `int_keys` did not exist before `d4635d2`.
But `d4635d2` postdates #007's close and was unreviewed — this is the
audit catching it.
**Fix**: build the integer table only from **canonical** keys — a key
qualifies iff `key.parse::<u64>()` succeeds *and*
`parsed.to_string() == key` (i.e. the key is exactly what
stringification would produce). Any non-canonical or non-numeric key
falls back to the string path (`int_keys: None`), which every consumer
already agrees on. No accepted schema's *reachable* behavior changed:
for fully-canonical mappings the numeric dispatch behaves identically
to stringified matching (the numeric value's `to_string()` equals the
key), and for non-canonical keys all consumers now reject exactly as
before `d4635d2`. The perf win (no per-read stringify) is preserved
for every mapping that was unambiguous to begin with.
**Resolution (2026-09-07):** exactly that — `compile_int_keys` now
requires `v.to_string() == *key` for the table to carry the entry;
any miss returns `Ok(None)` (string fallback). Doc comment states the
canonicality rule and why. Tests in `read_plan.rs`:
`r8_non_canonical_mapping_key_disables_int_dispatch` (key `"01"` →
`int_keys` is `None`) and `r8_canonical_mapping_keys_keep_int_dispatch`
(keys `"1"`,`"2"` → table `[(1,0),(2,1)]`). Probe output after the fix:
both `validate_bytes` and the reader reject disc 1 under key `"01"`
with the same error class — agreement restored.
### F2. `Vec::with_capacity(count)` reintroduced on both array materializers — ~477 MB simultaneous allocation from a ~1 KB schema
**Files**: `src/materialize.rs:247` (`materialize_plan_array` — the
`validate_bytes` packed path), `src/materialize.rs:650`
(`materialize_array_packed` — the legacy walker, reachable via the
aligned record arm)
**Problem**: `d4635d2`'s "materialize: with_capacity for bytes arrays,
arrays, and struct objects" item reintroduced
`Vec::with_capacity(count)` at two of the three sites H1's layer-1 fix
had converted to `Vec::new()` + push. `count` is now compile-capped at
`MAX_ARRAY_ELEMENTS` (2^16), so H1's 1 TB `SIGABRT` does not return —
but the per-array cap does not bound *nesting*:
```
PROBE validate peak bytes allocated simultaneously: 476780249
```
A schema of one 100-level nested array chain (each `count: 65535`,
innermost elements empty structs — all legal: the depth cap is 128 and
stride-0 chains evade `MAX_ARRAY_BYTES`, which only checks stride
products) peaks at **~477 MB of simultaneous allocation** on
`validate_bytes(&[])` from a ~1 KB schema and an *empty* buffer. Each
level's `with_capacity(65535 × sizeof(Value))` stays live across its
element walk, so the sizes multiply across ~127 legal depth levels
(the innermost zero-progress rejection fires only after the whole
chain has descended). On wasm32 — which this crate explicitly targets
— the same shape aborts the wasm heap well below 477 MB. H1's layer-1
rule ("no count-sized prealloc on untrusted input; the per-element
walk dominates") is exactly the invariant this violates; the perf
commit's own bench evidence doesn't need the prealloc either (see
below).
The other `with_capacity` additions in the commit are fine: byte-array
capacity from `b.len()` (a read slice), struct-object capacity from
`plan.fields().len()`, and the plan-compiler's from `fields.len()` —
all bounded by data/plan already in hand, not by declared counts.
**Fix**: restore H1's layer-1 shape at both sites — `Vec::new()` +
push loop (the loops already push `count` elements; the zero-progress
guard bounds honest progress per element). Optionally cap
preallocation at a small constant, but plain `Vec::new()` matches H1's
shipped behavior.
**Resolution (2026-09-07):** both sites restored to `Vec::new()` +
push. Bench before/after (criterion `--quick`, this session): packet
read 246 → 220 µs, chunk read 76/68 µs — the revert costs nothing
measurable on the bench shapes (small arrays; the materializer's
per-element work dominates), and the union/struct preallocs stay.
Locking test in `materialize.rs`:
`r8_deeply_nested_stride0_array_rejects_before_bulk_prealloc` (the
100-level chain still rejects cleanly with the zero-progress error at
the innermost level; the allocation shape itself is documented here —
in-tree cannot cheaply assert peak allocation, and the #008 probe was
deleted per the no-reproducer rule).
### N1. `d4635d2`'s fixed-size fast paths are sound (classified, no action)
**Files**: `src/read_plan.rs` (`fixed_size`, `fixed_plan_size`), `src/sequential_reader.rs`
The fixed-size struct fast path replaces the cursor size walk with one
bounds check; `fixed_plan_size` already returned `Result<Option>` with
clean overflow errors (L1's shape), and every new error arm formats
paths lazily on the error path only. The union-variant fast path
(`plan_variant_fixed_size`) applies only to struct variants and checks
bounds before use. No issue found.
### N2. Bench port (`dea96f0`) is methodology-honest (classified, no action)
**Files**: `benches/wire_vs_bast.rs`
The port drops alktty's async I/O group (correctly — it measured a
different stack) and adds a parity check before measurement so the
stream loop can't drift. The historical `read_chunk_stream` numbers
stay comparable by construction. No issue found.
### N3. N3a dispositions (review #007's deferred items)
**Files**: `src/bast.rs` (`source()` accessors, `BastField::synthetic`,
`resolve_typeref_as_def`'s inline arms)
- **`source()` accessors (7 sites)**: public API on `BastStruct`/
`BastUnion`/`BastField`/etc. Removal is a semver decision and
AGENTS.md's semver exception requires an explicit ask — **kept**.
They are one-line accessors over parsed source nodes, harmless, and
plausibly useful to downstream codegen (the announced consumer).
- **`BastField::synthetic` (`#[allow(dead_code)]`, zero callers)**:
`pub(crate)`, not public API — **deleted** (2026-09-07). No semver
impact; the `#[allow(dead_code)]` suppression is gone with it.
- **`resolve_typeref_as_def`'s inline struct/union/enum arms**: the
review-#007 suspicion ("plausibly dead after H3") was wrong —
probe-verified reachable: the meta-schema's
`mapping.additionalProperties: TypeRef` accepts inline struct
variants, and `LayoutBuilder`'s byte-disc and field-disc arms call
`resolve_typeref_as_def` on every union variant. The H3 parse rules
forbid variant *re-declaration of shared fields*, not inline variant
bodies. **Kept**, reachable.
### N4. Stale test-count references in review #006's resolution log
**Files**: `docs/reviews/006-implementation-review-030.md`
The bookkeeping note ("static count at `2eb086f` is 542 + 2 ignored")
and per-commit counts are accurate as written; no fix needed. Recorded
here so the review trail stays honest about what was re-checked
during this session's doc sweep. **Resolution (2026-09-07):** no code
change; superseded the "Fix" entry — this is the classification
record.
---
## What's Good
- The perf commit's core ideas are sound and survived review: the
compile-time `fixed_size` cache is computed through the existing
`Result`-returning sizer (no `unwrap_or_default` regression), and
the int-dispatch table's design was right — it just needed the
canonicality gate.
- The counting-allocator probe took 15 minutes and converted a
"probably too big" into a precise number (476,780,249 bytes) — the
same probe pattern the earlier reviews used, applied to allocation
instead of verdicts.
- `cargo bench --quick` before/after the fixes is the right tool for
guarding perf-fix reverts: packet read 220 µs post-fix vs 246 µs
baseline confirms the `Vec::new()` restore is free.
## Recommended Order
1. ~~**F1** — canonical-key gate on `compile_int_keys`~~ **fixed
2026-09-07**.
2. ~~**F2** — restore H1's no-prealloc rule at both array sites~~
**fixed 2026-09-07**.
3. ~~**N4** — `BastField::synthetic` deletion~~ **fixed 2026-09-07**.
4. **N3 source() accessors** — revisit only if/when the codegen
consumer confirms it does not want them (removal needs an explicit
ask per AGENTS.md).
## Notes
- Probe tests were run as `tests/zzz_probe*.rs` in-tree during the
session and deleted before any commit (the #006 pattern). None was a
crash hazard; the amplification probe allocates ~477 MB transiently
and completes in ~40 ms.
- Benches are not run in CI and are excluded from the publish (the
`[bench]` target ships — that is fine; benches don't affect the
library's API or its wasm compatibility).
- The 0.3.0 publish proceeds after these fixes: 0.1.0 and 0.2.0 are
on crates.io; this is the first 0.3.0 publish, so F1/F2's
behavior changes (both "previously-divergent, now-agreed" shapes)
land inside the version's first release — no semver bump implied.
+1049
View File
File diff suppressed because it is too large. Load diff
+52
View File
@@ -0,0 +1,52 @@
[package]
name = "alktype-fuzz"
version = "0.0.0"
publish = false
edition = "2021"
[package.metadata]
cargo-fuzz = true
[dependencies]
libfuzzer-sys = "0.4"
alktype-fuzz-shared = { path = "shared" }
[dependencies.alktype]
path = ".."
[[bin]]
name = "bast_compile"
path = "fuzz_targets/bast_compile.rs"
test = false
doc = false
bench = false
[[bin]]
name = "data_access"
path = "fuzz_targets/data_access.rs"
test = false
doc = false
bench = false
[[bin]]
name = "read_opseq"
path = "fuzz_targets/read_opseq.rs"
test = false
doc = false
bench = false
[[bin]]
name = "layout_build"
path = "fuzz_targets/layout_build.rs"
test = false
doc = false
bench = false
[[bin]]
name = "validate_pair"
path = "fuzz_targets/validate_pair.rs"
test = false
doc = false
bench = false
[workspace]
+67
View File
@@ -0,0 +1,67 @@
# alktype fuzzing
cargo-fuzz targets for the binary struct engine's untrusted-input
surfaces. The design and operating rules live in
`docs/plans/fuzzing.md` (adopted from alkhttp's
`docs/plans/fuzzing.md`; rationale in alkcall's
`docs/research/fuzzing.md`) — this README is the operational
cheat-sheet.
## Layout
- `fuzz_targets/` — nightly-only `fuzz_target!` binaries (thin wrappers).
- `shared/` — stable-toolchain library holding the invariant logic; the
corpus replay tests run here on plain `cargo test`.
- `corpus/<target>/` — committed seeds (regenerate with
`python3 fuzz/gen_fuzz_seeds.py`).
- `artifacts/` — gitignored crash/oom/timeout artifacts + campaign logs.
## Targets
| Target | Drives |
|---|---|
| `bast_compile` | `AlkTypeEngine::compile` in both layout modes over attacker-shaped BAST JSON (the whole schema side through one choke point) + `validate_bast_doc` + `build_validator` |
| `data_access` | the hand-rolled decode core (`src/data_access.rs`): fixed-width kinds, bool strictness, length-prefixed and indirect string/bytes, enums — over raw bytes with attacker-chosen offsets and endianness |
| `read_opseq` | the stateful `SequentialReader` (packed read side): op sequences (Next / NextBorrowed / Field / Reset / End) over hostile buffers under the compiled plan — cursor discipline, plan-order walks, the record-count spin bound, ADR-007 reader independence |
| `layout_build` | the packed write side (`LayoutBuilder::build`) with adversarial `var_sizes` over a five-schema menu — position disjointness/bounds, failed-write buffer-untouched contracts, write→read pair round trip |
## Running a campaign — always detached
Agent sessions must never run fuzzing in the foreground (an OOM in a
target can take down the session host; see docs/plans/fuzzing.md §2).
Use the detached runner:
```bash
fuzz/run-detached.sh bast_compile
# poll:
tail -n 50 fuzz/artifacts/bast_compile-*.log
ls fuzz/artifacts/bast_compile/
pgrep -f "cargo fuzz run bast_compile"
```
`FUZZ_RUNTIME_SECS=1800 fuzz/run-detached.sh bast_compile` for a longer
campaign. The runner pins `-fork=1 -rss_limit_mb=2048
-malloc_limit_mb=2048 -timeout=25` and detaches via `setsid` + `nohup`.
## Corpus replay (the standing fuzz gate)
```bash
cargo test --manifest-path fuzz/shared/Cargo.toml
```
replays every committed seed through the same invariant functions the
fuzz targets run — on stable, without nightly, no cargo-fuzz. Part of
the release verification checklist (AGENTS.md).
## Toolchain
`fuzz/rust-toolchain.toml` pins nightly (+ `llvm-tools-preview`) for
this subtree only; the main crate stays stable at MSRV 1.85. `cargo
fuzz build` works from any CWD inside `fuzz/` (rustup resolves the
toolchain per directory). Build:
```bash
cd fuzz && cargo fuzz build
# or from the repo root — the toolchain file is picked up by path:
cargo fuzz build -D
```
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}]}, "root": "S"}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "endian": "big", "fields": [{"name": "x", "kind": "uint8"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": []}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "f0", "kind": "int8"}, {"name": "f1", "kind": "int16"}, {"name": "f2", "kind": "int32"}, {"name": "f3", "kind": "int64"}, {"name": "f4", "kind": "uint8"}, {"name": "f5", "kind": "uint16"}, {"name": "f6", "kind": "uint32"}, {"name": "f7", "kind": "uint64"}, {"name": "f8", "kind": "float32"}, {"name": "f9", "kind": "float64"}, {"name": "f10", "kind": "bool"}, {"name": "f11", "kind": "string"}, {"name": "f12", "kind": "bytes"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "endian": "little", "fields": [{"name": "a", "kind": "uint32", "endian": "big"}, {"name": "b", "kind": "string", "encoding": "length-prefixed", "maxLength": 64}, {"name": "c", "kind": "bytes", "encoding": "offset-indirect", "maxLength": 128}, {"name": "d", "kind": "string", "encoding": "offset-indirect"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "child", "kind": {"$ref": "#/$defs/Nested"}}, {"name": "arr", "kind": {"kind": "array", "element": "uint32", "count": 3}}, {"name": "arr0", "kind": {"kind": "array", "element": "uint8", "count": 0}}, {"name": "rec", "kind": {"kind": "record", "values": "string"}}, {"name": "en", "kind": {"$ref": "#/$defs/E"}}, {"name": "un", "kind": {"$ref": "#/$defs/U1"}}, {"name": "un2", "kind": {"$ref": "#/$defs/U2"}}]}, "Nested": {"kind": "struct", "fields": [{"name": "y", "kind": "int16"}]}, "E": {"kind": "enum", "values": ["a", "b", "c"]}, "U1": {"kind": "union", "discriminator": {"kind": "byte", "offset": 0, "type": "uint8"}, "mapping": {"0": "uint8", "1": {"$ref": "#/$defs/Nested"}}}, "U2": {"kind": "union", "discriminator": {"kind": "field", "name": "tag"}, "fields": [{"name": "tag", "kind": "uint32"}], "mapping": {"0": "uint8", "7": "string"}}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "inl", "kind": {"kind": "struct", "fields": [{"name": "z", "kind": "uint8"}]}}, {"name": "inle", "kind": {"kind": "enum", "values": ["x"]}}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "l", "kind": {"$ref": "#/$defs/N"}}, {"name": "r", "kind": {"$ref": "#/$defs/N"}}]}, "N": {"kind": "struct", "fields": [{"name": "v", "kind": "uint8"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}], "align": 4096}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": "uint32", "align": 16}, {"name": "b", "kind": "uint8", "align": 1}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "s", "kind": "string", "maxLength": 67108864}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "s", "kind": "string", "maxLength": 0}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "align": 4097, "fields": []}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "align": 65536, "fields": []}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "align": 18446744073709551615, "fields": []}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"kind": "array", "element": "uint8", "count": 65537}}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"kind": "array", "element": "uint8", "count": 18446744073709551615}}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "s", "kind": "string", "maxLength": 67108865}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "nonsense"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "uint8", "fields": []}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct"}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"$ref": "#/$defs/S"}}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"$ref": "#/$defs/Missing"}}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {}}
+1
View File
@@ -0,0 +1 @@
[1, 2, 3]
+1
View File
@@ -0,0 +1 @@
{}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": 123}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "9bad", "kind": "uint8"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": "uint8", "bogus": true}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "union", "discriminator": {"kind": "byte", "offset": 0, "type": "uint8"}}}}
+1
View File
@@ -0,0 +1 @@
not json at all
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S":
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n",Line truncated
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n",Line truncated
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}]}, "D100": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/S"}}]}, "D99": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D100"}}]}, "D98": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D99"}}]}, "D97": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D98"}}]}, "D96": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D97"}}]}, "D95": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D96"}}]}, "D94": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D95"}}]}, "D93": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D94"}}]}, "D92": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D93"}}]}, "D91": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D92"}}]}, "D90": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D91"}}]}, "D89": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D90"}}]}, "D88": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D89"}}]}, "D87": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D88"}}]}, "D86": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D87"}}]}, "D85": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D86"}}]}, "D84": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D85"}}]}, "D83": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D84"}}]}, "D82": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D83"}}]}, "D81": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D82"}}]}, "D80": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D81"}}]}, "D79": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D80"}}]}, "D78": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D79"}}]}, "D77": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D78"}}]}, "D76": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D77"}}]}, "D75": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D76"}}]}, "D74": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D75"}}]}, "D73": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D74"}}]}, "D72": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D73"}}]}, "D71": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D72"}}]}, "D70": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D71"}}]}, "D69": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D70"}}]}, "D68": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D69"}}]}, "D67": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D68"}}]}, "D66": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D67"}}]}, "D65": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D66"}}]}, "D64": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D65"}}]}, "D63": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D64"}}]}, "D62": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D63"}}]}, "D61": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D62"}}]}, "D60": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D61"}}]}, "D59": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D60"}}]}, "D58": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D59"}}]}, "D57": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D58"}}]}, "D56": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D57"}}]}, "D55": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D56"}}]}, "D54": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D55"}}]}, "D53": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D54"}}]}, "D52": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D53"}}]}, "D51": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D52"}}]}, "D50": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D51"}}]}, "D49": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D50"}}]}, "D48": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D49"}}]}, "D47": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D48"}}]}, "D46": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D47"}}]}, "D45": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D46"}}]}, "D44": {"kind": "struct", "fields": [{"name": "Line truncated
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": "uint8"}, {"name": "a", "kind": "uint16"}]}}}
+1
View File
@@ -0,0 +1 @@
{"$defs": {"S": {"kind": "enum", "values": ["only"]}, "root": "S"}}
+1
View File
@@ -0,0 +1 @@

Binary file not shown.
+1
View File
@@ -0,0 +1 @@

+1
View File
@@ -0,0 +1 @@
�
+1
View File
@@ -0,0 +1 @@
�
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+1
View File
@@ -0,0 +1 @@
4
+1
View File
@@ -0,0 +1 @@
4
Binary file not shown.
Binary file not shown.
+1
View File
@@ -0,0 +1 @@
��
+1
View File
@@ -0,0 +1 @@
��
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+1
View File
@@ -0,0 +1 @@
D3"
+1
View File
@@ -0,0 +1 @@
"3D
Loaded 100 of 337 files, more files were not shown because too many files have changed in this diff. Show more