1 Commits
Author SHA1 Message Date
glm-5.2 f7c71da9e5 POC: ReadPlan shape derisking for ADR-011
Standalone workspace member at poc/readplan/ that depends on alktype
via path and exercises the ReadPlan/CompositePlan/ReadKind/
DiscriminatorPlan shape from ADR-011 against every BastType arm in the
current read loop.

Result: 28/30 tests pass. 2 deliberately ignored, both with documented
findings:

- Field-name-discriminator union read shape is a TODO (compile shape
  is correct; the read-side stub surfaces the work for implementation
  step 1 rather than hiding it).
- Existing SequentialReader returns element_stride=0 for fixed-size
  struct arrays (pre-existing limitation at sequential_reader.rs:567,
  not a plan-shape gap; the POC plan correctly computes the stride).

Coverage confirms every BastType arm compiles to the expected
ReadKind/CompositePlan. Equivalence tests confirm plan-driven read
produces identical (FieldValue, position) to the existing reader for
all covered cases. ReadPlan: Send + Sync confirmed.

Green light for ADR-011 implementation. See poc/readplan/FINDINGS.md
for the full writeup.

This branch is a derisking POC, not meant to merge to main (mirrors
the bast-validator-poc branch pattern). Cargo.toml gains a workspace
section that includes poc/readplan; that section is POC-only and would
be dropped if these files ever merged to main.
2026-08-18 09:37:33 +00:00
321 changed files with 3264 additions and 19584 deletions

No files matched your search

+1 -10
View File
@@ -1,12 +1,3 @@
target/
node_modules/
.worktrees/
fuzz/target/
fuzz/artifacts/
fuzz/coverage/
# Grown corpora: the hash-named files the campaigns drop into
# fuzz/corpus/<target>/ are gitignored (the quinn/h2 policy); the
# committed seeds are the seed-* files, kept via the per-dir
# .gitignore un-ignores.
fuzz/corpus/*/[0-9a-f][0-9a-f]*
.worktrees/
-33
View File
@@ -5,29 +5,6 @@ auto-loads this file as instructions, overriding the built-in defaults for
this project. Custom agents in `.opencode/agents/` inherit these rules
unless their own prompts say otherwise.
## Session Continuity (keep the agent loop alive)
opencode ends the turn whenever an assistant message contains no tool
call — including messages that are pure analysis. Long reasoning bursts
are welcome in this repo (they pre-catch errors and self-correct), but a
burst that ends as analysis-only text silently stops the session
mid-task. Past sessions documented this repeatedly ("the session
stalled"; the working fix discovered there: "call tools frequently, keep
thinking bursts short"). Keep the depth; change where the burst ends:
1. **Never end a turn with analysis-only text.** Every visible message
must either issue a tool call or be a final report for a genuinely
completed phase/task. When a thinking burst converges on a decision,
act on it (read, edit, bash) in the same turn.
2. **Land work incrementally.** Once a design decision is settled, write
the code before analyzing the next one. Do not emit full-design
essays in a single chat message; reasoning belongs in thinking tokens
or committed docs, not in the transcript.
3. **On a silent turn end, resume without re-deriving.** If the turn
ended after an analysis-only message and the task is incomplete,
pick up from the last settled decision — do not redo the analysis
and do not ask the user whether to continue.
## Git Workflow
**Commit and push when reasonable.** When a change is complete and
@@ -165,18 +142,8 @@ cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # if docs changed
cargo build --target wasm32-unknown-unknown --release # if layout/wasm-relevant code changed
cargo publish --dry-run --allow-dirty # before a release
cargo test --manifest-path fuzz/shared/Cargo.toml # fuzz corpus replay (the fuzz gate)
```
The corpus replay is the standing fuzz gate (the alkcall
`docs/research/fuzzing.md` §7.9 posture): it replays every committed
seed through the same invariant functions the fuzz targets run, on
stable, without nightly. Campaigns (nightly, cargo-fuzz) run manually
via `fuzz/run-detached.sh` — never as a foreground child of an agent
session — before releases, after touching `src/data_access.rs`,
`src/schema.rs`, the compile walks, or the sequential reader. See
`docs/plans/fuzzing.md` and `fuzz/README.md`.
## Architecture Context
- `docs/architecture/` — the authoritative spec. Read it before
-151
View File
@@ -4,156 +4,6 @@ All notable changes to this crate are documented here. The format is
based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and
this crate adheres to [Semantic Versioning](https://semver.org/).
## [0.4.0] - 2026-09-30
The fuzzing release. The fuzz/ workspace (docs/plans/fuzzing.md) put
five libFuzzer targets on the engine — bast_compile, data_access,
read_opseq, layout_build, validate_pair — with release-budget campaigns
across all five. Two genuine engine bugs found and fixed same-day,
plus one upstream pin (docs/plans/fuzzing.md §6, §5 campaign numbers).
### Breaking changes
- **`VariableEncoding` gains `MaxLengthReserved`** (finding W3-3):
the aligned-mode `maxLength` reservation (ADR-003 strategy 2, the
`VARCHAR(N)` pattern) is now recorded as its own encoding variant
instead of masquerading as `LengthPrefixed`. Code matching on
`VariableEncoding` exhaustively must add an arm; in the document
form the strategy is still expressed via `maxLength`, never as an
`encoding` value.
- **`read_field`/`write_field` fix for aligned `maxLength`
reservations** (finding W3-3, commit `a0dd3d2`): an aligned
String/Bytes leaf with a declared `maxLength` is now read as a raw
zero-padded, NUL-trimmed window and written zero-padded — previously
the first four raw bytes of the reservation window were misparsed as
a u32 length prefix, breaking the validate_bytes ⇒ read_field
agreement for every aligned schema declaring `maxLength`.
- **`plan_read_array` bounds check** (finding W2-1, wave 2): a
fixed-stride array whose declared window (count × stride) extends
past the buffer now returns an `Access` error naming the array
field instead of reporting success and deferring the failure to the
next field read (or masking it entirely when the array was last).
### Additions
- **`data_access::read_reservation` / `read_reservation_string` /
`write_reservation`** — the single source of truth for the
`maxLength` reservation leaf semantics, shared with the aligned
materializer.
- **`fuzz/` subtree** — five libFuzzer targets, 257 committed seeds,
the corpus-replay gate (`cargo test --manifest-path
fuzz/shared/Cargo.toml`, AGENTS.md verification checklist), the
detached campaign runner; nightly confined to `fuzz/`, `fuzz/`
excluded from the package. All targets, seeds, and the replay gate
are stable-toolchain safe.
## [0.3.0] - 2026-09-07
The compiled-forms release. The packed read path — the hot path for
stream parsing — is driven by a compile-once `ReadPlan` instead of a
per-call walk of the BAST typed tree; byte validation runs a compiled
`ValidationPlan`; plans and offset maps fingerprint to a stable hash
(ADR-011/ADR-012). Reads of SFTP-shaped packet streams went from
~189× hand-rolled Rust to ~74× (~2.4× faster), fixed-stride chunk
reads from ~18× to ~11×, and `SequentialReader::read_next_borrowed`
makes the per-field hot loop allocation-free.
### Breaking changes
- **`Bast*` types are owned.** `BastDoc`/`BastStruct`/`BastField`/… no
longer borrow from the source `serde_json::Value`; all v0.2.0
lifetimes are gone. `BastDoc::new` parses the root eagerly; `$ref`s
resolve lazily.
- **`OffsetMap::get` / `PackedLayout::get`** return `&OffsetEntry`
(was `Option<OffsetEntry>` by value), backed by an O(log n)
`BTreeMap` path→index (first-occurrence-wins for duplicate names).
- **`SequentialReader::new`** takes the compiled plan; construct via
`AlkTypeEngine::sequential_reader()` (packed mode only).
- **`materialize_packed` / `materialize_aligned`** take the compiled
plan / `(&BastDoc, &OffsetMap)` pair respectively.
- **Field-name-discriminator union wire convention** (ADR-011
addendum): the builder lays out the union's declared `fields`
(shared) first, then the variant's own fields. Variants must not
re-declare the discriminator or any shared field, and the
discriminator field must be the first entry in `fields` — all
enforced at parse with clean `Schema` errors. Schemas relying on
0.2.0's variant-only layout are rejected (they produced
reader↔builder-disagreeing bytes).
- **`maxLength` is string/bytes-only** — rejected at parse on every
other kind (it was silently unenforced there).
- **Aligned-mode `Record` fields reject `offset-indirect`** (the
materializer always walks the inline count-prefixed form — the
annotated shape was never readable).
- **Schema input bounds** (untrusted-schema hardening, AGENTS.md §3):
array `count` ≤ 2^16 and `count × stride` ≤ 2^26 bytes; `align` ≤
4096; `maxLength` ≤ 2^26; cyclic `$ref` graphs and >128-deep nesting
are rejected by every public walker (`OffsetMap::compute`,
`LayoutBuilder::new`, `materialize_aligned` included), not just the
engine.
### Additions
- **`ReadPlan`** (ADR-011) — the compiled packed-read plan, re-exported
with `CompositePlan`/`FieldPlan`/`ReadKind`/`DiscriminatorPlan`.
`ReadPlan::compile` is untrusted-input-safe standalone (depth cap +
cycle set). `fixed_size()` exposes the compile-time-known byte size
for fixed structs.
- **`ValidationPlan`** (ADR-012 §3) — the compiled `validate_bytes`
walker, with `ValidNode`/`ValidVariant` sub-types.
- **`fingerprint()`** on `ReadPlan`/`OffsetMap`/`ValidationPlan` +
`Hash`/`Eq` derives on the plan types (ADR-012 §1/§4) — plan
identity for cache-keying across processes.
- **`OffsetMap` `LeafMeta`** — each entry records whether it is
fixed/length-prefixed/offset-indirect so `read_field`/`write_field`
dispatch without re-walking the schema; `OffsetEntry` type re-exported.
- **`SequentialReader::read_next_borrowed`** — zero-allocation variant
of `read_next` (field name borrowed from the plan).
- **`AlkTypeEngine::validate_bytes`** now runs the compiled
`ValidationPlan` (was an interpretive BAST walk in 0.2.0).
### Fixes (post-release-commit hardening — reviews #006, #007, #008)
All found and fixed before the first crates.io publish of 0.3.0, so
no published version ever exhibited them.
- **Untrusted-input crashes removed.** A huge declared array count
OOM-aborted the process (`Vec::with_capacity(count)` before reading
a byte) — now compile-capped and walked with push-only growth.
Cyclic `$ref` graphs stack-overflowed the three standalone layout
walkers — now guarded by a shared reference-graph check. Deeply
nested stride-0 arrays briefly allowed ~477 MB of simultaneous
allocation from a ~1 KB schema — restored to incremental growth.
- **Cross-consumer divergences closed.** Builder, reader,
materializer, tunion, and the validation plan now agree on
field-disc union layout (shared-then-variant), on the discriminator
field's position (must be first), and on union mapping-key matching
(numeric fast-path dispatch only for canonical keys like `"2"`;
`"01"`/`"+1"` fall back to the string comparison all consumers
share). The legacy BAST walker's field-disc union arm walks shared
fields before the variant (it previously materialized variant fields
from shared fields' bytes).
- **Silently-corrupt layouts rejected.** Aligned record fields with
`maxLength`/`offset-indirect`; non-final inline length-prefixed
fields (records included — the ADR-006 check now sees them);
aligned-mode `maxLength`/`offset-indirect` on records; unions in
aligned mode (pre-existing, now tested).
- **Coverage**: 90.67% lines / 86.32% functions at review #007's
audit, 91.66% after its fixes; every uncovered region outside test
modules read and classified in-tree (docs/reviews/007).
### Non-breaking improvements
- Engine compile is one-shot and allocation-tidy; plans are
`Send + Sync` (statically asserted) and fingerprintable.
- Zero-progress array-element guard on all three array walkers (a
zero-size element makes the declared count unbounded on the wire).
- WASM-clean unchanged: two dependencies (`jsonschema`
default-features off, `serde_json` with `preserve_order`), no
`async`, no `unsafe`, no feature flags.
- Benches (`benches/wire_vs_bast.rs`): read/write chunk streams, an
SFTP-shaped union packet stream, and `validate_bytes` per buffer —
the numbers quoted above and in ADR-007/ADR-011.
## [0.2.0] - 2026-08-17
A breaking release that replaces the v0.1.0 `AlkType:*` custom-keyword
@@ -265,6 +115,5 @@ Initial crates.io release. Custom-keyword JSON Schema format
`AlkTypeEngine` with packed/aligned layout modes, builder API producing
`serde_json::Value`.
[0.3.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.3.0
[0.2.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.2.0
[0.1.0]: https://git.alk.dev/alkdev/alktype/releases/tag/v0.1.0
Generated
+9 -188
View File
@@ -27,9 +27,8 @@ dependencies = [
[[package]]
name = "alktype"
version = "0.4.0"
version = "0.2.0"
dependencies = [
"criterion",
"jsonschema",
"serde_json",
]
@@ -40,18 +39,6 @@ version = "0.2.21"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "683d7910e743518b0e34f1186f92494becacb047c7b6bf616c96772180fef923"
[[package]]
name = "anes"
version = "0.1.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "4b46cbb362ab8752921c97e041f5e366ee6297bd428a31275b9fcf1e380f7299"
[[package]]
name = "anstyle"
version = "1.0.14"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "940b3a0ca603d1eade50a4846a2afffd5ef57a9feac2c0e2ec2e14f9ead76000"
[[package]]
name = "autocfg"
version = "1.5.1"
@@ -97,107 +84,12 @@ version = "0.6.9"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "175812e0be2bccb6abe50bb8d566126198344f707e304f45c648fd8f2cc0365e"
[[package]]
name = "cast"
version = "0.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "37b2a672a2cb129a2e41c10b1224bb368f9f37a2b16b612598138befd7b37eb5"
[[package]]
name = "cfg-if"
version = "1.0.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801"
[[package]]
name = "ciborium"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "42e69ffd6f0917f5c029256a24d0161db17cea3997d185db0d35926308770f0e"
dependencies = [
"ciborium-io",
"ciborium-ll",
"serde",
]
[[package]]
name = "ciborium-io"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "05afea1e0a06c9be33d539b876f1ce3692f4afea2cb41f740e7743225ed1c757"
[[package]]
name = "ciborium-ll"
version = "0.2.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "57663b653d948a338bfb3eeba9bb2fd5fcfaecb9e199e87e1eda4d9e8b240fd9"
dependencies = [
"ciborium-io",
"half",
]
[[package]]
name = "clap"
version = "4.6.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "473c7e07f409a8d772161724aa8db6a765a2532a70f9667eeb7b49d3d02fbdca"
dependencies = [
"clap_builder",
]
[[package]]
name = "clap_builder"
version = "4.6.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "7b48fea5a88e9ae728a2dcbedbfc0e730f7d60da42e1cb049a83c9fb8b789889"
dependencies = [
"anstyle",
"clap_lex",
]
[[package]]
name = "clap_lex"
version = "1.1.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9"
[[package]]
name = "criterion"
version = "0.7.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "e1c047a62b0cc3e145fa84415a3191f628e980b194c2755aa12300a4e6cbd928"
dependencies = [
"anes",
"cast",
"ciborium",
"clap",
"criterion-plot",
"itertools",
"num-traits",
"oorandom",
"regex",
"serde",
"serde_json",
"tinytemplate",
"walkdir",
]
[[package]]
name = "criterion-plot"
version = "0.6.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9b1bcc0dc7dfae599d84ad0b1a55f80cde8af3725da8313b528da95ef783e338"
dependencies = [
"cast",
"itertools",
]
[[package]]
name = "crunchy"
version = "0.2.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "460fbee9c2c2f33933d720630a6a0bac33ba7053db5344fac858d4b8952d77d5"
[[package]]
name = "data-encoding"
version = "2.11.0"
@@ -215,12 +107,6 @@ dependencies = [
"syn 2.0.119",
]
[[package]]
name = "either"
version = "1.18.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "252afb9ae5eaa683babdc6a068b3f5726eb19e05070c731f9b2a23a7c3e8ed34"
[[package]]
name = "email_address"
version = "0.2.9"
@@ -288,17 +174,6 @@ dependencies = [
"wasm-bindgen",
]
[[package]]
name = "half"
version = "2.7.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "6ea2d84b969582b4b1864a92dc5d27cd2b77b622a8d79306834f1be5ba20d84b"
dependencies = [
"cfg-if",
"crunchy",
"zerocopy",
]
[[package]]
name = "hashbrown"
version = "0.16.1"
@@ -429,15 +304,6 @@ dependencies = [
"hashbrown 0.17.1",
]
[[package]]
name = "itertools"
version = "0.13.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "413ee7dfc52ee1a4949ceeb7dbc8a33f2d6c088194d9f922fb8318faf1f01186"
dependencies = [
"either",
]
[[package]]
name = "itoa"
version = "1.0.18"
@@ -613,12 +479,6 @@ version = "1.21.4"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50"
[[package]]
name = "oorandom"
version = "11.1.5"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
[[package]]
name = "outref"
version = "0.5.2"
@@ -687,6 +547,14 @@ version = "5.3.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "69cdb34c158ceb288df11e18b4bd39de994f6657d83847bdffdbd7f346754b0f"
[[package]]
name = "readplan-poc"
version = "0.0.0"
dependencies = [
"alktype",
"serde_json",
]
[[package]]
name = "redox_syscall"
version = "0.5.18"
@@ -768,15 +636,6 @@ version = "1.0.23"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "cf54715a573b99ac80df0bc206da022bcd442c974952c7b9720069370852e21f"
[[package]]
name = "same-file"
version = "1.0.6"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "93fc1dc3aaa9bfed95e02e6eadabb4baf7e3078b0bd1b4d7b6b0b68378900502"
dependencies = [
"winapi-util",
]
[[package]]
name = "scopeguard"
version = "1.2.0"
@@ -882,16 +741,6 @@ dependencies = [
"zerovec",
]
[[package]]
name = "tinytemplate"
version = "1.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "be4d6b5f19ff7664e8c98d03e2139cb510db9b0a60b55f8e8709b689d939b6bc"
dependencies = [
"serde",
"serde_json",
]
[[package]]
name = "unicode-general-category"
version = "1.1.0"
@@ -932,16 +781,6 @@ version = "0.8.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "5c3082ca00d5a5ef149bb8b555a72ae84c9c59f7250f013ac822ac2e49b19c64"
[[package]]
name = "walkdir"
version = "2.5.0"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "29790946404f91d9c5d06f9874efddea1dc06c5efe94541a7d6863108e3a5e4b"
dependencies = [
"same-file",
"winapi-util",
]
[[package]]
name = "wasip2"
version = "1.0.4+wasi-0.2.12"
@@ -996,30 +835,12 @@ dependencies = [
"unicode-ident",
]
[[package]]
name = "winapi-util"
version = "0.1.11"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "c2a7b1c03c876122aa43f3020e6c3c3ee5c05081c9a00739faf7503aeba10d22"
dependencies = [
"windows-sys",
]
[[package]]
name = "windows-link"
version = "0.2.1"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5"
[[package]]
name = "windows-sys"
version = "0.61.2"
source = "registry+https://github.com/rust-lang/crates.io-index"
checksum = "ae137229bcbd6cdf0f7b80a31df61766145077ddf49416a728b02cb3921ff3fc"
dependencies = [
"windows-link",
]
[[package]]
name = "wit-bindgen"
version = "0.57.1"
+4 -13
View File
@@ -1,6 +1,6 @@
[package]
name = "alktype"
version = "0.4.0"
version = "0.2.0"
edition = "2021"
rust-version = "1.85"
license = "MIT OR Apache-2.0"
@@ -9,15 +9,11 @@ repository = "https://git.alk.dev/alkdev/alktype"
readme = "README.md"
keywords = ["binary", "jsonschema", "wire-format", "serialization", "layout"]
categories = ["encoding", "data-structures", "parsing"]
exclude = [".opencode/", "docs/reviews/", "docs/research/", "docs/sdd_process.md", "Cargo.lock", "AGENTS.md", "fuzz/"]
exclude = [".opencode/", "docs/reviews/", "docs/research/", "docs/sdd_process.md", "Cargo.lock", "AGENTS.md"]
[lib]
name = "alktype"
[workspace]
members = ["."]
exclude = ["fuzz"]
[features]
default = []
@@ -25,10 +21,5 @@ default = []
jsonschema = { version = "0.46", default-features = false }
serde_json = { version = "1", features = ["preserve_order"] }
[dev-dependencies]
serde_json = "1"
criterion = { version = "0.7", default-features = false }
[[bench]]
name = "wire_vs_bast"
harness = false
[workspace]
members = ["poc/readplan"]
+8 -24
View File
@@ -24,11 +24,10 @@ A BAST document serves three roles simultaneously:
| Role | Mechanism | When |
|------|-----------|------|
| **Validation spec (bytes)** | Compiled `ValidationPlan` walk over the materialized `Value` (ADR-012) | Access time (`validate_bytes`) |
| **Validation spec (bytes)** | BAST-native validator (recursive walker over the BAST type tree) | Access time (`validate_bytes`) |
| **Validation spec (JSON)** | Standard `jsonschema::Validator` from a consumer-provided JSON Schema | Load time (build validator), access time (`validate_json`) |
| **Layout spec** | Offset computation from type sizes + field order | Load time (build offset map / packed layout) |
| **Data access** | Read/write at computed offsets | Access time (read field, write field) |
| **Wire access (packed)** | Compiled `ReadPlan` (ADR-011) — compile-once, no per-read schema walk | Access time (`SequentialReader`) |
No separate format definition, no separate parser, no separate
validator. The BAST document is the single source of truth for the
@@ -160,16 +159,11 @@ same BAST document can be compiled in either mode. Decided in ADR-002.
- **Byte-offset** — a fixed-size integer (`uint8`/`uint16`/`uint32`) at
a known byte offset. The SFTP `Packet` pattern: byte 0 is the type
byte, bytes 1..N are the variant struct. Mapping keys are stringified
integers. With all-canonical numeric keys the compiled reader
dispatches on the raw integer (no per-read stringification).
integers.
- **Field-name** — a named field within the union. The TypeBox
`typedef.ts` pattern. Mapping keys are string values matching the
discriminator field's value. The `fields` array declares the
discriminator field (D-BAST-005), which must be its first entry; the
variant must not re-declare it or any shared field. The builder lays
out the declared `fields` first, then the variant's own fields
(ADR-011 addendum) — builder, reader, materializer, and validator all
agree on that convention.
discriminator field (D-BAST-005).
Variant `$ref`s are resolved lazily — no compile-time inlining step.
@@ -188,10 +182,10 @@ types (ADR-VAL-SPLIT):
- `validate_bytes(&[u8])` — for raw byte buffers (channels' chunk
header, SFTP packets). Materializes a `Value` tree from the bytes via
the layout engine, then runs the compiled **`ValidationPlan`** (0.2.0
used an interpretive BAST walker; 0.3.0 compiles the value-domain
constraints — integer ranges, `maxLength`, enum index bounds, union
variant dispatch — once at compile time). No
the layout engine, then runs the **BAST-native validator** — a
recursive walker over the BAST type tree that checks the value-domain
constraints the materializer doesn't (integer ranges, `maxLength`,
enum index bounds, union variant constraints). No
`jsonschema` involvement; the BAST document is the complete
validation spec for bytes (D-BAST-006).
- `validate_json(&Value)` / `is_valid_json(&Value)` — for already-parsed
@@ -250,15 +244,6 @@ code was converted to `Err` ahead of v0.1.0 (review #002, L2); the
BAST parser preserves this invariant — overflow-safe arithmetic
(`checked_add`, `usize::try_from`) on all offset/count casts.
0.3.0 adds compile-time bounds for adversarial schemas: array counts
≤ 2^16 elements, computed array sizes ≤ 2^26 bytes, `align` ≤ 4096,
`maxLength` ≤ 2^26, and a shared reference-graph guard that rejects
cyclic `$ref`s and >128-deep nesting in every public schema walker.
Adversarial buffers fail with `Access` errors at read time — the
materializers never preallocate from declared counts. Reviews #006,
#007, and #008 document the audit trail
([docs/reviews/](docs/reviews/)).
## Documentation
Architecture documentation lives under [`docs/architecture/`](docs/architecture/):
@@ -284,8 +269,7 @@ Architecture documentation lives under [`docs/architecture/`](docs/architecture/
(ADR-003), error handling (ADR-004), int64/uint64 kinds (ADR-005),
non-final inline variable fields (ADR-006), packed-mode read factory
(ADR-007), TUnion in aligned mode (ADR-008), builder API (ADR-009),
`validate_bytes` (ADR-010), compiled read plan (ADR-011), plan
fingerprinting + `ValidationPlan` (ADR-012)
`validate_bytes` (ADR-010)
## License
-645
View File
@@ -1,645 +0,0 @@
//! Informal speed comparison: hand-rolled codec logic vs alktype-driven
//! codec over the same wire shapes.
//!
//! History: this bench originated (uncommitted) in `alktty` as the
//! curiosity probe that surfaced review #004's 400x read gap — the
//! finding that drove the 0.3.0 compiled-forms release (ADR-011/012).
//! It now lives here so alktype owns its perf story. The alktty-only
//! async roundtrip group (tokio `ChunkReader`/`ChunkWriter` over a
//! duplex pipe) was dropped — that measures alktty's I/O stack, not
//! this engine.
//!
//! The alktype engine / layout / plans are built **once outside** the
//! measured routine, per the "build cost is paid once" framing.
//!
//! Shapes:
//!
//! - **ChunkHeader** — a 5-byte header (`stream_type: uint8`,
//! `length: uint32` big-endian). The original shape, kept so numbers
//! stay comparable with the historical series (review #004: 400x →
//! 0.3.0: ~18x on read p64).
//! - **Read** — the hand-rolled path mirrors
//! `ChunkReader::read_chunk_after_peek` minus the tokio I/O
//! (identical overhead on both sides): validate `stream_type <= 4`,
//! parse `u32::from_be_bytes`, slice the payload. The alktype path
//! drives `SequentialReader::read_next` over the `ChunkHeader`
//! struct, then slices the payload at the parsed length. Both
//! return a `&[u8]` payload view — no allocation in either measured
//! path.
//! - **Write** — serialize the 5-byte header. Hand-rolled mirrors
//! `ChunkWriter::write_chunk`'s header writes; alktype uses
//! `PackedLayout` offsets (built once) and
//! `data_access::write_u8`/`write_u32` at those offsets.
//! - **Packet** — a byte-offset-discriminator union
//! (`Read {handle, length}` / `Write {handle, length, data: bytes}`),
//! the SFTP-shaped case ADR-011's framing argument was about:
//! exercises `CompositePlan::Union` dispatch, variant walks, and
//! length-prefixed variable reads. The alktype consumer pattern is
//! the documented one: `read_next` on the root yields
//! `FieldValue::Union { discriminator, variant_start }`, the consumer
//! selects the pre-built reader for that variant and walks it over
//! `&buf[variant_start..]`.
//! - **validate_bytes** — `engine.validate_bytes` per buffer
//! (materialize + `ValidationPlan` walk, ADR-010/ADR-012 §3): the
//! read+validate-on-untrusted-stream shape `alkcall` cares about. No
//! hand comparator: a hand-rolled codec validates inline during the
//! (already measured) parse, while `validate_bytes` additionally
//! materializes a `Value` tree per buffer — the honest reading is the
//! absolute per-chunk cost.
//!
//! One-shots (paid once at startup, not per chunk):
//! `alktype_engine_compile` (dominated by BAST meta-schema
//! validation), `alktype_sequential_reader_new` (an `Arc::clone`),
//! `alktype_layout_build`.
//!
//! Two payload sizes (64 B, 4 KiB) so per-chunk fixed overhead is
//! visible separately from payload-copy cost.
//!
//! Run: `cargo bench --bench wire_vs_bast`
use criterion::{criterion_group, criterion_main, BenchmarkId, Criterion};
use std::hint::black_box;
use alktype::{
data_access, AlkTypeEngine, Endian, FieldValue, LayoutBuilder, LayoutMode, PackedLayout,
ReadPlan, SequentialReader,
};
/// Mirrors `alktty::wire::MAX_CHUNK_LEN` — the hand-rolled comparator
/// validates against the same cap the real codec enforces.
const MAX_CHUNK_LEN: u32 = 16 * 1024 * 1024;
/// The `ChunkHeader` BAST definition. The `StreamType` enum is
/// intentionally NOT used — BAST enums encode as `u32`, but the
/// on-wire `stream_type` is a `uint8`; both sides read it as `uint8`.
const CHUNK_HEADER_BAST: &str = r#"{
"$schema": "https://alk.dev/bast/v1/schema",
"$defs": {
"ChunkHeader": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "stream_type", "kind": "uint8" },
{ "name": "length", "kind": "uint32" }
]
}
}
}"#;
/// SFTP-shaped byte-discriminator union: one byte selects the variant,
/// `Write` carries a trailing length-prefixed `bytes` field. The root
/// struct wraps the union (`AlkTypeEngine::compile` requires a struct
/// root); mapping keys are the stringified `uint8` discriminator
/// values.
const PACKET_BAST: &str = r##"{
"$schema": "https://alk.dev/bast/v1/schema",
"$defs": {
"Packet": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "event", "kind": { "$ref": "#/$defs/Event" } }
]
},
"Event": {
"kind": "union",
"discriminator": { "kind": "byte", "offset": 0, "type": "uint8" },
"mapping": {
"5": { "$ref": "#/$defs/Read" },
"6": { "$ref": "#/$defs/Write" }
}
},
"Read": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" }
]
},
"Write": {
"kind": "struct",
"endian": "big",
"fields": [
{ "name": "handle", "kind": "uint32" },
{ "name": "length", "kind": "uint32" },
{ "name": "data", "kind": "bytes" }
]
}
}
}"##;
// ---------------------------------------------------------------------------
// ChunkHeader fixtures
// ---------------------------------------------------------------------------
/// One chunk's worth of bytes on the wire: 5-byte header + payload.
fn make_chunk_bytes(stream_type: u8, payload: &[u8]) -> Vec<u8> {
let mut buf = Vec::with_capacity(5 + payload.len());
buf.push(stream_type);
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
buf.extend_from_slice(payload);
buf
}
/// Concatenate `n` chunks into one buffer, each with `payload_len` bytes.
fn make_chunk_stream(n: usize, payload_len: usize) -> Vec<u8> {
let payload = vec![0xA5u8; payload_len];
let mut buf = Vec::with_capacity(n * (5 + payload_len));
for i in 0..n {
let st = (i % 5) as u8;
buf.extend_from_slice(&make_chunk_bytes(st, &payload));
}
buf
}
// ---------------------------------------------------------------------------
// Hand-rolled chunk read: mirrors ChunkReader::read_chunk_after_peek minus
// the tokio I/O. Returns (stream_type, payload) so the compiler can't
// elide the work. Validates stream_type <= 4 and length <= MAX_CHUNK_LEN.
// ---------------------------------------------------------------------------
#[inline]
fn hand_read_header(buf: &[u8]) -> Option<(u8, u32)> {
if buf.len() < 5 {
return None;
}
let stream_type = buf[0];
if stream_type > 4 {
return None;
}
let length = u32::from_be_bytes([buf[1], buf[2], buf[3], buf[4]]);
if length > MAX_CHUNK_LEN {
return None;
}
Some((stream_type, length))
}
#[inline]
fn hand_read_chunk(buf: &[u8]) -> Option<(u8, &[u8])> {
let (st, len) = hand_read_header(buf)?;
let end = 5usize.checked_add(len as usize)?;
if buf.len() < end {
return None;
}
Some((st, &buf[5..end]))
}
/// Drive `hand_read_chunk` across `n` contiguous chunks in `buf`.
/// Returns the total payload bytes consumed (so the loop body is
/// meaningfully used and not optimized away).
fn hand_read_stream(buf: &[u8], n: usize) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
let (st, payload) = match hand_read_chunk(&buf[pos..]) {
Some(v) => v,
None => break,
};
total += payload.len();
pos += 5 + payload.len();
black_box(st);
}
black_box(total)
}
// ---------------------------------------------------------------------------
// alktype chunk read: SequentialReader over ChunkHeader. The reader is
// constructed once per benchmark group and reset() between chunks. After
// the header read, the payload is sliced at the parsed length — same as
// the hand-rolled path. We do NOT re-read a length prefix for the payload
// (that would be the double-prefix problem).
// ---------------------------------------------------------------------------
fn alktype_read_stream(buf: &[u8], n: usize, reader: &mut SequentialReader) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
reader.reset();
let st = match reader.read_next_borrowed(&buf[pos..]) {
Ok(Some((_, FieldValue::U8(v)))) => v,
_ => break,
};
let len = match reader.read_next_borrowed(&buf[pos..]) {
Ok(Some((_, FieldValue::U32(v)))) => v,
_ => break,
};
if len > MAX_CHUNK_LEN {
break;
}
let end = match 5usize.checked_add(len as usize) {
Some(e) if e <= buf.len() - pos => e,
_ => break,
};
let payload = &buf[pos + 5..pos + end];
total += payload.len();
pos += end;
black_box(st);
black_box(payload.as_ptr());
}
black_box(total)
}
// ---------------------------------------------------------------------------
// Hand-rolled chunk write: mirrors ChunkWriter::write_chunk's header
// writes into a caller-provided buffer. Writes `n` contiguous chunks.
// ---------------------------------------------------------------------------
fn hand_write_stream(out: &mut Vec<u8>, n: usize, payload_len: usize) {
let payload = vec![0xA5u8; payload_len];
for i in 0..n {
let st = (i % 5) as u8;
let start = out.len();
out.resize(start + 5 + payload_len, 0);
out[start] = st;
out[start + 1..start + 5].copy_from_slice(&(payload_len as u32).to_be_bytes());
out[start + 5..start + 5 + payload_len].copy_from_slice(&payload);
}
black_box(out.len());
}
// ---------------------------------------------------------------------------
// alktype chunk write: data_access::write_u8 / write_u32 at the
// PackedLayout offsets. The layout is built once per group and reused.
// Payload bytes are copied with the same slice copy as the hand-rolled
// path so the comparison isolates the header-encoding overhead.
// ---------------------------------------------------------------------------
fn alktype_write_stream(out: &mut Vec<u8>, n: usize, payload_len: usize, layout: &PackedLayout) {
let payload = vec![0xA5u8; payload_len];
let st_pos = layout.get("stream_type").expect("stream_type field").offset;
let len_pos = layout.get("length").expect("length field").offset;
for i in 0..n {
let start = out.len();
out.resize(start + 5 + payload_len, 0);
let _ = data_access::write_u8(out, start + st_pos, (i % 5) as u8, "stream_type");
let _ = data_access::write_u32(
out,
start + len_pos,
payload_len as u32,
"length",
Endian::Big,
);
out[start + 5..start + 5 + payload_len].copy_from_slice(&payload);
}
black_box(out.len());
}
// ---------------------------------------------------------------------------
// Packet fixtures: byte-disc union stream, alternating Read/Write
// variants. Wire layout per packet (packed, big-endian):
// Read: disc(1) + handle(4) + length(4) = 9 bytes
// Write: disc(1) + handle(4) + length(4) + len(4)+data = 13 + payload
// ---------------------------------------------------------------------------
fn make_packet_bytes(disc: u8, payload: &[u8]) -> Vec<u8> {
let mut buf = Vec::with_capacity(13 + 4 + payload.len());
buf.push(disc);
buf.extend_from_slice(&0x0102_0304u32.to_be_bytes());
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
if disc == 6 {
buf.extend_from_slice(&(payload.len() as u32).to_be_bytes());
buf.extend_from_slice(payload);
}
buf
}
fn make_packet_stream(n: usize, payload_len: usize) -> Vec<u8> {
let payload = vec![0xA5u8; payload_len];
let mut buf = Vec::new();
for i in 0..n {
let disc = if i % 2 == 0 { 5u8 } else { 6u8 };
buf.extend_from_slice(&make_packet_bytes(disc, &payload));
}
buf
}
// ---------------------------------------------------------------------------
// Hand-rolled packet read: read the discriminator byte, match the
// variant, parse its fields directly. Returns bytes consumed.
// ---------------------------------------------------------------------------
fn hand_read_packet(buf: &[u8]) -> Option<usize> {
let disc = *buf.first()?;
match disc {
5 => {
if buf.len() < 9 {
return None;
}
let handle = u32::from_be_bytes(buf[1..5].try_into().ok()?);
let length = u32::from_be_bytes(buf[5..9].try_into().ok()?);
black_box((handle, length));
Some(9)
}
6 => {
if buf.len() < 13 {
return None;
}
let handle = u32::from_be_bytes(buf[1..5].try_into().ok()?);
let length = u32::from_be_bytes(buf[5..9].try_into().ok()?);
let data_len = u32::from_be_bytes(buf[9..13].try_into().ok()?);
let end = 13usize.checked_add(data_len as usize)?;
if buf.len() < end {
return None;
}
black_box((handle, length));
black_box(&buf[13..end].as_ptr());
Some(end)
}
_ => None,
}
}
fn hand_read_packet_stream(buf: &[u8], n: usize) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
let Some(consumed) = hand_read_packet(&buf[pos..]) else {
break;
};
total += consumed;
pos += consumed;
}
black_box(total)
}
// ---------------------------------------------------------------------------
// alktype packet read: the documented union consumer contract. The root
// reader walks the wrapping struct; `read_next` returns
// `FieldValue::Union { discriminator, variant_start }`; the consumer
// selects the pre-built reader for that variant and walks it over
// `&buf[pos + variant_start..]` until exhausted.
// ---------------------------------------------------------------------------
/// Walk one variant's fields to exhaustion; returns bytes consumed.
/// Uses `read_next_borrowed` — the zero-alloc hot-loop pattern for
/// consumers that match or discard the field name.
fn alktype_walk_variant(reader: &mut SequentialReader, buf: &[u8]) -> Option<usize> {
reader.reset();
loop {
match reader.read_next_borrowed(buf) {
Ok(Some((name, value))) => {
black_box(name);
black_box(&value);
}
Ok(None) => return Some(reader.position()),
Err(_) => return None,
}
}
}
fn alktype_read_packet_stream(
buf: &[u8],
n: usize,
packet: &mut SequentialReader,
read: &mut SequentialReader,
write: &mut SequentialReader,
) -> usize {
let mut pos = 0usize;
let mut total = 0usize;
for _ in 0..n {
packet.reset();
let disc = match packet.read_next_borrowed(&buf[pos..]) {
Ok(Some((_, FieldValue::Union {
discriminator,
variant_start,
}))) => {
pos += variant_start;
discriminator
}
_ => break,
};
let vbuf = &buf[pos..];
let consumed = match disc.as_str() {
"5" => alktype_walk_variant(read, vbuf),
"6" => alktype_walk_variant(write, vbuf),
_ => break,
};
let Some(consumed) = consumed else {
break;
};
total += consumed;
pos += consumed;
}
black_box(total)
}
/// One-time sanity check (outside the measured loops): the union
/// consumer pattern the stream loop relies on — root reader reports the
/// mapping key and the variant start; the variant reader's walk to
/// exhaustion reports exactly the variant's byte size, so
/// `variant_start + consumed` lands on the next packet.
fn assert_packet_reader_parity(
payload_len: usize,
packet: &mut SequentialReader,
read: &mut SequentialReader,
write: &mut SequentialReader,
) {
let payload = vec![0u8; payload_len];
let read_pkt = make_packet_bytes(5, &payload);
packet.reset();
match packet.read_next_borrowed(&read_pkt) {
Ok(Some((_, FieldValue::Union {
discriminator,
variant_start,
}))) => {
assert_eq!(discriminator, "5");
assert_eq!(variant_start, 1, "variant starts after the 1-byte disc");
}
_ => panic!("expected union value for Read packet"),
}
let consumed = alktype_walk_variant(read, &read_pkt[1..]).expect("read variant walk");
assert_eq!(consumed, 8, "Read = handle(4) + length(4)");
assert_eq!(1 + consumed, read_pkt.len(), "Read packet fully consumed");
let write_pkt = make_packet_bytes(6, &payload);
packet.reset();
match packet.read_next_borrowed(&write_pkt) {
Ok(Some((_, FieldValue::Union {
discriminator,
variant_start,
}))) => {
assert_eq!(discriminator, "6");
assert_eq!(variant_start, 1);
}
_ => panic!("expected union value for Write packet"),
}
let consumed = alktype_walk_variant(write, &write_pkt[1..]).expect("write variant walk");
assert_eq!(
consumed,
12 + payload_len,
"Write = handle(4) + length(4) + len-prefix(4) + data"
);
assert_eq!(1 + consumed, write_pkt.len(), "Write packet fully consumed");
}
// ---------------------------------------------------------------------------
// Benchmarks
// ---------------------------------------------------------------------------
fn bench_read(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
let engine =
AlkTypeEngine::compile(&bast, "ChunkHeader", LayoutMode::Packed, None).expect("compile");
let mut reader = engine.sequential_reader().expect("packed reader");
let mut group = c.benchmark_group("read_chunk_stream");
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let n = 1024;
let buf = make_chunk_stream(n, payload_len);
group.bench_with_input(BenchmarkId::new("hand_rolled", label), &n, |b, &n| {
b.iter(|| hand_read_stream(black_box(&buf), n));
});
group.bench_with_input(BenchmarkId::new("alktype", label), &n, |b, &n| {
b.iter(|| alktype_read_stream(black_box(&buf), n, &mut reader));
});
}
group.finish();
}
fn bench_write(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
let builder = LayoutBuilder::new(&bast, "ChunkHeader").expect("builder");
let layout = builder
.build(&std::collections::HashMap::new())
.expect("layout");
let mut group = c.benchmark_group("write_chunk_stream");
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let n = 1024;
group.bench_with_input(
BenchmarkId::new("hand_rolled", label),
&(n, payload_len),
|b, &(n, pl)| {
b.iter(|| {
let mut out = Vec::with_capacity(n * (5 + pl));
hand_write_stream(&mut out, n, pl);
});
},
);
group.bench_with_input(
BenchmarkId::new("alktype", label),
&(n, payload_len),
|b, &(n, pl)| {
b.iter(|| {
let mut out = Vec::with_capacity(n * (5 + pl));
alktype_write_stream(&mut out, n, pl, &layout);
});
},
);
}
group.finish();
}
fn bench_packet_read(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(PACKET_BAST).expect("bast json");
let engine =
AlkTypeEngine::compile(&bast, "Packet", LayoutMode::Packed, None).expect("compile");
let mut packet_reader = engine.sequential_reader().expect("packed reader");
let read_plan = std::sync::Arc::new(ReadPlan::compile(&bast, "Read").expect("read plan"));
let write_plan = std::sync::Arc::new(ReadPlan::compile(&bast, "Write").expect("write plan"));
let mut read_reader = SequentialReader::new(read_plan);
let mut write_reader = SequentialReader::new(write_plan);
// One-time parity check of the union consumer pattern (not measured).
assert_packet_reader_parity(64, &mut packet_reader, &mut read_reader, &mut write_reader);
let mut group = c.benchmark_group("read_packet_stream");
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let n = 1024;
let buf = make_packet_stream(n, payload_len);
group.bench_with_input(BenchmarkId::new("hand_rolled", label), &n, |b, &n| {
b.iter(|| hand_read_packet_stream(black_box(&buf), n));
});
group.bench_with_input(BenchmarkId::new("alktype", label), &n, |b, &n| {
b.iter(|| {
alktype_read_packet_stream(
black_box(&buf),
n,
&mut packet_reader,
&mut read_reader,
&mut write_reader,
)
});
});
}
group.finish();
}
fn bench_validate(c: &mut Criterion) {
let header_bast: serde_json::Value =
serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
let header_engine = AlkTypeEngine::compile(&header_bast, "ChunkHeader", LayoutMode::Packed, None)
.expect("compile");
let packet_bast: serde_json::Value = serde_json::from_str(PACKET_BAST).expect("bast json");
let packet_engine =
AlkTypeEngine::compile(&packet_bast, "Packet", LayoutMode::Packed, None).expect("compile");
let mut group = c.benchmark_group("validate_stream");
let n = 1024;
let headers: Vec<Vec<u8>> = (0..n)
.map(|i| make_chunk_bytes((i % 5) as u8, &[0xA5u8; 64]))
.collect();
group.bench_function("alktype_chunk_header", |b| {
b.iter(|| {
for h in &headers {
header_engine.validate_bytes(black_box(h)).expect("validate");
}
})
});
for (payload_len, label) in [(64usize, "p64"), (4096usize, "p4k")] {
let payload = vec![0xA5u8; payload_len];
let packets: Vec<Vec<u8>> = (0..n)
.map(|i| make_packet_bytes(if i % 2 == 0 { 5 } else { 6 }, &payload))
.collect();
group.bench_with_input(
BenchmarkId::new("alktype_packet", label),
&packets,
|b, packets| {
b.iter(|| {
for p in packets {
packet_engine.validate_bytes(black_box(p)).expect("validate");
}
})
},
);
}
group.finish();
}
/// One-shot costs paid once at startup, not per chunk.
fn bench_oneshot(c: &mut Criterion) {
let bast: serde_json::Value = serde_json::from_str(CHUNK_HEADER_BAST).expect("bast json");
c.bench_function("alktype_engine_compile", |b| {
b.iter(|| {
let _ =
AlkTypeEngine::compile(black_box(&bast), "ChunkHeader", LayoutMode::Packed, None)
.expect("compile");
});
});
c.bench_function("alktype_sequential_reader_new", |b| {
let engine = AlkTypeEngine::compile(&bast, "ChunkHeader", LayoutMode::Packed, None)
.expect("compile");
b.iter(|| engine.sequential_reader());
});
c.bench_function("alktype_layout_build", |b| {
let builder = LayoutBuilder::new(&bast, "ChunkHeader").expect("builder");
b.iter(|| builder.build(&std::collections::HashMap::new()));
});
}
criterion_group!(
benches,
bench_read,
bench_write,
bench_packet_read,
bench_validate,
bench_oneshot
);
criterion_main!(benches);
+1 -2
View File
@@ -45,8 +45,7 @@ format definition; the engine is generic.
| [008](decisions/008-reject-tunion-in-aligned-mode.md) | Reject TUnion in Aligned Mode for v1 | Unions are the protocol pattern; aligned-mode union semantics were broken |
| [009](decisions/009-builder-api.md) | Builder API for Schema Construction | Fluent Rust API producing `serde_json::Value`; covers BAST kinds + standard JSON Schema; resolves OQ-003. *Output format amended to BAST / standard JSON Schema by ADR-BAST.* |
| [010](decisions/010-generalized-validation-validate-bytes.md) | Generalized Validation — `validate_bytes` on `AlkTypeEngine` | Single-call binary-buffer validation; materialize `Value` from bytes, then validate. *Validation step amended to the BAST-native validator by ADR-VAL-SPLIT.* |
| [011](decisions/011-compiled-read-plan-for-packed-mode.md) | Compiled Read Plan for Packed Mode | `ReadPlan` — the packed read-side compiled form, symmetric to `OffsetMap` (aligned) and `PackedLayout` (packed write). Closes review #004's 400x read-path gap; retires ADR-007's "re-parse on demand" framing. *Accepted — implemented in 0.3.0 (phases 1–2).* |
| [012](decisions/012-plan-fingerprinting-and-m1-closure.md) | Plan Fingerprinting, ValidationPlan, and Closing the Deferred M1 Sites in 0.3.0 | `ReadPlan`/`OffsetMap`/`ValidationPlan` `Hash + Eq` + `fingerprint()`; owned `BastDoc` (lifetime removal); `OffsetMap` carries `LeafMeta` to close the aligned-side M1 sites; `ValidationPlan` retires the interpretive `bast_validation` walk (review #005 M3 reversed the original deferral). Bundles with ADR-011 into one 0.3.0 breaking release. *Accepted — fully implemented in 0.3.0 (fingerprinting, owned `BastDoc`, `LeafMeta`, `ValidationPlan`).* |
| [011](decisions/011-compiled-read-plan-for-packed-mode.md) | Compiled Read Plan for Packed Mode | `ReadPlan` — the packed read-side compiled form, symmetric to `OffsetMap` (aligned) and `PackedLayout` (packed write). Closes review #004's 400x read-path gap; retires ADR-007's "re-parse on demand" framing. *Accepted.* |
## Relevant Open Questions
+22 -39
View File
@@ -128,12 +128,6 @@ These are different validators for different inputs.
"encoding": { "enum": ["length-prefixed", "offset-indirect"] },
"maxLength": { "type": "integer", "minimum": 0 }
},
"if": {
"properties": {
"kind": { "enum": ["string", "bytes"] }
}
},
"else": { "properties": { "maxLength": false } },
"required": ["name", "kind"],
"additionalProperties": false
},
@@ -275,11 +269,8 @@ These are different validators for different inputs.
- `align` (optional): field-level alignment (aligned mode only).
- `encoding` (optional): `"length-prefixed"` (default) or
`"offset-indirect"`. See [Variable-length encoding](#variable-length-encoding).
- `maxLength` (optional, `string`/`bytes` fields only): byte-length
cap. See [Variable-length encoding](#variable-length-encoding).
Rejected at parse on any other kind (review #006 N3: the annotation
was silently unenforced there — the validation plan bakes `maxLength`
into string/bytes leaves only).
- `maxLength` (optional): byte-length cap. See
[Variable-length encoding](#variable-length-encoding).
### TypeRef
@@ -465,11 +456,8 @@ override). In little-endian mode, `u32::from_le_bytes`; in big-endian
mode, `u32::from_be_bytes`. Ensures SFTP consumers (big-endian) have
consistent byte order for field values and length prefixes.
Applies to variable-length primitive types only: `string` and
`bytes`. The parser rejects `maxLength` (and the meta-schema forbids
it) on every other kind — including `record` (review #006 N3/M5: no
consumer honored it there, so the annotation was either silently
unenforced or, in aligned mode, silently corrupt).
Applies to all variable-length types: `string`, `bytes`,
`record`, and arrays of variable-length elements.
## Endianness
@@ -540,28 +528,24 @@ materializer iterates the field list), types are correct (`read_u32`
produces `Value::Number`), bounds are checked (via `data_access::
check_bounds`), UTF-8 is valid (via `from_utf8`), the discriminator is
in the mapping, and the boolean byte is 0 or 1. What the materializer
does NOT check — and what the validation half checks afterward — are
**value-domain constraints expressed in the BAST document**. Under
ADR-012 §3 these constraints are compiled once into a `ValidationPlan`
at engine-compile time (eager `$ref` resolution, cyclic-graph
rejection); each `validate_bytes` call walks the compiled constraint
tree against the `Value`. The plan's nodes enforce exactly these
constraints (the set is normative; the walker that enforced it
interpretively in 0.2.0 is retired):
does NOT check — and what the 19 v0.1.0 custom keyword validators check
afterward — are **value-domain constraints expressed in the BAST
document**. The BAST-native validator is a recursive walker over the
BAST type tree that checks exactly these:
| Constraint | `ValidNode` arm |
|------------|-----------------|
| Integer range (Int8..Uint32) | `Int { min, max }` / `Uint { max }` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `I64` / `U64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `Float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `Str { max_len }` — `maxLength` baked in from the owning field at compile time |
| Bytes `maxLength` (array length) | `Bytes { max_len }` — accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `Enum { count }` checks `idx < count` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `Union { variants }` reads `__discriminator`, dispatches on the compiled variant nodes |
| Struct fields | `Struct { fields }` requires each declared field present, recurses |
| Array count | `Array { count, element }` checks `arr.len() == count` and recurses per element |
| Record values | `Record { values }` recurses into each value |
| Boolean | `Bool` (materializer already rejects non-0/1 bytes) |
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `validate_float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `check_string` reads the field-level `maxLength` |
| Bytes `maxLength` (array length) | `check_bytes` accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `validate_enum` checks `idx < values.len()` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, looks up the variant, recurses via `validate_typeref` |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
@@ -600,8 +584,7 @@ The `jsonschema` crate **remains a direct dependency** for
meta-schema. The only thing removed is the custom keyword integration
path. The `validate_bytes` path no longer touches `jsonschema` — a
small wasm binary-size win in addition to the architecture
simplification. (Since ADR-012 §3, the interpretation step itself is
also compiled away: see the `ValidationPlan` above.)
simplification.
### `AlkTypeError::Validation` payload shape
@@ -136,9 +136,9 @@ reserving worst-case space.
- `true` is a shorthand for the default (length-prefixed). This keeps
the common case concise and the override explicit.
- The `encoding` annotation and `maxLength` apply to the variable-length
primitive types `AlkType:String` and `AlkType:Bytes`. (`maxLength` on
records was amended out by review #006 N3/M5 — see §3a.)
- The `encoding` annotation and `maxLength` apply to all variable-length
types: `AlkType:String`, `AlkType:Bytes`, `AlkType:Array`,
`AlkType:Record`, `AlkType:Timestamp`.
### 3a. TRecord value type
@@ -164,13 +164,8 @@ the `"values"` property in the schema:
the value's size is determined by its kind (fixed-size kinds have a
known size; variable-length kinds carry their own length prefix).
- The count and key-length prefixes respect the schema's endianness.
- ~~In aligned static mode with `maxLength`, the entire record is
reserved at `maxLength` bytes (zero-padded).~~ **Amended (review #006
N3/M5, 2026-09-02):** `maxLength` is rejected at parse on record
fields. The aligned materializer walks the record's inline
count-prefixed form and never honors the reservation (M5: silent
cross-field corruption), and no packed consumer enforced it either
(N3: silently unenforced). `maxLength` is `string`/`bytes`-only.
- In aligned static mode with `maxLength`, the entire record is reserved
at `maxLength` bytes (zero-padded).
### 4. TUnion discriminators
@@ -63,26 +63,23 @@ write-side.
### Cost
`SequentialReader::new(Arc<ReadPlan>)` is a refcount bump — 15.7 ns
(measured, alktty `wire_vs_bast` bench, 0.3.0). The reader shares the
engine's compiled [`ReadPlan`](011-compiled-read-plan-for-packed-mode.md)
(the packed read-side compiled form) via `Arc` instead of cloning
schema data; construction cost is negligible compared to reading a
buffer. The engine holds the owned `BastDoc` (ADR-012 §2a) for the
aligned materialize path and the one-shot `*::compile` paths.
`SequentialReader::new` clones the top-level struct's field schemas (a
`Vec<(String, Value)>` of the `properties` entries) and clones the
schema itself. This is cheap — a struct has a small number of fields
(SFTP's largest packet has 5). The construction cost is negligible
compared to the cost of reading a buffer.
> **Historical note**: the original 0.2.0 framing here ("re-parse on
> demand" — the read loop re-parsing `BastDoc::new` per field) was the
> root cause of the 400x read-path gap measured in
> **Note**: The "re-parse on demand" framing below (the read loop
> re-parsing `BastDoc::new` per field) is the root cause of the 400x
> read-path gap measured in
> [review #004](../../reviews/004-performance-review.md).
> [ADR-011](011-compiled-read-plan-for-packed-mode.md) (implemented,
> 0.3.0) retired it: the packed read loop walks `Arc<ReadPlan>` (2.27
> µs/chunk → 98 ns/chunk), `sequential_reader()` is an `Arc::clone`,
> and the owned `BastDoc` (ADR-012 §2a) removed the remaining
> per-access re-parse sites in `read_field`/`write_field`/`validate_bytes`.
> The factory decision itself (`sequential_reader() ->
> Option<SequentialReader>`, owned fresh reader, consumer-driven
> cursor) was retained unchanged.
> [ADR-011](011-compiled-read-plan-for-packed-mode.md) (Proposed)
> retires this framing by giving the packed read path a compiled
> `ReadPlan`; the "Cost" section here and the
> `src/engine.rs:112-115` doc comment will be updated in the
> implementation commit per ADR-011's recommended order. The factory
> decision itself (`sequential_reader() -> Option<SequentialReader>`,
> owned fresh reader, consumer-driven cursor) is retained.
## Consequences
@@ -2,53 +2,12 @@
## Status
Accepted. Implemented in 0.3.0 (phases 1–2, 2026-09-02). Closes review
#004 H1 + M1 (packed side) + L1 + L2;
Accepted. Closes review #004 H1 + M1 (packed side) + L1 + L2;
retires the "re-parse on demand" framing from ADR-007. A derisking
POC on branch `readplan-poc` confirmed the `ReadPlan` shape covers
every `BastType` arm in the current read loop before implementation
began (see "POC coverage" at the end).
**Refinements on ADR-012 acceptance (2026-08-20, review #005):** the
`CompositePlan::Union` shape was refined to carry `shared:
Option<Box<ReadPlan>>` (field-disc union shared fields, resolving POC
Finding 1 / review #005 H1) and to drop `VariantPlan`/`VariantKind`
in favor of `variants: Vec<(String, CompositePlan)>` (resolving review
#005 M1 — nested unions now work by `CompositePlan` recursion,
restoring the 0.2.0 capability the POC rejected). The "BastDoc
unchanged" scope statement stands as the ADR-011-only view; ADR-012
§2a subsequently makes `BastDoc` owned. The `ValidationPlan` this
ADR's "Out of scope" originally deferred indefinitely is now in
0.3.0 via ADR-012 §3 (review #005 M3 reversed the deferral). These
are pre-implementation refinements to types that do not yet exist on
`main`; the ADR-011 decision (a compiled `ReadPlan` for packed reads)
is unchanged.
**Addendum — field-disc union wire convention (2026-09-02, review #006
H3):** the packed-mode wire layout for a field-name-discriminator
TUnion is **shared-then-variant**: the union's declared `fields` (the
discriminator field + any shared fields) occupy the union's start
offset in declaration order, and the selected variant's fields follow
immediately after all shared fields. All three packed-mode consumers
now implement this one convention: the reader and materializer already
walked `shared` then the variant (the `shared` sub-plan shape above);
`LayoutBuilder` was corrected in the same pass — it previously laid out
only the selected variant, disagreeing with the read side on span and
field positions (review #006 H3 item 1). The convention requires that
a variant **must not re-declare** the discriminator field or any
shared field — `BastUnion::parse` enforces this at parse time (also:
the discriminator field must be declared in `fields`, and `fields`
must not contain duplicate names), so the shared walk and the variant
walk cover disjoint fields and the wire has exactly one copy of each
shared byte. Schemas whose variants redeclared shared fields were
ambiguous under the old split-convention behavior and are rejected
rather than given a silent meaning; this is a **breaking wire-format
constraint** for any 0.2.0-era schema that relied on re-declaration,
announced with the 0.3.x series. `DiscriminatorPlan::Field`'s disc
read is at the disc field's position within the shared walk (the
materializer's position-correct behavior, review #006 H3 item 2); the
reader's plan-walk reads it there too.
## Context
Review #004 (`docs/reviews/004-performance-review.md`) measured the
@@ -186,8 +145,7 @@ pub enum CompositePlan {
Struct(ReadPlan),
Union {
disc: DiscriminatorPlan,
shared: Option<Box<ReadPlan>>,
variants: Vec<(String, CompositePlan)>,
variants: Vec<(String, VariantPlan)>,
},
Array {
element: Box<CompositePlan>,
@@ -207,30 +165,12 @@ pub enum DiscriminatorPlan {
Nested structs share the `ReadPlan` shape (a struct field's `body` is
`CompositePlan::Struct(ReadPlan)`). Union variants are pre-resolved:
each `(key, CompositePlan)` entry carries the variant's compiled body,
so dispatch is a flat lookup + recurse — no `resolve_typeref_as_def` at
read time. A variant may itself be `CompositePlan::Union { ... }`, so
**nested unions** (a union variant that is itself a union, which the
0.2.0 reader supports via `resolve_and_walk_variant`'s `Union` arm) are
covered by ordinary recursion; no separate `VariantKind` enum is
needed. Array element strides are precomputed (`element_stride = 0`
each `(key, VariantPlan)` entry carries the variant's `ReadPlan`, so
dispatch is a flat lookup + recurse — no `resolve_typeref_as_def` at
read time. Array element strides are precomputed (`element_stride = 0`
signals variable-length elements, same convention as today's
`FieldValue::Array`).
`Union.shared` carries the union's declared `fields` (the discriminator
field + any shared fields) for the field-name-discriminator case —
`DiscriminatorPlan::Field.field_index` indexes into `shared`, and the
read loop walks `shared` first, then looks up and walks the selected
variant's `CompositePlan` starting after the shared fields. The
byte-offset-discriminator case has no shared fields (`shared: None`):
the discriminator byte is read at `disc.offset` and the variant starts
immediately after the discriminator size. The POC's `plan_read_union`
`Field` arm stub (Finding 1) is replaced by this `shared` sub-plan;
there is no separate `VariantPlan`/`VariantKind` type in the production
shape — the POC's `VariantPlan { kind, plan }` wrapper is dropped in
favor of recursing on `CompositePlan` directly, which is what makes
nested-union support fall out for free.
### Construction
```rust
@@ -275,16 +215,10 @@ plan being immutable owned data, but the implementation should add a
- `materialize_aligned` is unchanged (already takes `&OffsetMap`, a
compiled form).
- `ReadPlan` is a new public type, re-exported from `lib.rs`.
- `BastDoc` and the `Bast*` types are **unchanged by this ADR** — they
remain the validation-side typed tree, borrowed, as today. This is a
smaller breakage than review #004's Option A (which changed
`BastDoc<'a>` → `BastDoc` and every `Bast*` signature).
**Note (added on ADR-012 acceptance):** ADR-012 §2a subsequently
makes `BastDoc` owned, riding the same 0.3.0 bump. That is an
ADR-012 change, not an ADR-011 change; ADR-011's scope statement
stands as the ADR-011-only view. With ADR-012 §3 (ValidationPlan,
now in 0.3.0), `bast_validation` will also stop being the permanent
home of the `BastDoc` walk — see ADR-012.
- `BastDoc` and the `Bast*` types are **unchanged** — they remain the
validation-side typed tree, borrowed, as today. This is a smaller
breakage than review #004's Option A (which changed `BastDoc<'a>` →
`BastDoc` and every `Bast*` signature).
The crate is pre-1.0 with two in-house downstream consumers
(`alktty`, `alkcall`), both of which will be updated with the bump.
@@ -314,15 +248,12 @@ The crate is pre-1.0 with two in-house downstream consumers
- **`bast_validation`** — the BAST-native value-domain validator walks
`BastDoc` to check constraints (`maxLength`, enum string values,
union variant keys, integer ranges). These are value-domain checks,
not byte-position walks; they don't benefit from a *read* plan and
not byte-position walks; they don't benefit from a read plan and
would require a separate "validation plan" with a different shape.
**Not in scope for ADR-011** — but no longer deferred indefinitely:
ADR-012 §3 brings a `ValidationPlan` into 0.3.0. The "not a hot
loop" framing this paragraph originally relied on was re-evaluated
and rejected (see ADR-012 §3): read+validate on untrusted streams
makes validation hot in the same sense review #004 measured for
the read path. For 0.3.0 as accepted by ADR-011 alone,
`bast_validation` keeps walking `BastDoc`; ADR-012 §3 closes that.
Validation is not a per-chunk hot loop (AGENTS.md: validation is
opt-in per operation). A future `ValidationPlan` is a two-way door
if a bench motivates it; for now `bast_validation` keeps walking
`BastDoc`, which remains the validation-side typed tree.
- **`LayoutBuilder` / `PackedLayout`** — the packed write-side already
has a compiled form (`PackedLayout`). `LayoutBuilder::build`
(`src/layout_builder.rs:190`) re-parses `BastDoc::new` per `build()`
@@ -432,25 +363,19 @@ The crate is pre-1.0 with two in-house downstream consumers
## Scope Boundaries (What This Is Not)
- **Not a `BastDoc` replacement.** `BastDoc` stays as the
validation-side typed tree within ADR-011's scope (borrowed from
`&Value`, unchanged). The `Bast*` types and their signatures are not
touched by ADR-011. Validation (`bast_validation`), aligned one-shot
reads/writes (`engine.rs:334,467`), and `LayoutBuilder::build`
continue to walk `BastDoc` within ADR-011's scope. **ADR-012
subsequently revises two of these:** `BastDoc` becomes owned (§2a)
and `bast_validation` adopts a `ValidationPlan` (§3), both riding
the same 0.3.0 bump. ADR-011's scope statement is the ADR-011-only
view and is not re-litigated here.
validation-side typed tree, borrowed from `&Value`, unchanged. The
`Bast*` types and their signatures are not touched. Validation
(`bast_validation`), aligned one-shot reads/writes
(`engine.rs:334,467`), and `LayoutBuilder::build` continue to walk
`BastDoc` until they get their own compiled forms (additive, later).
- **Not a flat lookup table.** Packed positions are data-dependent;
the plan is a read program (instructions to walk), not a `(path,
offset)` table. This is inherent to packed sequential layout
(ADR-002), not a limitation of this design.
- **Not a validation plan.** `bast_validation`'s value-domain checks
(maxLength, enum values, union variant keys, integer ranges) are a
different concern and a different shape from `ReadPlan`. They are
out of ADR-011's scope; ADR-012 §3 adds a `ValidationPlan` in 0.3.0
rather than leaving validation on an interpretive `BastDoc` walk
indefinitely.
different concern and a different shape. They stay on `BastDoc`. A
future `ValidationPlan` is a two-way door.
- **Not the review's Option A or Option B.** It is the "compiled form"
path the review pointed at but did not name: Option A (make `BastDoc`
own its data) kills the re-parse but leaves the read loop as a
@@ -523,34 +448,23 @@ closes L2. L1 falls out at step 2. The aligned-side M1 paths
`validate_bytes` (packed) consumes the `ReadPlan` via
`materialize_packed`
## Future capabilities (in 0.3.0 via ADR-012)
## Future capabilities (not part of this ADR)
The deterministic-compile property of `ReadPlan` is a prerequisite for
several capabilities. ADR-012 ("Plan Fingerprinting, ValidationPlan,
and Closing the Deferred M1 Sites in 0.3.0") picks up all three items
below into the 0.3.0 release so they ship with this ADR's breaking
changes in one round of downstream churn, not two or three:
several capabilities that are explicitly out of scope here but worth
naming so a future ADR doesn't re-derive the prerequisite:
- **Fingerprinting the plan** (`#[derive(Hash)]` + a `fingerprint()`
method) for cross-run caching of compiled plans, disk-cached plans,
and `alkcall` schema-version handshakes. → **In 0.3.0 (ADR-012 §1).**
- **Closing the deferred M1 sites** via an owned `BastDoc` (lifetime
removal) for `LayoutBuilder` + extending `OffsetMap` with leaf
metadata for the aligned `read_field`/`write_field` paths. → **In
0.3.0 (ADR-012 §2).** Note: ADR-012 reframes the earlier "WritePlan"
candidate listed here as "not a new type — extend the existing
compiled forms (`PackedLayout`/`OffsetMap`) and cache the parse."
- A `ValidationPlan` that follows the same compile-once-walk-many
pattern for `bast_validation`. → **In 0.3.0 (ADR-012 §3).** Different
shape (value-domain, not byte-position) but the same class of
per-buffer re-walk cost on the read+validate-on-untrusted-input
common case. ADR-012 owns the shape decision and the implementation
plan scopes it. (Originally deferred by ADR-012 as "not a hot loop";
review #005 M3 reversed the deferral — see ADR-012 §3.)
- Fingerprinting the plan (e.g. `#[derive(Hash)]`) for cross-run
caching of compiled plans.
- Disk-cached compiled plans (skip `compile` on warm start).
- Schema-version handshakes for `alkcall`'s hub/spoke topology —
peers exchange plan fingerprints instead of full BAST documents.
- A `WritePlan` and/or `ValidationPlan` that follow the same
compile-once-walk-many pattern for the deferred M1 sites
(`engine.rs:334,467`, `layout_builder.rs:190`, `bast_validation`).
None of the in-0.3.0 items justify this ADR; the 400x read-path gap
does. They are listed here as forward references and to record that
the "WritePlan" candidate has been reframed out by ADR-012.
None of these justify this ADR; the 400x read-path gap does. They are
listed only as forward references.
## POC coverage
@@ -565,15 +479,7 @@ two spots where a plan arm could subtly miss a case:
`DiscriminatorPlan::Byte { offset, disc_type }` and
`DiscriminatorPlan::Field { name, field_index }`. The field-name case
pre-resolves the discriminator field's `ReadKind` so dispatch reads
it from the plan, not from a re-parsed `BastField`. The
field-disc union's declared `fields` (discriminator + any shared
fields) are carried as a sub-`ReadPlan` on `CompositePlan::Union`'s
`shared` field (a refinement of the POC shape, which stubbed the
`Field` arm — Finding 1); the production read loop walks `shared`
first, then the selected variant's `CompositePlan`. Nested-union
variants (a variant that is itself a union) are covered by ordinary
`CompositePlan` recursion; the POC rejected them, the 0.2.0 reader
accepts them, and the production shape restores parity.
it from the plan, not from a re-parsed `BastField`.
- **Array variable-element-stride (`element_stride = 0`).** Covered:
`CompositePlan::Array { element, count, element_stride }` preserves
the `0`-signals-variable convention, and the read loop walks
@@ -1,611 +0,0 @@
# ADR-012: Plan Fingerprinting, ValidationPlan, and Closing the Deferred M1 Sites in 0.3.0
## Status
Accepted. Implemented in 0.3.0 — §3's `ValidationPlan` (phase 7,
2026-08-31), §1's fingerprinting (phase 6), §2a's owned `BastDoc`
(phases 3–4), §2b's `LeafMeta` (phase 5); all shipped 2026-09-02.
Bundles three pieces of work into the 0.3.0 release so the
crate ships one round of breaking changes, not two (or three). The
three pieces: (a) fingerprinting `ReadPlan`/`OffsetMap`, (b) closing
the deferred M1 sites via an owned `BastDoc` + `OffsetMap` `LeafMeta`,
and (c) a `ValidationPlan` that retires the interpretive
`bast_validation` walk (added by reversing the original "defer
`ValidationPlan`" decision — see "ValidationPlan — in scope for
0.3.0" below). Companion to [ADR-011](011-compiled-read-plan-for-packed-mode.md)
(the `ReadPlan`) and the [0.3.0 implementation plan](../../plans/030-compiled-forms.md).
§3's concrete shape was scoped by the follow-on design session and is
implemented in `src/validation_plan.rs` — see §3a below.
## Context
ADR-011 accepted the `ReadPlan` as the packed read-side compiled form
and deferred two things to "future capabilities":
1. **Fingerprinting the plan** for cross-run caching, disk-cached
compiled plans, and `alkcall` hub/spoke schema-version handshakes.
2. **A `WritePlan` and/or `ValidationPlan`** following the same
compile-once-walk-many pattern for the deferred M1 sites and the
validation walk.
ADR-011 also explicitly deferred the aligned-side M1 sites
(`engine.rs:334,467` `read_field`/`write_field`;
`layout_builder.rs:190` `LayoutBuilder::build`) as "a deliberate
reversible bet that an aligned-mode hot loop won't emerge."
This ADR retires the deferrals in one release. The reasoning is
timing: 0.3.0 is already a breaking bump (ADR-011 changes
`SequentialReader::new` and `materialize_packed` signatures), and the
crate has no real downstream consumers yet (only `alktty`/`alkcall`,
both in-house). Doing all three pieces now costs one round of
downstream churn instead of two or three, and the fingerprinting work
cuts across both `ReadPlan` and `OffsetMap` — splitting would create
a cross-release dependency that's cleaner in one release. The
`ValidationPlan` inclusion follows the same logic applied to the
validation walk: deferring it would create a *second* breaking change
to `validate_bytes`/`bast_validation` after 0.3.0, which is exactly
the round of downstream churn this release exists to retire.
### Reframing "WritePlan"
ADR-011's "Future capabilities" section listed a `WritePlan` as a
candidate. On inspection, a new public `WritePlan` type is the wrong
shape for the deferred M1 sites, for two reasons:
1. **The packed write-side already has a compiled form: `PackedLayout`.**
`LayoutBuilder::build`'s M1 re-parse is the *builder* re-parsing
`BastDoc::new` on each `build()` call to get the typed tree it
walks. The fix is to cache the parsed tree on the builder at `new()`
time — internal, non-breaking, no new public type. The compiled
form (`PackedLayout`) is unchanged; only its construction stops
re-parsing.
2. **The aligned R/W side already has a compiled form: `OffsetMap`.**
`read_field`/`write_field`'s M1 re-parse is `lookup_leaf_field`
walking `BastDoc` to get leaf metadata (`kind`, `encoding`,
`endian`) that `OffsetMap` doesn't carry. The fix is to extend
`OffsetMap`'s entries with that metadata at `compute` time —
additive fields on an existing public type (breaking, but we're
bumping anyway). No new public type.
A new `WritePlan` type would overlap with `PackedLayout` (packed
write) and `OffsetMap` (aligned R/W) without a clean distinguishing
shape. The honest picture: the packed write-side compiled form is
`PackedLayout`; the aligned R/W compiled form is `OffsetMap`; the M1
fixes are "cache the parse" and "extend the compiled form with leaf
metadata," not "add a third compiled form." This serves the
minimal-public-API-changes goal better than a literal `WritePlan`.
### `ValidationPlan` — in scope for 0.3.0 (no longer deferred)
The BAST-native validator (`bast_validation`) walks `BastDoc` to
check value-domain constraints (enum value sets, integer ranges,
`maxLength` caps, union variant keys). This is a different shape
from `ReadPlan`/`OffsetMap` (value-domain, not byte-position), and
the earlier framing deferred it as "not a hot loop — validation is
opt-in per operation per AGENTS.md."
**That deferral is reversed.** The "not a hot loop" dismissal
under-counted the common case: **read + validate together on
untrusted input.** The downstream `alkcall` consumer accepts schemas
from arbitrary internet peers in a hub/spoke topology (AGENTS.md §3);
the common operation on an incoming frame is "read it, then validate
it before acting." `validate_bytes` (ADR-010) is therefore called
once per incoming buffer, and each call re-walks `BastDoc` for
validation even after ADR-011 makes the *read* half plan-fast. That
is the same class of per-buffer interpretive cost review #004 measured
for the read path (400x per chunk), on a different code path, on the
operation the untrusted-input discipline actually requires.
The cost-of-inaction framing that the original deferral relied on was
also wrong: a `ValidationPlan` introduced *after* 0.3.0 would be a
breaking change to `validate_bytes`'s contract and to the
`bast_validation` public surface, forcing rework of `alktty`/`alkcall`
— the exact downstream-churn this release is supposed to retire, not
create a second round of. Shipping it in 0.3.0 pays the cost once,
alongside the other breaking changes, while there are zero real
consumers. The cost of action now is a static, known quantity; the
cost of action later is the same work plus a second round of
downstream churn plus the risk of the interpretive path being the
one that gets used in the meantime on untrusted bytes.
**Decision: a `ValidationPlan` ships in 0.3.0 as §3 below.** The
shape is a compile-once-walk-many compiled form over the BAST
document's value-domain constraints, symmetric to `ReadPlan` (packed
read-side) and `OffsetMap` (aligned R/W). The concrete shape,
construction, and `validate_bytes` integration are scoped in the
0.3.0 implementation plan (a dedicated phase) and detailed in a
follow-on design session before implementation; this ADR commits the
*decision* (in 0.3.0, not deferred) and the *scope* (a compiled
validation form that retires the interpretive `BastDoc` walk in
`bast_validation`), so the deferral black hole is closed.
## Decision
### 1. Fingerprinting — `ReadPlan: Hash + Eq`, `OffsetMap: Hash + Eq`
Add `#[derive(Hash, Eq)]` (alongside the existing `Debug, Clone, PartialEq`)
to `ReadPlan` and `OffsetMap`, plus their public sub-types
(`FieldPlan`, `CompositePlan`, `ReadKind`, `DiscriminatorPlan`,
`ByteRange`, and the new `LeafMeta` — see §2). `VariantPlan`/
`VariantKind` are not in the production `ReadPlan` shape (ADR-011
was refined on acceptance to drop them — see ADR-011 status), so
they are not derived. `ValidationPlan` (§3) gets `Hash + Eq` + its
own `fingerprint()` as part of its public surface.
**`by_name` representation change.** `ReadPlan.by_name` is currently
`HashMap<String, usize>`. `HashMap` iteration order is non-deterministic
and `HashMap` does not implement `Hash`, which blocks `#[derive(Hash)]`
on `ReadPlan`. Switch `by_name` to `BTreeMap<String, usize>`. Lookup
cost at protocol-header N (~5 fields) is negligible (the `BTreeMap` is
only used by `read_field`'s name→index lookup, not by the sequential
`read_next` hot path). This makes the derived `Hash` cover the full
structural state of the plan.
**Fingerprint contract.** Two plans with equal `Hash` (or equal under
`PartialEq`) produce identical reads over identical bytes. Formally:
`plan1 == plan2 ⟹ ∀ buffer. read(plan1, buffer) == read(plan2, buffer)`.
This is the contract the downstream uses rely on:
- **Cross-run disk cache.** A consumer can hash a `ReadPlan`/
`OffsetMap` and cache the compiled plan keyed by the hash, skipping
`compile` on warm starts. Safe because the contract guarantees a
cache hit produces identical read behavior.
- **`alkcall` hub/spoke schema handshake.** Peers exchange plan
fingerprints instead of full BAST documents. A peer that receives a
fingerprint it has already compiled can skip re-transmitting the
schema. The contract guarantees fingerprint equality implies
behavioral equivalence, so the handshake is sound.
- **Schema-version diagnostics.** A consumer can log a plan
fingerprint alongside read results for reproducibility — two runs
over "the same schema" that produce different fingerprints reveal a
silent schema drift.
The contract is a *behavioral* equivalence, not a structural identity:
two plans with different `by_name` insertion order but the same
`fields` Vec produce the same reads, and after the `BTreeMap` change
they also produce the same `Hash`. The contract is documented on the
`Hash` impl and tested by a property-style test (compile the same
schema twice, assert `plan1 == plan2` and `plan1.hash() ==
plan2.hash()`).
**Fingerprint API.** No new public method is strictly needed —
consumers call `std::hash::Hash` directly. For ergonomics and to make
the contract visible, add a convenience method:
```rust
impl ReadPlan {
/// A stable 64-bit fingerprint of this plan's read behavior.
///
/// Two plans with the same fingerprint produce identical reads
/// over identical bytes (the fingerprint contract).
pub fn fingerprint(&self) -> u64;
}
impl OffsetMap {
/// A stable 64-bit fingerprint of this offset map's read/write
/// behavior. Same contract as `ReadPlan::fingerprint`.
pub fn fingerprint(&self) -> u64;
}
```
Implemented via `std::hash::DefaultHasher` (or a stable hasher like
`FxHasher` if we want cross-version stability — decision belongs to
the implementation step, called out in the plan). The fingerprint is
additive API, not breaking.
### 2. Closing the deferred M1 sites
#### 2a. `LayoutBuilder` — cache the parsed `BastDoc` at `new()`
`LayoutBuilder` currently stores `doc_value: Value` + `root_name: String`
and re-parses `BastDoc::new(&self.doc_value, &self.root_name)` on every
`build()` call (`layout_builder.rs:190`). The fix: store the parsed
typed tree at `new()` time and reuse it in `build()`.
This requires `BastDoc` to be owned (no lifetime borrowing from
`doc_value`). Two options:
- **Option α (smaller):** keep `BastDoc<'a>` borrowing, store
`doc_value: Value` + a *pre-resolved, owned* representation of just
what `build` needs (the field tree with `$ref`s resolved). This is
essentially a `WritePlan` by another name — rejected per the
reframing above.
- **Option β (cleaner):** make `BastDoc` own its data. This is
review #004's Option A, scoped to `LayoutBuilder` only. It's a
larger refactor but eliminates the lifetime entanglement for the
builder and is the prerequisite for any future owning consumer that
wants to cache the parsed tree.
**Decision: Option β, scoped to `LayoutBuilder`.** The `BastDoc<'a>` →
`BastDoc` (owned) refactor is the principled fix and is already
breaking (the `Bast*` types are re-exported from `lib.rs`), so it
rides the 0.3.0 bump. This does *not* change `SequentialReader` or
`materialize_packed` (those consume `ReadPlan` per ADR-011, not
`BastDoc`). It changes `LayoutBuilder::new` to parse once and `build`
to reuse. The `doc_value: Value` field is removed; the builder holds
the owned `BastDoc` directly.
**Note on `BastDoc` ownership scope:** ADR-011 left `BastDoc` borrowed
and unchanged ("the validation-side typed tree"). This ADR changes
that: `BastDoc` becomes owned. The validation-side (`bast_validation`)
and aligned-side (`OffsetMap::compute`, `materialize_aligned`)
consumers adapt to the owned `BastDoc` — they no longer need a
borrowed `&Value` kept alive alongside. This is a net simplification:
one typed-tree type, owned, used by all non-`ReadPlan` consumers. The
POC on `readplan-poc` confirmed `ReadPlan` doesn't need `BastDoc` to
be borrowed (it compiles from `&Value` once and discards the
`BastDoc`), so making `BastDoc` owned doesn't regress the read path.
#### 2b. `OffsetMap` — carry leaf metadata
`OffsetMap` currently stores `Vec<(String, ByteRange)>`. The
`read_field`/`write_field` M1 re-parse is `lookup_leaf_field` walking
`BastDoc` to get `LeafFieldInfo { kind, encoding, endian }`
(`engine.rs:556-600`). The fix: extend `OffsetMap`'s entries to carry
that metadata at `compute` time.
```rust
#[derive(Debug, Clone, Copy, PartialEq, Eq, Hash)]
pub struct LeafMeta {
pub kind: AlkTypeKind,
pub encoding: VariableEncoding,
pub endian: Endian,
}
pub struct OffsetMap {
fields: Vec<(String, ByteRange, LeafMeta)>, // was Vec<(String, ByteRange)>
total_size: usize,
}
```
`OffsetMap::compute` resolves each leaf field's `LeafMeta` during the
walk (it already walks the tree; it just doesn't currently record the
metadata). `read_field`/`write_field` drop the `BastDoc::new` +
`lookup_leaf_field` calls and read `LeafMeta` from the map. The
`LeafFieldInfo` struct in `engine.rs` is removed (replaced by
`OffsetMap`'s `LeafMeta`).
**Breaking changes:**
- `OffsetMap::get` return type: `Option<&ByteRange>` →
`Option<(&ByteRange, &LeafMeta)>` (or a small accessor struct).
Call sites in `alktty`/`alkcall` update with the bump.
- `ByteRange` is unchanged (still `Copy + Hash`).
- `LeafMeta` is a new public type, re-exported from `lib.rs`.
This is additive on the *capability* (the map now answers questions it
previously couldn't) but breaking on the *signature* (`get`'s return
type changes). Rides the 0.3.0 bump.
### 3. `ValidationPlan` — compile-once validation form
`bast_validation` currently walks `BastDoc` interpretively on every
`validate_bytes` call to check value-domain constraints (enum value
sets, integer ranges, `maxLength` caps, union variant keys). After
ADR-011, the *read* half of `validate_bytes` (packed) is plan-fast;
the *validation* half is still an interpretive `BastDoc` walk per
buffer. On the `alkcall` hub/spoke topology, `validate_bytes` is the
gate between "bytes arrived from an untrusted peer" and "act on the
decoded frame," so it runs once per incoming buffer and validation is
hot in the same sense review #004 measured for the read path.
**Decision: a `ValidationPlan` is a compiled form over the BAST
document's value-domain constraints, built once at `compile` time
(symmetric to `ReadPlan`/`OffsetMap`) and walked by
`bast_validation`/`validate_bytes` without re-touching `BastDoc`.**
The shape, construction, and `validate_bytes` integration are scoped
in the 0.3.0 implementation plan as a dedicated phase and detailed in
a follow-on design session before implementation begins. The
properties this ADR commits to (so the plan and any implementing agent
have a fixed contract):
- **Compile-once-walk-many.** `ValidationPlan::compile` walks `BastDoc`
once; `validate_bytes` (both modes) walks the `ValidationPlan` per
buffer, never `BastDoc`. This is the same pattern as `ReadPlan` and
`OffsetMap`; it is the structural reason the per-buffer
interpretive cost goes away.
- **Value-domain, not byte-position.** The plan carries constraint
descriptors (enum allowed-sets, integer range bounds, `maxLength`
caps, union variant keys, and any other value-domain checks
`bast_validation` performs today), keyed for dispatch against the
materialized `Value` tree, not byte offsets. The shape is therefore
different from `ReadPlan`/`OffsetMap`; the *pattern* (compiled form,
immutable, shared via `Arc`) is the same.
- **No new `BastDoc` walk in the hot path.** After this ADR, the only
consumers that walk `BastDoc` interpretively are the one-shot
`compile` paths (`ReadPlan::compile`, `OffsetMap::compute`,
`ValidationPlan::compile`, `LayoutBuilder::new`). The per-buffer
paths (`sequential_reader`, `materialize_packed`,
`materialize_aligned`, `validate_bytes`) all walk compiled forms.
This is the end state ADR-011 pointed at; this ADR closes it.
- **Semver.** `ValidationPlan` is a new public type, re-exported from
`lib.rs`. `validate_bytes`'s *signature* is unchanged (still
`(buffer) -> Result<(), AlkTypeError>`); the change is internal
(walks the plan instead of `BastDoc`). If the `ValidationPlan`
design surfaces a need to change `validate_bytes`'s signature, that
rides the 0.3.0 bump and is recorded in the plan's Semver Contract
table when the shape is scoped. `bast_validation`'s public surface
(`build_validator`, `validate_value`) is reviewed at shape-scope
time; additive changes ride the bump, removals/renames are avoided
unless the shape work shows they're necessary.
- **Fingerprinting.** `ValidationPlan` is `Hash + Eq` with a
`fingerprint()` method, same as `ReadPlan`/`OffsetMap` (§1/§4), so
the downstream uses (cross-run cache, `alkcall` handshake,
schema-version diagnostics) extend to the validation form without
new API. The fingerprint contract generalizes: two validation plans
with equal hashes accept/reject identical `(bytes)` identically.
**What this ADR does *not* decide** (left to the follow-on shape
session + plan phase): the concrete `ValidationPlan` struct/enum
shape, how `maxLength`/range/enum/union-key constraints are
represented, whether `bast_validation`'s `validate_value` is retired
or kept as a convenience wrapper over the plan, and whether the
`AlkTypeKind`-driven dispatch in `bast_validation` collapses into the
plan or stays a thin match over plan-carried descriptors. These are
shape questions, not decision questions; the decision (in 0.3.0,
compiled form, no per-buffer `BastDoc` walk) is fixed here.
### 3a. `ValidationPlan` shape — resolved by the design session
The follow-on design session (0.3.0 phase 7 predecessor) resolved the
open shape questions; implemented in `src/validation_plan.rs`:
- **Shape.** `ValidationPlan { root: ValidNode }`, a compiled
constraint tree — one `ValidNode` arm per value-domain check,
mirroring the interpretive walker's arms one-to-one:
`Int { min, max }` / `I64` / `Uint { max }` / `U64` / `Float` / `Bool`
/ `Str { max_len }` / `Bytes { max_len }` / `Enum { count }` /
`Struct { fields: Vec<ValidField> }` / `Union { variants:
Vec<ValidVariant> }` / `Array { count, element }` / `Record
{ values }`. `ValidField` carries `name + node`; `ValidVariant`
carries `key + node`. The nodes are public (diagnostics access via
`ValidationPlan::root()`); construction is only possible through
`compile`. Note the union node carries *only the variant nodes* — the
declared union `fields` (shared fields) are validated as part of the
variant walk, because the walker dispatches on the materialized
`__discriminator` and validates the whole object against the selected
variant (the `ValidNode::Union` doc comment records this; the
interpretive union arm recursed into the variant the same way).
- **Constraint representation.** Inline scalar fields on the node arms
(ranges as `i64`/`u64` pairs, `maxLength` as `Option<usize>`, enum
bound as `count: u64`, union keys as owned `String`s). `maxLength` is
resolved from the *owning field* at compile time and baked into the
`Str`/`Bytes` leaf — the walk never consults field annotations. It
never crosses a `$ref` (a `$ref` always targets a struct/union/enum
`$defs` entry, so the interpretive walk could never consult it
through one either).
- **`compile` signature.** `ValidationPlan::compile(&BastDoc) ->
Result<Self, AlkTypeError>` — the plan-table's `&str` root-name
parameter was vestigial (the doc already holds its root).
- **`validate_value` disposition.** Retained as a one-shot wrapper:
`compile(doc)` + `validate(value)`. The interpretive walker behind it
is *retired* (deleted) — the wrapper delegates to the plan, so there
is one constraint implementation, not two. `bast_validation.rs` keeps
the shared error helper (`validation_err`) and the
`__discriminator` key constant.
- **Error contract preserved.** The plan walk reproduces the
interpretive error messages byte-identically: a segment stack
(`field` / `[index]` / `[key]`) renders paths only on failure — zero
per-node allocation on the happy path. Numeric `__discriminator`
dispatch matches mapping keys without allocation for the u64/i64
forms (mapping keys are stringified integers; non-integer numbers
fall back to `Number::to_string`).
- **Compile-time rejection of adversarial graphs.** Eager `$ref`
resolution with a definition-level cycle set and a depth cap (128):
a cyclic or self-referential schema is `AlkTypeError::Schema` at
compile, not a stack overflow — the interpretive walker resolved
`$ref`s lazily with no guard and could overflow on recursion.
Diamond (shared, non-cyclic) refs compile fine; the cycle set is
path-scoped.
- **Engine integration.** `AlkTypeEngine` holds `Arc<ValidationPlan>`
built at `compile` time in *both* modes; the accessor
`validation_plan() -> &Arc<ValidationPlan>` is new public API.
`validate_bytes`'s signature is unchanged. The plan compile runs
*before* the layout build: it is the engine's reference-graph gate
(see Consequences).
- **Scope note.** phase-7's `Send + Sync` / `Hash + Eq` /
`fingerprint()` requirements are structural on the types above
(`#[derive(...)]` on plain owned data; the same `DefaultHasher`
fingerprint as phase 6).
### 4. Fingerprinting `OffsetMap` (bundled with §2b)
Since `OffsetMap` is getting new fields (`LeafMeta`) in §2b, its
`#[derive(Hash, Eq)]` (from §1) covers the new fields automatically.
The fingerprint contract for `OffsetMap` is the aligned-side analog
of `ReadPlan`'s: two offset maps with equal hashes produce identical
aligned reads/writes over identical bytes.
## Scope
### In scope
- `ReadPlan: Hash + Eq` + `fingerprint()` method (§1).
- `OffsetMap: Hash + Eq` + `fingerprint()` method (§1, §4).
- `BastDoc<'a>` → `BastDoc` (owned) refactor, scoped to the consumers
that currently hold `doc_value: Value` and re-parse: `LayoutBuilder`,
`bast_validation`, `materialize_aligned`, `OffsetMap::compute`
(§2a). `ReadPlan::compile` and the packed read path are unaffected
(they consume `&Value` once and discard `BastDoc`).
- `LayoutBuilder` caches the owned `BastDoc` at `new()`, `build()`
reuses it — no re-parse (§2a).
- `OffsetMap` carries `LeafMeta`; `read_field`/`write_field` drop
`BastDoc::new` + `lookup_leaf_field` (§2b).
- `LeafMeta` new public type (§2b).
- `BTreeMap` for `ReadPlan.by_name` (§1).
- `ValidationPlan` new public type + `compile` + `Hash + Eq` +
`fingerprint()` (§3). `validate_bytes` (both modes) walks the
`ValidationPlan` instead of re-walking `BastDoc`. `bast_validation`
adopts the plan; the public `validate_value`/`build_validator`
surface is reviewed at shape-scope time and rides the bump only if
the shape work shows a signature change is necessary.
- `ValidationPlan: Hash + Eq` + `fingerprint()` method (§3, §1) — the
fingerprint contract extends to the validation form.
### Out of scope
- Disk-cache or handshake *implementations* — the fingerprint
*contract* and method are in scope (§1, §3); the downstream uses
(cache format, wire protocol) are the consumers' problem, not this
ADR's.
- The `ValidationPlan` shape — **resolved** (§3a). Decided by the
design session and implemented in `src/validation_plan.rs`; the
decision (in 0.3.0, compiled form, no per-buffer `BastDoc` walk) was
fixed here.
- Cycle-guard hardening for the *layout* walkers' own recursion
(`LayoutBuilder`/`OffsetMap` struct recursion) beyond the engine-path
gate described in Consequences — if a non-engine entry point walking
those types on untrusted docs becomes a consumer pattern, the
guards get their own change (the `AlkTypeEngine::compile` gate
covers the supported path today).
- Cross-version fingerprint stability — the fingerprint is stable
within a crate version but may change across versions (a new
`AlkTypeKind` variant, for example, changes the hash). Cross-version
stability is a non-goal; consumers cache within a version. The
implementation step chooses a hasher and documents the stability
contract.
## Consequences
### Positive
- **One breaking release, not two (or three).** ADR-011's `ReadPlan`
+ this ADR's `BastDoc`-owned + `OffsetMap` extension +
`ValidationPlan` all ship together. The two in-house downstream
consumers (`alktty`, `alkcall`) update once.
- **Closes all deferred M1 sites.** `LayoutBuilder::build`
(`layout_builder.rs:190`), `read_field` (`engine.rs:334`),
`write_field` (`engine.rs:467`) all stop re-parsing. The packed-side
`validate_bytes` (`engine.rs:284`) was already closed by ADR-011;
this ADR closes the aligned-side equivalent.
- **Retires the interpretive validation walk.** `validate_bytes` on
untrusted streams (the `alkcall` common case) stops re-walking
`BastDoc` per buffer. This is the latent perf cliff review #005 M3
flagged: the read half was plan-fast after ADR-011, the validation
half was not. Closing it here — while there are zero real consumers
and one breaking bump already paying the downstream-churn cost —
avoids a second breaking change to `validate_bytes`/`bast_validation`
after 0.3.0. (Implemented: a spot benchmark of plan-validate on a
4-field mixed frame puts the validation half at ~0.2 µs/validate;
the compile-per-call one-shot it replaces runs ~2.7x slower before
the walk is even counted — and the full 0.2.0 per-buffer cost
included lazy `$ref` deep-clones that the one-shot no longer pays.
The materialize half, not validation, remains the dominant
`validate_bytes` cost.)
- **Compile-time rejection of cyclic `$ref` graphs.** A side effect of
eager plan compilation: a self-referential document is now a clean
`Schema` error instead of a stack overflow. The plan compile runs
*before* the layout build in `AlkTypeEngine::compile`, making it the
engine's reference-graph gate — `LayoutBuilder`/`OffsetMap`
struct-recursion had no cycle guard and previously could recurse
unboundedly on such a document (a pre-existing untrusted-schema
hazard, surfaced by the phase-7 `compile_rejects_cyclic_ref_graph`
test). (Resolved since: review #006 H2 added the shared
`walk_guard::check_ref_graph` guard at every standalone walker entry,
so the trust boundary no longer depends on the engine path.)
- **Fingerprinting enables downstream uses.** Cross-run plan caching,
`alkcall` schema handshake, and schema-version diagnostics all
become possible without further API work — across `ReadPlan`,
`OffsetMap`, and `ValidationPlan`.
- **`BastDoc` owned is a net simplification.** One typed-tree type,
owned, used by all `*::compile` paths. No more
lifetime-entanglement workarounds. The "re-parse on demand" framing
from ADR-007 is fully retired across read, write, and validation
paths.
- **`OffsetMap` extension is additive capability.** The map now
answers `kind`/`encoding`/`endian` questions it previously couldn't,
enabling future aligned-side tools without re-walking `BastDoc`.
### Negative
- **Breaking public-API changes (0.2.0 → 0.3.0).** `BastDoc<'a>` →
`BastDoc` (owned) changes every `Bast*` signature that took `&'a`.
`OffsetMap::get` return type changes. `LeafMeta` is new public.
`ReadPlan` is new public (from ADR-011). `ValidationPlan` (+ the
`ValidNode`/`ValidField`/`ValidVariant` node types) is new public
(§3a). All ride the bump.
- **`BastDoc` ownership refactor is broad.** Touches `bast.rs` (every
typed node: `&'a str` → `String`/`Arc<str>`, `&'a Value` →
`Value`/`Arc<Value>`) and every consumer (`layout_builder`,
`offset_map`, `materialize`, `bast_validation`, `engine`). This is
review #004's Option A, which ADR-011 deferred — this ADR picks it
up because the `LayoutBuilder` M1 fix requires it and we're bumping
anyway. The refactor is mechanical (lifetime removal, not logic
rewrites); the POC on `readplan-poc` confirmed the read path is
unaffected.
- **Interpretive `validate_value` is compile-per-call.** The retained
one-shot wrapper (`bast_validation::validate_value`) compiles a plan
then validates — fine for one-off/diagnostic use, wrong for per-
buffer use. Per-buffer callers must hold the engine (or a plan) —
the doc comments say so. The walker it replaced had the inverse
trade (no compile, but interpretive per call); the engine path
(compile once) is the one that matters.
- **Aligned-mode `maxLength`-reserved strings/bytes.** Materialization
emits the *full reserved* (zero-padded) data for these fields. A
`ValidationPlan` compiled from a document used in packed mode would
apply `maxLength` to trimmed length, matching packed semantics; the
aligned materializer's zero-padding means the value passed to
validation can carry trailing NULs. This is pre-existing
materialize behavior (not a plan artifact); consumers relying on
trimmed values already see it.
- **Fingerprint cross-version stability is not guaranteed.** A future
`AlkTypeKind` variant changes the hash. Documented as a within-
version contract. Consumers that need cross-version stability
serialize the BAST document and re-compile.
- **`BTreeMap` for `by_name` is a tiny lookup cost.** Negligible at
protocol-header N; irrelevant to the 400x fix.
## Scope Boundaries (What This Is Not)
- **Not a `WritePlan` type.** The packed write-side compiled form is
`PackedLayout`; the aligned R/W compiled form is `OffsetMap`. The
M1 fixes are "cache the parse" (§2a) and "extend the compiled form
with leaf metadata" (§2b), not "add a third compiled form."
- **Not a `ValidationPlan` deferral.** `ValidationPlan` is in scope
(§3) and implemented (§3a); the interpretive walker is retired.
- **Not cross-version fingerprint stability.** Within-version only.
- **Not a disk-cache or wire-protocol spec.** The fingerprint contract
and method are in scope; the downstream uses are the consumers'
concern.
## Recommended Order
See [the 0.3.0 implementation plan](../../plans/030-compiled-forms.md)
for the step-by-step execution order. The high-level grouping:
1. **`ReadPlan` (ADR-011 steps 1–5)** — the packed read-path fix. Closes
H1 + packed-side M1 + L1 + L2.
2. **`BastDoc` owned (§2a)** — the typed-tree ownership refactor. Prerequisite
for the `LayoutBuilder` M1 fix and for `ValidationPlan::compile`.
3. **`LayoutBuilder` M1 fix (§2a)** — cache the owned `BastDoc` at `new()`.
4. **`OffsetMap` extension (§2b)** — carry `LeafMeta`; close the
aligned-side `read_field`/`write_field` M1.
5. **`ValidationPlan` (§3)** — compiled validation form; `validate_bytes`
walks the plan instead of `BastDoc`. Requires the owned `BastDoc`
from step 2 for `ValidationPlan::compile`.
6. **Fingerprinting (§1, §4)** — `Hash + Eq` + `fingerprint()` on
`ReadPlan`, `OffsetMap`, and `ValidationPlan`. Rides on top of the
above.
7. **Public API bump (0.2.0 → 0.3.0)** — `lib.rs` re-exports, version
bump, update `alktty`/`alkcall`.
8. **Verification block** — full suite + wasm + bench.
## References
- [ADR-011](011-compiled-read-plan-for-packed-mode.md) — the
`ReadPlan` (packed read-side compiled form). This ADR extends the
0.3.0 release with fingerprinting, the deferred M1 fixes, and the
`ValidationPlan`.
- [Review #004](../../reviews/004-performance-review.md) — the
performance finding (H1, M1, L1, L2). ADR-011 closed H1 + packed
M1 + L1 + L2; this ADR closes the aligned-side M1.
- [Review #005](../../reviews/005-plan-review-030.md) — the 0.3.0 plan
review whose M3 finding reversed the `ValidationPlan` deferral.
- [ADR-007](007-packed-mode-read-factory.md) — the "re-parse on
demand" framing, retired across read, write, and validation paths
by ADR-011 + this ADR.
- [ADR-002](002-two-layout-modes-packed-vs-aligned.md) — the two
layout modes; `OffsetMap` is the aligned R/W compiled form extended
here with `LeafMeta`.
- [0.3.0 implementation plan](../../plans/030-compiled-forms.md) —
the step-by-step execution plan.
+8 -42
View File
@@ -25,7 +25,7 @@ protocols.
**Components:**
- **`LayoutBuilder`** — constructed via `LayoutBuilder::new(bast_doc, root_name)` (requires a `struct` at the root), then `builder.build(&var_sizes) -> Result<PackedLayout, AlkTypeError>` where `var_sizes: &HashMap<String, usize>` maps variable-length field paths (and TUnion discriminator/variant keys) to their actual byte sizes. Used at write time when the consumer knows the data sizes upfront. The builder computes positions only; the consumer writes data via the [`data_access`](data-access.md) functions at the computed positions.
- **`SequentialReader`** — constructed via `engine.sequential_reader()` (shares the engine's compiled `ReadPlan` via `Arc` — see [ADR-011](decisions/011-compiled-read-plan-for-packed-mode.md)), then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
- **`SequentialReader`** — constructed via `SequentialReader::new(bast_doc, root_name)`, then driven by `reader.read_next(&buffer) -> Result<Option<(String, FieldValue)>, AlkTypeError>` until `Ok(None)`, or `reader.read_field(&buffer, path)` to seek a single field (which walks all preceding fields to reach the target). `reader.reset()` rewinds to the start. Used at read time when the consumer is parsing an incoming frame.
**How it works:**
@@ -69,7 +69,7 @@ and safetensors.
**Component:**
- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result<Self, AlkTypeError>` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment, and resolves each leaf's `LeafMeta` (kind, encoding, effective endianness — ADR-012 §2b). The output is a flat table of `(field_path, OffsetEntry)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
- **`OffsetMap`** — constructed via `OffsetMap::compute(&doc) -> Result<Self, AlkTypeError>` (requires a `struct` at the root). Walks the BAST typed tree once, computes fixed byte positions for each field based on type sizes and alignment. The output is a flat table of `(field_path, byte_range)` pairs (see [Public Types](#public-types)). Used for both read and write at known offsets.
**How it works:**
@@ -113,11 +113,7 @@ with a `AlkTypeError::Offset` — the `OffsetMap` reserves only 4 bytes
(the length prefix), but `data_access::write_string` writes prefix +
data inline, which would clobber subsequent fields. Non-final variable
fields must use `maxLength` (fixed-size reservation) or
`"encoding": "offset-indirect"` — except `record` fields, for which
neither remedy is available (`maxLength` is rejected at parse — review
#006 N3 — and `offset-indirect` is rejected for records in aligned
mode — review #006 M5), so a non-final record field cannot be repaired
and must move to the last position. See
`"encoding": "offset-indirect"`. See
[ADR-006](decisions/006-reject-non-final-inline-length-prefixed-in-aligned-mode.md).
## Offset Computation Algorithm
@@ -198,12 +194,6 @@ annotation shapes).
only. The engine uses strategy 1 (inline length-prefixing) because
protocols don't benefit from fixed-size reservation.
`maxLength` applies to `string` and `bytes` fields only. The parser
rejects it on any other kind (review #006 N3): the validation plan
bakes it into string/bytes leaves only, so on a record (or any other
kind) the annotation did nothing — and in aligned mode a record
reservation was silently corrupt (review #006 M5).
**Strategy 3: Offset indirection (`"encoding": "offset-indirect"`).**
1. The field is a struct `{offset: u32, length: u32}`.
2. The `OffsetMap` records the position of this struct.
@@ -322,46 +312,22 @@ discriminators, the discriminator is recorded under the synthetic path
(schema `properties` order, with nested struct fields appearing inline
under their parent's path prefix).
### `LeafMeta` / `OffsetEntry` (aligned mode, 0.3.0)
```rust
pub struct LeafMeta {
pub kind: AlkTypeKind,
pub encoding: VariableEncoding,
pub endian: Endian,
}
pub struct OffsetEntry {
pub range: ByteRange,
pub meta: LeafMeta,
}
```
`OffsetMap::compute` resolves each leaf's read/write metadata (kind,
variable-length encoding, effective endianness — field override else
container default, propagated the aligned-materializer way) alongside
its byte range, so `read_field`/`write_field` dispatch on the entry
without re-walking the BAST tree per access (ADR-012 §2b).
### `OffsetMap` (aligned mode)
A flat table of `(field_path, OffsetEntry)` pairs computed from a schema.
A flat table of `(field_path, byte_range)` pairs computed from a schema.
```rust
impl OffsetMap {
pub fn compute(doc: &BastDoc) -> Result<Self, AlkTypeError>;
pub fn get(&self, field_path: &str) -> Option<&OffsetEntry>;
pub fn compute<'a>(doc: &'a BastDoc<'a>) -> Result<Self, AlkTypeError>;
pub fn get(&self, field_path: &str) -> Option<&ByteRange>;
pub fn total_size(&self) -> usize;
pub fn iter(&self) -> impl Iterator<Item = (&str, &OffsetEntry)>;
pub fn fingerprint(&self) -> u64;
pub fn iter(&self) -> impl Iterator<Item = &(String, ByteRange)>;
}
```
`compute` requires a `struct` at the root. `total_size`
includes trailing alignment padding. `iter` yields fields in the BAST
`fields` array order (nested struct fields appearing inline). The map
carries `Hash + Eq` (ADR-012 §1); `fingerprint()` is the
stable-within-version hash for caching and schema handshakes.
`fields` array order (nested struct fields appearing inline).
## Design Decisions
+1 -4
View File
@@ -230,10 +230,7 @@ type-level properties. The concrete BAST shapes are in
The `maxLength` keyword is *not* a BAST invention — it is the standard
JSON Schema `maxLength`, repurposed as a byte-length cap. In aligned
mode it reserves a fixed-size slot; in packed mode it is a validation
constraint only. It applies to `string`/`bytes` fields only: the parser
rejects it on any other kind (review #006 N3 — elsewhere it was
silently unenforced), and in aligned mode a record reservation was
silently corrupt (review #006 M5). See [bast-format.md §Variable-Length
constraint only. See [bast-format.md §Variable-Length
Encoding](bast-format.md#variable-length-encoding) and
[ADR-003](decisions/003-schema-annotations.md).
+36 -69
View File
@@ -1,6 +1,6 @@
---
status: accepted
last_updated: 2026-08-31
last_updated: 2026-08-15
---
# alktype — Validation
@@ -22,7 +22,7 @@ and recorded in [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md).
| Path | Input | Validator | Schema source |
|------|-------|-----------|---------------|
| `validate_bytes(&[u8])` | Raw bytes | Compiled `ValidationPlan` walk | The BAST document (binary layout + value constraints) |
| `validate_bytes(&[u8])` | Raw bytes | BAST-native validator (`bast_validation`) | The BAST document (binary layout + value constraints) |
| `validate_json(&Value)` | Parsed JSON `Value` | Standard `jsonschema::Validator` | A consumer-provided standard JSON Schema |
### `validate_bytes` — bytes in, BAST is the validator
@@ -35,39 +35,25 @@ produces `Value::Number`), bounds are checked (via
`data_access::check_bounds`), UTF-8 is valid (via `from_utf8`), the
discriminator is in the mapping, and the boolean byte is 0 or 1.
What the materializer does NOT check — and what the validation half
checks afterward — are **value-domain constraints expressed in the BAST
document**. Since ADR-012 §3 (0.3.0), those constraints are not walked
interpretively per buffer: they are **compiled once** into a
`ValidationPlan` ([`src/validation_plan.rs`](../../src/validation_plan.rs))
at `AlkTypeEngine::compile` time, and each `validate_bytes` call walks
the compiled constraint tree against the materialized `Value` — no
`$ref` re-resolution, no schema re-parse, no per-node path formatting
(error paths render only on failure). The plan's constraint nodes
implement exactly the table below (the constraint set is unchanged from
the retired interpretive walker):
What the materializer does NOT check — and what the BAST-native
validator checks afterward — are **value-domain constraints expressed
in the BAST document**. The BAST-native validator
(`src/bast_validation.rs`) is a recursive walker over the BAST typed
tree ([`crate::bast::BastDoc`]/[`BastType`]) that checks exactly these:
| Constraint | Plan node (`ValidNode`) |
|------------|-------------------------|
| Integer range (Int8..Uint32) | `Int { min, max }` / `Uint { max }` |
| Int64/Uint64 (full range) | `I64` / `U64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `Float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `Str { max_len }` — `maxLength` baked in from the owning field at compile time |
| Bytes `maxLength` (array length) | `Bytes { max_len }` — accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `Enum { count }` checks `idx < count` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `Union { variants }` reads `__discriminator`, dispatches on the compiled variant nodes |
| Struct fields | `Struct { fields }` requires each declared field present, recurses |
| Array count | `Array { count, element }` checks `arr.len() == count` and recurses per element |
| Record values | `Record { values }` recurses into each value |
| Boolean | `Bool` (materializer already rejects non-0/1 bytes) |
The plan is a public type (`ValidationPlan`, `Debug + Clone +
PartialEq + Eq + Hash + Send + Sync`): `engine.validation_plan()`
exposes it for consumers that validate their own materialized `Value`
trees or want its `fingerprint()` for caching / schema handshakes
(ADR-012 §1). The one-shot `bast_validation::validate_value(&doc,
&value)` remains as a convenience wrapper (compile + validate) for
callers holding a BAST document without an engine.
| Constraint | Validator arm |
|------------|---------------|
| Integer range (Int8..Uint64) | `validate_int`/`validate_uint` with `as_i64`/`as_u64` + range check |
| Int64/Uint64 (full range) | `validate_int64`/`validate_uint64` (JSON precision caveat per ADR-005) |
| Float finiteness (Float32/64) | `validate_float` with `as_f64().is_finite()` |
| String `maxLength` (byte length) | `check_string` reads the field-level `maxLength` |
| Bytes `maxLength` (array length) | `check_bytes` accepts the `Value::Array` form (the materializer emits bytes as an array of u8) |
| Enum index bounds | `validate_enum` checks `idx < values.len()` — **fixes the v0.1.0 dead constraint** |
| Union variant dispatch | `validate_union` reads `__discriminator`, looks up the variant, recurses via `validate_typeref` |
| Struct fields | `validate_struct` walks `fields`, requires each declared field present, recurses |
| Array count | `validate_array` checks `arr.len() == count` and recurses per element |
| Record values | `validate_record` recurses into each value's `values` type |
| Boolean | `validate_bool` (materializer already rejects non-0/1 bytes) |
No external JSON Schema is required for `validate_bytes`. The BAST
document is the complete specification of the binary format — it
@@ -136,14 +122,12 @@ simplification.
The strategy is decided in [ADR-004](decisions/004-error-handling-validation-strategy.md)
and refined by [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md):
1. **Load time:** Parse the BAST document into the typed tree, compile
the `ValidationPlan` (the value-domain constraint tree), compute
1. **Load time:** Parse the BAST document into the typed tree, compute
the layout engine, and (optionally) build the standard
`jsonschema::Validator` for the JSON-validation path. This is the
`AlkTypeEngine::compile` constructor.
2. **Access time:** Use the compiled engine for repeated read/write
operations. Validation is opt-in per operation: the validation half
walks the compiled `ValidationPlan`, never the BAST document.
operations. Validation is opt-in per operation.
### The `AlkTypeEngine` struct
@@ -154,7 +138,6 @@ supports both layout modes (ADR-002) via an internal `Layout` enum:
pub struct AlkTypeEngine {
layout: Layout, // packed or aligned (private enum)
json_validator: Option<jsonschema::Validator>, // None when no JSON Schema supplied
validation_plan: Arc<ValidationPlan>, // compiled value-domain constraints (ADR-012 §3)
endian: Endian, // parsed from the root struct's "endian"
bast_doc: Value, // retained for sequential_reader/read_field
root_name: String, // the selected $defs entry
@@ -192,7 +175,6 @@ impl AlkTypeEngine {
pub fn validate_json(&self, instance: &Value) -> Result<(), AlkTypeError>; // D-BAST-007
pub fn is_valid_json(&self, instance: &Value) -> bool; // D-BAST-007
pub fn validate_bytes(&self, buffer: &[u8]) -> Result<(), AlkTypeError>; // D-BAST-006
pub fn validation_plan(&self) -> &Arc<ValidationPlan>; // compiled constraints (ADR-012 §3)
}
```
@@ -291,23 +273,16 @@ The expensive work happens once at schema load time:
1. Parse the BAST document into the typed tree (`BastDoc::new`).
2. Parse the root struct's `"endian"` annotation.
3. Compile the `ValidationPlan` — the value-domain constraint tree,
with eager `$ref` resolution. Its compile walk rejects cyclic `$ref`
graphs with a clean `Schema` error *before* the layout computation.
(The layout walkers now also guard themselves — each standalone
entry point runs the shared reference-graph check
(`walk_guard::check_ref_graph`, review #006 H2) — so the plan-first
ordering is belt-and-suspenders at engine compile, and the trust
boundary no longer depends on the call path.)
4. Compute the layout (`LayoutBuilder` for packed, `OffsetMap` for
3. Compute the layout (`LayoutBuilder` for packed, `OffsetMap` for
aligned).
5. If `json_schema` is `Some`, build the standard
4. If `json_schema` is `Some`, build the standard
`jsonschema::Validator` via `validation::build_validator`.
The result is an `AlkTypeEngine` that can be used for repeated
operations. The validation half is pre-built: the engine holds an
`Arc<ValidationPlan>` and walks it per buffer without re-touching the
BAST document (ADR-012 §3).
operations. The BAST-native validator is not pre-built — it is a
recursive walker that runs on the materialized `Value` at access time,
re-using the `BastDoc` (re-parsed on demand from the retained
`bast_doc`).
### Access time: `engine.validate_bytes(&[u8])`
@@ -322,13 +297,10 @@ sequence (D-BAST-006):
→ dispatch then recurse; `Record` → object of key/value entries).
The read phase reuses the existing data-access functions and returns
`AlkTypeError::Access` (with field paths) on read failures.
2. **Validate the `Value` against the `ValidationPlan`.** The
materialized `Value` is walked against the compiled constraint tree
(`engine.validation_plan().validate(&value)`), producing
2. **Validate the `Value`.** The materialized `Value` is passed to
`bast_validation::validate_value(&doc, &value)`, producing
`AlkTypeError::Validation` on the first violated value-domain
constraint. This is the ADR-012 §3 end state: the only per-buffer
schema-touching step is the materialize half (the bytes must be
decoded against the tree); the validation half is plan-fast.
constraint.
Mode dispatch:
@@ -336,7 +308,7 @@ Mode dispatch:
- **Aligned mode** — uses the `OffsetMap` to read fields at their
computed offsets.
Both modes produce the same `Value` form; the validation plan is
Both modes produce the same `Value` form; the BAST-native validator is
mode-agnostic.
### Access time: `engine.validate_json(&Value)` / `engine.is_valid_json(&Value)`
@@ -407,10 +379,9 @@ then access the binary buffer.
| Decision | ADR | Summary |
|----------|-----|---------|
| Two-validator model (BAST-native + standard jsonschema) | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | `validate_bytes` uses the compiled `ValidationPlan`; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema; D-BAST-006/007/009 |
| Two-validator model (BAST-native + standard jsonschema) | [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) | `validate_bytes` uses the BAST-native validator; `validate_json` uses a standard `jsonschema::Validator` from a consumer-provided JSON Schema; D-BAST-006/007/009 |
| Error handling and validation strategy | [ADR-004](decisions/004-error-handling-validation-strategy.md) | `AlkTypeError` enum; load-time build, access-time check; field-path-carrying errors; jsonschema `ValidationError` wrapping |
| Generalized validation — `validate_bytes` | [ADR-010](decisions/010-generalized-validation-validate-bytes.md) | Single-call binary-buffer validation (materialize `Value` from bytes, then validate); two methods on one struct, not a trait |
| Compiled `ValidationPlan` | [ADR-012](decisions/012-plan-fingerprinting-and-m1-closure.md) | The value-domain constraint tree is compiled once at `compile` (eager `$ref` resolution, cycle rejection) and walked per buffer; `Hash + Eq` + `fingerprint()`; retires the interpretive `BastDoc` walk |
| BAST format | [ADR-BAST](decisions/bast-bast-format.md) | The BAST document is the complete binary-format spec (layout + value constraints) |
## Open Questions
@@ -426,19 +397,15 @@ see [builder.md](builder.md).
— the normative validation model
- [ADR-VAL-SPLIT](decisions/val-split-two-validator-model.md) — the
two-validator decision
- [ADR-012](decisions/012-plan-fingerprinting-and-m1-closure.md) — the
`ValidationPlan` decision (§3)
- [ADR-004](decisions/004-error-handling-validation-strategy.md) —
error handling and validation strategy
- [ADR-010](decisions/010-generalized-validation-validate-bytes.md) —
`validate_bytes` (the collapsed two-step dance)
- [schema-layer.md](schema-layer.md) — the BAST parser that the
plan compiler consumes
BAST-native validator walks
- [data-access.md](data-access.md) — read/write functions and the
materializer that produce the `Value` the plan checks
- [`src/validation_plan.rs`](../../src/validation_plan.rs) — the
compiled `ValidationPlan` implementation
materializer that produce the `Value` the validator checks
- [`src/bast_validation.rs`](../../src/bast_validation.rs) — the
one-shot wrapper (`validate_value`) and shared error helpers
BAST-native validator implementation
- [`src/validation.rs`](../../src/validation.rs) — the `build_validator`
helper
-983
View File
@@ -1,983 +0,0 @@
---
status: done
created: 2026-08-19
last_updated: 2026-09-02
adr: ADR-011, ADR-012
---
# 0.3.0 — Compiled Forms: ReadPlan, Owned BastDoc, OffsetMap LeafMeta, ValidationPlan, Fingerprinting
This is the execution plan for the 0.3.0 release: the compiled-form
rollup that closes review #004's 400x read-path gap (ADR-011) *and*
the deferred M1 sites (ADR-012) *and* retires the interpretive
validation walk via a `ValidationPlan` (ADR-012 §3, reversing the
original deferral per review #005 M3) *and* adds plan fingerprinting
(ADR-012 §1/§4) in one breaking bump. It is the **entry point** an
implementing agent reads first.
Companion documents:
- [ADR-011](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md)
— the `ReadPlan` decision (packed read-side compiled form).
- [ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md)
— fingerprinting + owned `BastDoc` + `OffsetMap` `LeafMeta` +
`ValidationPlan` (this release's other three pieces).
- [Review #004](../reviews/004-performance-review.md) — the
performance finding being closed.
- [POC findings](../../poc/readplan/FINDINGS.md) (branch `readplan-poc`)
— the derisking POC that confirmed the `ReadPlan` shape and surfaced
two findings (field-disc union read shape; struct-array stride).
**Working order:** read this plan top-to-bottom. The Semver Contract
section is the scope-creep guardrail — consult it before each step.
Each step links to its ADR and lists its verification gate. Implement
phases in order; within a phase, steps are ordered by dependency.
## Phases vs sessions
This plan is deliberately larger than one session's work. The eight
phases are the session boundaries — each phase is a coherent unit
that leaves the tree building and tests green, so any one session
can pick up a phase without needing context from the previous one.
Phase boundaries are also commit boundaries (and push boundaries per
AGENTS.md). If a phase is large enough to span sessions, the steps
within it are the sub-session boundaries.
## Semver Contract
The crate is on crates.io at 0.2.0 with zero real consumers (only
`alktty`/`alkcall`, both in-house path dev-deps). A breaking bump to
0.3.0 is free but the contract is explicit so the implementation
doesn't drift. Per AGENTS.md, the public surface is the `lib.rs`
re-exports.
| Public item (from `lib.rs` re-exports) | Class | Change |
|---|---|---|
| `AlkTypeEngine::compile` | **Breaking (internal)** | Signature unchanged `(bast_doc: &Value, root_name: &str, mode, json_schema) -> Result<Self, AlkTypeError>`. Internally builds a `ReadPlan` (packed) or extended `OffsetMap` (aligned) and stores it. The `bast_doc: Value` clone is retained (ADR-011 §Engine integration). |
| `AlkTypeEngine::sequential_reader` | **Breaking (return type)** | Returns `Option<SequentialReader>` (unchanged type), but the reader is now constructed from `Arc<ReadPlan>`, not from `&bast_doc`. The reader's public methods (`read_next`/`read_field`/`reset`/`position`/`endian`/`schema`) keep their signatures. `schema()` returns the `&Value` the plan was compiled from (retained on the engine). |
| `AlkTypeEngine::read_field` / `write_field` | **Unchanged (signature)** | Still `(buffer, field_path) -> Result<FieldValue, AlkTypeError>`. Internally reads `LeafMeta` from the extended `OffsetMap` instead of re-parsing `BastDoc`. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (signature)** | Packed mode calls `materialize_packed(&self.plan, buffer)` (ADR-011); aligned mode calls `materialize_aligned(&doc, buffer, &self.offset_map)` with the owned `BastDoc`. |
| `LayoutMode`, `AlkTypeEngine` | **Unchanged** | — |
| `BastDoc`, `BastDef`, `BastDefKind`, `BastStruct`, `BastField`, `BastType`, `BastUnion`, `BastDiscriminator`, `BastEnum`, `BastArray`, `BastRecord`, `BastRef` | **Breaking (lifetime removal)** | `BastDoc<'a>` → `BastDoc` (owned). Every `&'a str` → `String` (or `Arc<str>` — decision in phase 3). Every `&'a Value` → `Value` (or `Arc<Value>`). Every method signature that took/returned `&'a` changes. The `Bast*` types are re-exported from `lib.rs` so this is a public break. |
| `OffsetMap` | **Breaking (`get` return type)** | `get(field_path) -> Option<&ByteRange>` → `get(field_path) -> Option<&OffsetEntry>` where `OffsetEntry { range: ByteRange, meta: LeafMeta }` (or two accessors). Additive capability. |
| `ByteRange` | **Unchanged** | Still `Copy + PartialEq + Eq + Hash`. |
| `LeafMeta` | **New public type** | `{ kind: AlkTypeKind, encoding: VariableEncoding, endian: Endian }`, re-exported from `lib.rs`. `Copy + PartialEq + Eq + Hash`. |
| `ReadPlan` | **New public type** | From ADR-011. Re-exported from `lib.rs`. `Debug + Clone + PartialEq + Eq + Hash`. |
| `SequentialReader` | **Breaking (constructor + return type)** | `SequentialReader::new(&Value, &str) -> Result<Self, AlkTypeError>` → `SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible — just stores the `Arc`; the `BastDoc` parse moved to `ReadPlan::compile`). Public methods (`read_next`/`read_field`/`reset`/`position`/`endian`/`schema`) unchanged. `schema()` returns `&Value` retained on the plan (see phase 2 — the plan stores `Arc<Value>`, not `&Value`, to avoid the self-referential struct ADR-011 rejects). `engine.rs`'s `.ok()` on the old `Result` correspondingly goes away. |
| `FieldValue` | **Unchanged** | — |
| `materialize_packed` | **Breaking (signature)** | `materialize_packed(&BastDoc<'_>, &[u8])` → `materialize_packed(&ReadPlan, &[u8])`. |
| `materialize_aligned` | **Breaking (signature)** | `materialize_aligned(&BastDoc<'_>, &[u8], &OffsetMap)` → `materialize_aligned(&BastDoc, &[u8], &OffsetMap)` (owned `BastDoc`, no lifetime). |
| `ValidationPlan` | **New public type** | From ADR-012 §3. Re-exported from `lib.rs`. `Debug + Clone + PartialEq + Eq + Hash`. `compile(&BastDoc, &str) -> Result<Self, AlkTypeError>`, `fingerprint() -> u64`. Shape scoped in phase 7 (and a preceding design session); the contract is fixed in ADR-012 §3. |
| `AlkTypeEngine::validate_bytes` | **Unchanged (signature)** | Still `(buffer) -> Result<(), AlkTypeError>`. Internally walks `&ValidationPlan` (both modes) instead of re-walking `BastDoc` for value-domain checks. The `materialize` half is unchanged from the ADR-011/§2b work (packed: `materialize_packed(&self.plan, ...)`; aligned: `materialize_aligned(&self.doc, ..., &self.offset_map)`). |
| `bast_validation` (`build_validator`, `validate_value`) | **Additive (reviewed at phase 7)** | `validate_value` is expected to become a thin wrapper over `ValidationPlan` (or be retired if the shape work shows it's redundant). Additive changes ride the bump; renames/removals are avoided unless phase 7's shape work shows they're necessary. The BAST meta-schema validator (`build_validator`/`BAST_META_SCHEMA`) used at `compile` time is unaffected. |
| `LayoutBuilder`, `PackedLayout`, `FieldPosition` | **Unchanged (signature)** | `LayoutBuilder::new`/`build` signatures unchanged. Internally caches the owned `BastDoc` instead of re-parsing. |
| `AlkTypeKind`, `Endian`, `VariableEncoding` | **Unchanged** | — |
| `AlkTypeError` | **Unchanged** | — |
| `Schema`, `Definitions`, `Discriminator` builders | **Unchanged** | — |
| `UnionDispatch`, `build_validator`, `BAST_META_SCHEMA` | **Unchanged** | — |
| `data_access::*` functions | **Unchanged** | — |
**Net breaking surface:** `BastDoc` and all `Bast*` types (lifetime
removal), `OffsetMap::get` (return type), `SequentialReader::new`
(constructor + `Result` drop), `materialize_packed`/`materialize_aligned`
(signatures). **Net additive:** `ReadPlan`, `LeafMeta`, `ValidationPlan`,
`fingerprint()` methods, `Hash + Eq` derives on `ReadPlan`/`OffsetMap`/
`ValidationPlan`. **Net unchanged:** the builder, `AlkTypeEngine`
accessors (signatures), `FieldValue`, `AlkTypeKind`, `AlkTypeError`,
`data_access`, the `Schema`/`Definitions` builders, `validate_bytes`
(signature).
### Decisions deferred to their implementation phases
1. **`Arc<str>` vs `String` for owned `BastDoc` names** (phase 3): `Arc<str>`
shares allocation for repeated names (e.g. union variant keys appearing
in multiple places); `String` is simpler. The POC used `String`. Lean:
`String` unless a bench shows `Arc<str>` matters — the typed tree is
built once, not hot. Decided in phase 3.
2. **`OffsetMap::get` return shape** (phase 5): `Option<&OffsetEntry>` (a
new accessor struct) vs two methods `range(path) -> Option<&ByteRange>`
+ `meta(path) -> Option<&LeafMeta>`. The struct is fewer calls; the two
methods preserve back-compat shape for callers that only want the range.
Lean: struct — it's a breaking bump anyway and the struct is cleaner.
Decided in phase 5.
3. **Fingerprint hasher** (phase 6): `DefaultHasher` (std, stable within a
version) vs `FxHasher` (faster, also stable). Cross-version stability
is a non-goal (ADR-012). Lean: `DefaultHasher` — no new dep, the
fingerprint isn't hot. Decided in phase 6.
4. **Struct-array stride** (phase 2): the POC found the existing reader
returns `element_stride: 0` for fixed-size struct arrays
(`sequential_reader.rs:567`). The `ReadPlan` correctly computes the
stride. Decision: preserve the existing `0` behavior in `ReadPlan` for
back-compat with `SequentialReader`'s consumer contract, *or* fix it
and document the behavioral change. Lean: fix it — the `0` is a latent
bug, the stride is behaviorally observable, and we're bumping. The
plan step calls this out explicitly. Decided in phase 2.
## Phase 1 — `ReadPlan` type + `compile` (ADR-011 step 1) — **DONE (2026-09-02)**
> **Status: implemented.** `src/read_plan.rs` builds the refined
> `CompositePlan::Union` shape (`shared: Option<Box<ReadPlan>>` +
> `variants: Vec<(String, CompositePlan)>`, no `VariantPlan`/`VariantKind`)
> with eager `$ref` resolution, `BTreeMap` `by_name`, field-disc `shared`
> sub-plans, nested-union variant support, and true array strides
> (deferred decision 4 resolved: fixed struct arrays compute their real
> stride, not the 0.2.0 reader's `0`). Two parity notes recorded as
> compile-behavior locks in tests: (a) the plan propagates the
> *referring field's* effective endianness into nested structs/unions —
> exactly what the 0.2.0 packed reader/materializer do — rather than
> consulting nested containers' own `endian` annotations (the POC baked
> `s.endian()` there; its equivalence tests never covered a nested
> annotation, so the divergence was latent); (b) `compile` carries its
> own depth cap (128) + definition-level cycle set, so standalone
> `ReadPlan::compile` is untrusted-input-safe independent of the
> meta-schema and the engine's `ValidationPlan` gate.
**Goal:** Add the `ReadPlan`/`FieldPlan`/`CompositePlan`/`ReadKind`/
**ADR reference:** [ADR-011 §The `ReadPlan` shape](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#the-readplan-shape),
[ADR-011 §Construction](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#construction).
The `CompositePlan::Union` shape in ADR-011 was refined (vs the
accepted-at-POC shape) to carry the field-disc union's shared fields
and to drop `VariantPlan`/`VariantKind`; this phase implements the
refined shape.
**POC reference:** `poc/readplan/src/lib.rs` (branch `readplan-poc`)
is the reference scaffold. The production version lives in `src/` and
adds doc comments, clippy cleanliness, the field-name-discriminator
union read shape the POC stubbed (POC Finding 1), and nested-union
variant support the POC rejected but 0.2.0 accepts (POC "What this
POC does not cover" → nested unions). Both are resolved by the
refined `CompositePlan::Union` shape — see below.
**Files:** New `src/read_plan.rs`. Update `src/lib.rs` to add
`pub mod read_plan;` and re-export `ReadPlan` (and the plan sub-types
that are part of the public surface — `ReadKind`, `CompositePlan`,
etc. if the ADR's public-API section calls for them; the ADR lists
`ReadPlan` as the public type, sub-types can stay `pub` in-module if
consumers don't need to name them).
**Implementation notes:**
- `compile` walks `BastDoc` once (via the existing borrowed `BastDoc`,
which still exists at this phase — the owned-`BastDoc` refactor is
phase 3). Resolves all `$ref`s eagerly, computes effective endianness
at every node, inlines union variants. Malformed schemas surface as
`AlkTypeError::Schema` (AGENTS.md §3); overflow-safe arithmetic
(AGENTS.md §4).
- `by_name: BTreeMap<String, usize>` (not `HashMap` — ADR-012 §1
requires `Hash` on `ReadPlan`, and `HashMap` blocks derive). The
POC used `HashMap`; swap to `BTreeMap`.
- **Union shape (ADR-011 refined — resolves POC Finding 1 and the
nested-union gap):** `CompositePlan::Union { disc, shared,
variants: Vec<(String, CompositePlan)> }`.
- `shared: Option<Box<ReadPlan>>` — the union's declared `fields`
(the discriminator field + any shared fields) for the
field-name-discriminator case. `DiscriminatorPlan::Field.field_index`
indexes into `shared`. The read loop walks `shared` first, then
looks up the selected variant and walks its `CompositePlan`
starting after the shared fields. The byte-offset-discriminator
case sets `shared: None` (no shared fields; the variant starts
immediately after the discriminator size). This replaces the
POC's stubbed `plan_read_union` `Field` arm.
- `variants: Vec<(String, CompositePlan)>` — **not** the POC's
`Vec<(String, VariantPlan)>`. Dropping `VariantPlan`/`VariantKind`
means a variant's body is just a `CompositePlan`, so **nested
unions** (a variant that is itself a union, which the 0.2.0
reader supports via `resolve_and_walk_variant`'s `Union` arm at
`sequential_reader.rs:800`) work by ordinary `CompositePlan`
recursion — a variant can be `CompositePlan::Union { ... }`. No
separate `VariantKind::Union` arm, no behavioral drop vs 0.2.0,
no Semver Contract entry for a capability regression. The POC's
`VariantPlan { kind, plan }` wrapper is not carried forward.
- `compile_union` must reject a variant that is neither a struct
nor a union with `AlkTypeError::Schema` (mirroring
`resolve_and_walk_variant`'s `other => Err(...)` arm), so the
eager-resolution path keeps the untrusted-input discipline.
- Do *not* wire `ReadPlan` into `SequentialReader` or `materialize` yet
— that's phase 2. Phase 1 is the type + `compile` only, unit-tested
against the same BAST fixtures `bast.rs` uses (the existing `bast.rs`
tests are a ready source of fixtures).
**Verification:** `cargo test --release` (new unit tests for `compile`
covering every `BastType` arm — port the `cov_*` tests from
`poc/readplan/tests/coverage.rs`, **plus** a nested-union-variant
test asserting `compile_union` produces `CompositePlan::Union` whose
variant body is itself `CompositePlan::Union`, restoring the 0.2.0
capability the POC rejected). `cargo clippy --all-targets -- -D
warnings`. `cargo doc --no-deps` (new public type). `cargo build
--target wasm32-unknown-unknown --release` (`read_plan.rs` is
wasm-relevant). **Add a `fn read_plan_is_send_sync()` assertion
test** (a `const _: fn() = || { fn assert_send_sync<T: Send + Sync>()
{}; assert_send_sync::<ReadPlan>(); };` static-bound assertion, as
ADR-011 §"Engine integration" requires) to lock in `ReadPlan: Send +
Sync` so a future change can't break it silently — mirror the POC's
`readplan_is_send_sync` test.
---
## Phase 2 — `SequentialReader` + `materialize_packed` consume `ReadPlan` (ADR-011 steps 2–4) — **DONE (2026-09-02)**
> **Status: implemented.** The packed read loop walks `Arc<ReadPlan>`:
> `SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible; the old
> fallible constructor's work moved to `ReadPlan::compile`), the reader
> holds `plan: Arc<ReadPlan>` + cursor only, `schema()` returns the
> `Arc<Value>` retained on the plan (review #005 H2 closed — no
> self-referential struct), and a new `plan()` accessor exposes the
> shared plan. `materialize_packed(&ReadPlan, &[u8])` walks the same
> plan; the aligned materialize path keeps walking `BastDoc` with the
> retained `dummy_field_for`/`ty_source`/`materialize_typeref_packed`
> helpers (phase 5 Scope Boundary). Engine: `Layout::Packed` carries
> `Arc<ReadPlan>`; `sequential_reader()` is an `Arc::clone` (15.7 ns,
> was a whole-document `Value` clone); packed `validate_bytes` calls
> `materialize_packed(&self.plan, ...)`. The temporary validation
> bridge (reconstruct `BastDoc` for the validator) is still in place —
> phase 7 already retired it on `main`'s ValidationPlan; this phase's
> `validate_bytes` edit merged cleanly onto that state.
>
> **Stride (deferred decision 4):** fixed struct/nested-array elements
> now report their true stride through `FieldValue::Array`
> (0.2.0 returned `0`); doc comment updated; no existing test asserted
> the `0`, so no test needed changing — the plan-compile tests lock the
> new values.
>
> **Two parity subtleties found and preserved** (both invisible to the
> existing test suite, both now locked by tests or by construction):
> (a) the materializer unwraps the plan's anonymous single-field
> wrapper for primitive array elements/record values — without this,
> materialized records/arrays would nest each leaf under a synthetic
> object and `validate_bytes` would fail its own parity suite (caught
> by `materialize_record_packed_count_prefixed_pairs`); (b) the
> field-disc union's materialized key order keeps `__discriminator`
> first (matching 0.2.0's `Map` insertion order, observable under
> `preserve_order`).
>
> **Bench (alktty `wire_vs_bast`, 1024 chunks/stream):** read p64
> 2.27 µs/chunk (review #004) → **98 ns/chunk** (~23x; gap 400x →
> ~17x vs hand-rolled's 5.6 ns); read p4k → 100 ns/chunk. `engine.
> sequential_reader()` construction 15.7 ns (was a full `Value` clone).
> The residual gap is dominated by the per-field `String` allocation
> mandated by the unchanged `(String, FieldValue)` `read_next` return
> signature (2 allocs/chunk) plus `data_access` bounds checks — both
> outside this phase's scope (the signature is pinned by the Semver
> Contract).
**Goal:** Rewrite the packed read loop to walk `&ReadPlan` instead of
reconstructing `BastDoc`. `SequentialReader` stores `Arc<ReadPlan>` +
cursor state; `materialize_packed` takes `&ReadPlan`. Closes H1
(the 400x gap) + the packed-side M1 + L1.
**ADR reference:** [ADR-011 §Scope](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#scope),
[ADR-011 §Recommended Order](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#recommended-order)
steps 2–4.
**Files:** `src/sequential_reader.rs` (rewrite the read loop, change
`new`'s signature), `src/materialize.rs` (`materialize_packed` takes
`&ReadPlan`), `src/engine.rs` (`compile` builds `Arc<ReadPlan>` in
packed mode, `sequential_reader()` hands out `Arc::clone`, packed
`validate_bytes` calls `materialize_packed(&self.plan, ...)`).
**Implementation notes:**
- `SequentialReader::new(&Value, &str) -> Result<Self, AlkTypeError>` →
`SequentialReader::new(Arc<ReadPlan>) -> Self` (infallible — the
fallible `BastDoc` parse moved to `ReadPlan::compile` in phase 1;
`new` just stores the `Arc`). The reader stores `plan: Arc<ReadPlan>`,
`field_index: usize`, `position: usize`. `endian()` reads
`self.plan.endian()`.
- **`schema()` ownership (resolves review #005 H2):** `schema()`
returns `&Value` retained on the plan, but the plan stores an
**`Arc<Value>`**, not a `&Value`. ADR-011 §"Root cause" rejects the
self-referential struct pattern (a `ReadPlan` storing `&Value`
borrowing from the engine's `bast_doc: Value` would make the engine
self-referential — exactly the construction ADR-007 worked around
and ADR-011's `Arc<ReadPlan>` was meant to retire). The fix:
`ReadPlan` carries `schema: Arc<Value>`; `ReadPlan::compile` clones
the input `&Value` into `Arc<Value>` once; `schema()` returns
`&self.schema`. The engine stores `bast_doc: Arc<Value>` internally
(one allocation, shared via refcount between the engine and all
plans it builds) — this is an internal change, not a public
signature change (`compile` still takes `&Value`). `Arc<Value>`
implements `Hash + Eq` (`serde_json::Value: Hash + Eq` as of the
pinned `serde_json` with `preserve_order`; `Map::hash` sorts keys
for determinism), so phase 6's `#[derive(Hash)]` on `ReadPlan` is
not blocked by carrying the schema. **Note:** if a future
`serde_json` version regresses `Value: Hash`, phase 6 would need
`ReadPlan`'s hash to exclude the `schema` field; that is a phase-6
concern, not a phase-2 blocker.
- `read_field_at`/`read_field_value`/`read_typeref_value`/
`walk_struct_size`/`read_union_value`/`read_array_value`/
`read_record_value` are rewritten to take plan nodes
(`&FieldPlan`/`&CompositePlan`/`&ReadKind`) instead of
`&BastField`/`&BastType`/`&BastDoc`. Port `plan_read_field_at`/
`plan_walk_struct_size`/etc. from `poc/readplan/src/lib.rs` — they're
the reference implementations. The `read_union_value` rewrite
handles both discriminator kinds via the refined `CompositePlan::Union`
shape from phase 1: byte-disc reads the discriminator at
`disc.offset` then walks the variant (no `shared`); field-disc walks
`shared` first, reads the discriminator field at
`disc.field_index` within `shared`, looks up the variant, and walks
it starting after the shared fields. Nested unions (variant body is
itself `CompositePlan::Union`) recurse naturally — no special arm.
- **Struct-array stride (deferred decision 4):** the POC computes the
true fixed-struct stride; the existing reader returns `0`. The
production `ReadPlan::compile_array` should compute the true stride
(the POC's `fixed_struct_size` helper). Document the behavioral
change in the `FieldValue::Array` doc comment: `element_stride` is
now the true stride for fixed-size struct elements, not `0`. This is
a breaking behavioral change; rides the bump. Update the
`eq_array_ref_element`-style test to assert the new stride.
- **`materialize_packed` rewrite + packed/aligned split (resolves
review #005 L1 and L2):** `materialize_packed(&BastDoc<'_>, &[u8])`
→ `materialize_packed(&ReadPlan, &[u8])`. The materializer walks the
plan instead of `BastDoc`. **Scoped removal of helpers:** only the
*packed-side* call sites of `dummy_field_for`/`ty_source`
(`materialize.rs:249, 316, 351, 391` — the packed
`materialize_*_packed` paths) go away when packed-materialize moves
to the plan. The helpers themselves **stay**, because
`materialize_leaf_at` (`materialize.rs:631`, which calls
`dummy_field_for`) is on the **aligned** path — it's called by
`materialize_struct_aligned` (`:475`), `materialize_array_aligned`
(`:544`), `materialize_variable_aligned` (`:613`). Aligned
`materialize` keeps walking `BastDoc` through 0.3.0 (see the Scope
Boundary note in phase 5), so `dummy_field_for`/`ty_source` must
stay. **`materialize_typeref_packed` split:** this function is
currently shared by both packed and aligned paths (aligned's
`materialize_leaf_at` calls it to read leaves, and aligned's record
path at `:498-506` calls it directly). After phase 2,
packed-materialize gets a new plan-walking function;
`materialize_typeref_packed` stays for aligned's
`materialize_leaf_at` and the aligned record path (renamed or not —
implementer's choice; the function is private). This is two
mode-specific paths — the existing design — not a "parallel walker"
in the maintenance-tax sense ADR-011 §"Negative" cautions against
(that caution is about packed read-side `SequentialReader` +
`materialize_packed` sharing one plan, which this preserves).
- `AlkTypeEngine::compile` (packed branch): build `Arc<ReadPlan>` via
`ReadPlan::compile(bast_doc, root_name)`, store it in `Layout::Packed`.
`sequential_reader()` returns
`Some(SequentialReader::new(Arc::clone(&self.plan)))`.
`validate_bytes` (packed) calls
`materialize_packed(&self.plan, buffer)` then runs validation on the
materialized `Value`. **Validation path through phase 2:** until
phase 7 (ValidationPlan), `validate_bytes` reconstructs a `BastDoc`
for the validator only — the *read* path uses the plan, the
*validation* path uses `BastDoc`. This is a temporary bridge: phase 7
replaces it with a `ValidationPlan` walk (ADR-012 §3), retiring the
per-call `BastDoc` reconstruction. The bridge is acceptable for
phases 2–6 because the ValidationPlan work is committed in this
release (not deferred), so the bridge has a known removal point in
phase 7.
**Verification:** `cargo test --release` — the existing
`sequential_reader.rs` and `materialize.rs` tests drive `read_next`/
`read_field`/`reset`/`validate_bytes` through the public API, so they
validate the rewrite without modification. If any test breaks, the
rewrite diverged from the existing behavior — investigate before
patching the test. `cargo clippy --all-targets -- -D warnings`.
`cargo build --target wasm32-unknown-unknown --release`. **Re-run the
alktty `wire_vs_bast` bench** to confirm the 400x gap closes (this is
the headline result; record the before/after numbers in the commit
message).
---
## Phase 3 — Owned `BastDoc` (ADR-012 §2a) — **DONE (2026-09-02)**
> **Status: implemented.** Every `Bast*` type dropped its `<'a>`:
> `&'a str` → `String`, `&'a Value` → `Value` (deferred decision 1
> resolved: plain `String`/`Value` — the tree is built once, name
> sharing via `Arc<str>` needs a bench justification that doesn't
> exist). `BastDoc::new(&Value, &str)` still takes references in and
> clones into owned storage; `BastDoc` gained `Clone`. `resolve_*`
> return owned `BastDef`/`BastType`. Consumers adapted:
> `OffsetMap::compute(&BastDoc)`, `materialize_aligned(&BastDoc, ...)`
> (no lifetime), `BuildCtx`/`ComputeCtx` hold `&'d BastDoc`.
> **Engine ownership flip:** `AlkTypeEngine` now holds the owned
> `BastDoc` (replacing `bast_doc: Value` + `root_name: String` —
> `root_name()` delegates to the doc), killing its three per-call
> `BastDoc::new` re-parses (`validate_bytes` aligned path,
> `read_field`, `write_field` — review #004 M1 sites by construction;
> phase 5 removes the `lookup_leaf_field` walk itself). New public
> accessor `AlkTypeEngine::root_name()` (additive). **Bonus cleanup:**
> `materialize_typeref_packed`'s dead `_field` param dropped (phase 2
> left it dangling); under ownership it would have forced a deep
> `Value` clone per array element/record value/union variant via
> `dummy_field_for` — the param, `dummy_field_for`, and `ty_source`
> are gone (no behavior change; the `_field` arg was already
> ignored). `BastField::synthetic` retains an owned-signature
> `#[allow(dead_code)]` definition (no remaining callers; kept for the
> phase-4/5-aligned materializer helpers if they need it). Existing
> `BastDoc` consumers (`LayoutBuilder`'s `doc_value` re-parse cache)
> are unchanged pending phase 4. Engine stays `Send + Sync` with the
> owned doc — the engine's thread-share test now asserts it directly.
**Goal:** Make `BastDoc` own its data (drop the `<'a>` lifetime).
`&'a str` → `String`, `&'a Value` → `Value` (or `Arc<str>`/`Arc<Value>`
— deferred decision 1). This is the prerequisite for the `LayoutBuilder`
M1 fix (phase 4) and simplifies all owning consumers. Broad but
mechanical refactor.
**ADR reference:** [ADR-012 §2a](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2a-layoutbuilder--cache-the-parsed-bastdoc-at-new).
**Files:** `src/bast.rs` (every `Bast*` type), every consumer:
`src/layout_builder.rs`, `src/offset_map.rs`, `src/materialize.rs`,
`src/bast_validation.rs`, `src/engine.rs`, `src/tunion.rs`,
`src/bast_meta.rs` (if it walks `BastDoc`), `src/builder.rs` (if it
consumes `Bast*`). `src/lib.rs` re-exports (signatures change but
names stay).
**Implementation notes:**
- `BastDoc<'a>` → `BastDoc`. Fields: `root: Value` (was `&'a Value`),
`root_name: String` (was `&'a str`), `root_def: BastDef` (was
`BastDef<'a>`). `new(root: &Value, root_name: &str)` takes references
*in* (the caller still owns the input `Value`) but clones into owned
storage. The `&Value` → `Value` clone is the cost of ownership; it
happens once at `compile`/`new`, not per-field.
- `BastDef<'a>` → `BastDef`: `name: String`, `kind: BastDefKind`,
`source: Value`.
- `BastStruct<'a>` → `BastStruct`: `endian: Endian`, `align: Option<usize>`,
`fields: Vec<BastField>`, `source: Value`.
- `BastField<'a>` → `BastField`: `name: String`, `ty: BastType`,
`endian: Option<Endian>`, `align: Option<usize>`,
`encoding: VariableEncoding`, `max_length: Option<usize>`,
`source: Value`. `synthetic` constructor takes owned `BastType` +
`Value`.
- `BastType<'a>` → `BastType`: `Primitive(AlkTypeKind)`,
`Ref(BastRef)`, `Array(BastArray)`, `Record(BastRecord)`,
`Struct(BastStruct)`, `Union(BastUnion)`, `Enum(BastEnum)`.
- `BastUnion<'a>` → `BastUnion`: `endian: Endian`,
`discriminator: BastDiscriminator`, `fields: Vec<BastField>`,
`mapping: Vec<(String, BastType)>` (was `Vec<(&'a str, BastType)>`),
`source: Value`.
- `BastDiscriminator::Field { name: String }` (was `name: &'a str`).
- `BastEnum<'a>` → `BastEnum`: `values: Vec<String>` (was
`Vec<&'a str>`), `source: Value`.
- `BastArray<'a>` → `BastArray`: `element: Box<BastType>`, `count: usize`,
`source: Value`.
- `BastRecord<'a>` → `BastRecord`: `values: Box<BastType>`,
`source: Value`.
- `BastRef<'a>` → `BastRef`: `name: String`.
- **`resolve_typeref` / `resolve_ref` / `lookup_def`** now return owned
`BastType`/`BastDef`/`Value` instead of borrowed. The `clone()` in
the current `resolve_typeref` passthrough (`other => Ok(other.clone())`)
is no longer needed for the borrow case (everything is owned) but
the logic is unchanged — `BastType` is `Clone` either way.
- **Consumers adapt:** any code that held `&'a Value` alongside a
`BastDoc<'a>` (e.g. `LayoutBuilder.doc_value`, `AlkTypeEngine.bast_doc`,
`SequentialReader.doc_value` — though the reader is already on
`ReadPlan` after phase 2) drops the separate `Value` and holds the
owned `BastDoc` directly. `materialize_packed` is already on
`ReadPlan` (phase 2) and doesn't need `BastDoc` — unaffected.
`materialize_aligned` takes `&BastDoc` (owned, no lifetime).
- **`Arc<str>` vs `String` (deferred decision 1):** default to
`String`. The typed tree is built once; name sharing via `Arc<str>`
is a micro-optimization not justified without a bench. If phase 4's
`LayoutBuilder` work shows name allocation is measurable, revisit.
**Verification:** `cargo test --release` — the existing `bast.rs` tests
are the primary validation (they exercise every parser path). All
`Bast*`-consuming tests must pass unchanged (they go through public
APIs that still take `&Value`/`&str` in, just return owned types out).
`cargo clippy --all-targets -- -D warnings`. `cargo doc --no-deps`.
`cargo build --target wasm32-unknown-unknown --release` (`bast.rs` is
wasm-relevant).
---
## Phase 4 — `LayoutBuilder` caches the owned `BastDoc` (ADR-012 §2a) — **DONE (2026-09-02)**
> **Status: implemented.** `LayoutBuilder` stores `doc: BastDoc` +
> `endian` (the `doc_value: Value` + `root_name: String` cache is
> gone); `new` parses once, `build` walks `&self.doc` — the
> `layout_builder.rs` re-parse (M1) is retired. The `build`-time
> root-is-struct re-check replaced its `unreachable!()` with a clean
> `Schema` error (AGENTS.md §3 never-panic; the invariant is
> unchanged — `new` already rejects non-struct roots). Boxing fallout:
> the builder now holds the full owned tree, so `Layout::Packed`
> boxes it (`builder: Box<LayoutBuilder>`) to keep the engine's
> `Layout` enum variant sizes balanced (clippy
> `large_enum_variant`); `layout_builder()` still returns
> `Option<&LayoutBuilder>` via auto-deref, public API unchanged.
**Goal:** `LayoutBuilder::new` parses the owned `BastDoc` once and
stores it; `build` reuses it. Removes the `layout_builder.rs:190`
re-parse (M1). Non-breaking from the public API perspective
(`new`/`build` signatures unchanged); the change is internal.
**ADR reference:** [ADR-012 §2a](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2a-layoutbuilder--cache-the-parsed-bastdoc-at-new).
**Files:** `src/layout_builder.rs`.
**Implementation notes:**
- `LayoutBuilder` currently stores `doc_value: Value` + `root_name:
String` + `endian: Endian`. After phase 3, it stores `doc: BastDoc`
(owned) + `endian: Endian`. `new` calls `BastDoc::new` once;
`build(&self, var_sizes)` uses `&self.doc` directly — no
`BastDoc::new` call inside `build`.
- The `BuildCtx<'d>` struct (currently `doc: &'d BastDoc<'d>`) becomes
`doc: &BastDoc` (no lifetime, or a single lifetime for the borrow
from `&self`). The walk logic is unchanged.
- The `doc_value: Value` clone is removed; the builder holds the owned
`BastDoc` directly. `endian` is read from the doc at `new` time
(already is).
**Verification:** `cargo test --release` — the existing
`layout_builder.rs` tests pass unchanged (they go through
`LayoutBuilder::new` + `build`). `cargo clippy --all-targets -- -D
warnings`. `cargo build --target wasm32-unknown-unknown --release`.
---
## Phase 5 — `OffsetMap` carries `LeafMeta` (ADR-012 §2b) — **DONE (2026-09-02)**
> **Status: implemented.** Prerequisite first (review #005 M2):
> `Hash` added to `Endian`/`VariableEncoding` derives in `schema.rs`
> (additive; both are fieldless `Eq` enums). New public types
> `LeafMeta { kind, encoding, endian }` (`Copy + PartialEq + Eq +
> Hash`) and `OffsetEntry { range, meta }` (with `start()`/`end()`
> convenience accessors), both re-exported from `lib.rs`; `ByteRange`
> gained `Hash` (additive). Storage is `Vec<(String, OffsetEntry)>`
> (deferred decision 2 resolved: struct — `get` returns
> `Option<&OffsetEntry>`, `iter` yields `(&str, &OffsetEntry)`).
> `LeafMeta` is computed at `compute` time with **effective** endian
> threaded through the walk: container default → field override per
> field, propagated into nested-struct probes and array elements via
> the referring field (the same propagation the aligned materializer
> uses). **Parity note:** this replaces `engine.rs`'s
> `lookup_leaf_field` walk, which computed nested-struct defaults from
> the *nested struct's own* `endian` annotation — the two paths
> diverged whenever a nested struct declared `endian` and its
> referring field also declared one (the map now agrees with the
> aligned materializer and the packed `ReadPlan`; the old divergence
> was unreachable through `read_field` only when a nested annotation
> existed, and no test pinned it). `read_field`/`write_field` dispatch
> on the entry's `LeafMeta` — the `BastDoc` re-parse +
> `lookup_leaf_field`/`LeafFieldInfo` per access are gone (the last
> two M1 sites, engine.rs `read_field`/`write_field`). Behavior
> change: `read_field` on a path absent from the map (e.g. a
> whole-struct field path) now errors with `Offset` ("field not found
> in offset map") instead of `Access` ("does not support composite
> types") — the composite-path test already accepted either variant.
> `materialize_aligned`'s four `offset_map.get` call sites updated to
> `.range.start`. `alktty`/`alkcall` untouched (the bench never uses
> `OffsetMap::get`; alkcall has no dependency yet).
**Goal:** Extend `OffsetMap`'s entries with `LeafMeta { kind, encoding,
endian }` computed at `compute` time. `read_field`/`write_field` drop
the `BastDoc::new` + `lookup_leaf_field` calls (M1 aligned-side).
Closes the last two M1 sites (`engine.rs:334,467`).
**ADR reference:** [ADR-012 §2b](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#2b-offsetmap--carry-leaf-metadata).
**Files:** `src/offset_map.rs` (extend entries, compute `LeafMeta`),
`src/engine.rs` (rewrite `read_field`/`write_field` to use the map's
`LeafMeta`, remove `lookup_leaf_field` + `LeafFieldInfo`), `src/lib.rs`
(re-export `LeafMeta`).
**Implementation notes:**
- **`Hash` on `Endian`/`VariableEncoding` (resolves review #005 M2 —
do this first, it's a prerequisite):** `src/schema.rs:205` (`Endian`)
and `:212` (`VariableEncoding`) currently derive only `Debug, Clone,
Copy, PartialEq, Eq` — no `Hash`. `LeafMeta` (below) requires all
its fields to be `Hash` for `#[derive(Hash)]`, and phase 6's
`#[derive(Hash)]` on `ReadPlan`/`OffsetMap` requires `FieldPlan`'s
`endian: Endian` + `encoding: VariableEncoding` to be `Hash`. Add
`Hash` to both derives in `src/schema.rs`. Both are fieldless enums
already at `Eq + PartialEq`, so this is additive and semver-safe —
no behavioral change. Trivial, but it's an unstated prerequisite
the original plan omitted.
- New public type `LeafMeta { kind: AlkTypeKind, encoding:
VariableEncoding, endian: Endian }`. `Copy + PartialEq + Eq + Hash`
(all fields are `Copy + Hash` once the sub-step above is done —
`AlkTypeKind` already derives `Hash`; `Endian`/`VariableEncoding`
get it from the sub-step above).
- `OffsetMap` storage: `fields: Vec<(String, ByteRange, LeafMeta)>`
(was `Vec<(String, ByteRange)>`). The `compute` walk already resolves
each leaf's type; add the `LeafMeta` extraction at the point where
the leaf `ByteRange` is recorded.
- **`OffsetMap::get` return type (deferred decision 2):** change to
`get(field_path) -> Option<&OffsetEntry>` where `pub struct
OffsetEntry { range: ByteRange, meta: LeafMeta }`. Add
`OffsetEntry` to `lib.rs` re-exports. Callers that used
`map.get(path).unwrap().start` become
`map.get(path).unwrap().range.start`. Update `alktty`/`alkcall` call
sites (in-house).
- `engine.rs::read_field`/`write_field`: drop the
`BastDoc::new(&self.bast_doc, &self.root_name)?` +
`lookup_leaf_field(&doc, field_path)?` calls. Read `LeafMeta`
from `offset_map.get(field_path)?.meta`. The `kind`/`encoding`/
`endian` match arms in `read_field`/`write_field` are unchanged
(they already dispatch on `AlkTypeKind`/`VariableEncoding`/`Endian`).
- Remove `LeafFieldInfo` and `lookup_leaf_field` from `engine.rs`
(subsumed by `LeafMeta` on the map).
- `materialize_aligned` also uses `OffsetMap` — it currently calls
`offset_map.get(&path)?.start` for leaf reads. Update those call
sites to `.range.start`. The materializer's `resolve_typeref` calls
for composite walks stay (composites aren't in the offset map as
leaves; they're walked recursively). The `BastDoc` argument to
`materialize_aligned` is now owned (phase 3) — no signature change
beyond the lifetime drop.
- **Scope Boundary — aligned `materialize`'s `BastDoc` structure walk
(resolves review #005 L3):** `materialize_struct_aligned`
(`materialize.rs:451-521`) walks `BastDoc` to traverse
struct/array/record *structure*, using `OffsetMap` only for leaf
byte positions. This is the **permanent design for 0.3.0**, not a
deferral: after phase 3 the walk is over owned data (no re-parse,
not O(N²)), and aligned `validate_bytes` is one structure walk per
call (not per-field), so there is no perf driver analogous to review
#004's packed per-chunk gap. ADR-011 §"Out of scope" is half-true
here (aligned materialize takes `&OffsetMap` *and* `&BastDoc`) —
this note owns the decision: aligned materialize keeps walking owned
`BastDoc` for structure through 0.3.0. An `AlignedPlan` that
compiles the structure walk is **not** in scope; if a future bench
shows an aligned-mode hot loop, it gets its own ADR (tracked as an
open question, not a silent gap). Phase 7's `ValidationPlan` does
not change this — validation is value-domain, orthogonal to the
aligned structure walk.
**Verification:** `cargo test --release` — existing `offset_map.rs`
and `engine.rs` `read_field`/`write_field` tests pass (they go through
public APIs). `cargo clippy --all-targets -- -D warnings`. `cargo doc
--no-deps` (new public `LeafMeta`/`OffsetEntry`). `cargo build --target
wasm32-unknown-unknown --release`.
---
## Phase 6 — Fingerprinting `ReadPlan`/`OffsetMap` (ADR-012 §1, §4) — **DONE (2026-09-02)**
> **Status: implemented.** `#[derive(Hash, Eq)]` added to `ReadPlan`,
> `FieldPlan`, `CompositePlan`, `ReadKind`, `DiscriminatorPlan`
> (`ReadPlan`'s `schema: Arc<Value>` hashes fine — `serde_json::Value:
> Hash + Eq` under the pinned `preserve_order` serde_json) and to
> `OffsetMap` (`Clone` added alongside; its `LeafMeta`/`OffsetEntry`/
> `ByteRange` payload gained `Hash` in phase 5 / this phase). The
> POC's `VariantPlan`/`VariantKind` don't exist in the production
> shape (phase 1 dropped them). `fingerprint() -> u64` on both via
> `DefaultHasher` (deferred decision 3 resolved: std `DefaultHasher`,
> no new dep; the fingerprint isn't hot; cross-version stability is a
> non-goal per ADR-012). Fingerprint contract tests on both: same
> schema twice → equal `PartialEq` + equal fingerprint; field-kind
> change, field-order change, and endianness change each → different
> fingerprints; (ReadPlan) different root names over the same document
> → different fingerprints. `ValidationPlan` already carries its own
> `Hash + Eq` + `fingerprint` + contract test (phase 7).
**Goal:** Add `Hash + Eq` derives + `fingerprint() -> u64` to `ReadPlan`
and `OffsetMap`. Enables cross-run caching, `alkcall` schema handshake,
schema-version diagnostics. (`ValidationPlan` gets the same treatment
in phase 7, where it's built — it carries its own `Hash + Eq` +
`fingerprint()` as part of its public surface.)
**ADR reference:** [ADR-012 §1](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#1-fingerprinting--readplan-hash--eq-offsetmap-hash--eq),
[ADR-012 §4](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#4-fingerprinting-offsetmap-bundled-with-2b).
**Files:** `src/read_plan.rs` (derives + `fingerprint`), `src/offset_map.rs`
(derives + `fingerprint`), `src/lib.rs` (no new re-exports — `Hash`/`Eq`
are trait derives, `fingerprint` is an inherent method).
**Implementation notes:**
- `ReadPlan` already uses `BTreeMap` for `by_name` (phase 1), so
`#[derive(Hash, Eq)]` works. Add it alongside the existing
`Debug, Clone, PartialEq`. Same for `FieldPlan`, `CompositePlan`,
`ReadKind`, `DiscriminatorPlan`. (The POC's `VariantPlan`/`VariantKind`
are not in the production shape — phase 1 dropped them — so they are
not derived here.) `ReadPlan` also carries `schema: Arc<Value>` from
phase 2; `Arc<Value>: Hash + Eq` because `serde_json::Value: Hash +
Eq` (with `preserve_order`, `Map::hash` sorts keys deterministically),
so the `schema` field does not block the derive. If a future
`serde_json` version regresses `Value: Hash`, exclude `schema` from
the derived `Hash` via a manual `impl Hash for ReadPlan` that hashes
every field except `schema` — phase-6 concern, not a blocker.
- `OffsetMap` already carries `LeafMeta` (phase 5), and `LeafMeta` is
`Copy + Hash + Eq` (phase 5 added `Hash` to `Endian`/`VariableEncoding`).
Add `#[derive(Hash, Eq)]` to `OffsetMap`,
`OffsetEntry`, `ByteRange` (already `Eq + Hash`), `LeafMeta`.
- **Fingerprint hasher (deferred decision 3):** `DefaultHasher` (std,
no new dep). The fingerprint isn't hot; cross-version stability is a
non-goal. `fingerprint()`:
```rust
pub fn fingerprint(&self) -> u64 {
use std::hash::{Hash, Hasher};
let mut h = std::hash::DefaultHasher::new();
self.hash(&mut h);
h.finish()
}
```
- **Fingerprint contract test:** compile the same schema twice, assert
`plan1 == plan2` and `plan1.fingerprint() == plan2.fingerprint()`.
Compile a schema with one field changed, assert fingerprints differ.
This is the contract test for ADR-012 §1's "two plans with equal
hashes produce identical reads over identical bytes."
**Verification:** `cargo test --release` (new contract tests).
`cargo clippy --all-targets -- -D warnings`. `cargo doc --no-deps`.
---
## Phase 7 — `ValidationPlan` (ADR-012 §3) — **DONE (2026-08-31)**
> **Status: implemented.** The design session ran and the shape landed
> in `src/validation_plan.rs`. Summary of what was decided and built
> (full detail in ADR-012 §3a):
>
> - **Shape:** `ValidationPlan { root: ValidNode }` — a constraint tree
> with one arm per value-domain check (`Int`/`I64`/`Uint`/`U64`/
> `Float`/`Bool`/`Str`/`Bytes`/`Enum`/`Struct`/`Union`/`Array`/
> `Record`), `ValidField { name, node }`, `ValidVariant { key, node }`.
> `maxLength` baked into leaf nodes from the owning field at compile
> time. Unions compile to variant nodes only (the declared union
> fields are validated via the variant walk, matching the interpretive
> arm's dispatch-on-`__discriminator` semantics).
> - **`compile` signature:** `compile(&BastDoc) -> Result<Self,
> AlkTypeError>` (the plan table's `&str` param was vestigial).
> - **`bast_validation`:** the interpretive walker is *retired*
> (deleted, not just bypassed); `validate_value` survives as a
> compile-once-per-call wrapper over the plan (one-shot/diagnostic
> use); shared error helper retained.
> - **Engine:** `Arc<ValidationPlan>` built at `compile` in both modes;
> new accessor `validation_plan()`. `validate_bytes` walks the plan.
> **Bonus:** `ValidationPlan::compile` runs before the layout build and
> serves as the engine's cyclic-`$ref` gate. (As of the review #006 H2
> fix, the layout walkers also carry their own guard —
> `walk_guard::check_ref_graph` runs at each standalone entry — so this
> ordering is now belt-and-suspenders rather than the only defense; the
> plan text below predates that fix.)
> - **Remaining phase-7 bench work** (a `validate_bytes`-stream bench
> in alktty) moves with the bench work into phase 8; a spot check
> during development measured plan-validate at ~0.2 µs/call vs ~0.6
> µs for the compile-per-call one-shot it replaced.
**Goal:** Retire the interpretive `bast_validation` walk. Introduce a
`ValidationPlan` — a compile-once-walk-many compiled form over the
BAST document's value-domain constraints — built once at `compile`
time and walked by `validate_bytes` (both modes) per buffer instead
of re-walking `BastDoc`. Closes the latent perf cliff review #005 M3
flagged: after ADR-011 the *read* half of `validate_bytes` is
plan-fast, but the *validation* half still re-walks `BastDoc` per
buffer, which is hot on the `alkcall` read+validate-on-untrusted-stream
common case.
**ADR reference:** [ADR-012 §3](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md#3-validationplan--compile-once-validation-form).
**Predecessor for this phase:** a **design session** to scope the
concrete `ValidationPlan` shape (constraint representation,
`compile`/walk structure, `bast_validation` public-surface review)
**before** implementation begins. ADR-012 §3 fixes the decision (in
0.3.0, compiled form, no per-buffer `BastDoc` walk, `Hash + Eq` +
`fingerprint`) and lists what is *not* decided (the struct/enum
shape, the constraint descriptors, whether `validate_value` is
retired or kept as a wrapper). This phase implements whatever the
design session scopes; the contract below holds regardless of shape.
**Files:** New `src/validation_plan.rs` (the `ValidationPlan` type,
`compile`, walk entry points). `src/bast_validation.rs` (adopt the
plan; `validate_value` either becomes a thin wrapper over the plan
or is retired per the design session's call). `src/engine.rs`
(`compile` builds `Arc<ValidationPlan>` in both modes, stores it;
`validate_bytes` walks `&self.validation_plan` instead of
reconstructing a `BastDoc` for the validator — this removes the
temporary bridge from phase 2). `src/lib.rs` (re-export
`ValidationPlan`).
**Implementation contract (fixed by ADR-012 §3, independent of
shape):**
- **Compile-once-walk-many.** `ValidationPlan::compile` walks the
owned `BastDoc` once (phase 3 made it owned); `validate_bytes`
walks the `ValidationPlan` per buffer, never `BastDoc`. The only
consumers that walk `BastDoc` interpretively after this phase are
the one-shot `*::compile` paths (`ReadPlan::compile`,
`OffsetMap::compute`, `ValidationPlan::compile`,
`LayoutBuilder::new`).
- **Value-domain, not byte-position.** The plan carries constraint
descriptors (enum allowed-sets, integer range bounds, `maxLength`
caps, union variant keys, and any other value-domain checks
`bast_validation` performs today), keyed for dispatch against the
materialized `Value` tree. The shape is different from
`ReadPlan`/`OffsetMap`; the pattern (compiled form, immutable,
shared via `Arc`) is the same.
- **`Send + Sync`.** `ValidationPlan: Send + Sync` (immutable owned
data, no interior mutability) so `Arc<ValidationPlan>` shares from
the `Send + Sync` engine. Add a `static` bound assertion test
mirroring phase 1's `read_plan_is_send_sync`.
- **`Hash + Eq` + `fingerprint()`.** `ValidationPlan` derives
`Debug, Clone, PartialEq, Eq, Hash` and has
`fingerprint() -> u64` (same `DefaultHasher` implementation as
phase 6). The fingerprint contract generalizes: two validation
plans with equal hashes accept/reject identical `(bytes)`
identically. Add a fingerprint contract test (compile the same
schema twice, assert `plan1 == plan2` and `plan1.fingerprint() ==
plan2.fingerprint()`; change one constraint, assert fingerprints
differ).
- **Untrusted-input discipline.** `compile` surfaces malformed
schemas as `AlkTypeError::Schema` (AGENTS.md §3); overflow-safe
arithmetic (AGENTS.md §4). No `unsafe`, no `async`, no new deps,
wasm-clean (AGENTS.md §5–§11).
**What this phase does *not* include (shape-dependent, scoped by the
design session):** the concrete `ValidationPlan` struct/enum, the
constraint-descriptor representation, the `bast_validation`
public-surface decision (`validate_value` retire-vs-wrapper), and any
`AlkTypeError::Validation` variant changes. These are shape questions
the design session resolves; they are *not* a re-opening of the
"ship in 0.3.0" decision, which is fixed in ADR-012 §3.
**Verification:** `cargo test --release` — the existing
`bast_validation.rs` and `engine.rs` `validate_bytes` tests are the
primary validation (they drive validation through the public API and
must pass unchanged, confirming behavioral parity with the
interpretive walk). New unit tests for `ValidationPlan::compile`
covering every constraint kind. New `Send + Sync` assertion test.
New fingerprint contract tests. `cargo clippy --all-targets -- -D
warnings`. `cargo doc --no-deps` (new public type). `cargo build
--target wasm32-unknown-unknown --release` (`validation_plan.rs` is
wasm-relevant). **Re-run the alktty `wire_vs_bast` bench** and, if
the design session scopes one, a `validate_bytes`-on-untrusted-stream
bench alongside `wire_vs_bast` to confirm the validation half of
`validate_bytes` no longer dominates per-buffer.
---
## Phase 8 — Public API bump, docs, verification (ADR-011 step 6, ADR-012) — **DONE (2026-09-02)**
> **Status: implemented.** Version flipped 0.2.0 → 0.3.0; `lib.rs`
> re-exports complete (`ReadPlan` + sub-types, `LeafMeta`,
> `OffsetEntry`, `ValidationPlan` + sub-types from earlier phases).
> Docs: ADR-007 "Cost" section rewritten to the `Arc<ReadPlan>` cost
> (15.7 ns) with the old framing as a historical note (review #004 L2
> closed); ADR-011/012 status blocks flipped to implemented; the
> architecture README ADR table rows updated; `layout-engine.md`
> rewritten for the 0.3.0 surface (`SequentialReader` construction via
> the engine factory, `OffsetMap` `OffsetEntry`/`LeafMeta`/`fingerprint`
> public-types section, `OffsetMap::compute(&BastDoc)` owned signature);
> `SequentialReader` module doc now points at the engine factory.
> Reviews #004 and #005 status flipped to closed. Bench (alktty
> `wire_vs_bast`, re-run on the 0.3.0 tree): read p64 98 ns/chunk
> (hand-rolled 5.7 µs/stream — parity held from phase 2), read p4k
> unchanged, `alktype_layout_build` **180 ns** (was ~1.2 µs — the
> phase-4 owned-doc cache removed the per-build re-parse, ~7x),
> `sequential_reader_new` 15.7 ns (unchanged), write p64 −3%
> (37.9 µs), `engine_compile` 590 µs (unchanged; dominated by
> meta-schema validation). No `validate_bytes`-stream bench was added:
> the phase-7 spot check (~0.2 µs/call plan-validate vs ~0.6 µs
> compile-per-call) stands as the validation-half measurement; a
> dedicated bench remains a follow-up if `alkcall` profiling motivates
> it. Downstream: `alktty` compiles against the path dep unchanged
> (the bench uses `LayoutBuilder::new`/`build` and
> `engine.sequential_reader()` — no touched signatures);
> `alkcall` has no dependency yet.
**Goal:** Flip the version to 0.3.0, update `lib.rs` re-exports, update
the architecture docs (ADR-007 "Cost" rewrite, ADR-011/012 status flip
if not already, README ADR table), update in-house downstream
consumers, run the full verification block.
**ADR reference:** [ADR-011 §Public API change](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md#public-api-change-breaking--version-bump-to-030),
[ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md).
**Files:** `Cargo.toml` (version 0.2.0 → 0.3.0), `src/lib.rs`
(re-export `ReadPlan`, `LeafMeta`, `OffsetEntry`, `ValidationPlan`),
`docs/architecture/` (README ADR table, ADR-007 "Cost" section,
ADR-011/012 status), `docs/architecture/validation.md` /
`layout-engine.md` (mention `ReadPlan`/`LeafMeta`/`ValidationPlan`
where relevant), in-house downstream repos (`alktty`, `alkcall` —
update call sites for `OffsetMap::get`, `SequentialReader::new`,
`materialize_packed`, `BastDoc` owned, `validate_bytes` internal
change if any signature change surfaced in phase 7's shape work).
**Implementation notes:**
- **ADR-007 "Cost" section (L2 from review #004):** rewrite the
"re-parse on demand" paragraph to describe the `Arc<ReadPlan>` cost
and the owned-`BastDoc` cache. The factory decision itself stays
"Accepted." This is the last loose end from review #004.
- **`src/engine.rs:112-115` doc comment (L2):** rewrite the "re-parse
the typed tree on demand" comment to describe the compiled-form
architecture (`ReadPlan` for packed reads, `OffsetMap`+`LeafMeta`
for aligned, `ValidationPlan` for validation, owned `BastDoc` for
the builder and the `*::compile` paths).
- **`lib.rs` re-exports:** add `ReadPlan`, `LeafMeta`, `OffsetEntry`,
`ValidationPlan`. `BastDoc` and `Bast*` stay re-exported (signatures
changed in phase 3, names unchanged). `materialize_packed`/
`materialize_aligned` stay re-exported (signatures changed).
`SequentialReader` stays re-exported (`new` signature changed).
- **Downstream updates:** `alktty`'s bench (`benches/wire_vs_bast.rs`)
updates `SequentialReader::new` call + any `OffsetMap::get` usage.
`alkcall` updates similarly. Both are in-house path dev-deps; the
updates ride this release's commits (or follow-on commits in those
repos — they're separate repos, but the path dev-dep means a local
update is immediate).
- **`Cargo.toml` version bump:** `0.2.0` → `0.3.0`. The workspace
section added for the POC (`[workspace] members = ["poc/readplan"]`)
stays on the `readplan-poc` branch and is *not* merged to main — the
POC branch is derisking-only, like `bast-validator-poc`. If the POC
files ever merge to main, drop the workspace section (the POC is
disposable).
**Verification block (run all, all must pass):**
```bash
cargo test --release # full suite
cargo clippy --all-targets -- -D warnings
cargo doc --no-deps # new public types
cargo build --target wasm32-unknown-unknown --release # wasm-clean
cargo publish --dry-run --allow-dirty # before publish
```
Plus: **re-run the alktty `wire_vs_bast` bench** and record the
before/after numbers in the release commit message. The 400x gap
should close to within ~2–5x of hand-rolled (the `data_access` calls
are the same; the remaining gap is the `match` dispatch + `Arc` refcount
vs hand-rolled's direct calls). The SFTP-shaped union case (the one
ADR-011's framing argument cared about) should close further because
eager `$ref` resolution removes the `resolve_typeref_as_def` per-
variant dispatch cost.
---
## Cross-phase invariants
- **The tree builds and tests pass at every phase boundary.** No phase
leaves the crate in a non-compiling state. Phases 1 (add `ReadPlan`),
6 (add `Hash`/`Eq` derives to `ReadPlan`/`OffsetMap`), and 7 (add
`ValidationPlan`) are pure additions; phases 2–5 are rewrites that
must leave tests green; phase 8 is the bump/docs.
- **The POC on `readplan-poc` is the reference scaffold for phases 1–2.**
It is *not* merged to main; it stays on the branch as the derisking
record, like `bast-validator-poc`. If a phase 1–2 implementation
question arises about the plan shape, consult the POC. (The POC's
`VariantPlan`/`VariantKind` and its nested-union rejection are
**not** carried forward — phase 1's refined `CompositePlan::Union`
shape supersedes both; see phase 1.)
- **Review #004 is the closure target.** H1 → phase 2; packed M1 →
phase 2; aligned M1 → phases 4–5; L1 → phase 2 (falls out); L2 →
phase 8 (doc rewrite). The review's status flips to "closed" in the
phase 8 commit.
- **Review #005 is the closure target for the plan-spec issues.** H1
→ phase 1 (refined union shape in ADR-011 + plan); H2 → phase 2
(`Arc<Value>` on the plan); M1 → phase 1 (nested unions via
`CompositePlan` recursion, no behavioral drop); M2 → phase 5 (`Hash`
on `Endian`/`VariableEncoding`); M3 → ADR-012 §3 + phase 7
(`ValidationPlan` shipped in 0.3.0, deferral reversed); L1/L2/L3 →
phase 2 / phase 5 Scope Boundary; N1 → typo; N2 → phase 1
`Send + Sync` assertion test; N3 → Semver Contract table row.
- **No `unsafe`, no `async`, no new deps, no feature flags** (AGENTS.md
§5–§11). The owned-`BastDoc` refactor uses `String`/`Value`, not
`unsafe` self-referential tricks. `DefaultHasher` is std. Wasm-clean
throughout. `ValidationPlan` follows the same constraints.
- **`preserve_order` stays load-bearing** (AGENTS.md §8). The owned-
`BastDoc` refactor must not sort schema object keys anywhere; field
order in the `Value` still determines byte order in packed mode and
iteration order in both modes. `ValidationPlan::compile` inherits
this — value-domain checks that depend on field ordering (e.g. union
discriminator field lookup) respect `preserve_order`.
## What this plan is *not*
- **Not a disk-cache or wire-protocol spec.** The fingerprint contract
and method are in scope (phase 6 for `ReadPlan`/`OffsetMap`, phase 7
for `ValidationPlan`); downstream uses are the consumers' concern.
- **Not cross-version fingerprint stability.** Within-version only
(ADR-012). The fingerprint may change across versions if a new
`AlkTypeKind` variant is added; consumers cache within a version.
- **Not a perf bench.** The bench lives in alktty; this plan re-runs it
at phase 2 and phase 8 to confirm the gap closes. The plan itself
only asserts correctness/coverage.
- **Not an `AlignedPlan`.** Aligned `materialize`'s `BastDoc` structure
walk is the permanent 0.3.0 design (phase 5 Scope Boundary). An
`AlignedPlan` is out of scope; if a future bench motivates one, it
gets its own ADR.
-655
View File
@@ -1,655 +0,0 @@
# Plan: alktype fuzzing
Adopted from alkhttp's `docs/plans/fuzzing.md` (the pattern is operational
in six sibling crates: alkcall, alktty, alktunnels, alksocks, alkhttp —
each with the same `fuzz/` layout, the detached runner, and the
corpus-replay-as-plain-test gate). The rationale research lives in
alkcall's `docs/research/fuzzing.md` (tool landscape, comparable-crate
survey, the no-hosted-CI policy §7.9); this plan stays focused on what
alktype fuzzes and in what order.
Rationale for this crate in one paragraph (the detailed version applies
by reference from the two docs above): alktype is the binary engine the
alk* family consumes — alkcall's hub/spoke accepts BAST schema documents
from arbitrary internet peers, and those documents flow into this
crate's compile paths downstream; both untrusted-input shapes exist here
(attacker-shaped JSON BAST docs → `AlkTypeEngine::compile`, and
attacker-shaped byte buffers read according to a schema →
`validate_bytes` / `SequentialReader` / `tunion` dispatch /
`materialize`), and the byte side is hand-rolled decode
(`data_access`, `tunion` discriminators, indirect `{offset,length}`
pairs) — exactly the shapes where example tests miss off-by-one bugs.
The crate is fully synchronous, so targets are simpler than
alkcall/alkhttp's (no current-thread runtime shims anywhere).
**Status:** waves 1–3 implemented and verified (2026-09-30). The
pre-fuzzing inventory (§6) is verified against the code at 0.3.0.
All five targets and the full infrastructure are in-tree (commits
`ed41d77`, wave 1; `16b9023`, wave 2; `aef8d9f` + findings commits,
wave 3); the smoke campaigns ran clean (§5); the release-budget
campaigns across all five targets are complete (§5) — two targets
surfaced real findings: the wave-2 `read_opseq` packing bug (§6
candidate 6, fixed same-day) and three wave-3 `validate_pair`
findings (W3-1 harness pin, W3-2 upstream serde_json pin, W3-3 a real
engine bug — read_field/write_field misread aligned maxLength
reservations — fixed same-day with regression tests).
Progress log:
- **2026-09-30 — wave 3 landed.** Target 5 (`validate_pair`): the
two-input harness (10-lane schema menu incl. a raw-JSON-bytes lane
fused with the buffer), 48 committed seeds, decode-pin tests, and
the §3 target-5 invariants (mode agreement incl. the documented
aligned-rejection taxonomy, the materialize⇄validate_bytes verdict
lattice with verbatim error propagation, unknown-path echo,
non-finite-float Access pin, enum-Validation pin, record spin
bound). Pre-campaign hand-drives all held. **Release-budget
campaigns (§5):** 45–46 min across all five targets. bast_compile
818k execs / 19,894 edges (still growing at budget end);
data_access 11.3M execs / 684 edges (a value-profile run lifted the
saturated corpus from 379 to 684 edges); read_opseq 63k execs /
17,519 edges; layout_build 63k execs / 17,659 edges; validate_pair
37k execs / 10,850 edges. **Finding W3-1 (harness invariant
corrected — the fuzzer fired an over-assertion):** the harness
claimed validate_bytes Ok ⇒ every offset-map leaf's range.end ≤
buffer.len() — false for offset-indirect entries, whose pair points
absolutely into the buffer while a maxLength window may dwarf the
validated buffer; the wave-1 data_access bounds partition is the
real contract. Fixed the invariant, pinned the artifact bytes as
seed-044 + a named regression test (commit `9ca9922`).
**Finding W3-2 (upstream, pinned with slack — no alktype bug):**
serde_json's non-`float_roundtrip` parser drifts one ulp when
re-parsing its own emitted shortest repr of adversarial f64 values
(probe: 0x5bffffffffffffff emits 1.4536774485912136e+135 and
parses back one ulp low; std's parser and ryu's own float parse are
exact — the concise reparse is the drift). The harness replaced
bare `Value` equality with a one-ulp structural comparison; the
artifact bytes are seed-047 (commits `a8e955c`/`b7ead99`).
**Finding W3-3 (real engine bug — fixed, commit `a0dd3d2`):**
`AlkTypeEngine::read_field`/`write_field` treated an aligned
`maxLength` reservation (ADR-003 strategy 2, `VARCHAR(N)`: raw
zero-padded window, NUL-trimmed on read — exactly what the
materializer and `validate_bytes` implement) as length-prefixed,
parsing the window's first four raw bytes as a u32 length. Every
aligned schema declaring `maxLength` broke the
validate_bytes⇒read_field lattice whenever the reservation's first
bytes looked like a large prefix (validate Ok, read_field Access
with bogus bounds; write_field wrote prefix+data into a raw
window). Third crash artifact (W3-1's shape family: the campaign
re-found the disagreement space after W3-1's invariant was
corrected). Fix: `VariableEncoding` gains `MaxLengthReserved`
(additive variant, ADR-003 strategy 2); `OffsetMap::compute`
records it for maxLength fields with the default encoding
(`maxLength`+`offset-indirect` stays `OffsetIndirect`, preserving
W3-1's combination semantics); `read_field`/`write_field` dispatch
through new `data_access::read_reservation{,_string}/
write_reservation` (single source of truth with the materializer);
three engine regression tests + the W3-1/W3-2 artifacts as
committed corpus seeds. Post-fix restart: 30-min validate_pair
campaign clean to budget end (37k execs, 10,850 edges, exit 0).
- **2026-09-30 — wave 2 landed.** Targets 3 (`read_opseq`) and 4
(`layout_build`) with 73 committed seeds (58 + 15), decode-pin tests
for the arbitrary 1.4.2 derive encoding, and the plan §6 candidate-1
spin bound encoded as the explicit End-op assertion. Smoke campaigns:
see §5. **Finding W2-1 (fixed same session):** `plan_read_array`
returned `Ok` for a fixed-stride array whose declared window
(count × stride) extended past the buffer — the bounds check was
missing entirely from the array arm (struct/union arms had theirs).
A truncated array reported success with the failure deferred to the
*next* field read (wrong field path), or masked entirely when the
array was the last field. Found by hand-running the target-3 drive
before the campaign (the first semantic fixture); fixed in
`plan_read_array` with an end-vs-buffer bounds check + Access error
naming the array field, regression test
`array_truncated_below_fixed_stride_window_is_access_error_not_ok`.
**Finding W2-2 (pinned, not a bug):** packed mode compiles
`"encoding": "offset-indirect"` fields but the sequential reader
always reads them inline length-prefixed — this matches
bast-format.md §"Default strategy selection" ("Packed sequential
mode: always inline length-prefixing"), so the annotation is a
no-op in packed mode. Pinned as a corpus-replay invariant
(`packed_mode_is_always_inline_length_prefixed`) so any future
change to the packed reader's encoding awareness is deliberate. The
aligned materializer honors the encoding correctly. An upstream
question — should packed compile either honor the encoding or reject
the declaration — is recorded in §6 candidate 7.
- **2026-09-30 — wave 1 landed (commit `ed41d77`).** The `fuzz/`
workspace, 136 committed seeds (38 `bast_compile` + 98
`data_access`), the detached runner, the corpus-replay gate
(AGENTS.md checklist gains the line), nightly pinned subtree,
explicit root `[workspace]` exclusion, publish-exclude gain,
`json.dict`. `cargo fuzz build` clean; corpus replay 4/4 green;
main crate untouched (569 tests, clippy `-D warnings` clean). One
dict-format fix on the way: libFuzzer's dictionary parser does not
accept `\u` escapes (`"\u0000"` → the `\xAB` form) and needs fully
quoted lines — caught by the campaign launcher, not the fuzzer.
- **2026-09-30 — smoke campaigns clean (see §5).** `data_access`
saturated (pure decode core, the alkcall `chunk_header` profile);
`bast_compile` still discovering coverage at budget end (longer
campaigns keep paying). No crate findings — the §6 candidates
(record-count loops, indirect pairs) held under the parser-level
drives; both remain encoded as wave-2/3 invariants in the stateful
targets.
- **2026-09-30 — wave-2 smoke campaigns (see §5).** Results recorded
there alongside the wave-1 numbers.
---
## 1. Why alktype fuzzes (the if)
1. **Downstream of the trust boundary.** alktype is compiled against in
alkcall (the integration crate), whose peers are untrusted and whose
wire payloads carry schema-shaped JSON. A panic on a malicious BAST
doc or bytes read under one is the quinn-CVE class
(RUSTSEC-2026-0037) at one further hop: the alk* stack parses
documents it never vetted, and alktype is where they get walked.
2. **Both input shapes, one crate.** Sibling crates each had mostly one
parse shape (wire bytes); alktype has the schema-JSON shape *and*
the raw-buffer shape, plus two-input combined paths
(`read_field`/`write_field`, `materialize_aligned` are doc+bytes).
3. **Infrastructure is proven and cheap; the crate is the simplest
consumer yet.** Six siblings run the layout; alktype is sync, has
zero `unsafe`, zero `unwrap`/`expect` outside tests, and no
allocation-from-wire-count anywhere (grep-verified inventory). The
marginal cost is target logic only.
**Honest caveat (alkcall §1's shape):** the code is already well
hardened — `checked_add`/`check_bounds` everywhere, parse-time caps
(`MAX_ARRAY_ELEMENTS`, `MAX_ARRAY_BYTES`, `MAX_ALIGN` = 4096,
`MAX_LENGTH`), `MAX_GRAPH_DEPTH`/`MAX_COMPILE_DEPTH` = 128, meta-schema
gate before any walker. Expected yield is low-moderate: the residual
candidates in §6 are the first things to probe; a clean first campaign
is the successful negative result — "we think the engine is robust"
converted into a demonstrated property.
## 2. Infrastructure (identical to the siblings)
Layout (copy of alkcall/alkhttp):
```
fuzz/
├── Cargo.toml alktype-fuzz (nightly-only bins; own [workspace])
├── rust-toolchain.toml pins nightly + llvm-tools for this subtree only
├── fuzz_targets/ thin fuzz_target! wrappers (3 lines each)
├── shared/ alktype-fuzz-shared — STABLE-toolchain library:
│ invariant logic + corpus-replay tests
├── corpus/<target>/ committed seeds (generated by gen_fuzz_seeds.py)
├── artifacts/ gitignored crash/oom/timeout artifacts + logs
├── gen_fuzz_seeds.py deterministic seed generator (quiche pattern)
├── json.dict JSON/BAST token dictionary (bast_compile)
├── run-detached.sh detached campaign runner (copied from the siblings)
└── README.md operational cheat-sheet
```
Load-bearing details (all six siblings hit these; alkcall's doc is the
deep reference):
- **Invariant logic lives in `fuzz/shared/`**, not the target binaries.
The stable-toolchain shared crate replays every committed seed through
the identical invariant functions as plain `cargo test` — the standing
fuzz gate (alkcall §7.9 tier-3 deliverable; no hosted CI in this repo
by policy). The `fuzz_target!` binaries are thin wrappers.
- **Root `Cargo.toml` needs an explicit `[workspace]` table**
(`members = ["."]`, `exclude = ["fuzz"]`); without it auto-discovery
pulls `fuzz/shared/` into the main workspace and the stable toolchain
builds nightly-consumed dev-deps. alkcall hit this trap.
- **`fuzz/rust-toolchain.toml` pins nightly** so `cargo fuzz build`
works from any CWD; nightly stays confined to `fuzz/`, MSRV 1.85
untouched here. `fuzz/` joins the publish `exclude` list.
- **`.gitignore` additions**: `fuzz/artifacts/`, grown-corpus dirs
(committed seeds stay).
- **Detached runner (non-negotiable operating rule).** Campaigns never
run as a foreground child of an agent session; the runner pins
`-fork=1 -rss_limit_mb=2048 -malloc_limit_mb=2048 -timeout=25` and
detaches via `setsid` + `nohup` + log redirect; the agent polls the
log and artifact directory, never waits. Copied from the siblings.
- **No feature-gating needed in `fuzz/shared`**: alktype has
`default = []` and no feature flags, so the shared crate rides the
main crate build unconditionally (unlike alkhttp's gated
`openapi`/`mcp`).
- **Exposure needs are minimal.** The inventory found every target
entry point already `pub` (`compile`, `data_access::*`,
`SequentialReader`, `LayoutBuilder`, `tunion::*`, `materialize::*`,
`validate_bast_doc`, `build_validator`). No `#[cfg(fuzzing)]` hub is
expected — the first choice remains a minimal `fuzzing` hub only if a
needed item turns out `pub(crate)`, per the sibling pattern (alkcall
never needed one).
**Verification-gate change:** `cargo test --manifest-path
fuzz/shared/Cargo.toml` (corpus replay) joins AGENTS.md's verification
checklist, as the siblings did.
## 3. Target inventory (5, in waves)
| # | Target | Drives | Input style | Status |
|---|---|---|---|---|
| 1 | `bast_compile` | `AlkTypeEngine::compile` both modes (via `bast_meta` → `BastDoc` → plans → layout → validator) | raw bytes → serde_json → BAST doc | implemented 2026-09-30 |
| 2 | `data_access` | the hand-rolled decode core (`src/data_access.rs`, read + write side) | raw `&[u8]` + chosen (offset, endian) | implemented 2026-09-30 |
| 3 | `read_opseq` | stateful `SequentialReader` op sequences over hostile bytes under a fixed plan | `#[derive(Arbitrary)]` op enum | implemented 2026-09-30 |
| 4 | `layout_build` | `LayoutBuilder::build` with adversarial `var_sizes` | `#[derive(Arbitrary)]` map shapes | implemented 2026-09-30 |
| 5 | `validate_pair` | two-input structured: compile a schema once per exec, hammer hostile bytes through `validate_bytes`/`read_field`/`materialize` | `#[derive(Arbitrary)]` (doc, bytes) pair | implemented 2026-09-30 |
### Target 1 — `bast_compile` (the whole schema side, one choke point)
`AlkTypeEngine::compile` (`src/engine.rs:127-176`) fans out through the
entire untrusted-JSON surface: `bast_meta::validate_bast_doc` →
`BastDoc::new` → `ValidationPlan::compile` → `LayoutBuilder::new` +
`ReadPlan::compile` (packed) or `OffsetMap::compute` (aligned) →
`validation::build_validator` (when `json_schema` is `Some`).
- **Drives:** raw bytes → `serde_json` → `compile(value, root_name,
mode, None)` in both modes; a second lane feeds `Some(schema)` with a
second attacker-shaped JSON value for the jsonschema-build path.
- **Invariants:**
- no-panic on any JSON document, both modes;
- compile is always `Result` — every rejection is a clean
`AlkTypeError` (`Schema`/`Offset`/`Validation` payload classes,
`src/error.rs:11-28`), never a panic or a silent bogus engine;
- meta-schema gate ordering: any doc that fails
`validate_bast_doc` must surface `Schema(...)` and must never reach
layout/plan walks (shape partition);
- parse-time caps hold exactly: `align > 4096`, arrays above
`MAX_ARRAY_ELEMENTS`/`MAX_ARRAY_BYTES`, `maxLength >
MAX_LENGTH`, depth > 128, and `$ref` cycles all reject at compile
with the documented error classes (the walk-guard
`check_ref_graph`, compile-depth, and cycle-`seen` machinery
pinned by adversarial corpus entries);
- if compile fails in packed it must also fail in aligned (mode
independence of the schema-gate layer — the parse layers are
shared; divergence means a mode-specific parse bug);
- a successfully compiled engine's `endian()` equals the root
struct's declared endianness.
- **Seeds:** the full BAST feature menu (each kind, endian ×2, TUnion
byte/field/enum discriminators, records, arrays, string/bytes
encodings, `$ref` diamond), each reject-class corpus entry, plus the
hostile menu in §4.
### Target 2 — `data_access` (the decode core)
Every byte-touching decode funnels through `read_array<const N: usize>`
(`src/data_access.rs:48-75`): `checked_add(N)` → `check_bounds` →
`.get(..)` → `try_into`. The widest attacker-influenced values in the
crate are `read_bytes_indirect`'s absolute `{offset,length}` pair
(`src/data_access.rs:328-350`).
- **Drives:** the `pub` read functions directly with the fuzzer
choosing buffer, offset (including far-past-end and huge values),
and endianness; lanes for `read_bytes`/`read_string` (u32 length
prefix), `read_bytes_indirect`/`read_string_indirect` (the
`{offset,length}` pair), `read_enum`, `read_bool` strictness, and
each fixed-width kind from the macro family.
- **Invariants:**
- no-panic for any (buffer, offset, endian) triple;
- `bool` accepts exactly 0x00/0x01 and rejects everything else
(`:134-144` — the strictness is contract, pin it);
- invalid UTF-8 in `read_string` errors (`Access`), never a lossy
silently-corrupting parse (`:195-208`);
- bounds partition: an error implies `checked_add`-overflow or
`end > buffer_len` with the offending `field_path` named; an Ok
implies the field sits fully inside the buffer;
- nothing before/end-of-buffer is read: the decode consumes
exactly its declared width (offset unchanged on error paths);
- `read_bytes_indirect`'s data region always satisfies
`data_offset + data_length ≤ buffer_len` on Ok, and neither
field can push arithmetic past the buffer without an error
(the two `u32` widening casts at `:334, :341` widening-only,
verified by the partition).
### Target 3 — `read_opseq` (stateful, wave 2)
`SequentialReader` is stateful against attacker bytes (mutable cursor:
`field_index`, `position`; `src/sequential_reader.rs:144-148`) and
fuzzer-reachable operations are `read_next`, `read_next_borrowed`,
`read_field` (out-of-order names), `reset` (`:157-300`).
- **Drives:** `#[derive(Arbitrary)]` op sequences (Next, Field(name
choice), Reset, End) against a compiled plan — the plan built once
per exec from a fixed small schema menu, bytes adversarial.
- **Invariants:**
- no-panic over any op interleaving and any buffer;
- cursor discipline: a failed read leaves the reader usable (a
subsequent `reset` restores the exact initial state; cursor never
exceeds the buffer);
- `read_next` returns fields exactly in plan order and `None`
exactly at plan end; interleaved `read_field` for any field at
any cursor state never panics and never mutates the sequential
cursor (its offset argument comes from the plan, not the reader);
- record-count spin bound: wire-controlled `count` loops
(`src/materialize.rs:439-442`, `:817-820`,
`src/sequential_reader.rs:985-988`) consume ≥ 4 verified bytes per
iteration, so iterations are bounded by
`remaining_bytes / 4` — a hostile count fails fast with `Access`
(encode as an explicit per-exec assertion, not just
no-panic/OOM);
- engine-issued readers are independent: two readers over the same
plan and buffer never observe each other's cursors
(ADR-007's owned-fresh-reader contract).
### Target 4 — `layout_build` (wave 2)
`LayoutBuilder::new` parses once (`src/layout_builder.rs:154-156`);
`build(&HashMap<String, usize>)` (`:189`) is repeatable with
attacker-shaped `var_sizes` driving write-position arithmetic in
`walk_struct`.
- **Drives:** a fixed schema menu containing every variable-width
encoding × `#[derive(Arbitrary)]` `var_sizes` maps and write values
(`FieldValue` shapes).
- **Invariants:**
- no-panic across adversarial size maps (zero, huge, mismatched
with `max_length`/`count` declarations);
- every failed write leaves the buffer untouched (byte-equal to the
pre-call snapshot) or documented-partial exactly where the
contract allows — pin the actual contract the code implements;
- field positions from a successful `build` are disjoint and
in-bounds for the reported total size;
- `data_offset/length` pairs written by
`write_string_indirect`/`write_bytes_indirect` always satisfy the
read-side `read_*_indirect` bounds partition above — the write
side and the read side of the pair are one contract
(round-trip pair; `:373-417` guards verified by
`:733-750`-style assertions).
### Target 5 — `validate_pair` (two-input structured, wave 3)
The integration target: schema and bytes are both adversarial.
- **Drives:** `#[derive(Arbitrary)]` (doc, bytes) — compile once per
exec with whichever mode the fuzzer picks, then drive
`validate_bytes`, `read_field` (arbitrary field paths, including
junk paths), `materialize_packed`/`materialize_aligned`, and
`read_next` under the compiled plan.
- **Invariants:**
- no-panic for any (doc, bytes) pair, either mode;
- validate/read/materialize agreement lattice: `validate_bytes` Ok
⇒ `materialize_*` Ok and every `read_field` over a declared path
Ok; `materialize_*` error ⇒ `validate_bytes` error on the same
buffer (exact agreement direction pinned per the code's actual
contract — determine the strict/loose ordering from the
`validate_bytes` implementation, don't assume);
- non-finite floats (NaN/Inf) surfaced by `materialize` are always
`Access` errors, never silently `Null`/`0.0`
(`src/materialize.rs:876-883`);
- unknown field-path strings always error with `Access` naming the
path, never panic, never index the map by substring drift;
- `Value` output is serde-safe: `materialize_*` output round-trips
through `serde_json::to_vec` and back to a structurally equal
`Value` (structural only — this crate's serde_json builds with
`preserve_order`, so object key order is preserved; byte-identity
round-trips are acceptable only where the docs say lossless,
per the alkcall §7.3 false-positive trap when they don't).
## 4. Corpus policy
Committed hand-made seeds per target, generated by
`fuzz/gen_fuzz_seeds.py` (deterministic, in-tree, quiche pattern);
grown corpora and artifacts gitignored. Seed menus:
- `bast_compile`: a minimal valid packed doc and aligned doc; every
`AlkTypeKind` once; each reject class (`align` 4097/65536/u32-max,
`count` over cap, `maxLength` over cap, depth-129 nesting both
inline-nested and via `$ref` chains, `$ref` cycle, `$ref` to
missing def, root not a struct, missing `type`, unknown kind
string, duplicate field names first-wins probe, non-object doc,
deeply-nested JSON at serde_json's own 128 limit); `json.dict`
carries the BAST token set.
- `data_access`: minimal valid encodings per kind per endianness;
truncation at every prefix length (1..N-1 for each width);
`len = 0` / `MAX_LENGTH` / `u32::MAX` prefixes; the indirect pair at
{0,0}, {len, big}, {big, 0}, {u32::MAX, u32::MAX}; offset one-past-
end, offset u32-magnitude; 0x02 bool byte; invalid UTF-8 in a
string; enum value out of range; NaN/Inf bytes.
- `read_opseq` / `layout_build` (wave 2, **done 2026-09-30**): the
semantic fixtures — full sequential walk, reset-mid-walk then full
walk again, failed read then reset, record loop with a hostile count
under a real buffer, junk field paths; zero/huge/mismatched
`var_sizes`; overwrite-everything write; indirect-pair overflow
write, plus the write-then-read pair fixture. `read_opseq` seeds
also encode truncation at every prefix of the full-walk buffer. The
hand-written seeds are byte-encoded against the pinned arbitrary
1.4.2 derive layout and *pinned by decode tests* (both targets carry
a `decode_lands_on_the_intended_variants` replay test, the alkhttp
target-4 pattern).
- `validate_pair` (wave 3, **done 2026-09-30**): hostile-schema/
valid-bytes, valid-schema/hostile-bytes, valid/valid — plus the §6
candidate shapes as pinned reproducers. 48 committed seeds incl.
the W3-1/W3-2 artifact bytes, the aligned maxLength fixtures, and
the mode-agreement pins (ADR-006/ADR-008 lanes).
Stateful `Arbitrary` seeds are hand-encoded against the `arbitrary`
1.4.x derive layout with per-element keep-going bytes, pinned by
seed-decode tests (alkhttp's target-4 pattern).
## 5. Campaign + gate policy
Identical to the siblings (alkcall §7.9 posture; no hosted CI in this
repo):
- **Corpus replay is the standing fuzz gate:** `cargo test
--manifest-path fuzz/shared/Cargo.toml` — joins AGENTS.md's
verification checklist.
- **Campaigns run detached** via `fuzz/run-detached.sh`; budget 10 min
per target for a smoke campaign, 30–45 min before a release or after
touching `src/data_access.rs`, `src/schema.rs`, the compile walks, or
the sequential reader.
- **Grown corpora stay gitignored** (hash-named files ignored via
pattern; committed `seed-*` files stay); merge worthy entries into
seeds only deliberately.
- libFuzzer flags worth pinning: `-rss_limit_mb=2048`,
`-malloc_limit_mb=2048`, `-timeout=25`, `-max_len=65536`,
`-use_value_profile=1`, and `-dict=json.dict` on the JSON targets
(alkcall §7.2's set, minus the CI-tier concerns).
- Nightly stays confined to `fuzz/`; the main crate's stable build,
MSRV, and wasm target are untouched — `cargo fuzz build` must never
be a prerequisite for `cargo test`/`clippy`/`build`.
Release-budget campaign results (2026-09-30, 45–46 min per target,
§5 policy, detached, `-use_value_profile=1`, `-dict=json.dict` on the
JSON lanes): all five exited 0.
- `bast_compile` — 818k execs at ~300–500/s (2,756 s), coverage
19,894 edges / 86,455 features / 4,439 in-memory corpus entries,
**still growing at budget end** (2× the wave-1 45-min smoke edge
count). Longest campaigns keep paying on the schema side.
- `data_access` — 11.3M execs at ~3.7–4.7k/s (2,761 s), coverage 684
edges / 6,208 features — value-profile lifted the "saturated" 379
edges to 684 (the wave-1 ceiling was the no-profile ceiling).
- `read_opseq` — 63k execs at ~23/s (2,737 s), coverage 17,519 edges /
65,059 features / 1,668 entries; the stateful search kept adding
features across the full budget (assertion-throughput-bound).
- `layout_build` — 63k execs at ~23/s (2,751 s), coverage 17,659
edges / 71,477 features / 1,873 entries; write-side contracts held.
- `validate_pair` — three crashes over two runs before and one clean
run after the W3-1/W3-2/W3-3 fixes (final post-fix campaign 30 min):
37k execs (final 1,820 s run), coverage 10,850 edges / 38,364
features / 1,544 entries, `oom/timeout/crash: 0/0/0` at budget end.
The two-input harness is the highest-yield target in the crate: 1
real engine bug + 2 pinned contracts in its first campaigns.
- Artifact totals: the three `validate_pair` crash artifacts are all
pinned as committed seeds/regression tests (W3-1 `seed-044`, W3-2
`seed-047`, W3-3 reproduced by the aligned maxLength fixtures);
`bast_compile`/`data_access`/`read_opseq`/`layout_build` artifacts
empty. No OOM, timeout, or leak on any fork job of any target.
Wave-2 smoke campaign results (2026-09-30, 10 min per target, §5
policy, detached):
- `read_opseq` — 52.1k+ execs at ~90/s, exit 0, empty artifact dir,
no oom/timeout/crash on any fork job. The typed-input decode plus
the per-op assertion work makes this the slowest target per exec in
the crate so far; coverage ~9,365 edges / 15,797 features / 293
in-memory corpus entries. The heavy semantic invariants
(replay-after-failure, full-walk spin bounds, reader independence,
plan-order assertion) run per op, not per exec — the campaign is
assertion-throughput-bound, not coverage-bound; the stateful search
(op × cursor × buffer shape) was still adding features at budget end.
- `layout_build` — 42.4k+ execs at ~84/s, exit 0, empty artifact dir,
no oom/timeout/crash on any fork job; coverage 9,432 edges / 19,204
features / 181 in-memory corpus entries. The adversarial `var_sizes`
map space over five schema menus exercised the missing-size /
unknown-discriminator / overflow rejection paths; the write-side
contracts (failed write leaves buffer byte-identical, positional
disjointness and bounds) held everywhere.
Wave-1 smoke campaign results (2026-09-30, 10 min per target, both
exited 0, artifact dirs empty — no crash/hang/OOM/leak):
- `bast_compile` — 543k execs at ~1.1k/s (each exec compiles two
engines through the whole fan-out — meta gate, parse, plans, layout
walks — in both modes, so per-exec work is heavy), coverage 10,830
edges / 25,326 features / 1,293 in-memory corpus entries, **still
growing at budget end** — longer campaigns keep paying.
- `data_access` — 3.5M+ execs at ~7–8k/s, coverage saturated at 379
edges / 574 features / 37 corpus entries (the pure-decode-core
ceiling is fully enumerated; the alkcall `chunk_header` profile).
- Both exited 0 with empty artifact directories; `oom/timeout/crash:
0/0/0` on every fork job.
- **Seeds regenerate deterministically:** `python3
fuzz/gen_fuzz_seeds.py`.
- Toolchain notes live in `fuzz/README.md`; nightly stays confined to
`fuzz/`.
## 6. Pre-fuzzing candidate findings (confirm or refute)
These are pre-fuzzing code-review findings from the 0.3.0 inventory,
verified against the code. They define what the targets must encode as
invariants and are the first corpus entries to add; the fuzzer
confirms or refutes them.
1. **Wire-controlled `Record` count loops** (`src/materialize.rs:439-
442`, `:817-820`, `src/sequential_reader.rs:985-988`): the only
buffer-derived loop counts (`read_u32(..)? as usize` then
`for i in 0..count`). Each iteration performs at least one
bounds-checked read, so a hostile count should fail fast with
`Access` — bounded by `remaining_bytes / 4`, no allocation, no
spin. Correct as designed *if and only if* that holds; the
`read_opseq` target encodes it as an explicit assertion (§3
target 3) so the fuzzer can break it the moment any per-entry
cost stops being `≥ 4 verified bytes`.
2. **Attacker-controlled absolute `{offset,length}` pairs**
(`read_bytes_indirect`/`read_string_indirect`,
`src/data_access.rs:307-350`): the widest attacker-influenced
values in the crate (each `u32`, up to 2³²−1). The pattern
(widening cast → `checked_add` pair-sum → full bounds check)
looks correct; campaign confirms the bounds partition on every
(buffer, pair) input, both sides of the write/read contract
(§3 target 2 / target 4).
3. **Duplicate field-path tolerance by design** (`OffsetMap::build_
index`, `src/offset_map.rs:199-205`: first-wins; `BastStruct::
parse` does not reject duplicates): ambiguous lookups are
documented behavior. Encode as an invariant — a successful
engine's `read_field` resolves duplicates deterministically
(first wins) — so a future "reject duplicates" change shows up
as a deliberate contract change, not silent drift.
4. **`align_up`/`round_up` plain `+` arithmetic** (`src/offset_map.rs:
618-639`): un-checked `+ align - rem` in a field of
schema-controlled values. Overflow-infeasible today because
align is capped at parse (`MAX_ALIGN` = 4096) and the running
offset is monotonically checked — a `bast_compile` corpus entry
with align at the cap pinning the boundary keeps it that way if
the cap ever moves.
5. **`bast_validation::validate_value` recompiles a `ValidationPlan`
per call** (`src/bast_validation.rs:64-65`): a repeated-op DoS
surface if any consumer compiles-per-call. Not a bug in this
crate's API; fuzz targets compile once per exec, and the doc
records the compile cost as the consumer's responsibility.
6. **Missing bounds check in `plan_read_array` — CONFIRMED and FIXED
(wave 2, finding W2-1, 2026-09-30).** The struct arm
(`plan_bounds_check`) and the union fixed-size arm both verify the
computed end against the buffer; `plan_read_array`
(`src/sequential_reader.rs:898` pre-fix) computed `end` and
returned `Ok` without the check. A truncated fixed-stride array
reported `Ok(Some(...))` with `element_start + count × stride` past
the buffer; the failure surfaced at the *next* field (naming the
wrong field in the error), or never when the array was the last
field — a completed walk returning `Ok` over a short buffer. Found
by hand-running the `read_opseq` drive before the campaign; fixed
with an end-vs-buffer check returning `Access` naming the array
field; regression test in `src/sequential_reader.rs`. The wave-2
seeds' array-truncation fixtures keep the boundary pinned.
7. **Packed-mode `encoding: offset-indirect` is a silent no-op —
PINNED, open design question (wave 2).** `BastField::parse`
records the annotation; the packed reader
(`plan_read_primitive`) ignores it and reads inline
length-prefixed — correct per bast-format.md's "Default strategy
selection" table, but nothing rejects the declaration in packed
mode (aligned mode consumes it; aligned rejects it on record
fields only). The wave-2 replay test
`packed_mode_is_always_inline_length_prefixed` pins the current
behavior. If the packed reader ever grows encoding awareness, the
pin flips deliberately; if the format instead wants the
declaration rejected in packed mode, that is a schema-gate change
with the mode-agreement invariant to re-verify.
8. **Aligned `maxLength` reservation read paths disagreed —
CONFIRMED as a real engine bug and FIXED (wave 3, finding W3-3,
found by the `validate_pair` campaign).** The materializer and
`validate_bytes` implement ADR-003 strategy 2 correctly (raw
zero-padded window, NUL-trimmed), but `AlkTypeEngine::read_field`
dispatched every aligned String/Bytes leaf through the
length-prefixed `data_access::read_string`/`read_bytes` — parsing
the reservation window's first four raw bytes as a u32 length —
and `write_field` wrote prefix+data into the raw window. Any
aligned schema declaring `maxLength` whose first reservation bytes
looked like a large prefix broke the validate_bytes ⇒ read_field
lattice (validate Ok, read Access with bogus bounds). Fixed by
recording the strategy in `LeafMeta`'s encoding
(`VariableEncoding::MaxLengthReserved`, additive variant) and
dispatching read/write through new
`data_access::read_reservation{,_string}`/`write_reservation`
sharing the materializer's exact semantics. Regression tests in
`src/engine.rs` + `data_access.rs`; the W3-1 artifact bytes are
the committed reproducer (`seed-044`).
None of these rises to the alkcall §6.2 / alkhttp FWD-20 class; they
are boundary-confirmations, which is exactly the expected profile of
this crate (§1 honest caveat). Smoke-campaign evidence: no panics or
OOM/timeouts on any of the 383k+ wave-1/2 combined execs; the
release-budget campaigns added 12.3M+ execs with the W3 findings
above. Candidates 1, 2, and 4 held under the parser-level drives —
candidates 1 and 2 got their explicit stateful assertions in waves
2–3 (targets 3–5), and 4 keeps its boundary corpus entry. Candidate 3
(duplicate first-wins) is pinned by the wave-2 `layout_build` index
assertions and the existing crate tests; candidate 6 was confirmed as
a genuine bug and fixed in wave 2; candidate 7 is pinned with an open
design question; candidate 8 was confirmed as a genuine bug and fixed
in wave 3.
## 7. Sequencing
1. ✅ **Wave 1 (2026-09-30, commit `ed41d77`)** — infra (`workspace`
exclude, toolchain pin, runner, seed generator, README,
`.gitignore`) + targets 1–2 + corpora (136 seeds) + corpus replay +
AGENTS.md gate + smoke campaigns (clean, see §5).
2. ✅ **Wave 2 (2026-09-30)** — targets 3–4 (stateful `read_opseq`,
`layout_build`) + 73 seeds + decode-pin tests + smoke campaigns.
One real bug found and fixed first-session (`plan_read_array`
bounds check, §6 candidate 6); packed-mode offset-indirect pinned
as a documented no-op (§6 candidate 7).
3. ✅ **Wave 3 (2026-09-30)** — target 5 `validate_pair` (the
two-input structured harness) + 48 seeds + decode-pin tests +
45-min release-budget campaigns across all five targets. Three
findings: W3-1 (harness over-assertion corrected + pinned), W3-2
(upstream serde_json one-ulp f64 parse drift pinned with slack),
W3-3 (real engine bug — aligned maxLength reservation misread by
`read_field`/`write_field` — fixed with `VariableEncoding::
MaxLengthReserved` + `data_access::read_reservation*`/
`write_reservation`, commit `a0dd3d2`). Post-fix
validate_pair campaign clean to budget end. Fuzzing complete per
§2 scope: corpus replay 30/30 is the standing gate.
4. Everything else inherited verbatim: no hosted CI, OSS-Fuzz out,
no Actions/workflow files anywhere in the repo (alkcall §7.9).
## 8. References
- alkcall `docs/research/fuzzing.md` — rationale, tool landscape,
campaign containment (§7.6), no-hosted-CI policy (§7.9)
- alkhttp `docs/plans/fuzzing.md` — the live pattern this plan copies
(waves, findings log, `fuzzing` hub convention)
- RUSTSEC-2026-0037 / CVE-2026-31812 (quinn-proto) — the remote-DoS
class this crate's peers are exposed to
- Internal: ADR-002 (layout modes), ADR-004 (error/validation
strategy), ADR-006 (aligned-mode variable-field rejection), ADR-008
(aligned-mode TUnion rejection), ADR-010 (`validate_bytes`,
materialize-then-validate)
+2 -2
View File
@@ -1,6 +1,6 @@
---
status: closed
last_updated: 2026-09-02
status: open
last_updated: 2026-08-17
reviewed_artifacts:
- src/sequential_reader.rs
- src/bast.rs
-659
View File
@@ -1,659 +0,0 @@
---
status: closed
last_updated: 2026-09-02
resolved_findings: 2026-08-20 (all 11 — see "Resolution" at the end)
reviewed_artifacts:
- docs/plans/030-compiled-forms.md
- docs/architecture/decisions/011-compiled-read-plan-for-packed-mode.md
- docs/architecture/decisions/012-plan-fingerprinting-and-m1-closure.md
- docs/reviews/004-performance-review.md
- src/lib.rs
- src/bast.rs
- src/engine.rs
- src/sequential_reader.rs
- src/materialize.rs
- src/offset_map.rs
- src/layout_builder.rs
- src/schema.rs
- poc/readplan/{src/lib.rs, FINDINGS.md} (branch readplan-poc)
tool: manual source read + plan-vs-codebase cross-check + POC branch inspection
reviewer: 0.3.0 implementation plan review (triggered before phase 1)
---
# Review #005 — 0.3.0 Plan Review: Compiled Forms
## Purpose
The 0.3.0 implementation plan
([`docs/plans/030-compiled-forms.md`](../plans/030-compiled-forms.md)) is
the entry point an implementing agent reads first. It rolls up
[ADR-011](../architecture/decisions/011-compiled-read-plan-for-packed-mode.md)
(the `ReadPlan` packed read-side compiled form),
[ADR-012](../architecture/decisions/012-plan-fingerprinting-and-m1-closure.md)
(fingerprinting + owned `BastDoc` + `OffsetMap` `LeafMeta`), and the
fingerprinting work into one breaking bump. The plan is deliberately
structured as seven phases so each can be picked up by a fresh session
without prior context.
This review's purpose is to find planning-spec mistakes — factual
errors, contradictions, undocumented behavioral changes, hedges into an
unplanned future — *before* a phase-by-phase implementation starts,
because fresh-session implementations are reliable precisely when the
spec is accurate. A spec that contradicts the code or an ADR forces the
agent to either reverse-engineer the actual intent or guess, and the
failure rate goes up.
The review explicitly scans for the "deferral black hole" pattern: a
plan or ADR puts work off into a "future version/phase/downstream" with
no concrete reactivation condition, the next agent inherits the gap,
and the gap festers until something forces an untangle. This is a
known LLM-planning quirk distinct from classic planning mistakes, and
a default scan for it is part of this review's methodology.
## Methodology
- Full read of the plan and its two companion ADRs (011, 012), the
performance review (#004) the plan closes, and the POC findings on
branch `readplan-poc`.
- Cross-check every line-number reference and `src/` claim in the plan
against the actual codebase at `main` (commit `2310f6c`, v0.2.0).
Verified: `engine.rs:112-115,284,334,467`; `sequential_reader.rs:567`;
`layout_builder.rs:190`; `bast.rs:51-55`; `lib.rs` re-export list;
`Cargo.toml` version; presence of `poc/` (absent on main, present on
`readplan-poc` as expected); presence of `alktty`/`alkcall` downstream
path dev-deps.
- Cross-check the POC's `ReadPlan`/`CompositePlan` shape against both
ADR-011's shape section and the plan's phase-1 description.
- Cross-check the plan's phase 2 rewrite claims (`dummy_field_for`/
`ty_source` "are removed") against `materialize.rs`'s actual call
sites across both packed and aligned paths.
- Verify the derives the plan relies on (`Hash` on `LeafMeta`,
`ReadPlan`, `OffsetMap`) are reachable from the derives on their
constituent types in `src/schema.rs`.
- Scan for the deferral pattern by flagging every "future/deferred/
later/downstream/if needed" occurrence and asking: (a) is there a
concrete reactivation trigger? (b) is the decision owned or silent?
(c) does inaction have a cost that the deferral framing hides?
## Verification Baseline
The plan and both ADRs were read at the tree state at commit `2310f6c`
("Propose ADR-012 + 0.3.0 implementation plan"), which is `main` HEAD.
The codebase is v0.2.0 (`Cargo.toml`); the POC lives on branch
`readplan-poc` and is not merged, as the plan states. All line-number
references in the plan were verified correct against this tree.
## Summary Statistics
| Severity | Count |
|----------|------:|
| High | 2 (H1, H2) |
| Medium | 3 (M1, M2, M3) |
| Low | 3 (L1, L2, L3) |
| Nit | 3 (N1, N2, N3) |
The two High findings are correctness/contradiction issues that would
block or mislead an implementing agent. The Mediums are either
undocumented behavioral drops, missing implementation prerequisites, or
a deferral worth re-evaluating. Lows and Nits are wording/typo-level.
---
## Findings
### H1. Phase 1's field-name-discriminator union shape exists in neither ADR-011 nor the POC
**File**: `docs/plans/030-compiled-forms.md:144-152`
**Problem**: The plan describes the field-name-discriminator union read
shape as:
> `CompositePlan::Union` carries the union's declared `fields` as a
> sub-`ReadPlan` (the discriminator field + any shared fields), and the
> variant plans are laid out *after* the shared fields.
But ADR-011 §"The `ReadPlan` shape"
(`011-compiled-read-plan-for-packed-mode.md:144-158`) defines:
```rust
pub enum CompositePlan {
Struct(ReadPlan),
Union {
disc: DiscriminatorPlan,
variants: Vec<(String, VariantPlan)>,
},
...
}
```
There is no field for shared/declared fields on the `Union` variant.
The POC (`readplan-poc:poc/readplan/src/lib.rs`) matches the ADR's
shape — `CompositePlan::Union { disc, variants }` only — and its
`compile_union` does not carry shared fields. The POC's `FINDINGS.md`
Finding 1 (the same one the plan cites at lines 143-152) explicitly
says:
> `plan_read_union`'s `Field` arm is a stub that returns an error.
> ... The plan needs a sub-struct for the union's declared fields,
> separate from the variant plans.
So the plan describes a shape that exists in **neither** the accepted
ADR **nor** the reference POC, and presents it as "the production
version must implement it" within the existing `CompositePlan::Union`
shape. An implementing agent reading ADR-011 + plan + POC gets three
different `CompositePlan::Union` shapes and no guidance on where the
shared-fields sub-`ReadPlan` goes (a new `shared: Option<Box<ReadPlan>>`
field? a wrapper enum? two-variant split?).
This is a shape extension to an accepted ADR's public type. The plan
either needs to flag it as an ADR-011 refinement (with the ADR updated
first) or specify the concrete shape the agent should build.
**Lift**: unblocks phase 1. Without resolution, the agent will either
guess the shape and likely diverge from intent, or stop and ask.
---
### H2. `schema()` returning `&Value` from a `&Value` "stored on the plan" is a self-referential struct
**File**: `docs/plans/030-compiled-forms.md:188-191`
**Problem**: Phase 2 says:
> `schema()` returns a `&Value` retained on the plan (the plan stores
> the `&Value` it was compiled from — see ADR-011 §Engine integration;
> the `&Value` outlives the plan because the engine owns both).
The "the plan stores the `&Value` it was compiled from" is the
self-referential struct pattern ADR-011 §"Root cause"
(`011-...md:50-58`) explicitly identifies as impossible in safe Rust
and rejects. The engine owns `bast_doc: Value` and `Arc<ReadPlan>`. If
`ReadPlan` stores `&Value` borrowing from the engine's `bast_doc`, the
engine is self-referential — exactly the construction ADR-007 worked
around with "re-parse on demand" and ADR-011's `Arc<ReadPlan>` was
meant to retire. ADR-011 line 204 specifies the plan is "immutable
**owned** data"; it does not say the plan stores a `&Value`.
`schema()`'s current contract (`src/sequential_reader.rs:247`) is to
return the raw BAST `Value` the reader was built from. To preserve
that contract on `Arc<ReadPlan>` without a self-referential borrow,
`ReadPlan` must store an `Arc<Value>` (engine builds `Arc<Value>` at
compile time, hands a clone to the plan) or an owned `Value`. Then
`schema()` returns `&self.plan.value`. The plan should specify which.
**Lift**: prevents an agent from getting stuck in phase 2 trying to
make `&Value` in `Arc<ReadPlan>` work, which the borrow checker will
reject.
---
### M1. Nested unions: POC rejects a schema 0.2.0 accepts — undocumented behavioral drop
**Files**: `poc/readplan/FINDINGS.md` ("What this POC does not cover"),
`docs/plans/030-compiled-forms.md` (silent), `src/sequential_reader.rs:789-815`
**Problem**: The POC's `compile_union` rejects a union variant that is
itself a union with `AlkTypeError::Schema`. The existing reader
supports this: `resolve_and_walk_variant` at
`src/sequential_reader.rs:800` has a live `BastDefKind::Union` arm
that recurses via `read_union_value`. So 0.2.0 accepts and reads
nested-union schemas; phase 1's `ReadPlan::compile` (per the POC the
plan cites as the reference scaffold) would reject the same schema.
The plan's phase 1 calls out two POC findings explicitly (field-disc
union shape → H1 above, struct-array stride → deferred decision 4)
and says "the production version must implement/decide these." It does
**not** call out the nested-union rejection. An agent following the
plan would inherit the POC's reject-nested-unions behavior by default,
silently dropping a 0.2.0 capability — a behavioral regression that
rides the 0.3.0 bump without being listed in the Semver Contract
table.
This is also the cleanest example of the deferral-black-hole pattern
in the plan: the POC says "if a real schema needs it, the
implementation step adds a `VariantKind::Union` read path. Not
blocking — no current schema exercises it." The "if needed" framing
has no trigger, no OQ, no tracking — it's a black hole. The next agent
inherits the gap.
**Lift**: either (a) add `VariantKind::Union` read path in phase 1
(small — mirrors the existing `resolve_and_walk_variant` Union arm,
~20 lines), or (b) list it in the Semver Contract table as a
behavioral drop with a one-line OQ tracking the deferral. Given the
plan says there are zero real consumers, (b) is defensible, but it
must be *stated*, not silent. (a) is cheap and avoids the regression.
---
### M2. `Endian` and `VariableEncoding` don't derive `Hash` — phases 5/6 will not compile
**Files**: `src/schema.rs:205,212`, `docs/plans/030-compiled-forms.md:62,357,411-415`
**Problem**: `src/schema.rs:205` (`Endian`) and `:212`
(`VariableEncoding`) both derive only `Debug, Clone, Copy, PartialEq,
Eq` — no `Hash`. The plan requires:
- Phase 5 (line 62, 357): `LeafMeta { kind, encoding, endian }` as
`Copy + PartialEq + Eq + Hash`.
- Phase 6 (lines 411-415): `#[derive(Hash, Eq)]` on `ReadPlan`/
`OffsetMap`, and `FieldPlan` carries `endian: Endian` + `encoding:
VariableEncoding`.
Both derives will fail to compile: `#[derive(Hash)]` on a struct
requires all fields to be `Hash`. The plan never mentions adding
`Hash` to these two enums. The fix is trivial (both are fieldless
enums, already `Eq + PartialEq`, so adding `Hash` is semver-safe —
additive, no behavioral change), but it's a prerequisite the plan
omits. An agent working phase 5 will hit a compile error and have to
diagnose why.
**Lift**: trivial. Add a sub-step to phase 5 (or 6): "Add `Hash` to
`Endian` and `VariableEncoding` derives in `src/schema.rs`." This is
additive and safe to do earlier if convenient.
---
### M3. `ValidationPlan` deferral worth re-evaluating — the read+validate common case
**Files**: `docs/architecture/decisions/012-...md:64-73` ("Deferring
`ValidationPlan`"), `docs/plans/030-compiled-forms.md:526-528`,
`docs/reviews/004-performance-review.md` (the read-path perf review)
**Problem**: ADR-012 defers a `ValidationPlan` as "different shape
(value-domain, not byte-position), not a hot loop, separate ADR if a
bench motivates it." The plan inherits this deferral ("Not a
`ValidationPlan`" at lines 526-528). The deferral framing is "if a
bench motivates it" — a concrete trigger exists, so this is not a
black-hole hedge in the M1 sense.
Flagged for re-evaluation, not because the shape argument is wrong
(it's correct — value-domain checks are structurally different from
byte-position walks), but because the *hot-loop* dismissal may under-
account a common case: **read + validate together on untrusted input.**
Review #004 found the packed read path was 400x slow per chunk due to
per-field `BastDoc` re-parse. ADR-011 closes that. But
`validate_bytes`'s packed path (ADR-010) is `materialize_packed` →
`bast_validation::validate_value` over the materialized `Value`.
After ADR-011, `materialize_packed` walks the `ReadPlan` (fast).
`bast_validation::validate_value` still walks `BastDoc` to check
value-domain constraints — once per `validate_bytes` call, over the
full tree, on every buffer.
For a stream of N untrusted buffers (the `alkcall` hub/spoke topology
accepts schemas from arbitrary internet peers — AGENTS.md §3 — and
the common case is "read incoming frame, validate it before acting"),
`validate_bytes` is called N times. Each call does one `BastDoc`
walk for validation. After ADR-011, the *read* half of `validate_bytes`
is plan-fast; the *validation* half is still a `BastDoc` walk per call.
If validation is the common companion to read on untrusted input,
then skipping validation is risky (accepting untrusted bytes
unchecked) and running it re-walks `BastDoc` per buffer — the same
class of cost review #004 measured for the read path, just on a
different code path.
The argument is not "ValidationPlan has the same shape as ReadPlan"
(it doesn't). The argument is: ADR-012's "not a hot loop" dismissal
may be incomplete, because read+validate on untrusted streams makes
validation hot in the same sense read was hot. The deferral's
trigger ("if a bench motivates it") should be sharpened: either (a)
add a `validate_bytes`-on-untrusted-stream bench to alktty alongside
`wire_vs_bast` and let the bench decide, or (b) reason from the
existing review #004 numbers that the validation walk is
non-trivial and should be planned, not deferred.
This is not a request to implement `ValidationPlan` in 0.3.0. It's a
request to *own the decision*: either the trigger fires (and a
follow-on ADR/phase is scoped, possibly 0.4.0) or it doesn't (and the
deferral stands with a sharper justification than "not a hot loop").
As written, the deferral leaves the cost in the superposition where
it can neither be confirmed nor dismissed.
**Lift**: removes a latent perf cliff for the read+validate-on-
untrusted-input case that 0.3.0 is supposed to make viable.
---
### L1. `dummy_field_for`/`ty_source` are used in aligned `materialize`, not just packed
**Files**: `docs/plans/030-compiled-forms.md:209-210`,
`src/materialize.rs:249,316,351,391,631,650-663`
**Problem**: Phase 2 says:
> The `dummy_field_for`/`ty_source` helpers in `materialize.rs` are
> removed (the plan carries everything).
This is factually wrong. `dummy_field_for` is called at
`src/materialize.rs:631` inside `materialize_leaf_at`, which is called
by the **aligned** path: `materialize_struct_aligned` (line 475),
`materialize_array_aligned` (line 544), `materialize_variable_aligned`
(line 613). Aligned `materialize` keeps walking `BastDoc` through
0.3.0 (plan lines 379-385 confirm), so `dummy_field_for`/`ty_source`
must stay. Only the packed-side call sites (lines 249, 316, 351, 391)
go away when packed-materialize moves to the plan.
**Lift**: doc accuracy. An agent following the plan literally would
remove the helpers and break aligned `materialize`.
---
### L2. `materialize_packed` rewrite scope underspecified — packed-vs-aligned split of `materialize_typeref_packed`
**Files**: `docs/plans/030-compiled-forms.md:207-210`,
`src/materialize.rs:122-200, 498-506, 619-637`
**Problem**: `materialize_typeref_packed` is shared by both packed and
aligned paths — aligned's `materialize_leaf_at` (line 619-637) calls
`materialize_typeref_packed` to read leaves, and aligned's record path
(line 498-506) calls it directly. Phase 2 says
`materialize_packed(&ReadPlan, &[u8])` walks the plan instead of
`BastDoc` but does not state what happens to
`materialize_typeref_packed`.
The honest resolution: packed-materialize gets a new plan-walking
function; aligned keeps `materialize_typeref_packed` via
`materialize_leaf_at`; the function stays (renamed or not) for aligned.
This is two mode-specific paths — the existing design — not a
"parallel walker" in the maintenance-tax sense ADR-011 §"Negative"
(cautioning against) discusses. ADR-011's "one walker" claim (lines
234-237) is specifically about packed read-side (`SequentialReader` +
`materialize_packed` sharing the plan), not packed-vs-aligned, so
there's no ADR contradiction — just an underspecification in the plan.
**Lift**: prevents the agent from having to discover the split
mid-rewrite. Add one line to phase 2: "packed-materialize gets a new
plan-walking function; `materialize_typeref_packed` stays for
aligned's `materialize_leaf_at` and the aligned record path."
---
### L3. `materialize_aligned`'s `BastDoc` structure walk is silent in the plan
**Files**: `docs/plans/030-compiled-forms.md` (silent on this),
`src/materialize.rs:451-521`, `docs/architecture/decisions/011-...md:264`
**Problem**: `materialize_struct_aligned` walks `BastDoc` to traverse
struct/array/record structure, using `OffsetMap` only for leaf byte
positions. ADR-011 §"Out of scope" says "aligned mode is unchanged;
`materialize_aligned` already takes `&OffsetMap`" — which is half
true: it takes `&OffsetMap` for positions but also `&BastDoc` for
structure. The plan inherits the half-truth silently: there's no
statement anywhere that aligned materialize keeps walking `BastDoc`
for structure.
After phase 3 (owned `BastDoc`) + phase 5 (`LeafMeta`), the walk is
over owned data, no re-parse, not O(N²), and aligned `validate_bytes`
is one walk per call (not per-field). There's no perf driver
analogous to review #004's packed per-chunk gap. But the absence of
a driver is not the same as a decision: leaving it silent is a
deferral-by-omission. An implementing agent or future reader can't
tell whether the silence is "this is the permanent design" or "we'll
fix this later."
The decision should be owned. Either (a) add a "Scope Boundary" note
that aligned materialize keeps walking owned `BastDoc` for structure
as the permanent design (with an OQ if a future bench motivates an
`AlignedPlan`), or (b) if a bench motivation is plausible, scope an
OQ to track it. (a) is recommended — no perf driver, and after phase
3 the walk is over owned data, so it's not the re-parse pattern.
**Lift**: removes a silent gap that future agents would otherwise
have to reverse-engineer.
---
### N1. Typo: "back-comat" → "back-compat"
**File**: `docs/plans/030-compiled-forms.md:104-105`
**Problem**: "back-comat" in deferred decision 4.
**Lift**: trivial.
---
### N2. Phase 1 verification omits the `Send + Sync` assertion test ADR-011 requires
**Files**: `docs/plans/030-compiled-forms.md:158-163`,
`docs/architecture/decisions/011-...md:204-206`
**Problem**: ADR-011 §"Engine integration" says "the implementation
should add a `static` bound assertion test to lock it in" for
`ReadPlan: Send + Sync`. Phase 1's verification block lists `cargo
test`, `clippy`, `doc`, `wasm` but no mention of adding the assertion
test. An agent following the plan literally won't add it; the
property is currently true by construction but not asserted, so a
future change could break it silently.
**Lift**: add "add a `fn read_plan_is_send_sync()` assertion test" to
phase 1's verification, mirroring the POC's
`readplan_is_send_sync` test.
---
### N3. `SequentialReader::new` return-type change (`Result` drop) undocumented
**Files**: `docs/plans/030-compiled-forms.md:64`,
`src/sequential_reader.rs:129`, `src/engine.rs:205`
**Problem**: Currently `new(&Value, &str) -> Result<Self,
AlkTypeError>` — fallible (BastDoc parse). After phase 2,
`new(Arc<ReadPlan>)` is infallible (just stores the Arc) → returns
`Self`, not `Result<Self>`. The Semver Contract table (line 64) lists
only the argument-type change, not the `Result` drop.
`engine.rs:205`'s `.ok()` call correspondingly goes away. Minor, but
it's a signature change beyond what's listed.
**Lift**: add a row to the Semver Contract table noting the `Result`
drop.
---
## Deferral-pattern scan (LLM-planning quirk)
As part of the methodology, every "future/deferred/later/downstream/if
needed" occurrence in the plan and its ADRs was flagged and tested
for: (a) concrete reactivation trigger, (b) decision owned or silent,
(c) hidden cost of inaction.
| Item | Trigger? | Owned? | Cost of inaction | Finding |
|---|---|---|---|---|
| `ValidationPlan` (ADR-012) | "if a bench motivates it" | Yes (ADR + plan "What this is not") | Possible perf cliff on read+validate untrusted streams | M3 above — sharpen the trigger |
| Nested-union `ReadPlan` support | "if a real schema needs it" (POC) | No (POC only, plan silent) | Silent 0.2.0 capability drop | M1 above — state it |
| `materialize_aligned` structure walk | None — silent | No (silent) | Future agent ambiguity | L3 above — own the decision |
| `Arc<str>` vs `String` (decision 1) | "if phase 4 shows it's measurable" | Yes (deferred decision 1) | None | OK — has trigger, decided in phase 3 |
| `OffsetMap::get` shape (decision 2) | "decided in phase 5" | Yes (deferred decision 2) | None | OK |
| Fingerprint hasher (decision 3) | "decided in phase 6" | Yes (deferred decision 3) | None | OK |
| Struct-array stride (decision 4) | "decided in phase 2" | Yes (deferred decision 4) | None | OK |
| `BastDoc` `Arc<Value>` vs `Value` | None — silent | No (plan doesn't address) | Agent gets stuck (H2) | H2 above |
| Field-disc union shape (POC Finding 1) | "production version must implement" | Yes (plan phase 1) | None, but shape is undefined | H1 above — shape not in ADR |
The four explicit "deferred decisions" in the plan (items 4-7) all
have concrete triggers and decision points — these are the *good*
pattern. The black-hole pattern appears where deferrals lack triggers
(items 1-3, 8-9): three of those became findings (M1, L3, H2), and M3
is a deferral worth sharpening even though it has a trigger.
The general signal: a deferral is healthy when it has a concrete
reactivation condition and is tracked (OQ, ADR, or in-plan deferred
decision). A deferral is a black hole when it has no trigger, no
tracking, and the next agent inherits the gap by default.
---
## What's Good
- **Line-number accuracy is perfect.** Every `src/` reference in the
plan (`engine.rs:112-115,284,334,467`;
`sequential_reader.rs:567`; `layout_builder.rs:190`;
`bast.rs:51-55`; `lib.rs` re-exports) checks out against the v0.2.0
tree. This is unusual for a plan of this length and worth noting.
- **The Semver Contract table is a strong scope-creep guardrail.**
Walking every public `lib.rs` re-export against the table, the
classifications (Breaking / Unchanged / New) are correct for every
item, with the exceptions noted in N3 (the `Result` drop on `new`)
and M1 (the nested-union behavioral drop not listed).
- **The four explicit "deferred decisions" are the right pattern.**
Each has a trigger and a decision point in a named phase. This is
what deferrals should look like.
- **Phases are coherent session boundaries.** Phases 1 (pure
addition), 6 (pure addition), 7 (docs/bump) are small and clean.
Phases 3 (broad but mechanical), 4 (single file), 5 (single file +
engine) are well-scoped. Phase 2 is the largest and the plan
sanctions sub-session splits at the step level (lines 40-42), which
is the right escape valve.
- **Cross-phase invariants are stated and checkable.** "Tree builds
and tests pass at every phase boundary" is the right invariant; the
POC-on-`readplan-poc`-only convention is clearly separated from
production code; AGENTS.md §5-§11 constraints (no `unsafe`, no
`async`, no new deps, `preserve_order` load-bearing) are
reaffirmed.
- **The plan honestly scopes what it is not.** "Not a `ValidationPlan`",
"Not cross-version fingerprint stability", "Not a perf bench" —
these boundaries are stated rather than left implicit, which helps
an implementing agent resist scope creep. (M3 above is about
sharpening one of these, not removing the boundary.)
- **The POC reference is disciplined.** The plan is explicit that the
POC is "not production code," lives only on the branch, and is the
reference scaffold for phases 1-2 only. This matches how
`bast-validator-poc` was handled and avoids the POC leaking into
`main`.
---
## Recommended Order
1. **H1 (field-disc union shape)** — update ADR-011's
`CompositePlan::Union` to include the shared-fields sub-`ReadPlan`
(or document the wrapper shape), then update the plan's phase 1 to
reference the corrected ADR shape. Do this before phase 1 starts;
otherwise the implementing agent has to guess.
2. **H2 (`schema()` `&Value` on `Arc<ReadPlan>`)** — edit the plan's
phase 2 to specify `ReadPlan` stores `Arc<Value>` (or owned
`Value`), and `schema()` borrows from that. One-line edit to the
plan; avoid a phase-2 stuck point.
3. **M1 (nested unions)** — decide (a) implement `VariantKind::Union`
in phase 1, or (b) list as behavioral drop + OQ. Edit the plan and
(if b) the Semver Contract table accordingly. Decide before phase
1.
4. **M2 (`Hash` on `Endian`/`VariableEncoding`)** — add a sub-step
to phase 5 or 6. Trivial.
5. **M3 (`ValidationPlan` re-evaluation)** — either add a
`validate_bytes`-on-untrusted-stream bench to alktty (alongside
`wire_vs_bast`) and let the bench decide, or sharpen ADR-012's
"not a hot loop" justification. Does not block 0.3.0; can be
resolved in parallel with phase 1-7 work. **Flagged for
re-evaluation, not for implementation in 0.3.0.**
6. **L1, L2, L3** — edit the plan's phase 2 to fix the
`dummy_field_for` wording (L1), state the packed-vs-aligned
materialize split (L2), and add a Scope Boundary note for
aligned-materialize's `BastDoc` structure walk (L3). All three are
phase-2 doc edits.
7. **N1, N2, N3** — typo, `Send + Sync` assertion test, `Result`-drop
Semver row. Minor plan edits.
Items 1-3 must be resolved before phase 1 starts (they affect the
`ReadPlan` shape or 0.2.0 behavioral surface). Items 4-7 can be
resolved any time before their phase begins. Item 5 (M3) is
non-blocking and can run in parallel.
---
## Notes
- All line numbers refer to the tree at commit `2310f6c` (the plan's
commit) for `src/` files, and to the plan/ADR markdown as committed
at the same tree.
- The POC on `readplan-poc` was inspected via
`git show readplan-poc:poc/readplan/{src/lib.rs,FINDINGS.md}`; it is
not merged to `main` and the plan correctly states this.
- `alktty` and `alkcall` downstream repos exist as path dev-deps
(`/workspace/@alkdev/alktty`, `/workspace/@alkdev/alkcall`); the
plan's claim that they're in-house and updated with the bump is
verifiable, though this review did not inspect their call sites
in detail.
- This review does not re-litigate ADR-011 or ADR-012's accepted
decisions. H1 and H2 are about the plan *contradicting* the ADRs or
being unsound, not about the ADR decisions themselves; M3 is about
sharpening a deferral, not about re-deciding it.
- The deferral-pattern scan is a methodology experiment: a
pre-declared scan for LLM-specific planning quirks (deferral black
holes) alongside classic planning mistakes. It surfaced M1 and L3
that a conventional severity-only review would have missed or
under-weighted. Worth retaining as a default scan for future plan
reviews.
---
## Resolution (2026-08-20)
All 11 findings resolved in one docs-only edit pass to ADR-011,
ADR-012, and the 0.3.0 plan. No source changed; the crate still
builds/tests at v0.2.0. The M3 deferral reversal is the one
substantive decision change (per user direction: ship ValidationPlan
in 0.3.0, no more hedging); the rest are spec corrections or
pre-implementation refinements to types that do not yet exist on
`main`.
- **H1 (union shape):** ADR-011 §"The `ReadPlan` shape" refined —
`CompositePlan::Union` now carries `shared: Option<Box<ReadPlan>>`
(field-disc shared fields) and `variants: Vec<(String,
CompositePlan)>` (dropping `VariantPlan`/`VariantKind`). Plan
phase 1 rewritten to implement the refined shape. The shape
refinement is pre-implementation (the types don't exist on `main`).
- **H2 (`schema()` `&Value`):** plan phase 2 rewritten — `ReadPlan`
stores `schema: Arc<Value>` (not `&Value`); `schema()` returns
`&self.schema`. Verified `serde_json::Value: Hash + Eq` holds with
`preserve_order` (`Map::hash` sorts keys deterministically), so
phase 6's `#[derive(Hash)]` on `ReadPlan` is not blocked.
- **M1 (nested unions):** resolved as the review's option (a) —
nested-union support falls out of the H1 shape refinement (a
variant can be `CompositePlan::Union`), so no behavioral drop vs
0.2.0 and no Semver Contract entry for a capability regression.
Plan phase 1 adds a nested-union-variant test.
- **M2 (`Hash` on `Endian`/`VariableEncoding`):** plan phase 5
rewritten with an explicit first sub-step to add `Hash` to both
derives in `src/schema.rs` (additive, semver-safe). The inaccurate
"all fields are `Copy + Hash`" parenthetical on `LeafMeta` is
corrected.
- **M3 (`ValidationPlan`):** deferral **reversed** per user
direction. ADR-012 §"Deferring `ValidationPlan`" rewritten as
"ValidationPlan — in scope for 0.3.0"; new ADR-012 §3 commits the
decision (compiled form, no per-buffer `BastDoc` walk, `Hash + Eq`
+ `fingerprint()`) and lists the shape questions deferred to a
follow-on design session + the plan's new phase 7. Plan gains a
new phase 7 (ValidationPlan); old phase 7 (bump) renumbered to
phase 8. ADR-011's "Out of scope" `bast_validation` bullet and
"Scope Boundaries" `Not a validation plan` bullet updated to point
at ADR-012 §3. Plan's "What this plan is *not*" first bullet
removed. The deferral-black-hole pattern this review's methodology
flagged is closed: the work is committed in the plan with a
concrete reactivation trigger (the shape session before phase 7),
not hedged into an unplanned future.
- **L1 (`dummy_field_for`/`ty_source`):** plan phase 2 rewritten —
only the packed-side call sites go away; the helpers stay for the
aligned `materialize_leaf_at` path.
- **L2 (`materialize_typeref_packed` split):** plan phase 2
rewritten — packed-materialize gets a new plan-walking function;
`materialize_typeref_packed` stays for aligned's
`materialize_leaf_at` and the aligned record path.
- **L3 (aligned-materialize `BastDoc` structure walk):** plan phase 5
gains a Scope Boundary note — the walk is the permanent 0.3.0
design; an `AlignedPlan` is out of scope, tracked as an open
question if a future bench motivates it.
- **N1 (typo):** "back-comat" → "back-compat" in deferred decision 4.
- **N2 (`Send + Sync` assertion test):** plan phase 1 verification
rewritten to add the `read_plan_is_send_sync` static-bound
assertion test ADR-011 §"Engine integration" requires.
- **N3 (`Result` drop on `SequentialReader::new`):** Semver Contract
table row updated to note the constructor return-type change
(`Result<Self, AlkTypeError>` → `Self`) alongside the argument-type
change.
The deferral-pattern scan's general signal (healthy deferrals have a
concrete reactivation condition + tracking; black holes have neither)
is reaffirmed by the M3 reversal: the original "if a bench motivates
it" trigger was a black hole because no bench was ever going to be
run against a path that didn't exist yet, and the cost of inaction
(a second breaking change to `validate_bytes`/`bast_validation` after
0.3.0) was hidden by the "not a hot loop" framing.
File diff suppressed because it is too large. Load diff
-454
View File
@@ -1,454 +0,0 @@
---
status: resolved (F1, F2, C1, C2, C3, L1, L2 resolved 2026-09-02; N1/N2a/N3a/N4a are classified-no-action / deferred-by-design)
last_updated: 2026-09-02
reviewed_artifacts:
- src/materialize.rs
- src/sequential_reader.rs
- src/read_plan.rs
- src/offset_map.rs
- src/layout_builder.rs
- src/engine.rs
- src/data_access.rs
- src/bast.rs
- src/builder.rs
- src/tunion.rs
- src/validation_plan.rs
- src/walk_guard.rs
- src/bast_meta.rs
- tests/poc_roundtrip.rs
- tests/tunion_dispatch.rs
- tests/error_paths.rs
- tests/engine_integration.rs
- docs/reviews/006-implementation-review-030.md (post-fix coverage re-check)
tool: cargo-llvm-cov 0.8.4 (--release, per-line text) + manual classification of every uncovered production line + disposable probe tests (run in-session, then deleted)
reviewer: post-review-#006 coverage audit (session request — check test coverage for weak spots, meaningful tests, non-happy-path posture)
---
# Review #007 — Post-#006 Coverage Audit
## Purpose
Review #006 closed every finding and its M4 coverage map, but the
session-level posture (M4's item: "fold a coverage check into each fix
session") had never been run as a *whole-tree* pass after all those
fixes landed. This audit re-measures coverage after the eleven #006
commits, reads every uncovered production line, and classifies it —
the same "untested-but-fine / load-bearing / unreachable" discipline
M4's map used. Two probes were run in disposable tests (deleted after
the session, per #006's no-reproducer rule; neither was a crash
hazard — both reproduce cleanly inside the default harness).
## Methodology
- `cargo llvm-cov --release` (0.8.4, same tool as #006): summary +
per-line text. TOTAL **90.67% lines / 86.32% functions** — stable
with #006's post-M4 numbers (90.60%), the expected drift after the
N3 fix sessions added parse gates + tests.
- Per-file (worst first): `materialize.rs` 85.72, `data_access.rs`
80.32, `sequential_reader.rs` 86.19, `bast.rs` 87.41,
`builder.rs` 91.28, `layout_builder.rs` 91.42, `tunion.rs` 92.02,
`offset_map.rs` 92.93, `read_plan.rs` 90.69, `engine.rs` 96.44,
`validation_plan.rs` 93.97, `walk_guard.rs` 98.04,
`bast_meta.rs` 98.92, `bast_validation.rs`/`error.rs`/`schema.rs`/
`macros.rs`/`validation.rs` 100.
- Every uncovered line *outside* `#[cfg(test)]` modules (806 raw
lines) was read and classified. Lines inside test modules (the
`panic!("expected X, got {other:?}")` helpers) were excluded — they
distort per-file numbers (e.g. `bast.rs`'s 87.41% is really ~96%
production once its 60 helper lines are excluded).
- Two suspicions were probe-verified with disposable tests:
the F1 cross-consumer divergence and the F2 unbounded-`maxLength`
compile. Probe transcripts quoted verbatim in the findings.
- Happy-path posture audit: cross-checked which *error arms* adjacent
to covered code are 0-execution, and which public surfaces have only
success-path tests.
## Baseline
Audited at `main` HEAD `bb28ba3` ("Resolve N3"), 0.3.0, working tree
clean. Full suite green (548 tests static + 2 ignored doctests, per
#006's bookkeeping).
## Summary Statistics
| Severity | Count | Status |
|----------|------:|--------|
| High | 2 (F1, F2) | both resolved 2026-09-02 |
| Medium | 3 (C1, C2, C3) | all resolved 2026-09-02 |
| Low | 2 (L1, L2) | all resolved 2026-09-02 |
| Info | 4 (N1, N2a, N3a, N4a) | classified: N1 artifact, N2a/N4a no-action, N3a deferred to pre-release review |
**Resolution log:**
- **F1 + F2 (2026-09-02):** resolved in one commit — see the
resolution blocks on each finding. 477 lib tests green (511
static + 2 ignored doctests across all targets), clippy
`-D warnings` clean, wasm build green.
- **C1 + C2 + C3 (2026-09-02):** resolved in one commit — see the
resolution blocks. 481 lib tests green, clippy `-D warnings` clean,
wasm build green; `compile_variant`'s cycle arm confirmed executed
in the post-fix coverage run.
- **L1 + L2 (2026-09-02):** resolved in one commit — see the
resolution blocks. 488 lib tests green, clippy `-D warnings` clean,
wasm build green.
---
## Findings
### F1. Zero-progress guard missing in the plan materializer — `validate_bytes` accepts what `SequentialReader` rejects (cross-consumer divergence)
**Files**: `src/materialize.rs:227-253` (`materialize_plan_array` — no
guard), contrast `src/sequential_reader.rs:772-795`
(`plan_walk_variable_array_size` — has the guard) and
`src/materialize.rs:636-663` (`materialize_array_packed` — has the
guard)
**Problem**: The H1 fix session added the zero-progress runtime guard
("array element consumed 0 bytes") to two of the three array walkers:
the compiled reader's variable-array size walk and the legacy BAST
walker's packed array arm. The *plan-based* packed materializer —
`materialize_plan_array`, the walker `validate_bytes` actually uses in
packed mode (engine.rs:310-312) — got no guard.
A stride-0 array whose elements consume 0 bytes (empty-struct elements
are legal: the meta-schema's `StructDef` has no `minItems` on
`fields`) compiles with `element_stride: 0` and loops `count` times
materializing empty objects without reading a single buffer byte:
```
PROBE validate_bytes([]): OK — zero-progress guard MISSING in plan materializer
PROBE reader.read_next([]): Err(access error at items[0]: array element 0 consumed 0 bytes;
a zero-size element makes the declared count unbounded on the wire)
```
Schema: `{ "items": { "kind": "array", "element": { "kind":
"struct", "fields": [] }, "count": 8 } }`, packed mode, empty buffer.
Same schema, same buffer, opposite verdicts — the exact
cross-consumer-disagreement shape review #006 existed for (H3, M6).
Severity High by #006's own keying (AGENTS.md §3): `validate_bytes`
is the flagship untrusted-input path, and it silently accepts a
buffer the same engine's reader rejects. The H1 resolution text
("plan_walk_variable_array_size (reader) and materialize_array_packed
(materializer) now error") lists only two of the three walkers — the
plan materializer was missed because it is *not* the legacy walker
that finding named.
**Not a #006 regression**: the H1 fix text itself specified only the
reader and legacy-walker sites; the plan materializer predates the
guard and was outside that fix's blast radius. But the divergence is
new information — the guard's *invariant* ("a zero-progress element
makes the declared count unbounded") belongs to the array-walk
concept, not to two specific functions.
**Fix**: hoist the same guard into `materialize_plan_array`'s loop
(compare `*offset` before/after `materialize_plan_composite`; error
with the same wording the other two walkers use so downstream
matching sees one shape). Add a locking test driving the same schema
through BOTH paths asserting the verdicts agree (both reject an empty
buffer; both accept a buffer where the elements make progress —
empty-struct elements never do, so the acceptance half needs a
non-empty variant struct alongside).
**Resolution (2026-09-02):** the guard, hoisted verbatim from the two
existing sites (`*offset == before` after the element walk, same
"array element {i} consumed 0 bytes…" wording so downstream matching
sees one shape). Tests (3, in `materialize.rs`):
`f1_zero_progress_array_rejected_by_all_three_walkers` (plan
materializer + the record-value fallback path, both asserting the
`Access` error with the guard's wording),
`f1_validate_bytes_and_reader_agree_on_zero_progress_array` (the
cross-consumer agreement the probe showed was missing —
`validate_bytes` and `SequentialReader::read_next` both reject the
same schema+buffer with the same error class),
`f1_nonempty_variant_struct_array_still_materializes` (the
false-positive check: elements that consume bytes still walk).
Verified: 477 lib tests green, clippy `-D warnings` clean, wasm build
green.
### F2. `maxLength` is unbounded — the N2 analog
**Files**: `src/bast_meta.rs:81` (`"maxLength": { "type": "integer",
"minimum": 0 }` — no maximum), `src/bast.rs:1038-1043`
(`parse_max_length` — no cap, and silently drops non-`usize` values),
contrast `src/schema.rs` `MAX_ALIGN`/`parse_align` (the N2 pattern)
**Problem**: N2 bounded `align` at 4096 with a clean parse error plus
a meta-schema `"maximum"`. `maxLength` has the identical shape and
was not covered by that fix:
```
PROBE aligned maxLength 1e12 compiles; total_size = 1099511627776
```
A one-field schema declares a 1 TiB reservation; `total_size` in that
range is meaningless output the consumer may act on (N2's argument
(a)). Unlike align, no `Access` error follows at read time (an empty
buffer still fails buffer bounds first), so this is layout-meaningless
output, not a crash — the exact severity N2 recorded. Additionally,
`parse_max_length` returns `Option` and silently *drops* values that
overflow `usize` (`.and_then(|n| usize::try_from(n).ok())`) — on a
32-bit target a 5 GiB `maxLength` becomes "no maxLength", changing
layout semantics without telling the consumer.
**Fix**: the N2 playbook verbatim. A `MAX_LENGTH` cap in
`schema.rs` (value TBD — `align`'s 4096 is page granularity; a
reservation cap in the tens-of-megabytes range fits honest layouts;
suggest `2^26 = 67_108_864`, matching `MAX_ARRAY_BYTES`'s rationale),
enforced in `parse_max_length` (converted to `Result<Option<usize>>`,
clean `Schema` error naming the path/value/maximum — no silent drop),
plus `"maximum": 67108864` in the meta-schema's `maxLength` property
so the published contract matches the parser (the N2 dual-layer
pattern).
**Resolution (2026-09-02):** the N2 playbook, cap = `MAX_LENGTH`
(2^26 = 67_108_864, matching `MAX_ARRAY_BYTES`'s rationale: a single
fixed reservation no larger than the largest legal array):
1. `MAX_LENGTH` added to `schema.rs`, documented with the F2 probe
arithmetic.
2. `parse_max_length` converted to `Result<Option<usize>>`: non-integer
→ clean `Schema` error; `usize` overflow → clean `Schema` error (the
silent `.and_then(try_from().ok())` drop is gone); over-cap → clean
`Schema` error naming the path, value, and maximum.
3. Meta-schema `maxLength` property gains `"maximum": 67108864` — the
published contract matches the parser.
4. Tests (5, in `offset_map.rs`, mirroring the `n2_` family):
above-cap rejection naming value+maximum (bytes and string),
at-cap acceptance (`total_size == 67108864`), u64::MAX-scale value
rejected-not-silently-dropped (cap arm on 64-bit, overflow arm on
32-bit — one test covers whichever fires), and the meta-schema
dual-layer check (above-cap rejected, at-cap accepted).
Verified with F1's commit: 477 lib tests green, clippy clean, wasm
green.
### C1. Packed `validate_bytes` has never decoded a wide primitive
**Files**: `src/materialize.rs:121-169` (`materialize_plan_primitive`'s
Int16/Int32/Int64/Uint64/Float64/Boolean arms — all 0-execution),
`src/sequential_reader.rs:345-383` (the reader's same arms are covered
via `read_next` tests, but the materializer's are not)
**Problem**: every packed `validate_bytes` test feeds u8/uint32/
string-shaped data. The i16/i32/i64/u64/f64/bool arms of the plan
materializer — the code every untrusted packed wire buffer flows
through — have never executed through any test. Probe (in-session)
confirmed the BE i16/bool path works; the arms are correct, just
unexercised. This is the flagship decode path for `alkcall`'s packed
frames; one battery test closes it (mirror the aligned
`read_field` battery, tests/engine_integration.rs:130-190, which
already covers all twelve primitive kinds on the aligned side).
**Resolution (2026-09-02):** two tests in `engine.rs`:
`c1_validate_bytes_packed_decodes_all_twelve_primitives_le` (the full
eleven-field battery — i8..bool — plus a corrupted-bool rejection arm)
and `c1_validate_bytes_packed_decodes_big_endian_subset` (BE i16/u64/
f64 through the same public path). Both green.
### C2. Aligned `validate_bytes` never exercises the default inline encoding for string/bytes
**Files**: `src/materialize.rs:1037-1039`
(`materialize_variable_aligned`'s `LengthPrefixed`-else branch —
0-exec through the public path)
**Problem**: the aligned `validate_bytes` tests use records, unions,
maxLength reservations, and offset-indirect encodings. The *default*
encoding — an inline length-prefixed string or bytes field, the most
common real shape — reaches `read_field` (engine_integration.rs:192)
but never `validate_bytes`. The aligned `validate_bytes` surface has
thus never decoded the single most likely field kind through its
public path.
**Fix**: one aligned `validate_bytes` test with a trailing inline
string (and a bytes variant or arm), asserting acceptance plus a
short-buffer rejection.
**Resolution (2026-09-02):**
`c2_validate_bytes_aligned_inline_string_and_bytes_default_encoding`
in `engine.rs`. One wrinkle the test wrote itself into: ADR-006
allows an inline length-prefixed variable field only in the final
position, so the string and bytes shapes get separate one-field
schemas (string after a fixed `id`; bytes as a lone field). Asserts
acceptance for both plus a short-buffer rejection for the string.
### C3. `ReadPlan::compile`'s union-variant cycle arm is untested standalone
**Files**: `src/read_plan.rs:509-514` (`compile_variant`'s
`cycle_err` arm — 0-exec)
**Problem**: the H2 test family exercises `check_ref_graph` (walk
guard) via `OffsetMap::compute`/`LayoutBuilder::new`/
`materialize_aligned`, and `ValidationPlan::compile`'s cycle arm is
covered (`validation_plan.rs:297` shows executions, via the
`shared_refs_compile_without_false_cycle`/cycle tests). But
`ReadPlan::compile`'s own cycle rejection — the defense the *packed
read plan* relies on when driven standalone (its doc explicitly
promises untrusted-input safety) — has no test driving a two-def
cycle through it. The depth cap is tested
(`deep_nesting_beyond_depth_cap_is_schema_error`); the cycle arm is
shadowed in every engine-path test by the ValidationPlan gate running
first (engine.rs:153).
**Fix**: a `read_plan_compile_two_def_cycle_rejected` test calling
`ReadPlan::compile` directly on a two-def cycle, mirroring
`validation_plan.rs`'s existing standalone cycle test.
**Resolution (2026-09-02):**
`c3_cycle_through_union_mapping_variant_is_schema_error` in
`read_plan.rs`. Writing the test sharpened the finding: the
*field-level* cycle arm (`compile_typeref`, :362) was already covered
(2 execs) by `cyclic_ref_through_two_defs_is_schema_error`; the
0-exec arm was `compile_variant`'s own check (:513), reachable only
when the cycle closes through a **union mapping entry**. The new
test's shape (`A → B → U(mapping: "1" → $ref B)`) trips exactly that
arm — verified post-fix at the line level (1 execution).
### L1. `builder.rs`'s JSON-Schema conveniences are entirely untested
**Files**: `src/builder.rs:268-290` (`array()`, `number()`,
`boolean_()`, `null()`), `:465-479` (`items()`,
`additional_properties()`), `:508-565` (`maximum()`, `min_length()`,
`min_items()`, `max_items()`, `format()`, `title()`,
`description()`), `:440` (the `field()`-on-standard-repr path)
**Problem**: only the BAST-side builders have tests. The standard
JSON-Schema side feeds `jsonschema::build_validator` (the
`json_schema` parameter of `AlkTypeEngine::compile`), so a typo'd or
misplaced keyword would ship silently — the builder emits the JSON,
`jsonschema` interprets it, and nothing checks the translation. One
table-style test asserting each convenience produces the expected
JSON key/value closes the surface cheaply.
**Resolution (2026-09-02):** five tests in `builder.rs`:
`l1_standard_type_constructors_produce_type_keyword` (all eight
standard constructors, exact-JSON assertions),
`l1_field_on_standard_object_builds_properties` (`field()` on the
standard repr + `required()`),
`l1_items_and_additional_properties_on_standard_types`,
`l1_constraint_keywords_emit_expected_json_keys` (minimum/maximum/
minLength/minItems/maxItems/format/title/description, each asserted
on its exact keyword), and
`l1_standard_built_schema_compiles_as_json_validator` (the end of the
translation chain: the emitted JSON builds a `jsonschema` validator
and the constraints actually bite — valid passes, over-maximum/
missing-required/below-minimum fail).
### L2. `tunion::read_field_discriminator`'s enum arm is 0-exec
**Files**: `src/tunion.rs:166-169`
**Problem**: N1's resolution extended tunion to match the reader's
kind set and added uint16/uint32 tests both endians — but skipped the
enum arm, which is in the documented kind set
(tunion.rs:106-112 names "string / uint8 / uint16 / uint32 / enum").
The reader's enum arm is tested (`m4_field_disc_enum_dispatches_on_index`);
tunion's is not. One test locks parity on the last arm.
**Resolution (2026-09-02):** `l2_read_field_discriminator_enum_
dispatches_on_index` and `l2_read_field_discriminator_enum_big_endian`
in `tunion.rs` — enum index 0 (LE) and 1 (BE) dispatch with
`variant_offset == 4` / `discriminator_size == 4`.
### N1. `OffsetEntry::start()`/`end()` 0-execution in the combined run is a merge artifact, not a hole
**Files**: `src/offset_map.rs:79-86`
The combined `cargo llvm-cov --release` run reports these 0-exec;
`tests/poc_roundtrip.rs:187-189` calls `start()` (and the
`big_endian_round_trip_via_offset_map` test calls `end()`). Per-test
coverage confirms both execute (32/2 calls respectively in a
poc_roundtrip-only run). llvm-cov's profile merge does not attribute
integration-test-binary executions to the library in every run
configuration. Recorded so nobody "fixes" this by deleting the
accessors or writing a redundant in-module test. (Caveat for future
audits: when a combined run shows 0-exec on something an integration
test visibly calls, re-run per-test-target before classifying.)
### N2a. `data_access.rs`'s remaining uncovered lines are the documented >4 GiB guards — fine to leave
**Files**: `src/data_access.rs:54-95, 223-291, 336-414`
All are `checked_add` overflow arms and u32-truncation guards needing
multi-GiB slices or near-`usize::MAX` offsets — already documented as
defensively-unreachable on 64-bit test hardware in #006 M4 item 3's
resolution. (The `read_array`/`write_array` arms at :54-95 are
additionally unreachable-after-`check_bounds` belt-and-suspenders.)
No action.
### N3a. `bast.rs` dead-or-orphaned surface — flag for the pre-release review
**Files**: `src/bast.rs:212-214, 305-307, 404-406, 506-508, 741-743,
924-926, 973-975` (`source()` accessors — zero callers anywhere in
src or tests), `:353-363` (`BastField::synthetic`,
`#[allow(dead_code)]`, zero callers), `:149-167`
(`resolve_typeref_as_def`'s inline struct/union/enum arms — both call
sites pass `$ref`-only variants since H3's parse rules forbid
re-declaration; plausibly dead now)
Three small deletions-or-justifications. Not fixed this session (the
`source()` accessors are public API — removal is a semver decision
for the pre-release review, and AGENTS.md's semver exception list
says renames/removals need an explicit ask). Recorded so the
pre-release review session has the list.
### N4a. Internal-shape error arms are structurally unreachable — fine to leave
**Files**: `src/materialize.rs:241,309,422,858`,
`src/sequential_reader.rs:295,473,533,832`, `src/offset_map.rs:352`,
`src/layout_builder.rs:194,282`, `src/read_plan.rs:402`
The `"internal: …"` arms that dispatch on a `match` the caller
already narrowed (e.g. "union body at X is not CompositePlan::Union"
inside a function only reachable from a `Union` match arm). They are
honest defensive code — deleting them would force `unwrap()` — and
forcing them in tests would require constructing mid-walk corruption.
Leave uncovered; the pattern is consistent across the codebase.
---
## What's Good
- The #006 fix sessions left the tree in genuinely good shape: 90.67%
lines with every high-traffic wire path (reader dispatch, plan
compiler, offset map, walk guard) in the mid-90s or better.
- The untrusted-input discipline is visible in the coverage: every
parse-level gate added in #006 (H1 caps, N2 align cap, N3
string/bytes-only maxLength, H2 cycle rejections at all three
standalone walkers) has both rejection and boundary tests.
- The `#[cfg(test)]` helper noise is the only thing making
`bast.rs`/`data_access.rs` look worse than they are — the
production coverage of both is materially higher than the raw
per-file number.
## Recommended Order
1. ~~**F1** — guard hoist + cross-consumer agreement test~~
**resolved 2026-09-02** (with F2).
2. ~~**F2** — `MAX_LENGTH` cap, N2's dual-layer playbook verbatim~~
**resolved 2026-09-02** (with F1).
3. ~~**C1 + C2 + C3** — one locking test each~~ **resolved 2026-09-02**.
4. ~~**L1 + L2** — posture tests~~ **resolved 2026-09-02**.
5. **N3a** — defer to the pre-release review (semver decision).
## Notes
- Probe tests were run as `tests/zzz_probe.rs` in-tree during the
session and deleted before any commit (the #006 pattern). Neither
probe was a crash hazard; both reproduce safely in the default
harness.
- Per-file numbers are from a single `cargo llvm-cov --release`
run; the N1 merge artifact means integration-test-only calls
(e.g. `OffsetEntry::start()`) can show 0-exec in the combined
report — the classification above already accounts for that.
- The coverage holes fixed this session (C1-C3, L1, L2) were chosen
because each is load-bearing *and* one-test-cheap; the remaining
uncovered mass is dominated by N2a/N3a/N4a, which are documented
rather than forced.
- Static test counts at the review-#007 commits: 477 (F1/F2),
481 (C1-C3), 488 (L1/L2) — +14 net from the pre-audit 474.
- Post-fix coverage (same tool, full run): TOTAL **91.66% lines /
87.64% functions** (from 90.67/86.32). Per-file movement:
`builder.rs` 91.28→99.33, `engine.rs` 96.44→96.76,
`materialize.rs` 85.72→87.74, `read_plan.rs` 90.69→90.94,
`tunion.rs` 92.02→92.86. The remaining mass is the documented
N2a/N3a/N4a classes.
-296
View File
@@ -1,296 +0,0 @@
---
status: resolved (F1, F2 fixed 2026-09-07; N1, N2, N3 classified; N4 fixed 2026-09-07)
last_updated: 2026-09-07
reviewed_artifacts:
- src/read_plan.rs
- src/sequential_reader.rs
- src/materialize.rs
- src/bast.rs
- benches/wire_vs_bast.rs
- docs/reviews/007-coverage-audit.md (N3a disposition)
tool: manual diff review of post-#007 commits (dea96f0, d4635d2) + disposable probe tests (run in-session, then deleted) + cargo bench + counting-allocator peak-RSS probe
reviewer: pre-publish review #008 (session request — audit the two post-#007 perf/bench commits, then gate the 0.3.0 publish)
---
# Review #008 — Pre-Publish Review: Post-#007 Perf Commits
## Purpose
0.3.0's release commit (`9949f91`) predates review #006 entirely; the
fix sessions for #006 and #007 landed eleven more commits, and *after*
#007 closed, two more commits landed unreviewed: `dea96f0` (bench port
from alktty) and `d4635d2` (the perf commit — fixed-size struct fast
path, integer union dispatch, `read_next_borrowed`). The perf commit
touches the flagship packed read path, which every untrusted wire
buffer flows through. This review audits those two commits before the
first crates.io publish of the 0.3.x line (0.1.0 and 0.2.0 are
published; 0.3.0 never was — every post-release fix can legally ride
inside the first published 0.3.0, no semver conflict).
It also disposes of review #007's N3a — the one finding explicitly
deferred to "the pre-release review", which this session is.
## Methodology
- Full diff read of `d4635d2` (perf) and `dea96f0` (bench port),
cross-checked against the invariants the earlier reviews established:
cross-consumer dispatch agreement (#006 H3, #007 F1), the H1
no-count-sized-prealloc rule, and the N2/F2 dual-layer cap pattern.
- Disposable probe tests (`tests/zzz_probe*.rs`, deleted after the
session; none was a crash hazard) to confirm/deny the three
behaviors code reading flagged: the `int_keys` non-canonical-key
divergence, the engine's nested-array acceptance envelope, and the
`with_capacity` amplification.
- A counting-`GlobalAlloc` probe (peak-bytes metric) to measure the
worst-case simultaneous allocation of the amplification shape
precisely — RSS timing proved too noisy to separate the two test
cases.
- `cargo bench --quick` before/after the fixes to confirm the perf
commit's wins survive.
- N3a dispositions probed where cheap (inline-struct union variants).
## Baseline
Audited at `main` HEAD `d4635d2`, 0.3.0, working tree clean. 566
tests green (488 lib + 17 + 34 + 15 + 12, + 2 ignored doctests),
clippy `-D warnings` clean, wasm build green — per the perf commit's
verification block.
## Summary Statistics
| Severity | Count | Status |
|----------|------:|--------|
| High | 0 | — |
| Medium | 2 (F1, F2) | both fixed 2026-09-07 |
| Info | 3 (N1, N2, N3) | classified |
| Fix | 1 (N4) | fixed 2026-09-07 |
No Highs: both Mediums are probe-verified cross-consumer divergences
and resource-bound violations, but neither aborts the process
(`with_capacity` is now bounded per array by `MAX_ARRAY_ELEMENTS`, so
H1's 1 TB SIGABRT class does not return). Both were fixed in-session
because they violate AGENTS.md §3 (divergent verdicts on untrusted
input; unbounded-count-shaped allocation) — the publish gate.
**Resolution log:**
- **F1 + F2 (2026-09-07):** fixed in one commit — see the resolution
blocks. 569 tests green (491 lib + 17 + 34 + 15 + 12, + 2 ignored),
clippy `-D warnings` clean, doc 0 warnings, wasm green.
- **N4 (2026-09-07):** fixed with F1/F2 — see the block.
---
## Findings
### F1. `int_keys` integer dispatch breaks cross-consumer agreement on non-canonical mapping keys
**Files**: `src/read_plan.rs` (`compile_int_keys`, introduced by
`d4635d2`), contrast `src/materialize.rs` (`materialize_plan_union`'s
byte-disc arm — stringifies), `src/validation_plan.rs`
(`validate_union_numeric` — stringifies), `src/tunion.rs`
(`read_byte_discriminator` — stringifies)
**Problem**: `d4635d2` added a pre-parsed `(u64, variant_index)`
dispatch table for byte-discriminator unions: when every mapping key
parses as `u64`, the reader matches the raw discriminator integer
instead of stringifying per read. But the meta-schema does not
constrain mapping-key shape beyond "object property name", and
`key.parse::<u64>()` accepts **non-canonical** decimal strings:
```
PROBE1 reader: field=msg disc="01" (DISPATCHED)
PROBE1 validate_bytes: Err(access error at msg: union discriminator value 1 not in mapping)
```
With mapping key `"01"` (and discriminator `1` on the wire): the
reader's numeric dispatch **matches** (`"01".parse::<u64>() == 1`) and
dispatches — returning `discriminator == "01"` — while the
materializer (`1.to_string() == "1" ≠ "01"`), the validation plan, and
tunion all **reject** the identical buffer. Pre-`d4635d2`, all four
consumers stringified and all four rejected — agreement held (both
verdicts "reject", same error class). The perf commit flipped the
reader to accept-while-everyone-else-rejects: the exact
cross-consumer-divergence shape #006 H3 and #007 F1 exist for, on the
flagship path. `"+1"` parses as `u64` too (Rust's `from_str_radix`
accepts a leading `+`) — same class. The returned key string also
became schema-quirk-dependent: the reader reports `"01"` where the
materializer's `__discriminator` for a *matched* key would report the
stringified form.
**Not a #007 regression**: `int_keys` did not exist before `d4635d2`.
But `d4635d2` postdates #007's close and was unreviewed — this is the
audit catching it.
**Fix**: build the integer table only from **canonical** keys — a key
qualifies iff `key.parse::<u64>()` succeeds *and*
`parsed.to_string() == key` (i.e. the key is exactly what
stringification would produce). Any non-canonical or non-numeric key
falls back to the string path (`int_keys: None`), which every consumer
already agrees on. No accepted schema's *reachable* behavior changed:
for fully-canonical mappings the numeric dispatch behaves identically
to stringified matching (the numeric value's `to_string()` equals the
key), and for non-canonical keys all consumers now reject exactly as
before `d4635d2`. The perf win (no per-read stringify) is preserved
for every mapping that was unambiguous to begin with.
**Resolution (2026-09-07):** exactly that — `compile_int_keys` now
requires `v.to_string() == *key` for the table to carry the entry;
any miss returns `Ok(None)` (string fallback). Doc comment states the
canonicality rule and why. Tests in `read_plan.rs`:
`r8_non_canonical_mapping_key_disables_int_dispatch` (key `"01"` →
`int_keys` is `None`) and `r8_canonical_mapping_keys_keep_int_dispatch`
(keys `"1"`,`"2"` → table `[(1,0),(2,1)]`). Probe output after the fix:
both `validate_bytes` and the reader reject disc 1 under key `"01"`
with the same error class — agreement restored.
### F2. `Vec::with_capacity(count)` reintroduced on both array materializers — ~477 MB simultaneous allocation from a ~1 KB schema
**Files**: `src/materialize.rs:247` (`materialize_plan_array` — the
`validate_bytes` packed path), `src/materialize.rs:650`
(`materialize_array_packed` — the legacy walker, reachable via the
aligned record arm)
**Problem**: `d4635d2`'s "materialize: with_capacity for bytes arrays,
arrays, and struct objects" item reintroduced
`Vec::with_capacity(count)` at two of the three sites H1's layer-1 fix
had converted to `Vec::new()` + push. `count` is now compile-capped at
`MAX_ARRAY_ELEMENTS` (2^16), so H1's 1 TB `SIGABRT` does not return —
but the per-array cap does not bound *nesting*:
```
PROBE validate peak bytes allocated simultaneously: 476780249
```
A schema of one 100-level nested array chain (each `count: 65535`,
innermost elements empty structs — all legal: the depth cap is 128 and
stride-0 chains evade `MAX_ARRAY_BYTES`, which only checks stride
products) peaks at **~477 MB of simultaneous allocation** on
`validate_bytes(&[])` from a ~1 KB schema and an *empty* buffer. Each
level's `with_capacity(65535 × sizeof(Value))` stays live across its
element walk, so the sizes multiply across ~127 legal depth levels
(the innermost zero-progress rejection fires only after the whole
chain has descended). On wasm32 — which this crate explicitly targets
— the same shape aborts the wasm heap well below 477 MB. H1's layer-1
rule ("no count-sized prealloc on untrusted input; the per-element
walk dominates") is exactly the invariant this violates; the perf
commit's own bench evidence doesn't need the prealloc either (see
below).
The other `with_capacity` additions in the commit are fine: byte-array
capacity from `b.len()` (a read slice), struct-object capacity from
`plan.fields().len()`, and the plan-compiler's from `fields.len()` —
all bounded by data/plan already in hand, not by declared counts.
**Fix**: restore H1's layer-1 shape at both sites — `Vec::new()` +
push loop (the loops already push `count` elements; the zero-progress
guard bounds honest progress per element). Optionally cap
preallocation at a small constant, but plain `Vec::new()` matches H1's
shipped behavior.
**Resolution (2026-09-07):** both sites restored to `Vec::new()` +
push. Bench before/after (criterion `--quick`, this session): packet
read 246 → 220 µs, chunk read 76/68 µs — the revert costs nothing
measurable on the bench shapes (small arrays; the materializer's
per-element work dominates), and the union/struct preallocs stay.
Locking test in `materialize.rs`:
`r8_deeply_nested_stride0_array_rejects_before_bulk_prealloc` (the
100-level chain still rejects cleanly with the zero-progress error at
the innermost level; the allocation shape itself is documented here —
in-tree cannot cheaply assert peak allocation, and the #008 probe was
deleted per the no-reproducer rule).
### N1. `d4635d2`'s fixed-size fast paths are sound (classified, no action)
**Files**: `src/read_plan.rs` (`fixed_size`, `fixed_plan_size`), `src/sequential_reader.rs`
The fixed-size struct fast path replaces the cursor size walk with one
bounds check; `fixed_plan_size` already returned `Result<Option>` with
clean overflow errors (L1's shape), and every new error arm formats
paths lazily on the error path only. The union-variant fast path
(`plan_variant_fixed_size`) applies only to struct variants and checks
bounds before use. No issue found.
### N2. Bench port (`dea96f0`) is methodology-honest (classified, no action)
**Files**: `benches/wire_vs_bast.rs`
The port drops alktty's async I/O group (correctly — it measured a
different stack) and adds a parity check before measurement so the
stream loop can't drift. The historical `read_chunk_stream` numbers
stay comparable by construction. No issue found.
### N3. N3a dispositions (review #007's deferred items)
**Files**: `src/bast.rs` (`source()` accessors, `BastField::synthetic`,
`resolve_typeref_as_def`'s inline arms)
- **`source()` accessors (7 sites)**: public API on `BastStruct`/
`BastUnion`/`BastField`/etc. Removal is a semver decision and
AGENTS.md's semver exception requires an explicit ask — **kept**.
They are one-line accessors over parsed source nodes, harmless, and
plausibly useful to downstream codegen (the announced consumer).
- **`BastField::synthetic` (`#[allow(dead_code)]`, zero callers)**:
`pub(crate)`, not public API — **deleted** (2026-09-07). No semver
impact; the `#[allow(dead_code)]` suppression is gone with it.
- **`resolve_typeref_as_def`'s inline struct/union/enum arms**: the
review-#007 suspicion ("plausibly dead after H3") was wrong —
probe-verified reachable: the meta-schema's
`mapping.additionalProperties: TypeRef` accepts inline struct
variants, and `LayoutBuilder`'s byte-disc and field-disc arms call
`resolve_typeref_as_def` on every union variant. The H3 parse rules
forbid variant *re-declaration of shared fields*, not inline variant
bodies. **Kept**, reachable.
### N4. Stale test-count references in review #006's resolution log
**Files**: `docs/reviews/006-implementation-review-030.md`
The bookkeeping note ("static count at `2eb086f` is 542 + 2 ignored")
and per-commit counts are accurate as written; no fix needed. Recorded
here so the review trail stays honest about what was re-checked
during this session's doc sweep. **Resolution (2026-09-07):** no code
change; superseded the "Fix" entry — this is the classification
record.
---
## What's Good
- The perf commit's core ideas are sound and survived review: the
compile-time `fixed_size` cache is computed through the existing
`Result`-returning sizer (no `unwrap_or_default` regression), and
the int-dispatch table's design was right — it just needed the
canonicality gate.
- The counting-allocator probe took 15 minutes and converted a
"probably too big" into a precise number (476,780,249 bytes) — the
same probe pattern the earlier reviews used, applied to allocation
instead of verdicts.
- `cargo bench --quick` before/after the fixes is the right tool for
guarding perf-fix reverts: packet read 220 µs post-fix vs 246 µs
baseline confirms the `Vec::new()` restore is free.
## Recommended Order
1. ~~**F1** — canonical-key gate on `compile_int_keys`~~ **fixed
2026-09-07**.
2. ~~**F2** — restore H1's no-prealloc rule at both array sites~~
**fixed 2026-09-07**.
3. ~~**N4** — `BastField::synthetic` deletion~~ **fixed 2026-09-07**.
4. **N3 source() accessors** — revisit only if/when the codegen
consumer confirms it does not want them (removal needs an explicit
ask per AGENTS.md).
## Notes
- Probe tests were run as `tests/zzz_probe*.rs` in-tree during the
session and deleted before any commit (the #006 pattern). None was a
crash hazard; the amplification probe allocates ~477 MB transiently
and completes in ~40 ms.
- Benches are not run in CI and are excluded from the publish (the
`[bench]` target ships — that is fine; benches don't affect the
library's API or its wasm compatibility).
- The 0.3.0 publish proceeds after these fixes: 0.1.0 and 0.2.0 are
on crates.io; this is the first 0.3.0 publish, so F1/F2's
behavior changes (both "previously-divergent, now-agreed" shapes)
land inside the version's first release — no semver bump implied.
-1049
View File
File diff suppressed because it is too large. Load diff
-52
View File
@@ -1,52 +0,0 @@
[package]
name = "alktype-fuzz"
version = "0.0.0"
publish = false
edition = "2021"
[package.metadata]
cargo-fuzz = true
[dependencies]
libfuzzer-sys = "0.4"
alktype-fuzz-shared = { path = "shared" }
[dependencies.alktype]
path = ".."
[[bin]]
name = "bast_compile"
path = "fuzz_targets/bast_compile.rs"
test = false
doc = false
bench = false
[[bin]]
name = "data_access"
path = "fuzz_targets/data_access.rs"
test = false
doc = false
bench = false
[[bin]]
name = "read_opseq"
path = "fuzz_targets/read_opseq.rs"
test = false
doc = false
bench = false
[[bin]]
name = "layout_build"
path = "fuzz_targets/layout_build.rs"
test = false
doc = false
bench = false
[[bin]]
name = "validate_pair"
path = "fuzz_targets/validate_pair.rs"
test = false
doc = false
bench = false
[workspace]
-67
View File
@@ -1,67 +0,0 @@
# alktype fuzzing
cargo-fuzz targets for the binary struct engine's untrusted-input
surfaces. The design and operating rules live in
`docs/plans/fuzzing.md` (adopted from alkhttp's
`docs/plans/fuzzing.md`; rationale in alkcall's
`docs/research/fuzzing.md`) — this README is the operational
cheat-sheet.
## Layout
- `fuzz_targets/` — nightly-only `fuzz_target!` binaries (thin wrappers).
- `shared/` — stable-toolchain library holding the invariant logic; the
corpus replay tests run here on plain `cargo test`.
- `corpus/<target>/` — committed seeds (regenerate with
`python3 fuzz/gen_fuzz_seeds.py`).
- `artifacts/` — gitignored crash/oom/timeout artifacts + campaign logs.
## Targets
| Target | Drives |
|---|---|
| `bast_compile` | `AlkTypeEngine::compile` in both layout modes over attacker-shaped BAST JSON (the whole schema side through one choke point) + `validate_bast_doc` + `build_validator` |
| `data_access` | the hand-rolled decode core (`src/data_access.rs`): fixed-width kinds, bool strictness, length-prefixed and indirect string/bytes, enums — over raw bytes with attacker-chosen offsets and endianness |
| `read_opseq` | the stateful `SequentialReader` (packed read side): op sequences (Next / NextBorrowed / Field / Reset / End) over hostile buffers under the compiled plan — cursor discipline, plan-order walks, the record-count spin bound, ADR-007 reader independence |
| `layout_build` | the packed write side (`LayoutBuilder::build`) with adversarial `var_sizes` over a five-schema menu — position disjointness/bounds, failed-write buffer-untouched contracts, write→read pair round trip |
## Running a campaign — always detached
Agent sessions must never run fuzzing in the foreground (an OOM in a
target can take down the session host; see docs/plans/fuzzing.md §2).
Use the detached runner:
```bash
fuzz/run-detached.sh bast_compile
# poll:
tail -n 50 fuzz/artifacts/bast_compile-*.log
ls fuzz/artifacts/bast_compile/
pgrep -f "cargo fuzz run bast_compile"
```
`FUZZ_RUNTIME_SECS=1800 fuzz/run-detached.sh bast_compile` for a longer
campaign. The runner pins `-fork=1 -rss_limit_mb=2048
-malloc_limit_mb=2048 -timeout=25` and detaches via `setsid` + `nohup`.
## Corpus replay (the standing fuzz gate)
```bash
cargo test --manifest-path fuzz/shared/Cargo.toml
```
replays every committed seed through the same invariant functions the
fuzz targets run — on stable, without nightly, no cargo-fuzz. Part of
the release verification checklist (AGENTS.md).
## Toolchain
`fuzz/rust-toolchain.toml` pins nightly (+ `llvm-tools-preview`) for
this subtree only; the main crate stays stable at MSRV 1.85. `cargo
fuzz build` works from any CWD inside `fuzz/` (rustup resolves the
toolchain per directory). Build:
```bash
cd fuzz && cargo fuzz build
# or from the repo root — the toolchain file is picked up by path:
cargo fuzz build -D
```
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}]}, "root": "S"}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "endian": "big", "fields": [{"name": "x", "kind": "uint8"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": []}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "f0", "kind": "int8"}, {"name": "f1", "kind": "int16"}, {"name": "f2", "kind": "int32"}, {"name": "f3", "kind": "int64"}, {"name": "f4", "kind": "uint8"}, {"name": "f5", "kind": "uint16"}, {"name": "f6", "kind": "uint32"}, {"name": "f7", "kind": "uint64"}, {"name": "f8", "kind": "float32"}, {"name": "f9", "kind": "float64"}, {"name": "f10", "kind": "bool"}, {"name": "f11", "kind": "string"}, {"name": "f12", "kind": "bytes"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "endian": "little", "fields": [{"name": "a", "kind": "uint32", "endian": "big"}, {"name": "b", "kind": "string", "encoding": "length-prefixed", "maxLength": 64}, {"name": "c", "kind": "bytes", "encoding": "offset-indirect", "maxLength": 128}, {"name": "d", "kind": "string", "encoding": "offset-indirect"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "child", "kind": {"$ref": "#/$defs/Nested"}}, {"name": "arr", "kind": {"kind": "array", "element": "uint32", "count": 3}}, {"name": "arr0", "kind": {"kind": "array", "element": "uint8", "count": 0}}, {"name": "rec", "kind": {"kind": "record", "values": "string"}}, {"name": "en", "kind": {"$ref": "#/$defs/E"}}, {"name": "un", "kind": {"$ref": "#/$defs/U1"}}, {"name": "un2", "kind": {"$ref": "#/$defs/U2"}}]}, "Nested": {"kind": "struct", "fields": [{"name": "y", "kind": "int16"}]}, "E": {"kind": "enum", "values": ["a", "b", "c"]}, "U1": {"kind": "union", "discriminator": {"kind": "byte", "offset": 0, "type": "uint8"}, "mapping": {"0": "uint8", "1": {"$ref": "#/$defs/Nested"}}}, "U2": {"kind": "union", "discriminator": {"kind": "field", "name": "tag"}, "fields": [{"name": "tag", "kind": "uint32"}], "mapping": {"0": "uint8", "7": "string"}}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "inl", "kind": {"kind": "struct", "fields": [{"name": "z", "kind": "uint8"}]}}, {"name": "inle", "kind": {"kind": "enum", "values": ["x"]}}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "l", "kind": {"$ref": "#/$defs/N"}}, {"name": "r", "kind": {"$ref": "#/$defs/N"}}]}, "N": {"kind": "struct", "fields": [{"name": "v", "kind": "uint8"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}], "align": 4096}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": "uint32", "align": 16}, {"name": "b", "kind": "uint8", "align": 1}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "s", "kind": "string", "maxLength": 67108864}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "s", "kind": "string", "maxLength": 0}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "align": 4097, "fields": []}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "align": 65536, "fields": []}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "align": 18446744073709551615, "fields": []}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"kind": "array", "element": "uint8", "count": 65537}}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"kind": "array", "element": "uint8", "count": 18446744073709551615}}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "s", "kind": "string", "maxLength": 67108865}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "nonsense"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "uint8", "fields": []}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct"}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"$ref": "#/$defs/S"}}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": {"$ref": "#/$defs/Missing"}}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {}}
-1
View File
@@ -1 +0,0 @@
[1, 2, 3]
-1
View File
@@ -1 +0,0 @@
{}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": 123}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "9bad", "kind": "uint8"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": "uint8", "bogus": true}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "union", "discriminator": {"kind": "byte", "offset": 0, "type": "uint8"}}}}
-1
View File
@@ -1 +0,0 @@
not json at all
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S":
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n",Line truncated
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n", "kind": {"kind": "struct", "fields": [{"name": "n",Line truncated
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "x", "kind": "uint8"}]}, "D100": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/S"}}]}, "D99": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D100"}}]}, "D98": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D99"}}]}, "D97": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D98"}}]}, "D96": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D97"}}]}, "D95": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D96"}}]}, "D94": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D95"}}]}, "D93": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D94"}}]}, "D92": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D93"}}]}, "D91": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D92"}}]}, "D90": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D91"}}]}, "D89": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D90"}}]}, "D88": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D89"}}]}, "D87": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D88"}}]}, "D86": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D87"}}]}, "D85": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D86"}}]}, "D84": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D85"}}]}, "D83": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D84"}}]}, "D82": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D83"}}]}, "D81": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D82"}}]}, "D80": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D81"}}]}, "D79": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D80"}}]}, "D78": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D79"}}]}, "D77": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D78"}}]}, "D76": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D77"}}]}, "D75": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D76"}}]}, "D74": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D75"}}]}, "D73": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D74"}}]}, "D72": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D73"}}]}, "D71": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D72"}}]}, "D70": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D71"}}]}, "D69": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D70"}}]}, "D68": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D69"}}]}, "D67": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D68"}}]}, "D66": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D67"}}]}, "D65": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D66"}}]}, "D64": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D65"}}]}, "D63": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D64"}}]}, "D62": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D63"}}]}, "D61": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D62"}}]}, "D60": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D61"}}]}, "D59": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D60"}}]}, "D58": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D59"}}]}, "D57": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D58"}}]}, "D56": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D57"}}]}, "D55": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D56"}}]}, "D54": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D55"}}]}, "D53": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D54"}}]}, "D52": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D53"}}]}, "D51": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D52"}}]}, "D50": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D51"}}]}, "D49": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D50"}}]}, "D48": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D49"}}]}, "D47": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D48"}}]}, "D46": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D47"}}]}, "D45": {"kind": "struct", "fields": [{"name": "n", "kind": {"$ref": "#/$defs/D46"}}]}, "D44": {"kind": "struct", "fields": [{"name": "Line truncated
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "struct", "fields": [{"name": "a", "kind": "uint8"}, {"name": "a", "kind": "uint16"}]}}}
-1
View File
@@ -1 +0,0 @@
{"$defs": {"S": {"kind": "enum", "values": ["only"]}, "root": "S"}}
-1
View File
@@ -1 +0,0 @@

Binary file not shown.
-1
View File
@@ -1 +0,0 @@

-1
View File
@@ -1 +0,0 @@
�
-1
View File
@@ -1 +0,0 @@
�
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
-1
View File
@@ -1 +0,0 @@
4
-1
View File
@@ -1 +0,0 @@
4
Binary file not shown.
Binary file not shown.
-1
View File
@@ -1 +0,0 @@
��
-1
View File
@@ -1 +0,0 @@
��
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
-1
View File
@@ -1 +0,0 @@
D3"
-1
View File
@@ -1 +0,0 @@
"3D
Binary file not shown.
Binary file not shown.
-1
View File
@@ -1 +0,0 @@
����
-1
View File
@@ -1 +0,0 @@
����
-1
View File
@@ -1 +0,0 @@
�fUD3"
-1
View File
@@ -1 +0,0 @@
"3DUfwˆ
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Loaded 100 of 321 files, more files were not shown because too many files have changed in this diff. Show more